Source measurement unit acceleration method and system for bidirectional circulation attention transformation
By adopting the acceleration method of bidirectional cyclic attention transformation in the source measurement unit, integrating the BiRNN layer and the attention discrete cosine transformation module, the problem of measurement speed drop in high-precision measurement is solved, and more efficient measurement performance is achieved.
Patent Information
- Application Number
- CN202510514570.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing source measurement unit acceleration methods have problems such as decreasing measurement speed, low measurement accuracy and high system complexity in high accuracy measurement.
Using the acceleration method of bidirectional cyclic attention transformation, the encoder structure is improved, the time-domain and frequency-domain features are extracted, and data prediction is achieved through the decoder integration.
While ensuring measurement accuracy, the working efficiency and performance of the source measurement unit are significantly improved, and half of the measurement time can be saved without increasing the amount of data.
Smart Images

Figure CN120030394A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a source measurement unit acceleration method and system for bidirectional cyclic attention transformation, belonging to the technical field of time series signal trend prediction. Background Art
[0002] With the rapid development of the semiconductor industry, the demand for high-speed testing in large-scale production of semiconductor devices and integrated circuit production lines continues to grow. As one of the key tools for semiconductor device performance testing and integrated circuit testing, source measurement units (SMUs) play a vital role in high-precision transient response test scenarios. In high-precision measurement scenarios, multiple measurements and averaging are usually required to improve measurement accuracy, but this will reduce measurement speed.
[0003] There are many source measurement acceleration methods, but they all have some shortcomings. The first is the hardware parallelization acceleration method. The integrated multi-channel SMU system increases the multi-channel SMU hardware or adopts a parallel measurement architecture to perform multiple measurement tasks at the same time to improve the overall throughput. However, the timing synchronization and signal interference between channels during parallel measurement may affect the measurement accuracy, especially in high-speed transient testing. The second is fast sampling and data compression technology. By increasing the sampling rate or using data compression algorithms (such as wavelet transform and sparse representation) to reduce the amount of transmitted data, the measurement time is shortened. However, the compression algorithm may lose high-frequency details or introduce noise, affecting the reliability of high-precision testing. In addition, high-speed sampling has extremely high performance requirements for ADC and memory, which may increase the complexity of the system. The third is the prediction method based on traditional machine learning. Use historical data to train regression models (such as support vector machines and random forests) to predict subsequent measurement values and reduce the actual number of measurements. However, traditional machine learning has limited fitting capabilities for nonlinear and high-noise time series, especially in semiconductor testing, where device differences may cause model failure. In addition, traditional algorithms have limited processing capabilities for non-stationary signals (such as transient responses) and lack real-time adjustment capabilities. These limitations highlight the need for new acceleration methods to enable more efficient SMU measurements without sacrificing accuracy. Summary of the invention
[0004] In view of the deficiencies of the prior art, the present invention provides a source measurement unit acceleration method and system for bidirectional cyclic attention transformation; The present invention solves the problem of measurement speed reduction in high-precision measurement. On the one hand, the algorithm improves the encoder structure. Traditional time series prediction models may have limitations in processing long-term dependencies and trend extraction. BiRNN is an extended version of recurrent neural network (RNN) and consists of two directional RNNs. One forward RNN processes data from the beginning to the end of the sequence, while the other reverse RNN processes data from the end to the beginning of the sequence. Finally, the outputs of the two directions are merged at each time step. The present invention integrates BiRNN into the encoder structure, and the algorithm can process input sequences in both positive and negative directions at the same time, thereby more effectively capturing the bidirectional contextual information of the data and significantly enhancing the model's ability to capture time series trends. This improvement enables the model to more accurately grasp the overall trend of data changes during the prediction process, thereby improving the prediction accuracy. On the other hand, an attention discrete cosine transform (ADCT) module is introduced between the encoder and the decoder to convert the time domain signal into a frequency domain representation. This not only reveals the spectral characteristics of the signal, but also reduces data redundancy and improves the efficiency of subsequent processing by combining the attention mechanism. Finally, the algorithm performance is analyzed by analyzing the output characteristic curves of loads of different properties. The prediction method and the combined measurement and prediction method proposed in the present invention can save half of the measurement time while ensuring that the amount of data obtained remains unchanged by combining measurement and prediction, thereby effectively improving the working efficiency and performance of the source measurement unit.
[0005] To achieve the above object, the present invention provides the following technical solutions: A source measurement unit acceleration method and system for bidirectional cyclic attention transformation, comprising: The input data is preprocessed and then input into the trained acceleration model for acceleration processing; The acceleration model includes an encoder, an attention discrete cosine transform module (ADCT block) and a decoder; the BiRNN layer is integrated into the encoder, and an attention discrete cosine transform module (ADCT block) is introduced between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
[0006] Preferably, according to the present invention, preprocessing refers to embedding and encoding.
[0007] According to the preferred embodiment of the present invention, the input data Embed and encode the input data Convert to size Vector ; The calculation formula is as follows (1): (1);
[0008] The parameter α acts as a balancing factor to adjust the size between scalar projection and local / global embedding. , is a learnable global timestamp embedding, ,and, , is the embedding dimension of the acceleration model, which indicates the dimension of the high-dimensional space to which the time series data is mapped; is the length of the input sequence, which represents the number of time steps of the time series data; Represents a one-dimensional convolution operation, extracting local features from the input sequence and generating a high-dimensional representation; is the number of components into which the global timestamp is embedded.
[0009] According to the preferred embodiment of the present invention, the preprocessed data is processed by a BiRNN layer, and the BiRNN layer is The signal trend information is extracted from the , which is then used to calculate the self-attention weight; as shown in formula (2): (2);
[0010] in, The BiRNN layer is Extracted signal trend information, is the feature dimension, is a value vector containing the actual feature information, Represents the output features weighted by the attention weights.
[0011] First, the self-attention mechanism calculates the correlation between elements in the input sequence; then, it generates attention weights; and then, it uses the attention weights to weight the values ( ); Finally, the final output value is obtained by summing and weighting.
[0012] Preferably, according to the present invention, frequency domain features are extracted by an attention discrete cosine transform module; comprising: For time series , ,in, , time series The definition of one-dimensional attention discrete cosine transform is shown in formula (3): (3);
[0013] in, ; First, the ADCT block divides the input feature map, i.e., the multi-channel temporal feature matrix generated by the encoder, into several subsets along the channel dimension, as shown below: (4);
[0014] Among them, each subset represents a part of a channel, , ,n is the number of channels N; The multi-head attention weight matrix is multiplied by the frequency domain feature matrix to obtain the final sparse frequency domain feature matrix; the multi-head attention weight matrix refers to the dynamic weight matrix generated by the multi-head attention mechanism, denoted as ; The frequency domain feature matrix refers to the frequency domain representation generated by the ADCT module through discrete cosine transform (DCT), denoted as ; The sparse frequency domain feature matrix is calculated according to formulas (5) and (6): (5); (6);
[0015] in, represents the output of the encoder, yes The discrete cosine transform of Represents the multi-head attention weight parameter matrix; represents the number of attention heads, and the double sum represents the weighted feature accumulation of all heads; is the sparse frequency domain feature (after attention filtering), is the frequency domain feature matrix (DCT output); Specifically, for each head , first, calculate the weighted features of all positions , then, the weighted features of all positions Add; the accumulated features are passed through the fully connected layer ( ) to further extract features; then, through the normalization layer ( ) Normalize the features; Finally, the encoder output is fused with the time domain features and frequency domain features through the cross attention layer to obtain the final output; the final fused features are calculated as follows: (7); (8); (9); (10); (11);
[0016] Formula (7) converts the time domain features output by the encoder into a query vector, which is used to retrieve the key features of the frequency domain information; Formulas (8) and (9) respectively perform linear transformations on the frequency domain features generated by the ADCT module to generate a key vector and a value vector, where the key vector is used as a retrieval index and the value vector retains the original information in the frequency domain; Formula (10) generates the attention distribution of the time domain to the frequency domain through the similarity matrix between the query vector and the key vector, combined with the scaling factor to prevent gradient explosion, and dynamically allocates the dependence weights of each time point on different frequency components; Formula (11) uses the attention weight to weight the sum of the value vector, and outputs the fused features through the projection matrix, which has both time domain dynamics and frequency domain globality; in, is through the projection matrix Processing input sequence The resulting value matrix is ; and are respectively through the projection matrix and Processing sequence and After getting the query and key matrix, ; is the projection matrix, which is used to transform the input sequence Convert to Value Matrix , ; is the projection matrix, which is used to transform the input sequence Convert to Key Matrix , ; is the projection matrix, which is used to transform the input sequence Convert to Query Matrix , .
[0017] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, a source measurement unit acceleration method and system steps of a bidirectional cyclic attention transformation are implemented.
[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a source measurement unit acceleration method and system for bidirectional cyclic attention transformation.
[0019] A source measurement unit acceleration system for a bidirectional recurrent attention transform Informer, comprising: The data preprocessing module is configured to: perform data preprocessing on the input; The acceleration module is configured to: pre-process the input data and input it into the trained acceleration model for acceleration processing; the acceleration model includes an encoder, an attention discrete cosine transform module (ADCT block) and a decoder; integrate the BiRNN layer into the encoder, and introduce an attention discrete cosine transform module (ADCT block) between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. In the present invention, the encoder structure is improved based on the BiRNN algorithm. Traditional time series prediction models may have limitations in processing long-term dependencies and trend extraction. Compared with traditional unidirectional RNNs, BiRNN can more effectively process data with forward and backward dependencies. The present invention integrates BiRNN into the encoder structure. The algorithm can process input sequences in both positive and negative directions at the same time, effectively limiting the bidirectional contextual information of the data, and greatly enhancing the model's ability to capture time series trends. The algorithm can process input sequences in both positive and negative directions at the same time, thereby more effectively capturing the bidirectional contextual information of the data and significantly enhancing the model's ability to capture time series trends. This improvement enables the model to more accurately grasp the overall trend of data changes during the prediction process, thereby improving prediction accuracy.
[0021] 2. In the present invention, an attention discrete cosine transform (ADCT) module is introduced between the encoder and the decoder to convert the time domain signal into a frequency domain representation. This can not only reveal the spectral characteristics of the signal, but also reduce data redundancy and improve the efficiency of subsequent processing by combining the attention mechanism. Finally, the algorithm performance is analyzed by analyzing the output characteristic curves of loads of different properties. The prediction algorithm and the method combining measurement and prediction proposed in the present invention combine measurement and prediction while ensuring that the same amount of data is obtained, saving half of the measurement time.
[0022] 3. The present invention is applicable to various application scenarios of semiconductor device performance testing and integrated circuit testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic diagram of a source measurement unit acceleration method and system for bidirectional recurrent attention transformation; Figure 2 A schematic diagram of the workflow of a source measurement unit acceleration method and system for bidirectional recurrent attention transformation; Figure 3 Schematic diagram of the improved encoder (left) and ADCT block (right); Figure 4This is a schematic diagram of the overall architecture of the acceleration model; Figure 5 This is a comparison chart of various indicators of different models on the load data set Resistor; Figure 6 This is a comparison chart of various indicators of different models on the load data set Capacitor; Figure 7 This is a comparison chart of various indicators of different models on the load data set Inductor; Figure 8 The following are the prediction results under three load data sets. DETAILED DESCRIPTION
[0024] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0025] Terminology explanation: 1. Autoformer: An improved Transformer model based on the autocorrelation mechanism, focusing on long-term time series prediction. It replaces the traditional Transformer's self-attention with the auto-correlation mechanism to mine the periodic patterns of the sequence and improve prediction efficiency.
[0026] 2. LightTs: A lightweight time series prediction model that aims to reduce computational cost while maintaining high accuracy. It optimizes computational efficiency through hierarchical feature extraction and sparse attention.
[0027] 3. MSGNet: A multi-scale time series prediction model that captures dependencies of different time granularities through gating mechanism and multi-scale feature fusion.
[0028] 4. Transformer: The classic Transformer is based on the self-attention mechanism, which captures long-range dependencies by calculating the associated weights of all positions in the sequence.
[0029] 5. Informer: Informer (Improved Transformer) is a Transformer variant optimized for long sequence prediction. It proposes Prob sparse self-attention and distillation encoder to significantly reduce computational overhead.
[0030] 6. ETSformer: ETSformer combines the classic time series decomposition method (ETS: Error-Trend-Seasonality) with Transformer to improve prediction performance through interpretable time series decomposition.
[0031] 7. FEDformer: FEDformer (Frequency Enhanced Decomposed Transformer) improves prediction efficiency through frequency domain attention and hybrid decomposition, and is good at capturing frequency domain features.
[0032] Example 1 A source measurement unit acceleration method and system for bidirectional cyclic attention transformation, comprising: After preprocessing, the input data is input into the trained acceleration model for acceleration processing; if the input data is in units of 1PLC, the first half cycle of 1PLC, i.e., the first 10 sample points, can be used as input data; like Figure 4 As shown in FIG. 1 , the acceleration model includes an encoder, an attention discrete cosine transform module, namely an ADCT block, and a decoder; the BiRNN layer is integrated into the encoder, and an attention discrete cosine transform module, namely an ADCT block, is introduced between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
[0033] See also Figure 1 , the present invention proposes an acceleration model of source measurement unit combining bidirectional recurrent neural network and attention discrete cosine transform Informer, which is specially used to optimize SMU (source measurement unit) measurement algorithm. This framework is based on Transformer, and significantly improves processing efficiency and measurement accuracy by integrating bidirectional recurrent neural network (BiRNN) and introducing attention discrete cosine transform (ADCT) module. Especially when processing complex time series data, this architecture performs well and can more effectively capture long-term dependencies and trend changes. Figure 1 The overall structure of the framework is shown.
[0034] Extract time domain features through encoder; including: The encoder's time-domain feature extraction achieves robust modeling of signal trends through multi-stage collaborative processing.
[0035] First, the input sample points are passed through a one-dimensional convolutional layer to extract local transient features, and the learnable timestamp embedding is synchronously fused to encode millisecond-level sampling timing information. The bidirectional recurrent neural network (BiRNN) then performs bidirectional sequence modeling on the embedded features, capturing the dynamic characteristics of the rising and falling edges of the signal and generating hidden states containing bidirectional contextual information. On this basis, the multi-head self-attention mechanism dynamically assigns weights to each time step, highlighting key nodes in the transient response (such as voltage peaks or current transition points) while suppressing high-frequency noise interference. Finally, the feature distribution is stabilized through residual connections and layer normalization operations, and a time domain representation with both local details and global trends is output.
[0036] Extract frequency domain features through attention discrete cosine transform module; including: The ADCT module significantly improves the model's ability to characterize short-term high-noise signals through innovative frequency domain analysis methods. The module first divides the time domain features output by the encoder into channels, decomposes the multidimensional feature graph into several subsets, and independently performs discrete cosine transform (DCT) processing on each subset. Compared with the traditional Fourier transform, DCT effectively avoids the boundary noise interference caused by the Gibbs phenomenon, and achieves more efficient energy compression through real domain operations, ensuring the complete retention of the main frequency domain components of the signal. In order to overcome the frequency domain redundancy problem that may be introduced by fixed-weight DCT, the module innovatively introduces an attention mechanism to dynamically calibrate the transformed frequency domain features to suppress interference from irrelevant frequency bands.
[0037] like Figure 3 The figure shows a computational model involving a trend actuator and a cross-attention mechanism. Figure 3 The DCT (Discrete Cosine Transform) mentioned in is used to transform data from the time domain to the frequency domain in order to extract frequency features. The cross-attention mechanism and the self-attention mechanism are the core components of the model, which are used to capture key information and dependencies in the input data. Matrix multiplication plays an important role in these mechanisms to calculate the weights and associations between different data points. Overall, the model aims to effectively process and identify trends and patterns in complex data by combining the attention mechanism and frequency features.
[0038] Decode time domain features and frequency domain features; including: The cross-attention decoding layer enables the model to refer to the global frequency domain pattern of the signal while retaining the time domain dynamic details during the decoding process. This fusion strategy shows excellent adaptability in SMU measured data and effectively avoids the phase distortion problem of traditional methods in the transient response stage.
[0039] Data prediction: Specifically, taking 1PLC as a unit, the first half cycle of 1PLC, i.e. the first 10 sample points, is used as input data, and the remaining 10 sample points are predicted through the algorithm, so that the measurement task that originally required 1PLC can be completed with only 0.5PLC. This method combines measurement and prediction, saving half of the measurement time while ensuring that the amount of data obtained remains unchanged, and achieves acceleration effect through prediction.
[0040] In the present invention, the acceleration model uses the transformer model of Informer instead of Transformer. Informer is a transformer model designed for long sequence time series prediction (LSTF). It uses probabilistic sampling to select only a small number of important key-value pairs for calculation, which effectively captures the long-term dependencies in the sequence while greatly reducing the computational complexity and memory usage. In order to handle very long input sequences, Informer uses self-attention distillation technology. This process further improves computational efficiency by reducing the sequence length layer by layer, retaining important information and removing redundant parts.
[0041] BA-Informer is an improved Transformer architecture designed for high-precision and fast measurement of source measurement units (SMUs). Its core innovation lies in the deep integration of bidirectional recurrent neural networks (BiRNNs) and attention discrete cosine transform (ADCT) modules with traditional Transformer structures. The acceleration model first extracts features and performs spatiotemporal encoding of the input sequence through a one-dimensional convolutional layer (Conv1D) and timestamp embedding to generate a high-dimensional time series representation. The encoder part uses a bidirectional LSTM network (BiRNN) for bidirectional time series modeling. By splicing hidden states in both the forward and reverse directions, it effectively captures the rising and falling edge features in the signal, significantly improving the model's trend extraction capability in a high-noise environment. The self-attention layer then dynamically weights the BiRNN output to focus on the transient response features at key time points. The ADCT module innovatively uses discrete cosine transforms to replace traditional Fourier transforms, and achieves efficient frequency domain feature extraction through channel division and channel-by-channel DCT processing. It combines the multi-head attention mechanism to adaptively weight the frequency domain components, effectively suppressing the boundary noise interference caused by the Gibbs phenomenon. The decoder part deeply integrates the time domain features with the frequency domain features through the cross-attention mechanism, using the time domain features as the query (Query) and the frequency domain features as the key (Key) and value (Value), to achieve complementary enhancement of the time and frequency domain information. Finally, the prediction sequence is output at one time through the fully connected layer, and the non-autoregressive generation method is used to significantly improve the inference efficiency.
[0042] First, the input data is embedded and encoded, and then used as the input of the self-attention mechanism through BiRNN. After the encoding calculation, the output is used as the time domain feature. The output obtained by discrete cosine transform is combined with the weight parameters of the self-attention mechanism as the frequency domain feature. The cross-attention mechanism is used to fuse the features of the time domain and frequency domain scales, and the fused feature output is used as the input of the decoder.
[0043] Example 2 The source measurement unit acceleration method for bidirectional cyclic attention transformation according to embodiment 1 is different in that: See also Figure 2 , this is the overall workflow diagram of the source measurement unit acceleration method of bidirectional cyclic attention transformation. The method is mainly divided into two parts: data processing process and data prediction process. In high-precision measurement, the sampling of DC voltage and current is easily interfered by AC noise, resulting in a decrease in measurement accuracy and resolution. Increasing the sampling time (such as increasing the NPLC value) can reduce the impact of noise and improve measurement accuracy, but this will greatly increase the measurement time. In order to maintain high accuracy while shortening the measurement time, the present invention finds a method to strike a balance between accuracy and speed, combines a deep learning prediction algorithm, and entrusts some measurement tasks to the algorithm, thereby shortening the measurement time while maintaining or improving the measurement accuracy.
[0044] Data processing flow The data processing flow includes normalizing the collected raw data and dividing it into training and test sets as the input of the model.
[0045] Preprocessing refers to embedding and encoding. It increases the detail and accuracy of the encoder processing.
[0046] For input data Embed and encode the input data Convert to size Vector ; The calculation formula is as follows (1): (1);
[0047] The parameter α acts as a balancing factor to adjust the size between scalar projection and local / global embedding. , is a learnable global timestamp embedding, ,and, , is the embedding dimension of the acceleration model, which indicates the dimension of the high-dimensional space to which the time series data is mapped; is the length of the input sequence, which represents the number of time steps of the time series data; Represents a one-dimensional convolution operation, extracting local features from the input sequence and generating a high-dimensional representation; is the number of components into which the global timestamp is embedded.
[0048] Faced with complex signal processing tasks, traditional encoder structures perform poorly when processing load output characteristic curves affected by high excitation amplitude noise, and the prediction difficulty increases significantly with the increase of noise amplitude. To address this problem, the encoder of Informer is improved by introducing BiRNN to enhance its ability to capture signal trends. The introduction of BiRNN enables the encoder to analyze time series data from two directions and effectively capture past and future information, which is crucial for processing signals in high-noise environments. In this way, BRNN can provide a more comprehensive view of signal trends and help the self-attention mechanism to more accurately locate and interpret important features in the signal.
[0049] The time series forecasting model BiRNN is a variant of RNN that allows information to propagate in both forward and backward directions within the sequence. RNN-based models use autoregression for sequence modeling, but the cyclic structure may produce long-term modeling dependencies. Using BiRNN means that at any time step, it can utilize past and future contextual information, which is particularly useful for capturing trends because trends often involve long-term dependencies in the data.
[0050] The preprocessed data is processed by the BiRNN layer, which The signal trend information is extracted from the signal, which is then used to calculate the self-attention weight. The introduction of the self-attention mechanism further improves the model's ability to identify key features in the signal, as shown in formula (2): (2);
[0051] in, The BiRNN layer is Extracted signal trend information, is the feature dimension, is a value vector containing the actual feature information, Represents the output features weighted by the attention weights.
[0052] This self-attention mechanism enables the model to take into account the information of the entire sequence when processing each data point, thereby greatly improving the accuracy and robustness of the prediction. First, the self-attention mechanism calculates the correlation between elements in the input sequence; then, it generates attention weights; then, it uses the attention weights to weight the values ( ); Finally, the final output value is obtained by summing and weighting. The specific implementation process is as follows:
[0053] First, the correlation is calculated, and the input sequence is mapped into a triple representation of query vector (Q), key vector (K), and value vector (V) through linear transformation. The product of the transposed matrices of Q and K is calculated to obtain the original attention score matrix, which quantifies the degree of correlation of each element in the sequence with all other elements. In BA-Informer, Q and K are derived from the bidirectional temporal features extracted by the BiRNN layer, ensuring the integrity of trend information. Then weight normalization is performed, and the original attention scores are scaled and normalized by applying the Softmax function. This step converts the scores into probability distributions so that the sum of the weights at each time step is 1, while alleviating the gradient vanishing problem. Then feature weighting is performed, and the normalized attention weight matrix is multiplied by the value vector (V) to achieve dynamic recalibration of the original features. In this process, value vectors that are highly relevant to the current query will receive larger weights, while noisy or irrelevant features are weakened. Finally, the weighted feature vectors are summed along the sequence dimension to generate the final self-attention output. This mechanism enables the model to dynamically focus on the most important parts of the input sequence, thereby improving the performance and interpretability of the model.
[0054] Data prediction process: The training set data is used as input, and the time domain features are extracted through the encoder. The frequency domain features are extracted through the discrete cosine transform module, and the two scale features are decoded to complete the data prediction task. The weight parameters are continuously adjusted through the MSE loss function. The present invention adopts the ADCT module, which effectively improves the performance of the model by combining the frequency domain and time domain features of the time signal. The discrete cosine transform fundamentally avoids the Gibbs phenomenon caused by the periodicity of the discrete Fourier transform and the inverse discrete Fourier transform.
[0055] On this basis, the discrete cosine transform is combined with the attention mechanism in the encoder, and an ADCT block is designed between the encoder and the decoder to more effectively extract the frequency domain information of the signal.
[0056] Extract frequency domain features through attention discrete cosine transform module; including: When processing extremely short time series data, due to insufficient sample points, the traditional informer encoder structure often finds it difficult to fully learn the characteristics of the signal. In order to overcome this limitation and improve the performance of the model in processing such data, the present invention introduces the ADCT module, which effectively improves the performance of the model by combining the frequency domain and time domain characteristics of the time signal. The current mainstream frequency characteristic extraction method is based on Fourier transform, but due to the Gibbs phenomenon, the signal boundary will introduce high-frequency noise, which is not conducive to signal prediction. The discrete cosine transform fundamentally avoids the Gibbs phenomenon caused by the periodicity of the discrete Fourier transform and the inverse discrete Fourier transform, and has a more effective energy compression effect than the Fourier transform. On this basis, the discrete cosine transform is combined with the attention mechanism in the encoder, and an ADCT block is designed between the encoder and the decoder; to more effectively extract the frequency domain information of the signal.
[0057] For time series , ,in, , time series The definition of one-dimensional attention discrete cosine transform is shown in formula (3): (3);
[0058] in, ; Discrete cosine transform is actually a kind of DFT, and the input signal is a real even function.
[0059] First, the ADCT block divides the input feature map, i.e., the multi-channel temporal feature matrix generated by the encoder, into several subsets along the channel dimension, as shown below: (4);
[0060] Among them, each subset represents a part of a channel, , ,n is the number of channels N; In order to obtain more time series information from the feature map, DCT is used to extract more frequency domain features. However, the weight of DCT is constant, which means that there may be some redundant features in the extracted frequency domain features. These redundant features may not have a significant impact on the performance improvement of the prediction task. Therefore, the present invention introduces an attention mechanism to adjust the importance of different frequency domain features extracted by DCT through attention weights, so that the model can extract more representative feature information.
[0061] The multi-head attention weight matrix is multiplied by the frequency domain feature matrix to obtain the final sparse frequency domain feature matrix; the multi-head attention weight matrix refers to the dynamic weight matrix generated by the multi-head attention mechanism, which is used to filter the importance of frequency domain features and is recorded as ; The frequency domain feature matrix refers to the frequency domain representation generated by the ADCT module through discrete cosine transform (DCT), denoted as ; The sparse frequency domain feature matrix is calculated according to formulas (5) and (6): (5); (6);
[0062] in, represents the output of the encoder, yes The discrete cosine transform of Represents the multi-head attention weight parameter matrix; these weights are used to weight frequency features. represents the number of attention heads, and the double sum represents the weighted feature accumulation of all heads;
[0063] is the sparse frequency domain feature (after attention filtering), is the frequency domain feature matrix (DCT output); Specifically, for each head , first, calculate the weighted features of all positions , then, the weighted features of all positions Add; the accumulated features are passed through the fully connected layer ( ) to further extract features; then, through the normalization layer ( ) Normalize the features to stabilize the training process and improve the generalization ability of the model.
[0064] Finally, the encoder output is fused with the time domain features and frequency domain features through the cross-attention layer to obtain the final output; the cross-attention layer is the core fusion module of the decoder, which realizes the adaptive integration of time domain and frequency domain features through the multimodal attention mechanism. This layer uses the time domain features (BiRNN enhanced temporal dynamics) output by the encoder as the query (query vector), and the sparse frequency domain features generated by the ADCT module (DCT transformed and filtered by attention) as the key (key vector) and value (value vector), calculates the correlation weights between the time domain time points and the frequency domain components, and dynamically allocates the contribution of the frequency domain energy to the time domain features. The final fusion feature calculation is as follows:
[0065] (7); (8); (9); (10); (11);
[0066] Formulas (7) to (11) describe the cross-attention fusion process of time-domain and frequency-domain features in the BA-Informer decoder. Its core purpose is to adaptively fuse the time-domain features extracted by the encoder with the frequency-domain features generated by the ADCT module, and finally output a prediction result that has the advantages of both time and frequency domains.
[0067] Formula (7) converts the time domain features output by the encoder into a query vector, which is used to retrieve the key features of the frequency domain information; Formulas (8) and (9) respectively perform linear transformations on the frequency domain features generated by the ADCT module to generate a key vector and a value vector, where the key vector is used as a retrieval index and the value vector retains the original information in the frequency domain; Formula (10) generates the attention distribution of the time domain to the frequency domain through the similarity matrix between the query vector and the key vector, combined with the scaling factor to prevent gradient explosion, and dynamically allocates the dependence weights of each time point on different frequency components; Formula (11) uses the attention weight to weight the sum of the value vector, and outputs the fused features through the projection matrix, which has both time domain dynamics and frequency domain globality; in, is through the projection matrix Processing input sequence The resulting value matrix is ; and are respectively through the projection matrix and Processing sequence and After getting the query and key matrix, ; is the projection matrix, which is used to transform the input sequence Convert to Value Matrix , ; is the projection matrix, which is used to transform the input sequence Convert to Key Matrix , ; is the projection matrix, which is used to transform the input sequence Convert to Query Matrix , .
[0068] Design model performance comparison method: In order to test the performance of the model proposed in the present invention, the present invention compares it with other methods, including Autoformer, LightTs, MSGNet, Transformer, Informer, ETSformer, and FEDformer. In order to test the performance of the model, the present invention uses six performance indicators such as mean absolute percentage error (MAPE), and the mean absolute percentage error (MAPE) is shown in formula (12):
[0069] (12);
[0070] The normalized root mean square error (NRMSE) is shown in formula (13): (13);
[0071] Coefficient of determination: As shown in formula (14): (14);
[0072] in, is the actual value, is the predicted value, is the number of observations, represents the variance of all actual values, is the average of all actual values.
[0073] The mean square error (MSE) is shown in formula (15): (15);
[0074] The mean absolute error (MAE) is shown in formula (16): (16);
[0075] in, represent The actual value of the observation point (also called the true value or target value), represent -th observation point’s model prediction value, where n represents the total number of observation points (i.e., sample size).
[0076] The relative square error (RSE) is shown in formula (17): (17);
[0077] in is the average of the actual values.
[0078] The data set of this embodiment is based on the source measurement unit, and is obtained by applying voltages of different amplitudes to loads of different properties and measuring and recording their output characteristic curves. The performance of the model of the present invention is better than other models on each data set. The main performance indicators evaluated include mean absolute percentage error (MAPE), normalized root mean square error (NRMSE), mean square error (MSE), mean absolute error (MAE), relative square error (RSE) and coefficient of determination (R²). Figure 5 It can be seen that the model proposed in this paper outperforms other methods in all indicators, especially in MAPE and NRMSE indicators, showing significant accuracy improvement. The R² value is very high, indicating a strong correlation between the predicted value and the actual value and superior prediction performance. Figure 6 It can be seen that in the capacitance data set, the model proposed in this invention once again demonstrated its advantages, achieving lower MAPE, NRMSE and MSE values. The R² value of the model is the highest, indicating excellent predictive ability. Figure 7 It can be seen that the results of the inductance dataset further confirm the effectiveness of the proposed method. The model shows the lowest error rate (MAPE, NRMSE, MSE) and the highest R² value, proving its strong generalization ability on different data types.
[0079] Figure 8 The prediction results of the three load data sets used in this experiment. The model proposed in this invention shows significant adaptability and robustness on different data sets (such as resistance, capacitance, and inductance). By introducing BiRNN and ADCT in the encoder structure, the model's ability to capture complex time dependencies and frequency domain features is enhanced, thereby significantly improving the accuracy of the prediction, which is verified in the lower MAPE and NRMSE values. The introduction of ADCT enables the model to effectively extract and utilize frequency domain information, reduce data redundancy, and improve overall prediction performance. This is especially beneficial for processing short-time series data, while traditional methods may perform poorly in this scenario.
[0080] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a source measurement unit acceleration method for a bidirectional recurrent attention transformation described in Example 1 or 2 are implemented.
[0081] Example 4 A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a source measurement unit acceleration method of a bidirectional recurrent attention transform Informer described in Example 1 or 2 are implemented.
[0082] Example 5 A source measurement unit acceleration system for bidirectional recurrent attention transformation, comprising: The data preprocessing module is configured to: perform data preprocessing on the input; The acceleration module is configured to: pre-process the input data and input it into the trained acceleration model for acceleration processing; the acceleration model includes an encoder, an attention discrete cosine transform module (ADCT block) and a decoder; integrate the BiRNN layer into the encoder, and introduce an attention discrete cosine transform module (ADCT block) between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
Claims
1. A source measurement unit acceleration method for bidirectional recurrent attention transformation, characterized in that: include: The input data is preprocessed and then input into the trained acceleration model for acceleration processing; The acceleration model includes an encoder, an attention discrete cosine transform module (ADCT block) and a decoder; the BiRNN layer is integrated into the encoder, and an attention discrete cosine transform module (ADCT block) is introduced between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
2. The source measurement unit acceleration method of a bidirectional recurrent attention transformation according to claim 1 is characterized in that: Preprocessing refers to embedding and encoding.
3. The source measurement unit acceleration method of a bidirectional recurrent attention transformation according to claim 2 is characterized in that: For input data Embed and encode the input data Convert to size Vector ; The calculation formula is as follows (1): (1); The parameter α acts as a balancing factor to adjust the size between scalar projection and local / global embedding. , is a learnable global timestamp embedding, ,and, , is the embedding dimension of the acceleration model, which indicates the dimension of the high-dimensional space to which the time series data is mapped; is the length of the input sequence, which represents the number of time steps of the time series data; Represents a one-dimensional convolution operation, extracting local features from the input sequence and generating a high-dimensional representation; is the number of components into which the global timestamp is embedded.
4. The source measurement unit acceleration method of a bidirectional recurrent attention transformation according to claim 1 is characterized in that: The preprocessed data is processed by the BiRNN layer, which The signal trend information is extracted from the signal, which is then used to calculate the self-attention weight; As shown in formula (2): (2); in, The BiRNN layer is Extracted signal trend information, is the feature dimension, is a value vector containing the actual feature information, Represents the output features weighted by attention weights; First, the self-attention mechanism calculates the correlation between elements in the input sequence; then, it generates attention weights; then, it uses the attention weights to weight the values; finally, it obtains the final output value by summing and weighting.
5. The source measurement unit acceleration method of a bidirectional recurrent attention transformation according to claim 1 is characterized in that: Extract frequency domain features through attention discrete cosine transform module; including: For time series , ,in, , time series The definition of one-dimensional attention discrete cosine transform is shown in formula (3): (3); in, ; The ADCT block divides the input feature map, i.e., the multi-channel temporal feature matrix generated by the encoder, into several subsets along the channel dimension, as shown below: (4); Among them, each subset represents a part of a channel, , ,n is the number of channels N; The multi-head attention weight matrix is multiplied by the frequency domain feature matrix to obtain the final sparse frequency domain feature matrix; the multi-head attention weight matrix refers to the dynamic weight matrix generated by the multi-head attention mechanism, denoted as ; The frequency domain feature matrix refers to the frequency domain representation generated by the ADCT module through discrete cosine transform, denoted as ; The sparse frequency domain feature matrix is calculated according to formulas (5) and (6): (5); (6); in, represents the output of the encoder, yes The discrete cosine transform of Represents the multi-head attention weight parameter matrix; represents the number of attention heads, and the double sum represents the weighted feature accumulation of all heads; is the sparse frequency domain feature, is the frequency domain feature matrix; Specifically include: For each head , first, calculate the weighted features of all positions , then, the weighted features of all positions The accumulated features are transformed through the fully connected layer to further extract features. Then, the features are normalized through the normalization layer. Finally, the encoder output is fused with the time domain features and frequency domain features through the cross attention layer to obtain the final output. The final fused features are calculated as follows: (7); (8); (9); (10); (11); Formula (7) converts the time domain features output by the encoder into a query vector, which is used to retrieve the key features of the frequency domain information; Formulas (8) and (9) respectively perform linear transformations on the frequency domain features generated by the ADCT module to generate a key vector and a value vector, where the key vector is used as a retrieval index and the value vector retains the original information in the frequency domain; Formula (10) generates the attention distribution of the time domain to the frequency domain through the similarity matrix between the query vector and the key vector, combined with the scaling factor to prevent gradient explosion, and dynamically allocates the dependence weights of each time point on different frequency components; Formula (11) uses the attention weight to weight the sum of the value vector, and outputs the fused features through the projection matrix, which has both time domain dynamics and frequency domain globality; in, is through the projection matrix Processing input sequence The resulting value matrix is ; and are respectively through the projection matrix and Processing sequence and After getting the query and key matrix, ; is the projection matrix, which is used to transform the input sequence Convert to value matrix , ; is the projection matrix, which is used to transform the input sequence Convert to key matrix , ; is the projection matrix, which is used to transform the input sequence Convert to query matrix , .
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of a source measurement unit acceleration method of a bidirectional recurrent attention transform Informer as described in any one of claims 1-5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a source measurement unit acceleration method of a bidirectional recurrent attention transform Informer described in any one of claims 1-5 are implemented.
8. A source measurement unit acceleration system for bidirectional recurrent attention transformation, characterized in that: include: The data preprocessing module is configured to: perform data preprocessing on the input; The acceleration module is configured to: pre-process the input data and input it into the trained acceleration model for acceleration processing; the acceleration model includes an encoder, an attention discrete cosine transform module (ADCT block) and a decoder; integrate the BiRNN layer into the encoder, and introduce an attention discrete cosine transform module (ADCT block) between the encoder and the decoder; The time domain features are extracted through the encoder, the frequency domain features are extracted through the attention discrete cosine transform module, and the time domain features and frequency domain features are decoded to complete the data prediction task.
Citation Information
Patent Citations
Hierarchical text abstract acquisition method and based on discourse structure, system, terminal equipment and readable storage medium
CN113157907A
Refrigeration system load prediction method based on sequence-to-sequence model
CN115982567A
Method and system for increasing sampling rate of source measurement unit through interpolation based on time sequence prediction
CN117851823A
Video monitoring object counterfeiting detection system
CN118447376A
Cited By
Multi-input-output Transform model chip architecture and calculation method
CN121072630A
Self-adaptive modulation identification method based on time-frequency double flow and domain adversarial learning
CN121356955A
An adaptive modulation recognition method based on time-frequency dual stream and domain adversarial learning
CN121356955B
Robust lightweight time sequence prediction method based on bidirectional filling and geometric attention
CN121435109A
A robust light-weight time series forecasting method based on bidirectional padding and geometric attention
CN121435109B