Time series data analysis method, device and equipment based on improved Transform and medium

By improving the Transformer's time series data analysis method and using adversarial networks, adaptive position encoding, and multi-head attention mechanisms for data enhancement and training, the accuracy and interpretability issues of time series data analysis are solved, and efficient and comprehensive data analysis is achieved.

CN120653923APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510727829.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing time series data analysis methods are insufficient in accuracy and efficiency, and lack interpretability. In particular, when faced with complex and massive data, classic algorithms such as Random Forest, XGBoost, and RNN perform poorly.

Method used

A time series data analysis method based on an improved Transformer is adopted. Data enhancement is performed through an adversarial network. Combined with adaptive position encoding and multi-head attention mechanism, multi-task joint training and adversarial training are carried out. Interpretability analysis is performed using attention distribution features and intermediate layer output features.

Benefits of technology

It improves the accuracy and efficiency of time series data analysis, enhances the model's adaptability to different data situations, improves interpretability, and provides more comprehensive data analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653923A_ABST
    Figure CN120653923A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of finance, medical health and artificial intelligence, and provides an improved Transform-based time series data analysis method, device and equipment and a medium, which can utilize an adversarial network to perform data enhancement and effectively expand a training data set. The Transform model has the ability to perceive a sequence based on adaptive position coding, the multi-head attention mechanism based on a dynamic weight adjustment strategy enables the model to improve the ability to capture key information, and multi-head parallel computing can integrate attention results of different view angles to comprehensively mine dependency relationships in data; multi-task joint training and adversarial training are performed on the model, so that the task execution capability of the model can be improved on the premise of ensuring the training effect; the attention distribution features and the intermediate layer output features are utilized to perform interpretability analysis on the time series data analysis result, the interpretability of the analysis result can be improved, and a more comprehensive data analysis result can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of finance, medical health, and artificial intelligence technology, and in particular to a time series data analysis method, device, equipment, and medium based on an improved Transformer. Background Art

[0002] Nowadays, analyzing time series data is a common task in various fields to facilitate their development. For example, in finance, analyzing time series data can be used to predict stock prices, assess credit risk, monitor market risks, and analyze customer behavior. In healthcare, analyzing time series data can assist in disease prediction, condition monitoring, treatment effectiveness evaluation, and medical equipment failure prediction.

[0003] However, in the current data analysis field, while classic algorithms such as Random Forest and XGBoost (eXtreme Gradient Boosting) are widely used, their limitations are becoming increasingly apparent in the face of increasingly complex and massive data volumes. RNNs (Recurrent Neural Networks) and their various variants also perform poorly in large-scale data analysis scenarios due to long-term dependencies and limited parallel computing capabilities when processing sequential data, and they also lack interpretability. Summary of the Invention

[0004] In view of the above, it is necessary to provide a time series data analysis method, device, equipment and medium based on an improved Transformer, aiming to solve the problems of low accuracy, low efficiency and lack of interpretability in the time series data analysis process.

[0005] A time series data analysis method based on an improved Transformer, the time series data analysis method based on the improved Transformer comprising:

[0006] Collecting initial time series data, and performing data enhancement on the initial time series data using an adversarial network to obtain training samples;

[0007] An initial Transformer model is constructed based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy;

[0008] Performing adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model;

[0009] In response to an analysis instruction for target time series data, inputting the target time series data into the time series data analysis model to obtain a time series data analysis result;

[0010] Performing an interpretability analysis on the time series data analysis results using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain an interpretability analysis result;

[0011] Generate an analysis report of the target time series data based on the time series data analysis results and the interpretability analysis results.

[0012] A time series data analysis device based on an improved Transformer, comprising:

[0013] A data enhancement unit is used to collect initial time series data and perform data enhancement on the initial time series data using an adversarial network to obtain training samples;

[0014] A construction unit for constructing an initial Transformer model based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy;

[0015] A training unit, configured to perform adversarial training on the initial Transformer model based on the training samples and a multi-task joint training strategy to obtain a time series data analysis model;

[0016] an input unit, configured to input the target time series data into the time series data analysis model in response to an analysis instruction for the target time series data, and obtain a time series data analysis result;

[0017] An analysis unit, configured to perform an interpretability analysis on the time series data analysis result using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain an interpretability analysis result;

[0018] A generating unit is used to generate an analysis report of the target time series data based on the time series data analysis result and the interpretability analysis result.

[0019] A computer device, comprising:

[0020] a memory storing at least one instruction; and

[0021] A processor executes instructions stored in the memory to implement the time series data analysis method based on the improved Transformer.

[0022] A computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the time series data analysis method based on the improved Transformer.

[0023] It can be seen from the above technical solutions that the present invention can use the adversarial network to perform data enhancement on the initial time series data, effectively expand the training data set, and improve the model's adaptability to different data situations; based on the adaptive position coding and multi-head attention mechanism, the Transformer model is improved. The adaptive position coding can enable the model to have the ability to perceive the sequence order, and the multi-head attention mechanism based on the dynamic weight adjustment strategy can enable the model to pay more attention to important features and improve the ability to capture key information. The parallel computing of multiple heads can also comprehensively explore the dependencies in the data by integrating the attention results from different perspectives; multi-task joint training and adversarial training of the model can improve the task execution capability of the model while ensuring the training effect; the use of attention distribution characteristics and intermediate layer output characteristics to perform interpretable analysis on the time series data analysis results can improve the interpretability of the analysis results and provide more comprehensive data analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of a preferred embodiment of the time series data analysis method based on the improved Transformer of the present invention.

[0025] Figure 2 It is a functional module diagram of a preferred embodiment of the time series data analysis device based on the improved Transformer of the present invention.

[0026] Figure 3 It is a structural diagram of a computer device for implementing a preferred embodiment of the time series data analysis method based on the improved Transformer of the present invention. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] like Figure 1 FIG. 1 is a flow chart of a preferred embodiment of the time series data analysis method based on the improved Transformer of the present invention. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.

[0029] The improved Transformer-based time series data analysis method is applied to one or more computer devices, which are devices that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Their hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0030] The computer device may be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.

[0031] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0032] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0033] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0034] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0035] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0036] S10, collecting initial time series data, and using an adversarial network to perform data enhancement on the initial time series data to obtain training samples.

[0037] In this embodiment, the initial time series data may be time series data in a corresponding field.

[0038] For example, in the financial field, the initial time series data may be time series data such as stock prices and trading volumes at different time points in the past, or time series information such as financial data and credit history of an enterprise or individual over a period of time.

[0039] For example, in the field of medical health, the initial time series data can be time series data such as the patient's physiological parameters (such as body temperature, blood pressure, heart rate, blood sugar, etc.), symptom manifestations, etc. that change over time, and time series data such as the vital signs of hospitalized patients (such as electrocardiogram signals, blood oxygen saturation, etc.).

[0040] For another example, in the field of industrial monitoring, the initial time series data may be time series data of the temperature, pressure, etc. of a sensor that changes over time.

[0041] In this embodiment, the use of an adversarial network to perform data enhancement on the initial time series data to obtain training samples includes:

[0042] Cleaning the initial time series data to obtain first data;

[0043] Normalizing the first data to obtain second data;

[0044] performing sequence block processing on the second data to obtain third data;

[0045] Using random noise as input, using the generator in the adversarial network to generate a data block similar to the initial time series data, and using the discriminator in the adversarial network to distinguish whether the data block input to the discriminator is a real data block or a generated data block during the generation process;

[0046] The training sample is constructed using the third data and the generated data block.

[0047] Among them, data points with obvious errors or abnormalities in the initial time series data can be removed, such as records with too many missing values, outliers that exceed a reasonable range, etc. Mean filling, interpolation, etc. can also be used to process a small number of missing values ​​to achieve preliminary cleaning of the initial time series data.

[0048] The first data may be normalized in different ways according to the characteristics of different data types.

[0049] For example, for price series data in the financial field, Z-score normalization can be used; for trading volume in the financial field, Min-Max algorithm can be used for normalization.

[0050] Through normalization, data of different magnitudes can be brought to a unified scale, which is conducive to model training and convergence.

[0051] The second data may be divided into sequence blocks of fixed length.

[0052] By dividing the sequence into blocks and using them as input for subsequent models, the model can better capture the data characteristics and patterns within the local time range.

[0053] Among them, data enhancement through adversarial networks can increase data diversity and the generalization ability of the model. The block-based time series data is augmented with data based on generative adversarial networks (GANs). An adversarial network consisting of a generator and a discriminator is constructed. The generator takes random noise as input and generates data blocks similar to real time series data. The discriminator is responsible for distinguishing whether the input data blocks are real data or generated data. During the training process, the generator is continuously optimized to deceive the discriminator, and the discriminator is continuously optimized to accurately distinguish between true and false data. After multiple rounds of adversarial training, a large number of time series data blocks with similar distribution to the original data but different content are generated, effectively expanding the training data set and improving the model's adaptability to different data situations.

[0054] S11, constructing an initial Transformer model based on adaptive position encoding and multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy.

[0055] In this embodiment, the construction of the initial Transformer model based on adaptive position encoding and multi-head attention mechanism includes:

[0056] Calculate the data change between each position and the adjacent positions in the time series data sequence;

[0057] The following formula is used to position-code each position according to the data variation:

[0058]

[0059] Among them, PE (pos,2i) Indicates the position code with index value 2i; PE (pos,2i+1)Indicates the position code with index value (2i+1); pos indicates the position; d model represents the model dimension; β represents the adjustment coefficient; Δx pos represents the data change; i represents the index value.

[0060] The optimal value of the adjustment coefficient can be determined based on experiments.

[0061] The above embodiment solves the problem that the Transformer model itself lacks the ability to perceive sequence order. This embodiment can adaptively and dynamically generate position encodings based on the degree of local variation in the data. The position encodings are added to the data embedding vector to provide the model with time series information, enabling the model to distinguish data at different time steps.

[0062] In this embodiment, the construction of the initial Transformer model based on the adaptive position encoding and multi-head attention mechanism also includes:

[0063] For each attention head, linearly transform the input data to obtain the query, key, and value;

[0064] The attention score is adjusted based on the query, key, and value using the following formula:

[0065]

[0066] Among them, Attention new (Q h ,K h ,V h ) represents the adjusted attention score corresponding to the h-th attention head; Q h represents the query corresponding to the h-th attention head; K h represents the key corresponding to the h-th attention head; V h represents the value corresponding to the h-th attention head; d k represents the kth dimension; I represents the feature importance vector.

[0067] Through the above embodiments, a dynamic weight adjustment strategy is introduced on the basis of the traditional multi-head attention mechanism, so that the model can pay more attention to important features when calculating attention, improve the ability to capture key information, and different attention heads can be calculated in parallel, integrating the attention results from different perspectives to comprehensively explore the dependencies in the data.

[0068] S12: performing adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model.

[0069] In this embodiment, in order to fully mine the various information in the time series data, a multi-task joint training strategy is adopted for model training.

[0070] Specifically, the adversarial training of the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain the time series data analysis model includes:

[0071] Obtaining the weight of each task and constructing a total loss function based on the weight of each task; wherein each task shares the bottom feature extraction layer of the initial Transformer model, and each task corresponds to a different top output layer;

[0072] During adversarial training, the initial Transformer model is used as a generator, and a discriminator is used to determine whether the output of the initial Transformer model is true;

[0073] When the total loss function reaches convergence, the training is stopped and the currently obtained model is determined as the time series data analysis model.

[0074] For example: in addition to traditional prediction tasks, auxiliary tasks are added, such as data trend classification (rising, falling, stable), outlier detection, etc.

[0075] Among them, the weight of each task can be adjusted through experiments to balance the training of different tasks, so that the model can learn richer and more comprehensive feature representations, thereby improving the performance of the model on various data analysis tasks.

[0076] Among them, when conducting adversarial training, it is necessary to build an adversarial network. Specifically, this embodiment uses the Transformer model as the generator of the adversarial network, and the discriminator is used to determine whether the output of the Transformer model is true. During the training process, the Transformer model must not only minimize its own prediction loss, but also "cheat" the discriminator, while the discriminator must identify the output of the Transformer model as accurately as possible. Through this adversarial training method, the Transformer model is prompted to learn more robust and generalized feature representations, thereby improving the performance of the model in different data scenarios.

[0077] In this embodiment, after the time series data analysis model is obtained through training, the accuracy, mean square error and other indicators of the model can be calculated using a test data set to evaluate the trained model.

[0078] S13 , in response to an analysis instruction for the target time series data, input the target time series data into the time series data analysis model to obtain a time series data analysis result.

[0079] In this embodiment, the analysis instruction can be triggered by relevant staff according to actual analysis needs, or it can be triggered periodically, or automatically triggered when time series data is uploaded, to assist in more timely response.

[0080] In this embodiment, after the target time series data is input into the time series data analysis model, the data must first pass through the embedding layer to convert the data of each time step into a vector representation of fixed dimension for subsequent processing by the model.

[0081] Furthermore, the data processed by the multi-head attention layer enters the feedforward neural network layer. A gating unit is added between the output and input of the feedforward neural network layer. The gating unit is calculated through a linear transformation and a sigmoid activation function.

[0082] Through this gating mechanism, the model dynamically determines the ratio of the output information of the feedforward neural network layer to the original input information based on data characteristics, thereby enhancing its adaptability to different data patterns. This process performs a nonlinear transformation on the feature vector at each position to further extract data features.

[0083] The model also includes residual connections and layer normalization. Residual connections are used to add inputs directly to the outputs of the corresponding layer, helping to alleviate the vanishing gradient problem and enabling the model to learn more complex features. Layer normalization normalizes the outputs of each layer, accelerating model convergence and making model training more stable.

[0084] In this embodiment, inputting the target time series data into the time series data analysis model to obtain the time series data analysis results includes:

[0085] When the time series data analysis model performs a single-step task, the time series data analysis model is used to output a predicted value of the next time step according to the learned time series pattern to obtain the time series data analysis result; or

[0086] When the time series data analysis model performs a multi-step task, the time series data analysis model is used to take the prediction result of the previous step as the input of the next step prediction, and gradually generate a prediction sequence of multiple time steps to obtain the time series data analysis result.

[0087] Of course, when the time series data analysis model performs multi-step tasks, the model can also directly output prediction sequences of multiple time steps, which depends on the model's architectural design and training method.

[0088] S14, using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to perform interpretability analysis on the time series data analysis results to obtain interpretability analysis results.

[0089] In this embodiment, the interpretability analysis of the time series data analysis results is performed using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model, and the interpretability analysis results obtained include:

[0090] Determining the degree to which the time series data analysis model focuses on different features at different time steps according to the attention weight matrix of the time series data analysis model;

[0091] Performing cluster analysis on the output features of the intermediate layer to obtain feature patterns;

[0092] The interpretability analysis result is generated according to the degree of attention paid by the time series data analysis model to different features at different time steps and the feature pattern.

[0093] For example, by visualizing the attention weight matrix, we can intuitively demonstrate the model's focus on different features at different time steps, helping users understand the basis for the model's decisions. Furthermore, by leveraging the model's intermediate layer output features, combined with methods such as cluster analysis, we can conduct in-depth feature mining and pattern classification on the data. For example, by using intermediate layer feature vectors as sample points and employing clustering algorithms such as K-means, we can cluster data with similar characteristics into a single category, thereby discovering underlying patterns and regularities in the data and providing users with more comprehensive data analysis results.

[0094] In the above embodiment, by analyzing the attention distribution of the Transformer model and using the intermediate layer features for cluster analysis, it is possible not only to intuitively understand the key data parts that the model focuses on when performing data analysis and prediction, thereby improving the interpretability of the model, but also to provide users with more comprehensive and in-depth data analysis results, helping users to better understand the patterns and trends behind the data.

[0095] S15: Generate an analysis report of the target time series data according to the time series data analysis result and the interpretability analysis result.

[0096] In this embodiment, the analysis report may be sent to a triggerer of the data analysis instruction.

[0097] This embodiment uses technical improvements such as data augmentation, improved multi-head attention mechanism, adaptive position encoding, and multi-task joint training to enable the Transformer model to more accurately capture complex long-term dependencies and nonlinear patterns in time series data. Compared with traditional data analysis algorithms and unimproved Transformer models, it can greatly improve the accuracy of data analysis.

[0098] Furthermore, the combination of the parallel computing characteristics of the Transformer itself and the improvement measures of this embodiment can significantly shorten the computing time of the model when processing large-scale data, making it particularly suitable for data analysis scenarios with high real-time requirements.

[0099] For example, in the financial sector, by collecting time series data such as stock prices and trading volumes at different points in the past and applying models to analyze trends, cyclical patterns, and other patterns within the data, it is possible to predict future stock price trends and help investors formulate investment strategies. For example, by analyzing the fluctuation patterns of a stock's daily closing price, opening price, and trading volume over the past year, we can predict stock price fluctuations over the next week and help investors decide when to buy or sell.

[0100] Another example: In the field of healthcare, collecting time-series data on patients' physiological parameters (such as body temperature, blood pressure, heart rate, blood sugar, etc.) and symptoms over time, and using models to analyze data trends and anomalies, can help predict the risk of disease or assist in diagnosing the disease. For example, based on the time-series changes in blood sugar monitoring values ​​of diabetic patients over the past few months, combined with relevant records such as diet and exercise, it is possible to predict blood sugar fluctuation trends and assist in adjusting treatment plans; or by analyzing the time-series data of symptoms such as cough frequency, body temperature, and degree of dyspnea over a period of time, it can assist in determining whether the patient has a respiratory disease.

[0101] It can be seen from the above technical solutions that the present invention can use the adversarial network to perform data enhancement on the initial time series data, effectively expand the training data set, and improve the model's adaptability to different data situations; based on the adaptive position coding and multi-head attention mechanism, the Transformer model is improved. The adaptive position coding can enable the model to have the ability to perceive the sequence order, and the multi-head attention mechanism based on the dynamic weight adjustment strategy can enable the model to pay more attention to important features and improve the ability to capture key information. The parallel computing of multiple heads can also comprehensively explore the dependencies in the data by integrating the attention results from different perspectives; multi-task joint training and adversarial training of the model can improve the task execution capability of the model while ensuring the training effect; the use of attention distribution characteristics and intermediate layer output characteristics to perform interpretable analysis on the time series data analysis results can improve the interpretability of the analysis results and provide more comprehensive data analysis results.

[0102] like Figure 2, which is a functional module diagram of a preferred embodiment of the time series data analysis device based on the improved Transformer of the present invention. The time series data analysis device 11 based on the improved Transformer includes a data enhancement unit 110, a construction unit 111, a training unit 112, an input unit 113, an analysis unit 114, and a generation unit 115. The module / unit referred to in the present invention refers to a series of computer program segments that can be executed by a processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0103] The data enhancement unit 110 is used to collect initial time series data and perform data enhancement on the initial time series data using an adversarial network to obtain training samples.

[0104] In this embodiment, the initial time series data may be time series data in a corresponding field.

[0105] For example, in the financial field, the initial time series data may be time series data such as stock prices and trading volumes at different time points in the past, or time series information such as financial data and credit history of an enterprise or individual over a period of time.

[0106] For example, in the field of medical health, the initial time series data can be time series data such as the patient's physiological parameters (such as body temperature, blood pressure, heart rate, blood sugar, etc.), symptom manifestations, etc. that change over time, and time series data such as the vital signs of hospitalized patients (such as electrocardiogram signals, blood oxygen saturation, etc.).

[0107] For another example, in the field of industrial monitoring, the initial time series data may be time series data of the temperature, pressure, etc. of a sensor that changes over time.

[0108] In this embodiment, the data enhancement unit 110 uses an adversarial network to perform data enhancement on the initial time series data, and the obtained training samples include:

[0109] Cleaning the initial time series data to obtain first data;

[0110] Normalizing the first data to obtain second data;

[0111] performing sequence block processing on the second data to obtain third data;

[0112] Using random noise as input, using the generator in the adversarial network to generate a data block similar to the initial time series data, and using the discriminator in the adversarial network to distinguish whether the data block input to the discriminator is a real data block or a generated data block during the generation process;

[0113] The training sample is constructed using the third data and the generated data block.

[0114] Among them, data points with obvious errors or abnormalities in the initial time series data can be removed, such as records with too many missing values, outliers that exceed a reasonable range, etc. Mean filling, interpolation, etc. can also be used to process a small number of missing values ​​to achieve preliminary cleaning of the initial time series data.

[0115] The first data may be normalized in different ways according to the characteristics of different data types.

[0116] For example, for price series data in the financial field, Z-score normalization can be used; for trading volume in the financial field, Min-Max algorithm can be used for normalization.

[0117] Through normalization, data of different magnitudes can be brought to a unified scale, which is conducive to model training and convergence.

[0118] The second data may be divided into sequence blocks of fixed length.

[0119] By dividing the sequence into blocks and using them as input for subsequent models, the model can better capture the data characteristics and patterns within the local time range.

[0120] Among them, data enhancement through adversarial networks can increase data diversity and the generalization ability of the model. The block-based time series data is augmented with data based on generative adversarial networks (GANs). An adversarial network consisting of a generator and a discriminator is constructed. The generator takes random noise as input and generates data blocks similar to real time series data. The discriminator is responsible for distinguishing whether the input data blocks are real data or generated data. During the training process, the generator is continuously optimized to deceive the discriminator, and the discriminator is continuously optimized to accurately distinguish between true and false data. After multiple rounds of adversarial training, a large number of time series data blocks with similar distribution to the original data but different content are generated, effectively expanding the training data set and improving the model's adaptability to different data situations.

[0121] The construction unit 111 is used to construct an initial Transformer model based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy.

[0122] In this embodiment, the construction unit 111 constructs the initial Transformer model based on the adaptive position encoding and multi-head attention mechanism, including:

[0123] Calculate the data change between each position and the adjacent positions in the time series data sequence;

[0124] The following formula is used to position-code each position according to the data variation:

[0125]

[0126] Among them, PE (pos,2i) Indicates the position code with index value 2i; PE (pos,2i+1) Indicates the position code with index value (2i+1); pos indicates the position; d model represents the model dimension; β represents the adjustment coefficient; Δx pos represents the data change; i represents the index value.

[0127] The optimal value of the adjustment coefficient can be determined based on experiments.

[0128] The above embodiment solves the problem that the Transformer model itself lacks the ability to perceive sequence order. This embodiment can adaptively and dynamically generate position encodings based on the degree of local variation in the data. The position encodings are added to the data embedding vector to provide the model with time series information, enabling the model to distinguish data at different time steps.

[0129] In this embodiment, the construction unit 111 constructs the initial Transformer model based on the adaptive position encoding and multi-head attention mechanism, further comprising:

[0130] For each attention head, linearly transform the input data to obtain the query, key, and value;

[0131] The attention score is adjusted based on the query, key, and value using the following formula:

[0132]

[0133] Among them, Attention new (Q h ,K h ,V h ) represents the adjusted attention score corresponding to the h-th attention head; Q h represents the query corresponding to the h-th attention head; K h represents the key corresponding to the h-th attention head; V h represents the value corresponding to the h-th attention head; d k represents the kth dimension; I represents the feature importance vector.

[0134] Through the above embodiments, a dynamic weight adjustment strategy is introduced on the basis of the traditional multi-head attention mechanism, so that the model can pay more attention to important features when calculating attention, improve the ability to capture key information, and different attention heads can be calculated in parallel, integrating the attention results from different perspectives to comprehensively explore the dependencies in the data.

[0135] The training unit 112 is used to perform adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model.

[0136] In this embodiment, in order to fully mine the various information in the time series data, a multi-task joint training strategy is adopted for model training.

[0137] Specifically, the training unit 112 performs adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy, and obtains a time series data analysis model including:

[0138] Obtaining the weight of each task and constructing a total loss function based on the weight of each task; wherein each task shares the bottom feature extraction layer of the initial Transformer model, and each task corresponds to a different top output layer;

[0139] During adversarial training, the initial Transformer model is used as a generator, and a discriminator is used to determine whether the output of the initial Transformer model is true;

[0140] When the total loss function reaches convergence, the training is stopped and the currently obtained model is determined as the time series data analysis model.

[0141] For example: in addition to traditional prediction tasks, auxiliary tasks are added, such as data trend classification (rising, falling, stable), outlier detection, etc.

[0142] Among them, the weight of each task can be adjusted through experiments to balance the training of different tasks, so that the model can learn richer and more comprehensive feature representations, thereby improving the performance of the model on various data analysis tasks.

[0143] Among them, when conducting adversarial training, it is necessary to build an adversarial network. Specifically, this embodiment uses the Transformer model as the generator of the adversarial network, and the discriminator is used to determine whether the output of the Transformer model is true. During the training process, the Transformer model must not only minimize its own prediction loss, but also "cheat" the discriminator, while the discriminator must identify the output of the Transformer model as accurately as possible. Through this adversarial training method, the Transformer model is prompted to learn more robust and generalized feature representations, thereby improving the performance of the model in different data scenarios.

[0144] In this embodiment, after the time series data analysis model is obtained through training, the accuracy, mean square error and other indicators of the model can be calculated using a test data set to evaluate the trained model.

[0145] The input unit 113 is configured to input the target time series data into the time series data analysis model in response to an analysis instruction for the target time series data, and obtain a time series data analysis result.

[0146] In this embodiment, the analysis instruction can be triggered by relevant staff according to actual analysis needs, or it can be triggered periodically, or automatically triggered when time series data is uploaded, to assist in more timely response.

[0147] In this embodiment, after the target time series data is input into the time series data analysis model, the data must first pass through the embedding layer to convert the data of each time step into a vector representation of fixed dimension for subsequent processing by the model.

[0148] Furthermore, the data processed by the multi-head attention layer enters the feedforward neural network layer. A gating unit is added between the output and input of the feedforward neural network layer. The gating unit is calculated through a linear transformation and a sigmoid activation function.

[0149] Through this gating mechanism, the model dynamically determines the ratio of the output information of the feedforward neural network layer to the original input information based on data characteristics, thereby enhancing its adaptability to different data patterns. This process performs a nonlinear transformation on the feature vector at each position to further extract data features.

[0150] The model also includes residual connections and layer normalization. Residual connections are used to add inputs directly to the outputs of the corresponding layer, helping to alleviate the vanishing gradient problem and enabling the model to learn more complex features. Layer normalization normalizes the outputs of each layer, accelerating model convergence and making model training more stable.

[0151] In this embodiment, the input unit 113 inputs the target time series data into the time series data analysis model, and obtains the time series data analysis results including:

[0152] When the time series data analysis model performs a single-step task, the time series data analysis model is used to output a predicted value of the next time step according to the learned time series pattern to obtain the time series data analysis result; or

[0153] When the time series data analysis model performs a multi-step task, the time series data analysis model is used to take the prediction result of the previous step as the input of the next step prediction, and gradually generate a prediction sequence of multiple time steps to obtain the time series data analysis result.

[0154] Of course, when the time series data analysis model performs multi-step tasks, the model can also directly output prediction sequences of multiple time steps, which depends on the model's architectural design and training method.

[0155] The analysis unit 114 is used to perform interpretability analysis on the time series data analysis results using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain interpretability analysis results.

[0156] In this embodiment, the analysis unit 114 performs interpretability analysis on the time series data analysis results using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model, and the interpretability analysis results obtained include:

[0157] Determining the degree to which the time series data analysis model focuses on different features at different time steps according to the attention weight matrix of the time series data analysis model;

[0158] Performing cluster analysis on the output features of the intermediate layer to obtain feature patterns;

[0159] The interpretability analysis result is generated according to the degree of attention paid by the time series data analysis model to different features at different time steps and the feature pattern.

[0160] For example, by visualizing the attention weight matrix, we can intuitively demonstrate the model's focus on different features at different time steps, helping users understand the basis for the model's decisions. Furthermore, by leveraging the model's intermediate layer output features, combined with methods such as cluster analysis, we can conduct in-depth feature mining and pattern classification on the data. For example, by using intermediate layer feature vectors as sample points and employing clustering algorithms such as K-means, we can cluster data with similar characteristics into a single category, thereby discovering underlying patterns and regularities in the data and providing users with more comprehensive data analysis results.

[0161] In the above embodiment, by analyzing the attention distribution of the Transformer model and using the intermediate layer features for cluster analysis, it is possible not only to intuitively understand the key data parts that the model focuses on when performing data analysis and prediction, thereby improving the interpretability of the model, but also to provide users with more comprehensive and in-depth data analysis results, helping users to better understand the patterns and trends behind the data.

[0162] The generating unit 115 is configured to generate an analysis report of the target time series data according to the time series data analysis result and the interpretability analysis result.

[0163] In this embodiment, the analysis report may be sent to a triggerer of the data analysis instruction.

[0164] This embodiment uses technical improvements such as data augmentation, improved multi-head attention mechanism, adaptive position encoding, and multi-task joint training to enable the Transformer model to more accurately capture complex long-term dependencies and nonlinear patterns in time series data. Compared with traditional data analysis algorithms and unimproved Transformer models, it can greatly improve the accuracy of data analysis.

[0165] Furthermore, the combination of the parallel computing characteristics of the Transformer itself and the improvement measures of this embodiment can significantly shorten the computing time of the model when processing large-scale data, making it particularly suitable for data analysis scenarios with high real-time requirements.

[0166] For example, in the financial sector, by collecting time series data such as stock prices and trading volumes at different points in the past and applying models to analyze trends, cyclical patterns, and other patterns within the data, it is possible to predict future stock price trends and help investors formulate investment strategies. For example, by analyzing the fluctuation patterns of a stock's daily closing price, opening price, and trading volume over the past year, we can predict stock price fluctuations over the next week and help investors decide when to buy or sell.

[0167] Another example: In the field of healthcare, collecting time-series data on patients' physiological parameters (such as body temperature, blood pressure, heart rate, blood sugar, etc.) and symptoms over time, and using models to analyze data trends and anomalies, can help predict the risk of disease or assist in diagnosing the disease. For example, based on the time-series changes in blood sugar monitoring values ​​of diabetic patients over the past few months, combined with relevant records such as diet and exercise, it is possible to predict blood sugar fluctuation trends and assist in adjusting treatment plans; or by analyzing the time-series data of symptoms such as cough frequency, body temperature, and degree of dyspnea over a period of time, it can assist in determining whether the patient has a respiratory disease.

[0168] It can be seen from the above technical solutions that the present invention can use the adversarial network to perform data enhancement on the initial time series data, effectively expand the training data set, and improve the model's adaptability to different data situations; based on the adaptive position coding and multi-head attention mechanism, the Transformer model is improved. The adaptive position coding can enable the model to have the ability to perceive the sequence order, and the multi-head attention mechanism based on the dynamic weight adjustment strategy can enable the model to pay more attention to important features and improve the ability to capture key information. The parallel computing of multiple heads can also comprehensively explore the dependencies in the data by integrating the attention results from different perspectives; multi-task joint training and adversarial training of the model can improve the task execution capability of the model while ensuring the training effect; the use of attention distribution characteristics and intermediate layer output characteristics to perform interpretable analysis on the time series data analysis results can improve the interpretability of the analysis results and provide more comprehensive data analysis results.

[0169] like Figure 3 , which is a structural diagram of a computer device for implementing a preferred embodiment of the time series data analysis method based on the improved Transformer according to the present invention.

[0170] The computer device 1 may include a memory 12, a processor 13 and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a time series data analysis program based on an improved Transformer.

[0171] Those skilled in the art will understand that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may have either a bus structure or a star structure. The computer device 1 may also include more or less other hardware or software than shown in the figure, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.

[0172] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that are suitable for the present invention should also be included in the scope of protection of the present invention and included here by reference.

[0173] The memory 12 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 may be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 may also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 1. Furthermore, the memory 12 may include both an internal storage unit of the computer device 1 and an external storage device. The memory 12 can be used not only to store application software and various types of data installed in the computer device 1, such as the code of a time series data analysis program based on the improved Transformer, but also to temporarily store data that has been output or is about to be output.

[0174] In some embodiments, the processor 13 may be composed of an integrated circuit, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, connecting the various components of the entire computer device 1 using various interfaces and lines. It executes or executes programs or modules stored in the memory 12 (such as executing a time series data analysis program based on an improved Transformer) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.

[0175] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned embodiments of the time series data analysis method based on the improved Transformer, for example Figure 1 Steps shown.

[0176] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a data enhancement unit 110, a construction unit 111, a training unit 112, an input unit 113, an analysis unit 114, and a generation unit 115.

[0177] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to execute the portion of the time series data analysis method based on the improved Transformer described in various embodiments of the present invention.

[0178] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the processes in the above-mentioned method embodiments by instructing relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments.

[0179] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, etc.

[0180] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0181] Blockchain, as used in this article, refers to a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.

[0182] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The figure shows that only one straight line is used, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13.

[0183] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.

[0184] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.

[0185] Optionally, the computer device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the computer device 1 and to display a visual user interface.

[0186] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0187] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0188] Combine Figure 1 The memory 12 in the computer device 1 stores a plurality of instructions to implement a time series data analysis method based on an improved Transformer, and the processor 13 can execute the plurality of instructions to implement:

[0189] Collecting initial time series data, and performing data enhancement on the initial time series data using an adversarial network to obtain training samples;

[0190] An initial Transformer model is constructed based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy;

[0191] Performing adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model;

[0192] In response to an analysis instruction for target time series data, inputting the target time series data into the time series data analysis model to obtain a time series data analysis result;

[0193] Performing an interpretability analysis on the time series data analysis results using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain an interpretability analysis result;

[0194] Generate an analysis report of the target time series data based on the time series data analysis results and the interpretability analysis results.

[0195] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0196] It should be noted that the data involved in this case were all obtained legally. The software tools or components not produced by our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0197] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division, and actual implementation may employ other division methods.

[0198] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0199] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0200] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0201] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0202] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0203] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the present invention may also be implemented by a single unit or device through software or hardware. Terms such as first and second are used to indicate names and do not imply any particular order.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A time series data analysis method based on an improved Transformer, characterized in that: The time series data analysis method based on the improved Transformer includes: Collecting initial time series data, and performing data enhancement on the initial time series data using an adversarial network to obtain training samples; An initial Transformer model is constructed based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy; Performing adversarial training on the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model; In response to an analysis instruction for target time series data, inputting the target time series data into the time series data analysis model to obtain a time series data analysis result; Performing an interpretability analysis on the time series data analysis results using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain an interpretability analysis result; Generate an analysis report of the target time series data based on the time series data analysis results and the interpretability analysis results.

2. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: The use of the adversarial network to perform data enhancement on the initial time series data to obtain training samples includes: Cleaning the initial time series data to obtain first data; Normalizing the first data to obtain second data; performing sequence block processing on the second data to obtain third data; Using random noise as input, using the generator in the adversarial network to generate a data block similar to the initial time series data, and using the discriminator in the adversarial network to distinguish whether the data block input to the discriminator is a real data block or a generated data block during the generation process; The training sample is constructed using the third data and the generated data block.

3. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: The initial Transformer model based on adaptive position encoding and multi-head attention mechanism is constructed as follows: Calculate the data change between each position and the adjacent positions in the time series data sequence; The following formula is used to position-code each position according to the data variation: Among them, PE (pos,2i) Indicates the position code with index value 2i; PE (pos,2i+1) Indicates the position code with index value (2i+1); pos indicates the position; d model represents the model dimension; β represents the adjustment coefficient; Δx pos represents the data change; i represents the index value.

4. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: The initial Transformer model constructed based on adaptive position encoding and multi-head attention mechanism also includes: For each attention head, linearly transform the input data to obtain the query, key, and value; The attention score is adjusted based on the query, key, and value using the following formula: Among them, Attention new (Q h ,K h ,V h ) represents the adjusted attention score corresponding to the h-th attention head; Q h represents the query corresponding to the h-th attention head; K h represents the key corresponding to the h-th attention head; V h represents the value corresponding to the h-th attention head; d k represents the kth dimension; I represents the feature importance vector.

5. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: The adversarial training of the initial Transformer model based on the training samples and the multi-task joint training strategy to obtain a time series data analysis model includes: Obtaining the weight of each task and constructing a total loss function based on the weight of each task; wherein each task shares the bottom feature extraction layer of the initial Transformer model, and each task corresponds to a different top output layer; During adversarial training, the initial Transformer model is used as a generator, and a discriminator is used to determine whether the output of the initial Transformer model is true; When the total loss function reaches convergence, the training is stopped and the currently obtained model is determined as the time series data analysis model.

6. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: Inputting the target time series data into the time series data analysis model to obtain the time series data analysis results includes: When the time series data analysis model performs a single-step task, the time series data analysis model is used to output a predicted value of the next time step according to the learned time series pattern to obtain the time series data analysis result; or When the time series data analysis model performs a multi-step task, the time series data analysis model is used to take the prediction result of the previous step as the input of the next step prediction, and gradually generate a prediction sequence of multiple time steps to obtain the time series data analysis result.

7. The time series data analysis method based on the improved Transformer according to claim 1, characterized in that: The interpretability analysis of the time series data analysis results is performed using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model, and the interpretability analysis results obtained include: Determining the degree to which the time series data analysis model focuses on different features at different time steps according to the attention weight matrix of the time series data analysis model; Performing cluster analysis on the output features of the intermediate layer to obtain feature patterns; The interpretability analysis result is generated according to the degree of attention paid by the time series data analysis model to different features at different time steps and the feature pattern.

8. A time series data analysis device based on an improved Transformer, characterized in that: The time series data analysis device based on the improved Transformer includes: A data enhancement unit is used to collect initial time series data and perform data enhancement on the initial time series data using an adversarial network to obtain training samples; A construction unit for constructing an initial Transformer model based on adaptive position encoding and a multi-head attention mechanism; wherein the attention score under the multi-head attention mechanism is improved based on a dynamic weight adjustment strategy; A training unit, configured to perform adversarial training on the initial Transformer model based on the training samples and a multi-task joint training strategy to obtain a time series data analysis model; an input unit, configured to input the target time series data into the time series data analysis model in response to an analysis instruction for the target time series data, and obtain a time series data analysis result; An analysis unit, configured to perform an interpretability analysis on the time series data analysis result using the attention distribution characteristics and intermediate layer output characteristics of the time series data analysis model to obtain an interpretability analysis result; A generating unit is used to generate an analysis report of the target time series data based on the time series data analysis result and the interpretability analysis result.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the time series data analysis method based on the improved Transformer as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the time series data analysis method based on the improved Transformer as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Power system frequency safety assessment method and system based on improved GAN and Transformer

    CN121353019A

  • Power system frequency security assessment method and system based on improved gan and transformer

    CN121353019B