Power asset anomaly detection method based on time sequence prediction
By building a time series prediction model based on transformer, dynamically detecting power asset abnormalities, solving the problem of inefficient detection of power asset abnormalities in the existing technology, achieving efficient and accurate abnormal detection and real-time security policy implementation, and improving the safety of the power system.
Patent Information
- Application Number
- CN202510462236.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has low efficiency in power asset abnormality detection in power systems, making it difficult to achieve dynamic evaluation and real-time processing, and cannot discover potential abnormal data in a timely manner.
Using a time series prediction method, a time series prediction model based on power assets is constructed by collecting basic, business and security attribute data of power assets, and a transformer-based time series prediction model is optimized using forward propagation and backpropagation, and anomaly behavior is detected in combination with Adam optimizer and cosine similarity.
It realizes efficient and accurate detection of abnormalities of power assets, supports real-time execution of safety policies, and improves the safety protection capabilities of the power system.
Smart Images

Figure CN120387032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial control system security, and particularly to a method for detecting anomalies in power assets based on time series prediction. Background Art
[0002] Identifying anomalies in power assets is a crucial task in the power system. Its core objective is to timely detect and locate potential abnormal data to ensure the stable operation of the power system. Power assets exhibit typical big data characteristics, and the detection and location of abnormal data face great challenges.
[0003] Currently, the discovery of power abnormal data mainly relies on traditional passive rule verification and script detection. These methods are usually time-consuming, laborious and inefficient. Traditional anomaly detection methods take an average of 48 hours to discover some abnormal data. With the power system facing more and more security challenges, many scholars have proposed hierarchical analysis models, dynamic risk assessment schemes based on neural networks, and assessment schemes based on hidden Markov models. These methods can evaluate the security of power assets to a certain extent, but still have limitations such as high computational complexity and inconvenient update.
[0004] Patent CN115983250A discloses a method and system for locating the root cause of power abnormal data based on a knowledge graph; it sorts out and extracts knowledge from the data of the target power system, constructs a corresponding knowledge graph; uses parsing algorithms and regular expressions in natural language processing to intelligently analyze the source of abnormal data; adopts parallel computing technology to improve the efficiency of data analysis. However, it is unable to dynamically evaluate the anomalies of power assets in order to adopt real-time processing strategies such as immediately blocking communication. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for detecting anomalies in power assets based on time series prediction to achieve dynamic detection of abnormal behaviors of power assets.
[0006] To achieve the above purpose, the present invention adopts the following technical scheme: A method for detecting anomalies in power assets based on time series prediction, comprising the following steps:
[0007] Step 1: Collect the basic attribute data, service attribute data and security attribute data of power assets, and preprocess the collected data based on time characteristics;
[0008] Step 2: Construct a time series prediction model based on transformer, process the input data through forward propagation, calculate the gradient through backpropagation to update the model parameters, and then optimize the model based on Adam and evaluate the model performance and anomaly handling;
[0009] Step 3: Detect anomalies in power assets.
[0010] In a preferred embodiment, in step 1, the basic attribute data of power assets is collected to describe the hardware, operating system, temperature, and load basic information of power equipment, and to provide the operating health status and performance level of the assets. Specifically, it includes: collecting the asset name, model, operating system type, and manufacturer data to distinguish different types of equipment and be used to identify the equipment when establishing an asset model; collecting asset partition, the hierarchy, and usage scenario data to identify the role of the equipment in the network topology and physical hierarchy; collecting the power module status, fan temperature status, motherboard temperature status, memory utilization rate, and CPU utilization rate data of the asset to identify the operating health status of the equipment.
[0011] In a preferred embodiment, in step 1, the business attribute data of power assets is collected, including IP address, port information, and network traffic communication behavior data, to provide network data for analyzing whether there are abnormalities in the assets. Specifically, it includes: collecting IPv4 format addresses, IP network segments, MAC addresses, port numbers, network access types, transport layer protocols, and application layer protocol data to establish the network communication characteristics of the equipment; collecting service traffic volume, byte count, transport layer protocol used by the port, services opened by the port, and port online status data to identify the communication activity of the assets; collecting network card status, network packet loss rate, and network error packet rate data to measure the network quality and equipment health status.
[0012] In a preferred embodiment, in step 1, the security attribute data of power assets is collected, specifically including threat event number, threat event type, threat status, threat level, and detection of the cause of the threat event, to identify potential security risks.
[0013] In a preferred embodiment, in step 1, data preprocessing specifically includes extracting time series features, handling missing values and abnormal data in the data, and data standardization. First, time series features are extracted from the original data. For the basic attribute data of power assets, for asset temperature and utilization data, the current value, maximum value, mean value, and change rate are extracted. For power module data, a state sequence is extracted according to its state change characteristics. For asset business attribute data, for IP and port features, address, port, and network card status information are extracted, the network activity intensity of the device is calculated, and statistical analysis is performed on the data to generate a time series. For service and protocol features, behavior features at the network protocol layer are extracted. For security attribute data, the threat event number, threat level, and status information of security threat events are extracted, each threat event is marked with a timestamp and converted into a time series. The data is processed by time windowing, that is, the data within each time period is divided into windows of a fixed length. After extracting features for each window, missing value handling and anomaly detection are performed. When handling missing values, if a piece of data has many missing values, this piece of feature data is directly deleted. If the missing values are few, interpolation or mean value methods are used for data filling. For the detected outliers, statistical methods such as Z-score are used to detect and remove the outliers, and then standardization and normalization processing are performed. For the data that has completed the preprocessing operation, it is divided into a training set, a validation set, and a test set, accounting for 80%, 10%, and 10% respectively. The size of the preset time window is 24h, and the data is divided into multiple <input, output> pairs, that is, the historical data of the past 24h is selected as the input to predict the power asset data within the next 1h.
[0014] In a preferred embodiment, in the transformer sequence prediction model of step 2, the encoder is used to process time series data. The encoder consists of multiple self-attention layers and a feed-forward neural network, allowing the model to focus on other time steps in the sequence at each time step and capture long-term dependencies. During the process of processing the sequence, positional encoding is added to the input time series data to ensure that the model can perceive the time order of the data. The input of the model is a tensor with the shape of batch size, sequence length, and input feature dimension, and the output is a tensor with the shape of batch size and prediction range.
[0015] In a preferred embodiment, the transformer model in step 2 adopts a binary cross-entropy loss function, as shown in the following formula. The distribution of normal and abnormal data in the data is unbalanced, and different weights are set for different categories to make the model pay more attention to abnormal samples.
[0016]
[0017] where y i is the true label, is the probability predicted by the model, and N is the total number of data.
[0018] In a preferred embodiment, the Transformer model in step 2 calculates gradients through backpropagation and updates the model parameters; the goal of backpropagation is to calculate the gradient of the loss function with respect to each parameter through the chain rule, including the weights of the input embedding layer, the weight matrices in the self-attention mechanism, and the weight matrices in the feed-forward neural network; when calculating the gradients, first take the derivative of the loss function to obtain each predicted value the gradient of the loss function, that is, the predicted value deviates from the true label y i the degree; then, based on the reciprocal of the loss function, calculate the gradients of the loss function with respect to the parameters in turn, including the gradient with respect to the model output and the gradient with respect to the model weights; finally, apply the chain rule to calculate the gradients of the loss function with respect to the parameters of each layer; after completing the gradient calculation, use an optimization algorithm to update the model parameters, adopt the Adam optimizer, and dynamically adjust the learning rate based on the moving average of the gradients; first, calculate the first moment estimate, specifically as follows:
[0019]
[0020] where β1 is the decay rate controlling the momentum, m t is the first moment estimate at the current time t, m t-1 is the first moment estimate at the previous time t-1, L is the loss function, and θ is the parameter; then, calculate the second moment estimate, specifically as follows:
[0021]
[0022] where β2 is the decay rate controlling the adaptive learning rate, v t is the second moment estimate at the current time t, v t-1 the second moment estimate at the previous time t-1; next, calculate the correction bias, specifically as follows:
[0023]
[0024] where, is the corrected first moment estimate, is the corrected second moment estimate, is the power of β1 at time step t, similarly; finally, update the model parameters, specifically as follows:
[0025]
[0026] Among them, η is the learning rate, and ∈ is a small constant used to prevent division-by-zero errors. Through the above steps, the model parameters are updated according to the calculated gradients, enabling the model to better fit the data in the next round of training and gradually minimizing the loss function.
[0027] Based on the trained transformer model, the model performance is evaluated on the test set. The root mean square error (RMSE) method is used to evaluate the deviation between the predicted value and the true value, and it is sensitive to outliers in the data. The method is as follows:
[0028]
[0029] According to the constructed transformer time series prediction model, the regular state prediction function of power assets is constructed as follows:
[0030]
[0031] Among them, X t-p+1:t is the input time series data set, W out ∈R d×1 is the weight matrix, and b out is the bias term. The output of the last layer of the encoder, EncoderOutput t is used as the prediction input for time t + 1, as follows:
[0032] EncoderOutput = LaverNorm(MultiHead(Q, K, V) + X')
[0033] After passing through multiple self-attention layers and feed-forward neural network layers, where Q, K, and V are the query, key, and value matrices respectively, and X' is the input.
[0034] Power asset data is collected in real time, the time series data for future time is predicted, and it is compared with the real data to detect whether there is a large deviation in the current data and determine whether there is abnormal behavior in the power assets. The similarity calculation method uses cosine similarity. If the similarity is lower than the threshold, it indicates that the data is abnormal. The similarity S i The calculation method is as follows:
[0035]
[0036] Among them, n is the total number of data, and the threshold is an empirical parameter that is adaptively adjusted based on each detection result to continuously optimize the model and improve the detection accuracy.
[0037] In a preferred embodiment, the value of β1 is 0.9.
[0038] In a preferred embodiment, the value of β2 is 0.999.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] The present invention collects the basic attribute data, service attribute data and security attribute data of power assets, preprocesses the collected data based on time characteristics, and forms an anomaly detection training data set D, a model test data set D', and a real-time detection data set X. Constructs a time series prediction model based on transformer, processes the input data through forward propagation, calculates the gradient through backpropagation to update the model parameters, and updates the model parameters θ based on Adam. Calculates the model error RMSE based on the test data set D' and optimizes the model. Constructs a time series prediction function f Transformer (X t-p+1:t ) for the normal state of power assets based on the time series prediction model. Compares the deviation between the real-time detection data set X and the prediction result. If it is less than the threshold, it indicates that there is an abnormal behavior in the current asset, and the security policy is executed in real time to realize the dynamic detection of abnormal behaviors of power assets.
[0041] The power asset anomaly detection method proposed by the present invention uses a transformer model to achieve efficient and highly accurate detection of abnormal data. The threshold is adaptively adjusted to continuously optimize the model accuracy. By dynamically detecting abnormal behaviors of power assets, it supports the execution of real-time security policies and improves the security protection ability of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flow chart of a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] The present invention will be further described below with reference to the drawings and embodiments.
[0044] It should be noted that the following detailed description is illustrative and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0046] Such as Figure 1As shown, the present invention provides a method for detecting anomalies in power assets based on time series prediction, comprising the following steps:
[0047] In step (1), in the power asset data collection and preprocessing stage, sufficient raw data is obtained from the power assets to analyze the operating status, network status, and security events of the assets. First, collect the basic attribute data of the assets to describe the basic information such as the hardware, operating system, temperature, load, etc. of the power equipment, and provide the operating health status and performance level of the assets. Specifically, collect data such as the asset name, model, operating system type, manufacturer, etc. to distinguish different types of devices and be used to identify the devices when establishing the asset model; collect data such as the asset's partition, hierarchy, usage scenario, etc. to identify the role of the device in the network topology and physical hierarchy; collect data such as the power supply module status, fan temperature status, motherboard temperature status, memory utilization rate, CPU utilization rate, etc. of the asset to identify the operating health status of the device.
[0048] In step (2), collect the business attribute data of the assets, including communication behavior data such as IP addresses, port information, network traffic, etc., to provide network data for analyzing whether the assets are abnormal. Specifically, collect data such as IPv4 format addresses, IP network segments, MAC addresses, port numbers, network access types, transport layer protocols, application layer protocols, etc. to establish the network communication characteristics of the device; collect data such as business traffic volume, number of bytes, transport layer protocol used by the port, services provided by the open port, online status of the port, etc. to identify the communication activity of the asset; collect data such as network card status, network packet loss rate, network error packet rate, etc. to measure the network quality and device health status.
[0049] In step (3), collect the security attribute data of the assets, specifically including threat event numbers, threat event types, threat status, threat levels, detection of threat event causes, etc., to identify potential security risks.
[0050] Step (4) performs data preprocessing based on the power asset data in steps (1), (2), and (3), specifically including extracting time series features, handling missing values and abnormal data in the data, and data standardization. First, extract time series features from the original data. For the basic attribute data of power assets, for data such as asset temperature and utilization rate, extract the current value, maximum value, mean value, change rate, etc.; for power module data, extract the state sequence according to its state change characteristics. For the asset business attribute data, for IP and port characteristics, extract information such as address, port, network card status, etc., calculate the network activity intensity (traffic, number of bytes, number of connections, etc.) of the device, perform statistical analysis on these data, and generate time series; for service and protocol characteristics, extract the behavioral characteristics at the network protocol level. For security attribute data, extract information such as threat event number, threat level, status, etc. of security threat events, mark each threat event with a timestamp and convert it into a time series. In this step, perform time windowing on the data, that is, divide the data within each time period into windows of a fixed length, and extract features for each window;
[0051] Step (5) performs missing value handling and anomaly detection based on the data in step (4). When handling missing values, if a piece of data has too many missing values, directly delete this piece of feature data; if the missing values are few, use interpolation or mean method to fill the data. For the detected outliers, the reasons may be noise during the acquisition process or extreme anomalies of the device. Use statistical methods such as Z-score to detect and remove the outliers;
[0052] Step (6) performs standardization and normalization processing on the data in step (5) to make different features have similar scales and complete the construction of time series data. For the data that has completed the preprocessing operation, divide it into a training set, a validation set, and a test set, accounting for 80%, 10%, and 10% respectively. Preset the size of the time window to 24h, and divide the data into multiple <input, output> pairs, that is, select the historical data of the past 24h as the input to predict the power asset data within the next 1h;
[0053] Step (7) inputs the training data set into the transformer model according to the time series data in step (6), performs forward propagation, and calculates the predicted value. In the transformer sequence prediction model of the present invention, use the encoder to process the time series data. The encoder consists of multiple self-attention layers and a feed-forward neural network, allowing the model to focus on other time steps in the sequence at each time step and capture long-term dependencies. During the process of processing the sequence, add position encoding to the input time series data to ensure that the model can perceive the time order of the data. The model input is a tensor with the shape of (batch size, sequence length, input feature dimension), and the output is a tensor with the shape of (batch size, prediction range);
[0054] Step (8). For the Transformer model described in step (7), the binary cross-entropy loss function is adopted. As shown in the following formula, since the distribution of normal and abnormal data in the data is unbalanced, different weights are set for different categories to make the model pay more attention to abnormal samples;
[0055]
[0056] where y i is the true label, is the probability predicted by the model, and N is the total number of data;
[0057] Step (9). For the Transformer model described in step (7), calculate the gradient through backpropagation and update the model parameters. The goal of backpropagation is to calculate the gradient of the loss function with respect to each parameter through the chain rule, including the weights of the input embedding layer, the weight matrix in the self-attention mechanism, the weight matrix in the feed-forward neural network, etc. When calculating the gradient, first take the derivative of the loss function to obtain the gradient of each predicted value with respect to the loss function, that is, the degree to which the predicted value deviates from the true label y i . Then, based on the reciprocal of the loss function, calculate the gradient of the loss function with respect to the parameters in turn, including the gradient of the model output, the gradient of the model weights, etc. Finally, apply the chain rule to calculate the gradient of the loss function with respect to the parameters of each layer. After completing the gradient calculation, use an optimization algorithm to update the model parameters. The Adam optimizer is adopted, and the learning rate is dynamically adjusted based on the moving average of the gradient. First, calculate the first moment estimate as follows:
[0058]
[0059] where β1 is the decay rate controlling the momentum. Then, calculate the second moment estimate as follows:
[0060]
[0061] where β2 is the decay rate controlling the adaptive learning rate. Next, calculate the correction bias as follows:
[0062]
[0063] Finally, update the model parameters as follows:
[0064]
[0065] Among them, η is the learning rate, and ∈ is a small constant used to prevent division-by-zero errors. Through the above steps, the model parameters are updated according to the calculated gradients, enabling the model to better fit the data in the next round of training and gradually minimizing the loss function;
[0066] Step (10), based on the transformer model trained in step (9), evaluate the model performance on the test set. Using the root mean square error (RMSE) method, evaluate the deviation between the predicted value and the true value, and it is sensitive to outliers in the data. The method is as follows:
[0067]
[0068] where N is the total number of samples in the test set, y i is the true value at the i-th time point, is the predicted value at the i-th time point;
[0069] Step (11), based on the transformer time series prediction model constructed in step (9), construct the regular state prediction function of the power asset as follows:
[0070] f Transformer (X t-p+1:t ) = W out ·EncoderOutput t + b out
[0071] where, X t-p+1:t is the input time series data set, W out ∈ R d×1 , is the weight matrix, b out is the bias term. Use the output of the last layer of the encoder, EncoderOutput t as the predicted input for time t + 1, as follows:
[0072] EncoderOutput = LayerNorm(MultiHead(Q, K, V) + X')
[0073] After passing through multiple self-attention layers and feed-forward neural network layers, where Q, K, V are the query, key, and value matrices respectively, and X' is the input;
[0074] Step (12), collect power asset data in real time, predict the time series data of the equipment for future time based on step (11), and compare it with the real data to detect whether there is a large deviation in the current data and determine whether there is abnormal behavior in the power asset. The similarity calculation method uses cosine similarity. If the similarity is lower than the threshold, it indicates that the data is abnormal. The similarity calculation method is as follows:
[0075]
[0076] Among them, the threshold is an empirical parameter, which is adaptively adjusted based on the detection results each time, continuously optimizing the model to improve the detection accuracy.
[0077] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for detecting anomalies in power assets based on time series prediction, characterized in that, It includes the following steps: Step 1: Collect the basic attribute data, business attribute data, and security attribute data of power assets, and preprocess the collected data based on time characteristics; Step 2: Build a time series prediction model based on transformer, process the input data through forward propagation, calculate the gradient to update the model parameters through backward propagation, and then optimize the model based on Adam and evaluate the model performance and anomaly handling; Step 3: Power asset anomaly detection.
2. The method for detecting abnormal power assets based on time series prediction according to claim 1, wherein, In Step 1, collect the basic attribute data of power assets, describe the hardware, operating system, temperature, and load basic information of power equipment, and provide the operating health status and performance level of the assets. Specifically, it includes: collecting the asset name, model, operating system type, and manufacturer data to distinguish different types of equipment and identify the equipment when building an asset model; collecting asset partition, the hierarchy, and usage scenario data to identify the role of the equipment in the network topology and physical hierarchy; collecting the power module status, fan temperature status, motherboard temperature status, memory utilization rate, and CPU utilization rate data of the asset to identify the operating health status of the equipment.
3. A method for abnormal detection of power assets based on time series prediction according to claim 1, characterized in that, In Step 1, collect the business attribute data of power assets, including IP address, port information, and network traffic communication behavior data, to provide network data for analyzing whether the asset is abnormal. Specifically, it includes: collecting IPv4 format addresses, IP network segments, MAC addresses, port numbers, network access types, transport layer protocols, and application layer protocol data to establish the network communication characteristics of the device; collecting business traffic volume, byte count, transport layer protocol used by the port, services opened by the port, and port online status data to identify the communication activity of the asset; collecting network card status, network packet loss rate, and network error packet rate data to measure the network quality and equipment health status.
4. The method for detecting abnormal power assets based on time series prediction according to claim 1, wherein, In Step 1, collect the security attribute data of power assets, specifically including threat event number, threat event type, threat status, threat level, and detection of threat event reasons to identify potential security risks.
5. A method for abnormal detection of power assets based on time series prediction according to claim 1, characterized in that, In Step 1, data preprocessing specifically includes extracting time series features, handling missing values and abnormal data in the data, and data standardization. First, extract time series features from the original data; for the basic attribute data of power assets, for asset temperature and utilization rate data, extract the current value, maximum value, average value, and change rate; For power module data, extract the status sequence according to its status change characteristics; for the business attribute data of the asset, for IP and port characteristics, extract address, port, and network card status information, calculate the network activity intensity of the device, conduct statistical analysis on the data, and generate a time series; for service and protocol characteristics, extract the behavior characteristics at the network protocol level. For security attribute data, extract the threat event number, threat level, and status information of security threat events, mark each threat event with a timestamp and convert it into a time series; perform time windowing on the data, that is, divide the data within each time period into windows of a fixed length, extract features for each window, and then perform missing value processing and anomaly detection; during missing value processing, if a data entry has many missing values, directly delete this feature data; if the missing values are few, use interpolation or the mean method to fill in the data; for the detected outliers, use statistical methods such as Z-score to detect and remove the outliers, and then perform standardization and normalization processing; for the data that has completed the preprocessing operations, divide it into a training set, a validation set, and a test set, accounting for 80%, 10%, and 10% respectively; the size of the preset time window is 24h, and the data is divided into multiple <input, output> pairs, that is, select the historical data of the past 24h as the input to predict the power asset data within the next 1h.
6. A method for detecting abnormal power assets based on time series prediction according to claim 1, characterized in that, In the transformer sequence prediction model of step 2, use the encoder to process the time series data; the encoder consists of multiple self-attention layers and a feed-forward neural network, allowing the model to focus on other time steps in the sequence at each time step and capture long-term dependencies; During the process of processing the sequence, add positional encoding to the input time series data to ensure that the model can perceive the time order of the data; the model input is a tensor with the shape of batch size, sequence length, and input feature dimension, and the output is a tensor with the shape of batch size and prediction range.
7. A method for abnormal detection of power assets based on time series prediction according to claim 6, characterized in that, The transformer model in step 2 uses the binary cross-entropy loss function, as shown in the following formula. The normal and abnormal data distributions in the data are unbalanced, and different weights are set for different classes to make the model pay more attention to abnormal samples; where y i is the true label, is the probability predicted by the model, and N is the total number of data.
8. A method for abnormal detection of power assets based on time series prediction according to claim 7, characterized in that, The Transformer model in step 2 calculates the gradients through backpropagation and updates the model parameters; the goal of backpropagation is to calculate the gradients of the loss function with respect to each parameter through the chain rule, including the weights of the input embedding layer, the weight matrices in the self-attention mechanism, and the weight matrices in the feed-forward neural network; when calculating the gradients, first take the derivative of the loss function to obtain each predicted value the gradient of the loss function, that is, the predicted value deviates from the true label y i to the extent; then, based on the reciprocal of the loss function, calculate the gradients of the loss function with respect to the parameters in turn, including the gradients with respect to the model output and the gradients with respect to the model weights; finally, apply the chain rule to calculate the gradients of the loss function with respect to the parameters of each layer; after completing the gradient calculation, use an optimization algorithm to update the model parameters, adopt the Adam optimizer, and dynamically adjust the learning rate based on the moving average of the gradients; first, calculate the first-order moment estimate, specifically as follows: where β1 is the decay rate for controlling momentum, and m t is the first moment estimate at the current time t, and m t-1 is the first moment estimate at the previous time t-1, L is the loss function, and θ is the parameter; then, calculate the second moment estimate as follows: where β2 is the decay rate for controlling the adaptive learning rate, v t is the second moment estimate at the current time t, and v t-1 is the second moment estimate at the previous time t-1; Next, calculate the correction bias as follows: Among them, is the corrected first moment estimate, is the corrected second moment estimate, is the power of β1 at time step t, Similarly; finally, update the model parameters as follows: Among them, η is the learning rate, ∈ is a small constant used to prevent division-by-zero errors; through the above steps, update the model parameters according to the calculated gradients, so that the model can better fit the data in the next round of training and gradually minimize the loss function; According to the trained transformer model, evaluate the model performance on the test set, using the root mean square error (RMSE) method to evaluate the deviation between the predicted value and the true value, and it is sensitive to outliers in the data. The method is as follows: According to the constructed transformer time series prediction model, construct the regular state prediction function of the power asset as follows: f Transformer (X t-p+1:t ) = W out ·EncoderOutput t + b out Among them, X t-p+1:t is the input time series data set, W out ∈R d×1 is the weight matrix, and b out is the bias term; the output EncoderOutput of the last layer of the encoder t is used as the predicted input for time t+1, specifically as follows: EncoderOutput = LayerNorm(MultiHead(Q, K, V) + X') After passing through multiple self-attention layers and feed-forward neural network layers, where Q, K, and V are the query, key, and value matrices respectively, and X' is the input; Collect power asset data in real time, predict time series data for future time, compare it with real data, detect whether there are large deviations in the current data, and determine whether there are abnormal behaviors in power assets; the similarity calculation method uses cosine similarity. If the similarity is lower than the threshold, it indicates that the data is abnormal, and the similarity is S i The calculation method is as follows: Among them, n is the total number of data, the threshold is an empirical parameter, which is adaptively adjusted based on each detection result, continuously optimize the model, and improve the detection accuracy.
9. The method for detecting abnormal power assets based on time series prediction according to claim 8, characterized in that, The value of β1 is 0.
9.
10. A method for abnormal detection of power assets based on time series prediction according to claim 8, characterized in that, The value of β2 is 0.999.