Data processing method and related equipment
By generating adversarial networks for power load data processing, the problem of difficulty in capturing seasonal and daily changes in power loads is solved in the existing technology, and more accurate power load prediction and data understanding are achieved.
Patent Information
- Application Number
- CN202411892500.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively capture the seasonal and daily changes in power loads in power demand forecasting, resulting in insufficient prediction results.
Generative adversarial networks (GANs) are used for data processing. Through dimensionality reduction, restoration, game and other steps, the generative adversarial networks are iteratively trained, combined with weight adjustment, and finally a more accurate power load data set is obtained.
It improves the decomposition accuracy and robustness of power load data, improves the effect of seasonal-trend decomposition, enhances the understanding and modeling ability of power load data, and can more accurately predict future power demand.
Smart Images

Figure CN119988831A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing method and related equipment. Background Art
[0002] In the context of big data with the popularization of smart grids, it is increasingly important to predict electricity demand in future periods. Accurate and effective estimation of electricity demand can avoid costly mistakes, help develop electricity supply strategies, financing plans and electricity management, and have practical significance for maintaining the stability of production and life. Accurate analysis and prediction of power data has important guiding significance for grid planning and management decisions of economic departments, but most existing models are only studied on a single time scale. Because the power system has complex dynamics and variability, research on a single time scale may ignore seasonal changes, such as the large difference in electricity demand in summer and winter; it will ignore daily changes, and a single time scale may not be able to fully capture the fluctuations in power load within a day. For example, there may be obvious differences in electricity consumption patterns during the day and at night, and a single time scale may not be able to refine such changes. Summary of the invention
[0003] In view of this, the purpose of this application is to provide a data processing method and related equipment.
[0004] Based on the above purpose, the present application provides a data processing method, including:
[0005] Inputting the preprocessed first data into a generative adversarial network;
[0006] Performing dimensionality reduction processing on the first data to obtain second data;
[0007] Performing restoration processing on the second data to obtain third data;
[0008] generating fourth data based on random noise;
[0009] Performing a game based on the second data and the fourth data to obtain fifth data;
[0010] Calculate the loss based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively train the generative adversarial network; in response to the end of the generative adversarial network training, obtain sixth data;
[0011] The first data and the sixth data are combined based on a first weight to obtain a final data set.
[0012] In a possible implementation manner, the calculating the loss based on the first data, the second data, the third data, the fourth data, and the fifth data includes:
[0013] Calculate a first loss based on the first data, the second data, the third data and the fourth data;
[0014] Based on the first data and the fifth data, a second loss is constructed.
[0015] In a possible implementation manner, the calculating the first loss based on the first data, the second data, the third data, and the fourth data includes:
[0016] Calculate a third loss based on the first data and the third data;
[0017] Calculate a fourth loss based on the second data and the fourth data;
[0018] Based on the second weight, the third loss and the fourth loss are combined to obtain the first loss.
[0019] In a possible implementation, the iterative training of the generative adversarial network includes:
[0020] Training the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss;
[0021] Based on the second loss, the adversary is trained with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
[0022] In a possible implementation, the method further includes:
[0023] Adjusting the first weight according to the performance of the generator of the generative adversarial network;
[0024] And / or, adjusting the first weight according to a rate of decrease of the loss of the generative adversarial network.
[0025] In a possible implementation manner, the first weight corresponds to the sixth data;
[0026] The adjusting the first weight according to the performance of the generator includes:
[0027] In response to an improvement in performance of the generator, increasing the first weight;
[0028] In response to a decrease in performance of the generator, the first weight is decreased.
[0029] In a possible implementation manner, the first weight corresponds to the sixth data;
[0030] The adjusting the first weight according to the loss decreasing speed of the generative adversarial network includes:
[0031] In response to the loss decreasing speed increasing, increasing the first weight;
[0032] In response to the loss decreasing rate decreasing, the first weight is decreased.
[0033] In a possible implementation, the method further includes:
[0034] Decomposing the final data set based on trend to obtain trend components;
[0035] Decomposing the final data set based on seasonality to obtain periodic components;
[0036] Decomposing the final data set based on residuals to obtain residual components;
[0037] The first data is analyzed and processed based on the trend component, the period component and the residual component.
[0038] Based on the same inventive concept, the embodiment of the present application further provides a data processing device, including:
[0039] A generative adversarial network module is configured to input the preprocessed first data into the generative adversarial network; the generative adversarial network module includes a dimensionality reduction network, a restoration network, a generator and an adversary;
[0040] A dimension reduction network is configured to perform dimension reduction processing on the first data to obtain second data;
[0041] A restoration network is configured to perform restoration processing on the second data to obtain third data;
[0042] a generator configured to generate fourth data based on random noise;
[0043] An adversary configured to perform a game based on the second data and the fourth data to obtain fifth data;
[0044] A training module is configured to calculate a loss based on the first data, the second data, the third data, the fourth data, and the fifth data, and iteratively train the generative adversarial network; in response to the end of the generative adversarial network training, obtain sixth data;
[0045] The combining module is configured to combine the first data and the sixth data based on a first weight to obtain a final data set.
[0046] In a possible implementation, the training module is further configured to:
[0047] a calculation unit, configured to calculate a first loss based on the first data, the second data, the third data, and the fourth data;
[0048] The constructing unit is configured to construct a second loss based on the first data and the fifth data.
[0049] In a possible implementation, the computing unit is further configured to:
[0050] Calculate a third loss based on the first data and the third data;
[0051] Calculate a fourth loss based on the second data and the fourth data;
[0052] Based on the second weight, the third loss and the fourth loss are combined to obtain the first loss.
[0053] In a possible implementation, the training module is further configured to:
[0054] A first training unit is configured to train the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss;
[0055] The second training unit is configured to train the adversary based on the second loss with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
[0056] In a possible implementation manner, the device further includes:
[0057] An adjustment module, configured to adjust the first weight according to the performance of the generator of the generative adversarial network;
[0058] And / or, adjusting the first weight according to a rate of decrease of the loss of the generative adversarial network.
[0059] In a possible implementation manner, the first weight corresponds to the sixth data;
[0060] The adjustment module is further configured to:
[0061] In response to an improvement in performance of the generator, increasing the first weight;
[0062] In response to a decrease in performance of the generator, the first weight is decreased.
[0063] In a possible implementation manner, the first weight corresponds to the sixth data;
[0064] The adjustment module is further configured to:
[0065] In response to the loss decreasing speed increasing, increasing the first weight;
[0066] In response to the loss decreasing rate decreasing, the first weight is decreased.
[0067] In a possible implementation manner, the device further includes:
[0068] A first decomposition module is configured to decompose the final data set based on trend to obtain a trend component;
[0069] A second decomposition module is configured to decompose the final data set based on seasonality to obtain a periodic component;
[0070] A third decomposition module is configured to decompose the final data set based on residuals to obtain residual components;
[0071] The analysis module is configured to analyze and process the first data based on the trend component, the periodic component and the residual component.
[0072] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data processing method as described in any one of the above is implemented.
[0073] Based on the same inventive concept, an embodiment of the present application further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned data processing methods.
[0074] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, which includes computer program instructions, and the computer instructions are used to enable the computer program product to execute any of the above-mentioned data processing methods.
[0075] As can be seen from the above, the data processing method and related equipment provided by the present application are as follows: the first data after preprocessing is input into the generative adversarial network; the first data is subjected to dimensionality reduction processing to obtain the second data; the second data is subjected to restoration processing to obtain the third data; the fourth data is generated based on random noise; the fifth data is obtained based on the second data and the fourth data; the loss is calculated based on the first data, the second data, the third data, the fourth data and the fifth data, and the generative adversarial network is iteratively trained; in response to the end of the generative adversarial network training, the sixth data is obtained; the first data and the sixth data are combined based on the first weight to obtain the final data set. The embodiment of the present application can better adapt to the complex load time series data pattern through the generative adversarial network. The generative adversarial network can learn the nonlinear relationship and complex time dependence in the data, thereby improving the accuracy and robustness of the decomposition. It can improve the effect of seasonal-trend decomposition using local regression scatter smoothing method (Seasonal-Trend decomposition using Loess, STL) decomposition, thereby improving the understanding and modeling ability of power load data, and helping to more accurately predict future power demand. It can also learn online in real-time and dynamic environments to adapt to changes in load patterns. This makes it more adaptable and able to handle real-time changes and emergencies in the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0077] Figure 1 A schematic diagram of a data processing method flow chart of an embodiment of the present application;
[0078] Figure 2 This is a schematic diagram of the structure of a data processing device according to an embodiment of the present application;
[0079] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0080] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0081] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the usual meanings understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0082] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0083] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0084] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0085] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0086] As mentioned in the background technology section, in the context of big data with the popularization of smart grids, it has become increasingly important to predict electricity demand in future periods. Accurate and effective estimation of electricity demand can avoid costly mistakes, help develop electricity supply strategies, financing plans and electricity management, and have practical significance for maintaining the stability of production and life. Accurate analysis and prediction of power data has important guiding significance for grid planning and management decisions of economic departments, but most existing models are only studied on a single time scale. Because the power system has complex dynamics and variability, research on a single time scale may ignore seasonal changes, such as the large difference in electricity demand in summer and winter; it will ignore daily changes, and a single time scale may not be able to fully capture the fluctuations in power load within a day. For example, there may be obvious differences in electricity consumption patterns during the day and at night, and a single time scale may not be able to refine such changes.
[0087] Based on the above considerations, the embodiment of the present application proposes a data processing method, by inputting the pre-processed first data into a generative adversarial network; performing dimensionality reduction processing on the first data to obtain second data; restoring the second data to obtain third data; generating fourth data based on random noise; playing games based on the second data and the fourth data to obtain fifth data; calculating losses based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively training the generative adversarial network; in response to the end of the generative adversarial network training, obtaining sixth data; combining the first data and the sixth data based on the first weight to obtain the final data set. The embodiment of the present application can better adapt to complex load time series data patterns through generative adversarial networks. Generative adversarial networks can learn nonlinear relationships and complex time dependencies in data, thereby improving the accuracy and robustness of decomposition. It can improve the effect of seasonal-trend decomposition using local regression scatter point smoothing (Seasonal-Trenddecomposition using Loess, STL) decomposition, thereby improving the understanding and modeling capabilities of power load data, and helping to more accurately predict future power demand. It can also learn online in real-time and dynamic environments to adapt to changes in load patterns. This makes it more adaptable and able to handle real-time changes and emergencies in the power system.
[0088] The technical solutions of the embodiments of the present application are described in detail below through specific examples.
[0089] refer to Figure 1 The data processing method of the embodiment of the present application comprises the following steps:
[0090] Step S101, inputting the preprocessed first data into a generative adversarial network;
[0091] Step S102, performing dimensionality reduction processing on the first data to obtain second data;
[0092] Step S103, restoring the second data to obtain third data;
[0093] Step S104, generating fourth data based on random noise;
[0094] Step S105, performing a game based on the second data and the fourth data to obtain fifth data;
[0095] Step S106, calculating the loss based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively training the generative adversarial network; in response to the end of the generative adversarial network training, obtaining sixth data;
[0096] Step S107: combining the first data and the sixth data based on the first weight to obtain a final data set.
[0097] With respect to step S101, before inputting the first data into the generative adversarial network, the first data needs to be preprocessed.
[0098] In this embodiment, data collection is first performed to obtain raw power load time series data, which usually includes power load measurement values within hourly, daily or shorter time intervals. The source of the data may be actual measurements in the power system, sensor data or other reliable data sources. The correctness and consistency of the timestamp are ensured.
[0099] Check the timestamp format in your data, handle missing or outliers, and ensure that timestamps are in the correct chronological order.
[0100] Outlier handling, identifying and handling outliers, which may include incorrect measurements, missing data, or unreasonable extreme values. Outlier handling can be done by interpolation, smoothing, or deletion, depending on the actual situation.
[0101] Standardization and normalization, standardize or normalize the power load time series data to eliminate the differences between different measurement units and numerical ranges. This helps improve the stability of the model and the training effect.
[0102] Through the above steps, data preprocessing ensures the quality and consistency of the original power load time series data, providing reliable input for the subsequent generative adversarial network and STL load time series decomposition. Such a preprocessing process helps to improve the robustness of the model and ensure that the generated pseudo data can better reflect the characteristics of the original data.
[0103] Furthermore, the preprocessed first data is input into the generative adversarial network.
[0104] Steps S102-S106 of the present application are described based on the following content.
[0105] The generative adversarial network is quite special. The process of processing data is the training process of the generative adversarial network. Therefore, the following will combine training and data processing to illustrate the embodiments of the present application.
[0106] In some embodiments, the loss is calculated based on the first data, the second data, the third data, the fourth data and the fifth data, including: calculating a first loss based on the first data, the second data, the third data and the fourth data; and constructing a second loss based on the first data and the fifth data.
[0107] In some embodiments, the first loss is calculated based on the first data, the second data, the third data and the fourth data, including: calculating the third loss based on the first data and the third data; calculating the fourth loss based on the second data and the fourth data; and obtaining the first loss based on the second weight and combining the third loss and the fourth loss.
[0108] In some embodiments, the iterative training of the generative adversarial network includes: training the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss; and training the adversary based on the second loss with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
[0109] For the generative adversarial network, we first need to initialize the network parameters: initialize the parameters θ of the generator g , the parameters of the adversary θ d , the parameter θ of the dimension reduction network e .
[0110] Furthermore, the generator and dimensionality reduction network are trained.
[0111] Specifically, the training of the generator and the dimensionality reduction network is performed jointly. By minimizing the reconstruction loss L between the dimensionality reduced data and the original data R The temporal loss L between the data output by the generator and the data after dimensionality reduction U To update the parameters of the generator and dimensionality reduction network: (First loss). Among them, λ is a hyperparameter used to weigh the reconstruction loss and timing loss.
[0112] The number of neurons in each gated recurrent unit (GRU) in the hidden layer of the dimension reduction network is n (n is the dimension of the data after dimension reduction). Its input is the high-dimensional real data sequence P (first data), and the network sample batch is set to 128; the output is the dimension-reduced data l = [l1, l2, ..., l T ](second data), where the output at time t is l t =e(l t-1 , P t ), e(·) is the mapping function of the dimensionality reduction network.
[0113] The number of neurons in each GRU in the hidden layer of the restoration network is m (the number of electrical appliances in the factory). The number of electrical appliances in the factory here directly affects the complexity and dimension of the load time series data. Each appliance has a unique load pattern. By using the number of appliances as a reference for the number of GRU neurons, the restoration network can better learn and restore the load characteristics of each appliance, thereby more accurately capturing and representing the load time series data of each appliance in the factory, and improving the performance of the entire system in power load time series decomposition and prediction. Its input is the reduced-dimensional data l, and its output is (third data), wherein r(·) is the restoration network mapping function.
[0114] Construct the loss function L R To measure the performance of the dimensionality reduction network and the restoration network, the third loss is calculated by the following formula:
[0115]
[0116] Where: is the distribution function P data (P) expectation, where P data (P) is the probability distribution of P; L R It is used to calculate the loss of real data after dimensionality reduction and restoration. The training goal is to make L R Close to 0. t is the high-dimensional real data sequence at time t, Restore the output sequence of the network at time t.
[0117] Each GRU in the generator hidden layer contains 128 neurons. The input is a random noise sequence z (random noise), where P noise (z) represents the probability distribution of z, which is a Gaussian distribution. The output data is (Fourth data), where the output at time t is g(·) is the mapping function of the generator. In order to improve the generator’s learning ability of timing characteristics, a timing loss function L is established for it U, through L u The numerical feedback is used to supervise the generator's learning of the data time series features, and the fourth loss is calculated by the following formula:
[0118]
[0119] Where: L U Used to calculate the real low-dimensional data l t And the low-dimensional data output by the generator The training goal is to make L U Close to 0.
[0120] Training the adversary: The goal of the adversary is to maximize the adversarial loss function L V By iteratively updating the adversarial parameters θ d :
[0121] Each GRU in the hidden layer of the adversary contains 128 neurons. Its input is the output of the generator (corresponding to the probability distribution information of generated data) and the output l of the dimensionality reduction network (corresponding to the probability distribution information of real data), the output after the adversary game is (Fifth data), where d(·) is the adversary mapping function, When the adversary determines that the data read this time is real data close to 1, otherwise close to 0. Due to the change in input data, the network's loss function is no longer a measure of the difference in the conditional probability density of two one-dimensional random variables, but the difference in the joint probability density of two multi-dimensional random variables. Generate adversarial loss function L V (Second loss) is:
[0122]
[0123] Where: is the distribution function P noise (z), g represents the generator, and d represents the adversary. The adversarial loss function needs to measure the performance of both the generator and the adversary at the same time, and its ultimate training goal is to achieve a Nash equilibrium between the two, that is, L V Finally it converges to 0.5.
[0124] Update the generator and dimensionality reduction network again: After adversarial training, update the generator and dimensionality reduction network again to minimize the first loss L R +λL U .
[0125] Iterative training: Repeat the above training steps until the performance of the generator and adversary reaches a satisfactory level and the sixth data is obtained. This is an iterative process, and the hyperparameters and number of training times need to be adjusted according to the actual training situation and convergence.
[0126] The training process of GAN is to continuously optimize the generator, adversary and dimensionality reduction network so that they gradually reach a state of equilibrium in the game. The generator learns the distribution of real data, the dimensionality reduction network maps high-dimensional data to low-dimensional space, and the adversary distinguishes the difference between generated data and real data. During the training process, the parameters of these three networks are updated collaboratively to improve the generator's ability to generate realistic data and the dimensionality reduction network's ability to learn time series features, while enabling the adversary to more accurately distinguish generated data from real data. It should be noted that GAN training is a dynamic process, and problems such as mode collapse and gradient disappearance may be encountered during training. Some techniques (such as batch normalization, learning rate adjustment, etc.) need to be adopted to improve the stability of training.
[0127] Further, with respect to step S107, the first data and the sixth data are combined based on the first weight to obtain a final data set.
[0128] In some embodiments, the method further includes: adjusting the first weight according to the performance of the generator of the generative adversarial network; and / or adjusting the first weight according to the loss decrease rate of the generative adversarial network.
[0129] Steps to address data imbalance usually involve properly combining real and generated data during training of a generative adversarial network (GAN) to ensure that the trained model can better capture the characteristics of the data distribution.
[0130] Combine the generated pseudo data with the original real data to build a mixed data set. The generated pseudo data can be combined with the real data by simple superposition, weighted average, etc. in, is the mixed data set, and α is the weight coefficient, which controls the influence of the generated data in the final data.
[0131] Set weight balance: In the training of mixed data sets, the influence of generated pseudo data and real data can be balanced by setting appropriate weights. This can be done by setting different sample weights for different data sources, so that more emphasis is placed on those parts with insufficient data during training.
[0132] In this implementation, in the training of mixed data sets, in order to balance the impact of generated pseudo data and real data, high weights can be set for parts where data is insufficient or difficult to obtain, thereby improving the model's modeling ability for these scarce data; on the contrary, lower weights are set for parts where data is abundant or easy to obtain. At the same time, based on the performance of the model on the validation set, the weights of various types of data can be dynamically adjusted to ensure that the model can fully learn and capture the characteristics of the data during the training process, thereby improving the accuracy and robustness of the model.
[0133] In some embodiments, the first weight corresponds to the sixth data; and the adjusting the first weight according to the performance of the generator includes: increasing the first weight in response to an improvement in the performance of the generator; and decreasing the first weight in response to a decrease in the performance of the generator.
[0134] In this embodiment, the weight is dynamically adjusted according to the data distribution during the training process. If the performance of the generator improves, the weight of the generated data can be gradually increased to ensure that the influence of the generated data on the model training gradually increases. Conversely, if the performance of the generator decreases, the weight of the generated data can be reduced.
[0135] In some embodiments, the first weight corresponds to the sixth data; and adjusting the first weight according to the loss decrease rate of the generative adversarial network includes: increasing the first weight in response to an increase in the loss decrease rate; and decreasing the first weight in response to a decrease in the loss decrease rate.
[0136] In this embodiment, according to the training process of the generative adversarial network, the weight can be dynamically adjusted based on the change of the loss function. For example, when the loss of the generated data decreases rapidly, its weight can be appropriately increased to improve its influence in the training.
[0137] In this embodiment, the distribution of the mixed data set can also be monitored regularly to understand the relative contribution of the generated data and the real data in the training. If it is found that the imbalance of data distribution still exists, the weights and strategies can be adjusted according to the monitoring results.
[0138] A feedback mechanism can also be introduced to adjust the data imbalance handling strategy through the evaluation results of model performance. For example, at the end of each training iteration or cycle, the performance of the model on the validation data is evaluated, and the data weight and balance strategy are adjusted based on the performance.
[0139] In some embodiments, the method further includes: decomposing the final data set based on trend to obtain a trend component; decomposing the final data set based on seasonality to obtain a periodic component; decomposing the final data set based on residual to obtain a residual component; and analyzing and processing the first data based on the trend component, the periodic component and the residual component.
[0140] In this embodiment, the STL algorithm is used to decompose the load data according to three factors: trend, seasonal factors, and irregularity, and decompose it into trend component, periodic component and residual component. The decomposition form can be selected as multiplication form or addition form. This application uses the multiplication form.
[0141] Trend decomposition: First, the trend part is fitted by the Loess method. The trend component T(t) is obtained by the following multiplication form: Where G(·) is the generator, which is used to fit the trend; is the parameter of the trend generator; n is the number of generators, which can be set according to the specific situation; y(t) is the original time series data, where t is the time index; Artificial trend data generated for the generator.
[0142] Seasonal decomposition: The detrended data is y(t) / T(t). Then the seasonal component is fitted using a generative adversarial network to obtain the seasonal component S(t). Here, the seasonal decomposition is expressed in multiplication form: Where G(·) is the generator used to fit seasonality; is the parameter of the season generator; m is the number of generators, which can be set according to the specific situation.
[0143] Residual decomposition: Finally, the residual component is obtained by subtracting the trend and seasonal components from the original data: R(t) = y(t)-T(t)·S(t). In this way, the time series data y(t) is decomposed into three parts: trend T(t), season S(t) and residual R(t).
[0144] The advantage of using the multiplication form for seasonal decomposition in this embodiment is that it can more flexibly adapt to the amplitude changes in different seasons. Different from the addition form of seasonal decomposition, the multiplication form is more suitable for seasonality with relatively stable amplitude.
[0145] The above STL method can more accurately capture the seasonal changes in power load data by separating the seasonal, trend and residual components. This is very important for predicting seasonal peaks and valleys, such as peak electricity demand in summer and heating demand in winter.
[0146] The STL method can extract trends in load data, including long-term changes and trends. This is helpful for understanding the long-term development trend of power load and possible structural changes, such as those caused by population growth, economic changes or technological innovation.
[0147] By separating the residual part, the STL method can reduce the noise interference in the data, making the analysis clearer and more reliable. This helps to improve the accuracy of load forecasting, especially when faced with some random factors.
[0148] In addition, GAN can be used to improve the effect of STL decomposition. By introducing a generative model, it can better adapt to complex load time series data patterns. The generative model can learn nonlinear relationships and complex time dependencies in the data, thereby improving the accuracy and robustness of the decomposition.
[0149] GAN can also perform online learning in real-time and dynamic environments to adapt to changes in load patterns. This makes the method more adaptive and able to handle real-time changes and emergencies in the power system.
[0150] In summary, the GAN-based STL load time series data processing method can improve the understanding and modeling capabilities of power load data, help to more accurately predict future power demand, and improve the operating efficiency and reliability of the power system.
[0151] In some other feasible embodiments, the technical solution of the present application can also be used for data processing in other fields, which will not be introduced in detail here.
[0152] It can be seen from the above embodiments that the data processing method described in the embodiment of the present application is to input the pre-processed first data into the generative adversarial network; perform dimensionality reduction processing on the first data to obtain the second data; perform restoration processing on the second data to obtain the third data; generate the fourth data based on random noise; perform game based on the second data and the fourth data to obtain the fifth data; calculate the loss based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively train the generative adversarial network; in response to the end of the generative adversarial network training, obtain the sixth data; combine the first data and the sixth data based on the first weight to obtain the final data set. The embodiment of the present application can better adapt to the complex load time series data pattern through the generative adversarial network. The generative adversarial network can learn the nonlinear relationship and complex time dependence in the data, thereby improving the accuracy and robustness of the decomposition. It can improve the effect of seasonal-trend decomposition using local regression scatter point smoothing method (Seasonal-Trend decomposition using Loess, STL) decomposition, thereby improving the understanding and modeling ability of power load data, and helping to more accurately predict future power demand. Online learning can also be performed in real-time and dynamic environments to adapt to changes in load patterns. This makes it more adaptable and able to handle real-time changes and emergencies in the power system.
[0153] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.
[0154] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0155] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a data processing device.
[0156] refer to Figure 2 , the data processing device comprises:
[0157] A generative adversarial network 21 is configured to input the preprocessed first data into the generative adversarial network; the generative adversarial network module includes a dimensionality reduction network, a restoration network, a generator and an adversary;
[0158] A dimension reduction network 22 is configured to perform dimension reduction processing on the first data to obtain second data;
[0159] The restoration network 23 is configured to perform restoration processing on the second data to obtain third data;
[0160] A generator 24 configured to generate fourth data based on random noise;
[0161] The antagonist 25 is configured to perform a game based on the second data and the fourth data to obtain fifth data;
[0162] The training module 26 is configured to calculate the loss based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively train the generative adversarial network; in response to the end of the generative adversarial network training, obtain sixth data;
[0163] The combining module 27 is configured to combine the first data and the sixth data based on a first weight to obtain a final data set.
[0164] In some embodiments, the training module 26 is further configured to:
[0165] a calculation unit, configured to calculate a first loss based on the first data, the second data, the third data, and the fourth data;
[0166] The constructing unit is configured to construct a second loss based on the first data and the fifth data.
[0167] In some embodiments, the computing unit is further configured to:
[0168] Calculate a third loss based on the first data and the third data;
[0169] Calculate a fourth loss based on the second data and the fourth data;
[0170] Based on the second weight, the third loss and the fourth loss are combined to obtain the first loss.
[0171] In some embodiments, the training module 26 is further configured to:
[0172] A first training unit is configured to train the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss;
[0173] The second training unit is configured to train the adversary based on the second loss with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
[0174] In some embodiments, the apparatus further comprises:
[0175] An adjustment module, configured to adjust the first weight according to the performance of the generator of the generative adversarial network;
[0176] And / or, adjusting the first weight according to a rate of decrease of the loss of the generative adversarial network.
[0177] In some embodiments, the first weight corresponds to the sixth data;
[0178] The adjustment module is further configured to:
[0179] In response to an improvement in performance of the generator, increasing the first weight;
[0180] In response to a decrease in performance of the generator, the first weight is decreased.
[0181] In some embodiments, the first weight corresponds to the sixth data;
[0182] The adjustment module is further configured to:
[0183] In response to the loss decreasing speed increasing, increasing the first weight;
[0184] In response to the loss decreasing rate decreasing, the first weight is decreased.
[0185] In some embodiments, the apparatus further comprises:
[0186] A first decomposition module is configured to decompose the final data set based on trend to obtain a trend component;
[0187] A second decomposition module is configured to decompose the final data set based on seasonality to obtain a periodic component;
[0188] A third decomposition module is configured to decompose the final data set based on residuals to obtain residual components;
[0189] The analysis module is configured to analyze and process the first data based on the trend component, the periodic component and the residual component.
[0190] For the convenience of description, the above device is described in terms of functions divided into various modules. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0191] The device of the above embodiment is used to implement the corresponding data processing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0192] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data processing method described in any of the above embodiments is implemented.
[0193] Figure 3 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0194] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0195] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0196] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0197] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0198] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0199] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0200] The electronic device of the above embodiment is used to implement the corresponding data processing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0201] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the data processing method described in any of the above embodiments.
[0202] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0203] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the data processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0204] Based on the same inventive concept, corresponding to the data processing method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the data processing method. Corresponding to the execution subject corresponding to each step in each embodiment of the data processing method, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0205] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the data processing method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0206] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0207] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0208] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0209] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A data processing method, characterized in that: include: Inputting the preprocessed first data into a generative adversarial network; Performing dimensionality reduction processing on the first data to obtain second data; Performing restoration processing on the second data to obtain third data; generating fourth data based on random noise; Performing a game based on the second data and the fourth data to obtain fifth data; Calculate the loss based on the first data, the second data, the third data, the fourth data and the fifth data, and iteratively train the generative adversarial network; In response to the completion of the generative adversarial network training, obtaining sixth data; The first data and the sixth data are combined based on a first weight to obtain a final data set.
2. The method according to claim 1, characterized in that The calculating the loss based on the first data, the second data, the third data, the fourth data and the fifth data comprises: Calculate a first loss based on the first data, the second data, the third data and the fourth data; Based on the first data and the fifth data, a second loss is constructed.
3. The method according to claim 2, characterized in that The calculating the first loss based on the first data, the second data, the third data and the fourth data includes: Calculate a third loss based on the first data and the third data; Calculate a fourth loss based on the second data and the fourth data; Based on the second weight, the third loss and the fourth loss are combined to obtain the first loss.
4. The method according to claim 2, characterized in that: The iterative training of the generative adversarial network comprises: Training the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss; Based on the second loss, the adversary is trained with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
5. The method according to claim 1, characterized in that The method further comprises: Adjusting the first weight according to the performance of the generator of the generative adversarial network; And / or, adjusting the first weight according to a rate of decrease of the loss of the generative adversarial network.
6. The method according to claim 5, characterized in that The first weight corresponds to the sixth data; The adjusting the first weight according to the performance of the generator includes: In response to an improvement in performance of the generator, increasing the first weight; In response to a decrease in performance of the generator, the first weight is decreased.
7. The method according to claim 5, characterized in that The first weight corresponds to the sixth data; The adjusting the first weight according to the loss decreasing speed of the generative adversarial network includes: In response to the loss decreasing speed increasing, increasing the first weight; In response to the loss decreasing rate decreasing, the first weight is decreased.
8. The method according to claim 1, characterized in that: The method further comprises: Decomposing the final data set based on trend to obtain trend components; Decomposing the final data set based on seasonality to obtain periodic components; Decomposing the final data set based on residuals to obtain residual components; The first data is analyzed and processed based on the trend component, the period component and the residual component.
9. A data processing device, characterized in that: include: A generative adversarial network module, configured to input the preprocessed first data into the generative adversarial network; The generative adversarial network module includes a dimensionality reduction network, a restoration network, a generator and an adversary; A dimension reduction network is configured to perform dimension reduction processing on the first data to obtain second data; A restoration network is configured to perform restoration processing on the second data to obtain third data; a generator configured to generate fourth data based on random noise; An adversary configured to perform a game based on the second data and the fourth data to obtain fifth data; A training module, configured to calculate a loss based on the first data, the second data, the third data, the fourth data, and the fifth data, and iteratively train the generative adversarial network; In response to the completion of the generative adversarial network training, obtaining sixth data; The combining module is configured to combine the first data and the sixth data based on a first weight to obtain a final data set.
10. The device according to claim 9, characterized in that The training module is further configured to: a calculation unit, configured to calculate a first loss based on the first data, the second data, the third data, and the fourth data; The constructing unit is configured to construct a second loss based on the first data and the fifth data.
11. The device according to claim 10, characterized in that The computing unit is further configured to: Calculate a third loss based on the first data and the third data; Calculate a fourth loss based on the second data and the fourth data; Based on the second weight, the third loss and the fourth loss are combined to obtain the first loss.
12. The device according to claim 10, characterized in that The training module is further configured to: A first training unit is configured to train the dimension reduction network and the generator in the generative adversarial network with the goal of minimizing the first loss; The second training unit is configured to train the adversary based on the second loss with the goal of achieving Nash equilibrium between the performance of the generator and the performance of the adversary in the generative adversarial network.
13. The device according to claim 9, characterized in that The device also includes: An adjustment module, configured to adjust the first weight according to the performance of the generator of the generative adversarial network; And / or, adjusting the first weight according to a rate of decrease of the loss of the generative adversarial network.
14. The device according to claim 13, characterized in that The first weight corresponds to the sixth data; The adjustment module is further configured to: In response to an improvement in performance of the generator, increasing the first weight; In response to a decrease in performance of the generator, the first weight is decreased.
15. The device according to claim 13, characterized in that The first weight corresponds to the sixth data; The adjustment module is further configured to: In response to the loss decreasing speed increasing, increasing the first weight; In response to the loss decreasing rate decreasing, the first weight is decreased.
16. The device according to claim 9, characterized in that The device also includes: A first decomposition module is configured to decompose the final data set based on trend to obtain a trend component; A second decomposition module is configured to decompose the final data set based on seasonality to obtain a periodic component; A third decomposition module is configured to decompose the final data set based on residuals to obtain residual components; The analysis module is configured to analyze and process the first data based on the trend component, the periodic component and the residual component.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
18. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.
19. A computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 8.