Power data processing method, device and equipment based on sand table simulation deduction

CN122734484APending Publication Date: 2026-09-11CHINA DATANG TECH & ECONOMY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610989652.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]本发明提供了一种基于沙盘仿真推演的电力数据处理方法、装置及设备,解决了电力推演规则分离、时序处理差与协同难的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734484A_ABST
    Figure CN122734484A_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, and equipment for power data processing based on sand table simulation, belonging to the field of information processing technology. The method includes: acquiring multi-source heterogeneous data from the power industry; performing time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset; determining a subset of core factors based on the standard power time-series dataset; inputting the subset of core factors into a hybrid simulation model for processing to obtain initial indicator calculation results; determining the scenario game simulation results based on simulation scenario parameters, initial indicator calculation results, and multi-agent game simulation rules; determining risk classification and risk transmission path information based on the scenario game simulation results and preset risk assessment rules; and determining optimization strategy data and a simulation report based on the risk classification and risk transmission path information. This solution improves the accuracy and efficiency of power simulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer information processing technology, and in particular to a method, apparatus and equipment for processing power data based on sand table simulation. Background Technology

[0002] In recent years, time-of-use pricing mechanisms, the parallel operation of medium- and long-term and spot trading, fluctuations in renewable energy output, and carbon emission constraints have transformed power decision-making from planned dispatch under a single electricity price model to a complex problem involving multiple time scales, trading instruments, and market participants. Currently, the commonly used simulation methods in the industry mainly fall into three categories: manual calculation based on spreadsheets, relying on manual data processing and indicator calculation; general-purpose business simulation software, using built-in financial models for simulation; and general-purpose artificial intelligence analysis platforms, employing machine learning algorithms to model data and extrapolate trends. In addition, the power industry has also explored related areas such as system operation simulation and market transaction simulation, but these mainly focus on single business segments.

[0003] The aforementioned existing technologies rely on manual data cleaning, alignment, and aggregation, resulting in low efficiency and susceptibility to inaccuracies. Their computational capabilities are limited to linear operations, making it difficult to handle nonlinear relationships between complex time-series variables such as time-of-use load and electricity prices. Risk assessment depends on expert experience, lacking quantitative methods and traceability mechanisms. General-purpose commercial simulation software lacks industry-specific adaptations and cannot incorporate power business logic such as spot trading rules, carbon emission trading mechanisms, unit output physical constraints, and grid dispatch constraints, leading to significant discrepancies between simulation results and actual conditions. Its module data flow relies on manual configuration, failing to form an automated, closed-loop process. While general-purpose AI analysis platforms have advantages in data modeling, their algorithms do not fully consider the temporal characteristics of power data, lack differentiated screening mechanisms for long and short time-series factors, and limit fitting accuracy. In risk analysis, they can only provide qualitative suggestions and cannot quantify the transmission path and extent of risk between indicators. Summary of the Invention

[0004] This invention provides a method, apparatus, and equipment for power data processing based on sand table simulation, which solves the problems of power simulation rule separation, poor time-series processing, and difficulty in coordination.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: This invention provides a power data processing method based on sand table simulation, comprising: Acquire multi-source heterogeneous data from the power industry; The multi-source heterogeneous data is subjected to time-series alignment and standardization to obtain a standard power time-series dataset; Based on the aforementioned power time-series standard dataset, a subset of core factors is determined; The core factor subset is input into the hybrid inference model for processing to obtain the initial measurement results of the indicators. The hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. Obtain the simulation scenario parameters configured by the user; Based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules, the scenario game simulation results are determined. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Based on the game theory simulation results of the scenario, the risk transmission path information and risk classification judgment results are determined; Based on the risk classification results and risk transmission path information, optimize strategy data and simulation reports are determined.

[0006] Optionally, the multi-source heterogeneous data is subjected to time-series alignment and standardization to obtain a standard power time-series dataset, including: Time series data is aligned and deduplicated by performing time series alignment and sorting on each data source according to a uniform timestamp granularity to obtain aligned time series data. The missing values ​​in the aligned time series data are time-series filled in to obtain the filled time series data. Based on the rigid thresholds and time-series fluctuation characteristics of the power business, outlier identification and correction are performed on the completed time-series data to obtain corrected time-series data; The corrected time-series data is standardized to obtain a standard power time-series dataset.

[0007] Optionally, based on the aforementioned standard power time-series dataset, a subset of core factors is determined, including: The standard power time series dataset is traversed and extracted to obtain multiple candidate factors and corresponding indicators; Based on the candidate factors, a multidimensional feature dataset is determined, which includes the values ​​of each candidate factor at each time point. Based on the aforementioned indicators, an indicator label set is determined, which includes the values ​​of each indicator at each time point. The multidimensional feature dataset and indicator label set are paired and divided according to time points to obtain a first time series sample set and a second time series sample set. The first time series sample set and the second time series sample set include multiple time series samples, and each time series sample includes feature data and label data corresponding to the same time point. Based on the first time series sample set and the second time series sample set, determine the basic mutual information values ​​of each candidate factor and indicator; Based on the preset time-series weighting coefficients, the basic mutual information values ​​of each candidate factor and the index are weighted and fused to obtain the weighted mutual information value of each candidate factor. According to the numerical value of the weighted mutual information, the candidate factors are sorted, and a predetermined number of candidate factors at the top of the sorting are selected as the core factor subset.

[0008] Optionally, the training process of the hybrid inference model includes: Construct a first sub-model, the structure of which consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the core factor subset, and the hidden layer uses an activation function. Construct a second sub-model, where the number of decision trees in the first sub-model is a preset value; The first calculation result output by the first sub-model and the second calculation result output by the second sub-model are weighted and fused according to the preset fusion weight to obtain the initial inference model; Obtain historical electricity datasets; The historical power dataset is divided into training set, validation set and test set according to a preset ratio; The training set is input into the initial inference model for parameter learning to obtain the intermediate inference model; Based on the validation set, the intermediate inference model is hyperparameter-tuned to obtain the optimized inference model; The generalization ability of the optimized inference model is evaluated based on the test set, and the inference model that passes the evaluation is used as the hybrid inference model after training.

[0009] Optionally, based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules, the scenario game simulation results are determined, including: Based on the scenario parameters, the initial states and behavioral strategies of multiple intelligent agents are determined, and the initial strategy set of each intelligent agent is obtained. The multiple intelligent agents include the power generation entity intelligent agent, the power grid entity intelligent agent, the user entity intelligent agent, and the regulatory entity intelligent agent. Based on the initial strategy set and the preset power unit output constraint rules, determine the output boundary value of each unit in each simulation period; Based on the output boundary value and the preset time-sharing clearing weight rule, the corrected revenue value of each unit in each simulation period is determined; Based on the corrected payout value and the initial policy set of each agent, the agents are driven to conduct multiple rounds of game iterations until an equilibrium state is reached, and the data of each agent in the equilibrium state is used as the game inference result of the scenario.

[0010] Optionally, based on the results of the scenario game theory simulation, the risk transmission path information and risk classification judgment results are determined, including: Based on the real-time inferred values ​​of each indicator and their corresponding historical benchmark values ​​in the scenario game simulation results, the fluctuation deviation rate of each indicator is determined. The risk level of each indicator is determined based on the preset deviation range in which the volatility deviation rate of each indicator falls; Based on the risk level of each indicator, determine the target indicator pair; Based on the time-series sample values ​​of each target indicator pair in each simulation period of the game simulation results in the scenario, determine the risk transmission coefficient between each target indicator pair; Based on the risk transmission coefficient, determine the risk transmission path information between each pair of target indicators; Based on the risk transmission path information and the risk level of each indicator, the risk classification result is determined.

[0011] Optionally, based on the risk classification results and risk transmission path information, optimization strategy data and simulation reports are determined, including: Obtain the results of multiple simulation schemes; The results of the multiple simulation schemes are compared to generate multi-scheme benchmark difference data; Based on the risk classification and risk transmission path information, an optimization strategy corresponding to the current simulation condition is matched from the preset strategy rule base. Based on the comparison difference data of the multiple schemes and the optimization strategy, a simulation report is generated.

[0012] This invention also provides a power data processing device based on sand table simulation, comprising: The first acquisition module is used to acquire multi-source heterogeneous data from the power industry. The first processing module is used to perform time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset; determine a subset of core factors based on the standard power time-series dataset; input the subset of core factors into a hybrid inference model for processing to obtain initial calculation results of the indicators, wherein the hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. The second acquisition module is used to acquire the simulation scenario parameters configured by the user; The second processing module is used to determine the scenario game simulation results based on the simulation scenario parameters, initial indicator calculation results, and preset multi-agent game simulation rules. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Based on the scenario game simulation results, the module determines risk transmission path information and risk classification judgment results. Based on the risk classification judgment results and risk transmission path information, the module determines optimization strategy data and simulation report.

[0013] This invention also provides a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when run by the processor, executes the above-described method.

[0014] This invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described method.

[0015] The technical solution of the present invention has at least the following effects: The above-mentioned solution of the present invention acquires multi-source heterogeneous data from the power industry; performs time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset; determines a subset of core factors based on the standard power time-series dataset; inputs the subset of core factors into a hybrid inference model for processing to obtain initial indicator calculation results, wherein the hybrid inference model includes a first sub-model and a second sub-model fused in parallel; acquires user-configured inference scenario parameters; determines scenario game inference results based on the inference scenario parameters, initial indicator calculation results, and preset multi-agent game simulation rules, wherein the multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules; determines risk transmission path information and risk classification judgment results based on the scenario game inference results; and determines optimization strategy data and inference reports based on the risk classification judgment results and risk transmission path information, thereby improving the accuracy and efficiency of power inference. Attached Figure Description

[0016] Figure 1 This is a flowchart of the power data processing method based on sand table simulation provided in this embodiment of the invention; Figure 2 This is a structural diagram of the power data processing device based on sand table simulation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the computing device provided in an embodiment of the present invention. Detailed Implementation

[0017] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0018] like Figure 1 As shown, an embodiment of the present invention proposes a power data processing method based on sand table simulation, which may include: Step 11: Obtain multi-source heterogeneous data from the power industry; Step 12: Perform time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset; Step 13: Determine the core factor subset based on the power time series standard dataset; Step 14: Input the subset of core factors into the hybrid inference model for processing to obtain the initial calculation results of the indicators. The hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. Step 15: Obtain the simulation scenario parameters configured by the user; Step 16: Determine the scenario game simulation result based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Step 17: Based on the scenario game simulation results, determine the risk transmission path information and the risk classification judgment results; Step 18: Based on the risk classification results and risk transmission path information, determine the optimization strategy data and simulation report.

[0019] This invention proposes the above-mentioned technical solution, which establishes a fully automated processing architecture from data governance, factor screening, model calculation, game simulation to risk assessment. It performs time-series alignment and standardization governance on multi-source heterogeneous data to form a time-series standard dataset, constructs a hybrid inference model that integrates deep neural networks and weighted random forests in parallel to calculate indicators, embeds unit output constraints and time-of-use clearing weight rules into the multi-agent game simulation process, and quantifies risk levels and transmission paths based on the degree of fluctuation and deviation of indicators and time-series sample values. Data flow at each stage is traceable, avoiding the problems of chaotic data flow and inconsistent definitions in traditional discrete systems. Simultaneously, it improves the accuracy of indicator calculation compared to general algorithms, enabling simulation results to reflect the physical output boundaries of units and the differences in time-of-use market returns, transforming risk identification from qualitative judgment to quantitative analysis, and automating the comparison of multiple schemes and strategy generation. The overall inference cycle is shortened compared to manual methods.

[0020] In an optional embodiment of the present invention, step 12, which involves performing time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset, may include: Step 121: Perform time-series alignment and deduplication sorting on each data source according to a unified timestamp granularity to obtain aligned time-series data; Step 122: Perform time-series completion on the missing values ​​in the aligned time-series data to obtain the completed time-series data; Step 123: Based on the rigid threshold of power business and the time-series fluctuation characteristics, perform outlier identification and correction on the completed time-series data to obtain the corrected time-series data; Step 124: Standardize the corrected time series data to obtain a standard power time series dataset.

[0021] In step 121 of this embodiment, because the multi-source heterogeneous data in the power industry comes from different data sources such as thermal power or new energy production monitoring systems, power marketing systems, power grid EMS dispatching systems, carbon trading registration systems, commodity price systems, electricity price policy databases, and regional electricity load time series databases, the timestamp granularity of the data recorded by each system varies. For example, some systems record at the second level, while others record at the minute or hour level, and multiple duplicate records from different systems may appear at the same time point. To address these issues, the timestamps of all data sources are uniformly converted to a preset timestamp granularity, aligning all data on a unified timeline. Simultaneously, records appearing multiple times at the same time point are deduplicated, retaining only valid records, and all records are sorted according to chronological order to correct time series gaps and out-of-order issues caused during data acquisition or transmission, resulting in aligned time series data.

[0022] In step 122, the aligned time-series data may contain missing data due to reasons such as acquisition equipment failure, communication interruption, or data write failure. A hierarchical time-series completion strategy is adopted to handle different degrees of missing data. For slightly missing data points, linear interpolation is used for completion, that is, using the valid data values ​​of adjacent time points before and after the missing position, linear interpolation is performed according to the time distance to obtain the estimated value of the missing position. For moderately missing data segments, historical periodic weighted fitting is used for completion, that is, selecting historical periodic data with similar power operation periodic characteristics to the current time period, performing weighted averaging according to preset weights, and then filling in the missing position. For severe missing data or missing key indicators, a dual-data source cross-validation method is used for completion, that is, extracting records of the same indicator from two or more independent data sources, cross-referencing and cross-validating to determine the completed value, and obtaining the completed time-series data.

[0023] In step 123, the completed time-series data may still contain outliers caused by equipment malfunctions, data transmission interference, or other reasons. A dual verification process is performed, combining rigid thresholds for power business data with time-series fluctuation characteristics. Rigid threshold verification identifies records exceeding reasonable ranges as outliers based on inherent boundary conditions of power business, such as the rated output range of generating units, upper and lower limits of electricity prices, and typical load fluctuation intervals. Time-series fluctuation characteristic verification identifies data points deviating from normal fluctuation patterns by more than a preset multiple based on the periodic fluctuation patterns exhibited by indicators such as power load and electricity prices over similar time periods. The identified outliers are then corrected to remove acquisition noise, resulting in corrected time-series data.

[0024] In step 124, the corrected time-series data exhibits differences in physical units, statistical definitions, and dimensional dictionaries due to varying sources. For example, different systems may use different units of measurement for power, and their definitions of the same indicator may differ. To address these issues, all data is unified to the same physical units, statistical definitions, and dimensional dictionaries, eliminating data bias caused by different system sources. Simultaneously, in accordance with power data security requirements, fields containing sensitive information undergo tiered anonymization, ultimately generating a standard power time-series dataset. This dataset possesses a unified time granularity, consistent physical units, and consistent statistical definitions, directly supporting subsequent factor selection and model calculations.

[0025] In an optional embodiment of the present invention, step 13, determining the core factor subset based on the power time series standard dataset, may include: Step 131: Perform traversal extraction on the power time series standard dataset to obtain multiple candidate factors and corresponding indicators; Step 132: Determine a multidimensional feature dataset based on the candidate factors, wherein the multidimensional feature dataset includes the values ​​of each candidate factor at each time point; Step 133: Determine the indicator label set based on the indicator, wherein the indicator label set includes the value of each indicator at each time point; Step 134: Perform sample pairing and division on the multidimensional feature dataset and indicator label set according to time points to obtain a first time series sample set and a second time series sample set. The first time series sample set and the second time series sample set include multiple time series samples, and each time series sample includes feature data and label data corresponding to the same time point. Step 135: Determine the basic mutual information values ​​of each candidate factor and index based on the first time series sample set and the second time series sample set; Step 136: Based on the preset time-series weighting coefficients, the basic mutual information values ​​of each candidate factor and the index are weighted and fused to obtain the weighted mutual information value of each candidate factor. Step 137: Sort the candidate factors according to the weighted mutual information values, and select a preset number of candidate factors at the top of the sorting as the core factor subset.

[0026] In step 131 of this embodiment, the power time-series standard dataset contains multiple data fields. Some fields characterize the driving factors of power companies, such as grid-connected electricity volume, fuel cost per kilowatt-hour, time-of-use pricing, carbon price, and unit utilization hours. Other fields characterize results, such as company revenue, net profit, cash flow, cost per kilowatt-hour, unit carbon emissions, and unit utilization rate. All data fields of the power time-series standard dataset are traversed to extract candidate factors and corresponding indicators. The candidate factors are used to select the input features required for subsequent modeling, while the indicators serve as the target variables in the modeling process.

[0027] In step 132, a multidimensional feature dataset is constructed based on the extracted candidate factors. Each column of the multidimensional feature dataset corresponds to a candidate factor, each row corresponds to a time point, and the value at the intersection of the row and column is the specific value of the candidate factor at that time point. The number of feature dimensions is equal to the total number of candidate factors.

[0028] In step 133, an indicator label set is constructed based on the extracted indicators. Each column of the indicator label set corresponds to an indicator, each row corresponds to a time point, and the value at the intersection of the row and column is the specific value of the indicator at that time point.

[0029] In step 134, the multidimensional feature dataset and the indicator label set are paired at the same time point, so that each time series sample contains both feature data and label data at that time point. After pairing, all samples are divided into a first time series sample set and a second time series sample set according to the time granularity. The first time series sample set is a short time series sample set, with hours or days as the time unit, used to capture short-term dynamics such as spot transactions and intraday load fluctuations; the second time series sample set is a long time series sample set, with months or years as the time unit, used to reflect long-term trends such as annual strategic planning, electricity price policy trends, and changes in installed capacity.

[0030] In step 135, for the first and second time-series sample sets, the basic mutual information value between each candidate factor and each indicator is calculated. The mutual information value measures the degree of correlation between the values ​​of the candidate factors and the values ​​of the indicators; a higher mutual information value indicates a more significant influence of the candidate factor on the indicator. The formula for calculating the basic mutual information value is:

[0031] in, Candidate factors With indicators The basic mutual information value between them; Candidate factors The value of With indicators The value of The joint probability distribution of ; Candidate factors The value of The marginal probability distribution; As an indicator The value of The marginal probability distribution.

[0032] In step 136, the influence of the same candidate factor on the indicator varies at different time scales. The basic mutual information values ​​corresponding to the short-time series sample set and the long-time series sample set are weighted and fused according to a preset time-series weighting coefficient to obtain the weighted mutual information value of each candidate factor. The formula for calculating the weighted mutual information value is as follows:

[0033] in, Candidate factors With indicators The weighted mutual information value between them; Candidate factors for short time series samples With indicators Mutual information value; Candidate factors for long-term time series samples With indicators Mutual information value; This is the time-series weighting coefficient, and its value is determined based on the business scenario being simulated. In short-time-series simulation scenarios such as electricity spot trading and intraday bidding... In long-term strategic scenarios such as annual planning and expansion of new energy capacity In extreme operating conditions such as unit failure and market shutdown .

[0034] In step 137, the candidate factors are sorted in descending order of their weighted mutual information values. The larger the weighted mutual information value, the stronger the correlation between the candidate factor and the indicator. A predetermined number of candidate factors at the top of the ranking are selected as the core factor subset, and candidate factors with weaker correlations at the bottom of the ranking are removed. The selected core factor subset is used as the input features for the subsequent hybrid inference model.

[0035] In an optional embodiment of the present invention, the training process of the hybrid inference model may include: Step 141: Construct a first sub-model. The structure of the first sub-model consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the core factor subset. The hidden layer uses an activation function. Step 142: Construct the second sub-model, where the number of decision trees in the first sub-model is a preset value; Step 143: The first calculation result output by the first sub-model and the second calculation result output by the second sub-model are weighted and fused according to the preset fusion weight to obtain the initial inference model; Step 144: Obtain historical electricity dataset; Step 145: Divide the historical power dataset into training set, validation set and test set according to a preset ratio; Step 146: Input the training set into the initial inference model to learn the parameters and obtain the intermediate inference model; Step 147: Perform hyperparameter tuning on the intermediate inference model based on the validation set to obtain the optimized inference model; Step 148: Evaluate the generalization ability of the optimized inference model based on the test set, and use the inference model that passes the evaluation as the trained hybrid inference model.

[0036] In step 141 of this embodiment, the first sub-model adopts a deep neural network structure, which consists of an input layer, three hidden layers, and an output layer. The input layer has 68 neurons, consistent with the dimension of the core factor subset, and is used to receive the filtered 68-dimensional core factors. The three hidden layers have 128, 64, and 32 neurons respectively, and each hidden layer uses the ReLU activation function to introduce nonlinear transformation capability, enabling the network to fit the complex nonlinear coupling relationship between the core factors and the indicators. The output layer has 12 neurons and is used to output the measurement results of 12 indicators.

[0037] In step 142, the second sub-model adopts a weighted random forest structure with 100 decision trees. Multiple sub-models are constructed using the bootstrap sampling method, and the ensemble output of these decision trees provides stable prediction results under discrete conditions. The weighted random forest incorporates a time-series sample weighting mechanism during its construction, giving more weight to recent samples than to historical samples. Specifically, recent samples have a weight of 0.7, while historical samples have a weight of 0.3, allowing the model to pay closer attention to recent data trends during training.

[0038] In step 143, the first measurement result output by the deep neural network and the second measurement result output by the weighted random forest are weighted and fused according to preset fusion weights to obtain the initial inference model. The fusion calculation formula is:

[0039] in, The results of the fusion calculation of the initial inference model; This is the first measurement result output by the deep neural network. This is the second calculation result of the weighted random forest output; This is the fusion weight for the deep neural network, with a value of 0.6; The fusion weights for the weighted random forest are 0.4, and .

[0040] In step 144, the historical power dataset is obtained. This dataset consists of historical time period data from the standard power time series dataset, including the historical values ​​of each dimension of the core factor subset and the corresponding historical values ​​of the indicators.

[0041] In step 145, the historical power dataset is divided into a training set, a validation set, and a test set according to a preset ratio of 7:2:1. The training set accounts for 70% of the total data and is used for model parameter learning; the validation set accounts for 20% of the total data and is used for hyperparameter tuning and preventing overfitting; the test set accounts for 10% of the total data and is used for evaluating the model's generalization ability.

[0042] In step 146, the training set is input into the initial inference model for parameter learning. During training, the mean squared error loss function is used to calculate the error between the model output and the training set labels. The error calculation formula is as follows:

[0043] in, This represents the mean squared error loss value. For the first The true index value of each sample; For the first Model calculation values ​​for each sample; The sample size is denoted as . The Adam optimizer is used for parameter updates, with an initial learning rate set to 0.001. After 200 rounds of basic training, an intermediate inference model is obtained.

[0044] In step 147, the validation set is input into the intermediate inference model for hyperparameter tuning. By monitoring the trend of the loss value of the intermediate inference model on the validation set, the hyperparameter configuration is adjusted to achieve optimal performance on the validation set, resulting in the optimized inference model.

[0045] In step 148, the test set is input into the optimized inference model to evaluate its generalization ability. When the loss value of the optimized inference model on the test set is lower than the preset threshold, the inference model that passes the evaluation is used as the trained hybrid inference model for subsequent index calculation.

[0046] In step 141 of this embodiment, the first sub-model adopts a deep neural network structure, which consists of an input layer, three hidden layers, and an output layer. The input layer has 68 neurons, consistent with the dimension of the core factor subset, and is used to receive the filtered 68-dimensional core factors. The three hidden layers have 128, 64, and 32 neurons respectively, and each hidden layer uses the ReLU activation function to introduce nonlinear transformation capability, enabling the network to fit the complex nonlinear coupling relationship between the core factors and the indicators. The output layer has 12 neurons and is used to output the measurement results of 12 indicators.

[0047] In step 142, the second sub-model adopts a weighted random forest structure with 100 decision trees. Multiple sub-models are constructed using the bootstrap sampling method, and the ensemble output of these decision trees provides stable prediction results under discrete conditions. The weighted random forest incorporates a time-series sample weighting mechanism during its construction, giving more weight to recent samples than to historical samples. Specifically, recent samples have a weight of 0.7, while historical samples have a weight of 0.3, allowing the model to pay closer attention to recent data trends during training.

[0048] In step 143, the first measurement result output by the deep neural network and the second measurement result output by the weighted random forest are weighted and fused according to preset fusion weights to obtain the initial inference model. The fusion calculation formula is:

[0049] in, The results of the fusion calculation of the initial inference model; This is the first measurement result output by the deep neural network. This is the second calculation result of the weighted random forest output; This is the fusion weight for the deep neural network, with a value of 0.6; The fusion weights for the weighted random forest are 0.4, and .

[0050] In step 144, the historical power dataset is obtained. This dataset consists of historical time period data from the standard power time series dataset, including the historical values ​​of each dimension of the core factor subset and the corresponding historical values ​​of the indicators.

[0051] In step 145, the historical power dataset is divided into a training set, a validation set, and a test set according to a preset ratio of 7:2:1. The training set accounts for 70% of the total data and is used for model parameter learning; the validation set accounts for 20% of the total data and is used for hyperparameter tuning and preventing overfitting; the test set accounts for 10% of the total data and is used for evaluating the model's generalization ability.

[0052] In step 146, the training set is input into the initial inference model for parameter learning. During training, the mean squared error loss function is used to calculate the error between the model output and the training set labels. The error calculation formula is as follows:

[0053] in, This represents the mean squared error loss value. For the first The true index value of each sample; For the first Model calculation values ​​for each sample; The sample size is denoted as . The Adam optimizer is used for parameter updates, with an initial learning rate set to 0.001. After 200 rounds of basic training, an intermediate inference model is obtained.

[0054] In step 147, the validation set is input into the intermediate inference model for hyperparameter tuning. Hyperparameter tuning targets parameters affecting the model structure, such as the number of hidden layer neurons, learning rate, and regularization coefficient. A grid search method is used to traverse candidate combinations in a predefined hyperparameter combination space. The model performance under each combination is evaluated on the validation set, and the hyperparameter combination with the minimum loss value on the validation set is selected as the optimal configuration. Simultaneously, the trend between the training set loss value and the validation set loss value is monitored during training. When the validation set loss value stops decreasing while the training set loss value continues to decrease, an overfitting signal is identified. Training is terminated early, and the model weights corresponding to the iteration step with the lowest validation set loss value are backtracked to prevent the model from overfitting to the training set noise. After the above hyperparameter tuning and overfitting prevention measures, the optimized inference model is obtained.

[0055] In step 148, the test set is input into the optimized inference model for generalization capability evaluation. Specifically, the test set is input into the optimized inference model to perform forward propagation calculation, obtaining the calculation results for each test sample. The mean squared error loss function, the same as in step 146, is used to calculate the loss value on the test set. When the test set loss value is lower than a preset threshold, it indicates that the model has good generalization capability, and the inference model that passes the evaluation is used as the trained hybrid inference model. When the test set loss value is higher than the preset threshold, it indicates that the model's performance on unseen data has not met expectations. At this time, the hyperparameter configuration is readjusted, and the training and evaluation process from steps 146 to 148 is repeated until the test set loss value is lower than the preset threshold. Finally, the trained hybrid inference model is obtained for subsequent index calculation.

[0056] In an optional embodiment of the present invention, step 16, determining the scenario game simulation result based on the scenario parameters, the initial index calculation results, and the preset multi-agent game simulation rules, may include: Step 161: Based on the inferred scenario parameters, determine the initial state and behavior strategy of multiple intelligent agents to obtain the initial strategy set of each intelligent agent. The multiple intelligent agents include the power generation entity intelligent agent, the power grid entity intelligent agent, the user entity intelligent agent, and the regulatory entity intelligent agent. Step 162: Determine the output boundary value of each unit in each simulation period according to the initial strategy set and the preset power unit output constraint rules; Step 163: Determine the corrected revenue value of each unit in each simulation period based on the output boundary value and the preset time-sharing clearing weight rule; Step 164: Based on the corrected payout value and the initial strategy set of each agent, drive each agent to conduct multiple rounds of game iterations until an equilibrium state is reached, and use the data of each agent in the equilibrium state as the game inference result of the scenario.

[0057] In step 161 of this embodiment, the scenario parameters include scenario type, market environment parameters, unit technical parameters, fuel prices, carbon emission quotas, and other boundary conditions. Based on these scenario parameters, the initial states and behavioral strategies of the power generation entity, grid entity, user entity, and regulatory entity are determined respectively. The power generation entity is responsible for formulating bidding strategies and output plans for each unit; its initial state includes the available capacity, marginal cost, and signed medium- and long-term contracted electricity volume of each unit. The grid entity is responsible for executing transmission and distribution scheduling constraints and congestion management; its initial state includes the grid topology, transmission channel capacity, and system reserve capacity requirements. The user entity is responsible for responding to electricity price signals and adjusting electricity load demand; its initial state includes the electricity load curve and electricity price elasticity coefficient. The regulatory entity is responsible for executing market rules and policy boundary constraints; its initial state includes market price limits, total carbon emission quotas, and renewable energy consumption responsibility weights. Each entity generates its own initial behavioral strategy based on the above initial states, collectively forming the initial strategy set of each entity.

[0058] In step 162, within each simulation period, preset power unit output constraint rules are used to limit the simulated output of each unit, ensuring that all output results conform to the physical constraints of power production. The power unit output constraint rules are as follows:

[0059] in, For the first Simulated output value of the generator unit for the current time period; The minimum technical output limit of the unit is determined by the unit's stable combustion load, minimum start-up and shutdown time, and other technical conditions. The rated maximum output limit of the unit is determined by factors such as the unit's rated capacity and equipment health status. After the output schemes in the initial strategy set of each agent are verified by the above constraint rules, the output boundary values ​​of each unit in each simulation period are obtained.

[0060] In step 163, the preset time-of-use clearing weighting rule is used to weight and adjust the unit revenue for different time periods. Because the price formation mechanism and supply-demand relationship differ in the electricity market at different times, using a uniform electricity price for revenue calculation cannot reflect the true differences in time-of-use revenue. The time-of-use clearing weighting rule is as follows:

[0061] in, for Simulated revenue of generator units during specific time periods; for The simulated output value of the unit during the time period is constrained by the output boundary value determined in step 162 and should not exceed the output boundary value range; for Time-of-use market clearing electricity prices; for Time-of-use clearing weighting coefficient. The time-of-use clearing weighting coefficient is determined according to the typical time period rules of the power grid, with peak periods... During the flat period Low period Substituting the output boundary value and the time-sharing clearing weight coefficient into the above formula, the corrected revenue value of each unit in each simulation period is obtained.

[0062] In step 164, each agent aims to maximize its own profit during the game. The power generation agent adjusts its bidding strategy and output plan to strive for higher grid-connected electricity volume and transaction prices, while the user agent adjusts its load demand to reduce electricity costs. The actions of each agent influence each other in multiple iterations. After one agent adjusts its strategy, other agents adjust their strategies accordingly based on the changed market state. After repeated game interactions, the system gradually converges to an equilibrium state. At this equilibrium state, unilaterally changing one's strategy will not yield higher profits for any party. Once the system reaches equilibrium, the data of each agent in this state is output as the scenario game simulation result. This result includes data such as the output plans of each unit, the market clearing result, and the profits of each agent.

[0063] In an optional embodiment of the present invention, step 17, determining the risk transmission path information and risk classification result based on the scenario game simulation results, may include: Step 171: Determine the fluctuation deviation rate of each indicator based on the real-time inferred values ​​of each indicator and their corresponding historical benchmark values ​​in the scenario game simulation results. Step 172: Determine the risk level of each indicator based on the preset deviation range in which the fluctuation deviation rate of each indicator falls; Step 173: Determine the target indicator pair based on the risk level of each indicator; Step 174: Determine the risk transmission coefficient between each pair of target indicators based on the time-series sample values ​​of each target indicator pair in each simulation period of the scenario game simulation results. Step 175: Determine the risk transmission path information between each pair of target indicators based on the risk transmission coefficient; Step 176: Determine the risk classification result based on the risk transmission path information and the risk level of each indicator.

[0064] In step 171 of this embodiment, the scenario game simulation result includes the real-time simulation values ​​of each indicator at the end of the simulation. These indicators include cost per kilowatt-hour, cash flow, carbon allowance surplus, spot price deviation rate, and unit utilization hours. Each indicator has a corresponding historical benchmark value, which is taken from the value of the same indicator in the same historical period or the annual budget target value in the electricity time-series standard dataset. For each indicator, its real-time simulation value is compared with the historical benchmark value, and the fluctuation deviation rate is calculated. The formula for calculating the fluctuation deviation rate is:

[0065] in, The volatility deviation rate of the indicator is used to measure the degree of fluctuation of the current projected value of the indicator relative to the benchmark value. This refers to the real-time projected value of the indicator; The historical benchmark value or annual budget target value of the indicator.

[0066] In step 172, the preset deviation range includes three intervals, each corresponding to one of the three risk levels. The first preset deviation range is... The first, corresponding to a general risk level, indicates that the indicator deviates slightly from the benchmark value without affecting overall stability; the second preset deviation range is... The corresponding early warning risk level indicates that the indicator shows significant abnormalities and there is a potential risk of loss; the third preset deviation range is... The corresponding high-risk level indicates that the indicator has deviated significantly and has posed a substantial risk. The risk level of each indicator is determined based on the preset deviation range into which the fluctuation deviation rate of each indicator falls.

[0067] In step 173, after obtaining the risk levels of all indicators, the target indicator pairs that require risk transmission analysis are determined. Specifically, indicators whose risk levels reach the warning risk level or high-risk risk level are selected as risk source indicators. These risk source indicators are then paired with other indicators that may be affected by their transmission to form target indicator pairs. For example, abnormal fluctuations in the cost per kilowatt-hour may be transmitted to cash flow, carbon quota gaps may be transmitted to carbon emission compliance costs, and abnormal unit utilization hours may be transmitted to unit power generation costs.

[0068] In step 174, the scenario game simulation results include complete time-series sample values ​​of each indicator within each simulation period. For each pair of target indicators, the time-series sample values ​​of both indicators within each simulation period are extracted from the scenario game simulation results, and the risk transmission coefficient between the target indicator pair is calculated. The risk transmission coefficient combines the Pearson correlation coefficient and historical volatility multiples to quantify the strength of risk transmission between indicators. The formula for calculating the Pearson correlation coefficient is:

[0069] in, As an indicator With indicators The Pearson correlation coefficient between them; As an indicator The first in the target indicator pair One time-series sample value; As an indicator The first in the target indicator pair One time-series sample value; As an indicator The sample mean; As an indicator The sample mean. The formula for calculating historical volatility is:

[0070] in, As an indicator or indicators Historical volatility multiples; This represents the recent standard deviation of the indicator's volatility. This represents the historical benchmark standard deviation of the indicator. The formula for calculating the risk transmission coefficient is:

[0071] in, For target indicators to be aligned with indicators With indicators The risk transmission coefficient between indicators indicates that the larger the value, the stronger the ability of risk to be transmitted from one indicator to another, and the more significant the correlation effect.

[0072] In step 175, the risk transmission path information between each pair of target indicators is determined based on the risk transmission coefficient between them. All pairs of target indicators are sorted in descending order of risk transmission coefficient. The indicators in the top-ranked pairs are connected according to the transmission direction to form a directed transmission path from the high-risk source indicator to the affected indicators, thus constructing a risk transmission network between indicators.

[0073] In step 176, the risk transmission path information and the risk level of each indicator are comprehensively assessed. Specifically, the risk level of each indicator is used as the basic assessment criterion, and the risk transmission path information is used as the auxiliary assessment criterion. When an indicator is simultaneously at a high-risk level and there are multiple risk transmission paths pointing to that indicator, the risk classification result of that indicator is adjusted upwards. When an indicator is simultaneously at a general risk level and there are transmission paths pointing to that indicator, the risk classification result of that indicator is comprehensively adjusted in conjunction with the risk level of the source indicator in the transmission path, and finally the final risk classification result of each indicator is determined.

[0074] In an optional embodiment of the present invention, step 18, determining the optimization strategy data and simulation report based on the risk classification determination result and risk transmission path information, may include: Step 181: Obtain multiple sets of simulation results; Step 182: Compare the results of the multiple sets of simulation schemes to generate multi-scheme benchmark difference data; Step 183: Based on the risk classification judgment result and risk transmission path information, match the optimization strategy corresponding to the current simulation condition from the preset strategy rule base; Step 184: Generate a simulation report based on the multi-scheme benchmarking difference data and the optimization strategy.

[0075] In step 181 of this embodiment, the multiple sets of simulation results originate from a multi-level collaborative interactive simulation process. Users at each level adjust different simulation parameter configurations and execute the complete scenario game simulation process, resulting in multiple sets of scenario game simulation results. Each set of simulation results includes the indicator calculation value under that scheme, game simulation process data, risk classification judgment results, and risk transmission path information.

[0076] In step 182, a multi-dimensional comparative analysis is conducted on the obtained results of multiple simulation schemes. The comparison dimensions include revenue level, cost structure, risk level, carbon asset returns, and unit utilization rate.

[0077] In step 183, based on the risk classification results and risk transmission path information, an optimization strategy corresponding to the current simulated operating condition is matched from a pre-set strategy rule base. The strategy rule base includes four types of power-specific strategies: fuel procurement strategy, spot market time-of-use pricing strategy, carbon quota allocation strategy, and unit operation optimization strategy. The fuel procurement strategy addresses fuel price fluctuation risks and includes suggestions on procurement timing, procurement volume, and inventory management under different price ranges. The spot market time-of-use pricing strategy addresses spot market bidding violation risks and time-of-use revenue optimization, including pricing ranges and bidding strategies under different time periods and supply-demand relationships. The carbon quota allocation strategy addresses carbon quota shortfall risks and includes optimization schemes for carbon quota buying, selling, or inter-period allocation. The unit operation optimization strategy addresses abnormal unit utilization hours risks and includes unit start-up and shutdown optimization and load allocation optimization schemes. The matching process selects strategy entries corresponding to the risk characteristics from the strategy rule base based on the risk type, risk level, and risk transmission path in the current simulated operating condition, determining the optimization strategy applicable to the current simulated operating condition.

[0078] In step 184, the comparison data of multiple schemes and optimization strategies are used as input to generate a simulation report. The simulation report includes a summary of scenario parameters, simulation process data, indicator calculation results, risk assessment results, conclusions of multi-scheme comparison, and optimization strategy recommendations.

[0079] A specific embodiment of the power data processing method based on sand table simulation provided in this invention is as follows: Step 1: Acquire multi-source heterogeneous data from the power industry; The system is specifically designed to connect with power industry-specific systems such as thermal power / new energy production monitoring systems, power marketing systems, power grid EMS dispatching systems, carbon trading registration systems, commodity price systems, electricity price policy databases, and regional electricity load time series databases. It collects various types of heterogeneous data, including electricity time-of-use load, spot trading data, unit operation data, carbon price time series data, fuel price data, and electricity price policy data, as the raw data foundation for subsequent governance and analysis.

[0080] Step 2: Perform time-series alignment and standardization on the multi-source heterogeneous data to obtain a standard power time-series dataset; First, unified time-series alignment and deduplication sorting of multi-source data were completed, and standardized to Beijing timestamps with uniform granularity to fix time-series gaps and out-of-order issues. Based on the periodic characteristics of power data, a hierarchical time-series completion strategy was adopted, using linear interpolation, weighted fitting of historical data, and cross-verification of dual data sources to differentiate and repair mild, moderate, and severe missing data. Combined with the rigid thresholds of power business and time-series fluctuation characteristics, dual verification was performed to identify and correct erroneous data such as unit output exceeding limits, abnormal electricity prices, and abnormal load fluctuations, and invalid acquisition noise was eliminated. Abnormal data that could not be directly corrected was returned to the data access stage for reprocessing.

[0081] Based on this, the units, statistical standards and dimension dictionaries of the full data are unified to eliminate the differences and deviations between data from multiple systems. Then, the classified information is desensitized in accordance with the power data security specifications, and finally a standardized and accurate power time series standard dataset is generated. This dataset supports automatic iterative updates at the hourly and daily granular levels.

[0082] Step 3: Determine the core factor subset based on the power time series standard dataset; The system iterates through all candidate power factors, constructs a multidimensional feature dataset and an indicator label set, and automatically divides the sample sets into hourly / daily short-time series samples and monthly / grade long-time series samples according to time granularity, calculating the basic mutual information values ​​of the two types of sample sets respectively; and automatically matches the time series weight coefficients according to the current inference scenario. The final factor correlation degree is calculated using an improved weighted mutual information formula, which is as follows:

[0083] in, Candidate factors With indicators The weighted mutual information value between them; Candidate factors for short time series samples With indicators Mutual information value; Candidate factors for long-term time series samples With indicators The mutual information value.

[0084] Weighting coefficient Configuration follows fixed rules: Short-term simulation scenarios such as spot trading and intraday bidding are used. Annual planning, expansion of new energy capacity, and other long-term strategic scenarios Automatic locking in extreme conditions such as unit failure and market shutdown Only short-term factors are calculated; finally, the factors are sorted in descending order of weighted mutual information values, highly correlated core factors are selected, low-correlation redundant factors are removed, and the optimal feature subset for modeling is output.

[0085] Step 4: Input the core factor subset into the hybrid inference model for processing to obtain the initial calculation results of the indicators. The hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. The first sub-model is a deep neural network with a fixed structure of an input layer, three hidden layers, and an output layer. The input layer has 68 neurons corresponding to the selected core factors. The first to third hidden layers have 128, 64, and 32 neurons respectively, all using the ReLU activation function. The output layer has 12 neurons corresponding to power generation indices, used to fit the continuous nonlinear coupling relationship of the factors. The second sub-model is a weighted random forest model with 100 ensemble decision trees. A time-series sample weighting mechanism is introduced, setting the weight of recent samples to 0.7 and the weight of long-term historical samples to 0.3. Multiple sub-models are constructed through bootstrap sampling to fit the characteristics of discrete mutations and extreme operating condition fluctuations.

[0086] The outputs of the two sub-models are fused using fixed weights to obtain the final measurement result. The fusion formula is as follows:

[0087] in, Fixed weight , ; During model training, the dataset is divided into training, validation, and test sets in a 7:2:1 ratio. Mean Squared Error (MSE) is used as the global loss function, with the following formula:

[0088] With the Adam adaptive optimizer, the initial learning rate is 0.001, the basic training iterations are 200 rounds, and the incremental iterations are 50 rounds after adding new data every day to achieve dynamic model updates.

[0089] Step 5: Obtain the simulation scenario parameters configured by the user; The system allows users to select from eight preset core power scenarios, including conventional market-based operation, expansion of new energy units, adjustment of transmission and distribution prices, emergency power supply, bidding in the electricity spot market, carbon asset trading, response to fuel price fluctuations, and cross-regional power trading. It also allows users to customize scenario parameters and boundary conditions. The system is adapted to the three-tier organizational structure of the power group, namely "headquarters-regional subsidiaries-grassroots stations". Different levels of accounts have corresponding hierarchical operation permissions and parameter modification permissions. The parameters configured by lower-level units must not exceed the rigid indicators issued by the higher-level units.

[0090] The rigid constraints issued by the superior include core parameters such as total profit, carbon emission quota, total power generation, upper limit of cost per kilowatt-hour, and minimum utilization hours of generating units. If the parameters entered by the subordinate exceed the constraint boundary, the system will automatically intercept and save the parameters and prompt the risk of exceeding the boundary. Only after forced rebound correction can the configuration continue to ensure that the simulation plan is in line with the overall strategic goals of the group.

[0091] Step 6: Determine the scenario game simulation result based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. The system embeds the power unit output constraints into the game theory algorithm kernel, through the formula: Lock the effective output boundaries of the units for each time period, among which For the first Simulated output of the generator unit during the current period. , These represent the unit's minimum technical output and rated maximum output, respectively; a time-based clearing weighting factor is also introduced. Construct a time-weighted clearing calculation formula: ,in for Time-of-use unit simulation revenue, for Market-clearing electricity prices during peak hours During the flat period Low periods Adjust the power generation revenue and transaction weights for different time periods.

[0092] During the simulation, the system uses scenario parameters as boundaries to synchronously drive four types of intelligent agents—power generation entity, grid entity, user entity, and regulatory entity—to conduct collaborative game iterations. After multiple rounds of Nash equilibrium iterations and convergence, the system outputs the optimal returns, unit output schemes, market transaction results, and corresponding risk data for each scenario. When extreme conditions such as market shutdown or trading suspension are detected, the system automatically shuts down the multi-agent game iteration function and switches to a pure internal production calculation mode to ensure the stability and effectiveness of the simulation results.

[0093] Step 7: Based on the scenario game simulation results, determine the risk transmission path information and risk classification results; The system first calculates the indicator volatility deviation rate, using the following formula:

[0094] in, For the benchmark value of the indicator, These are real-time projected values; risk levels are divided into three levels based on the deviation rate: A risk level of [10%, 30%) indicates moderate risk, while a risk level of [30%, 50%) indicates a warning risk. The risk is classified as high-risk. At the same time, for five types of power-specific risks, namely, inverted cost per kilowatt-hour, cash flow gap, carbon quota gap, spot price violation, and abnormal unit utilization hours, special risk identification is completed in accordance with the corresponding judgment rules.

[0095] The system constructs a directed topological transmission graph of indicators and uses a two-factor joint algorithm to calculate the risk transmission coefficient: first, the linear correlation between indicators is calculated using the Pearson correlation coefficient formula. Then calculate the historical volatility ratio. Finally, the joint coefficient of risk transmission was obtained. This quantifies the intensity of risk transmission between indicators and traces the source and transmission path of risks; the system fixes a 150% indicator fluctuation as the risk circuit breaker threshold, triggering a circuit breaker when any core indicator deviates by a certain percentage. If the circuit breaker mechanism is triggered immediately, the current simulation will be suspended and an alarm and source tracing information will be sent.

[0096] Step 8: Based on the risk classification results and risk transmission path information, determine the optimization strategy data and simulation report.

[0097] The system employs power-specific visualization components to graphically render and display core projection indicators such as time-of-use load curves, time-series changes in cost per kilowatt-hour, carbon asset trends, and cash flow fluctuations. It supports parallel comparative analysis of multiple projection scenarios from multiple dimensions, including revenue, cost, risk, and carbon assets. The system intelligently matches the current projection conditions, risk points, and revenue structure with the built-in power industry case library and expert strategy library, providing targeted optimization suggestions for fuel procurement, spot time-of-use pricing, carbon quota allocation, and unit operation optimization.

[0098] The system has built-in standardized power simulation report and benchmarking report templates, automatically summarizing scenario parameters, simulation process data, indicator calculation results, risk assessment conclusions, multi-scheme comparison results and optimization strategies. It supports users to export formal simulation reports and benchmarking statistical reports with one click. At the same time, it supports multi-level accounts of the group to collaboratively adjust parameters and carry out multi-scheme iterative simulations, realize the results to be summarized and verified at each level, and ensure the consistency of calculation standards and strategic implementation across the entire group.

[0099] This invention proposes the above-mentioned technical solution, which forms a standardized power time-series dataset through time-series alignment, hierarchical completion, and standardized governance of multi-source heterogeneous power data; it employs an improved mutual information algorithm with dynamic adjustment of time-series weights to hierarchically screen factors, and completes indicator calculation using a hybrid model composed of deep neural networks and weighted random forests, achieving dynamic model updates through incremental iteration; it embeds unit output constraints and time-of-use clearing rules into the kernel of a multi-agent game algorithm for scenario simulation; it utilizes the transmission coefficient calculated jointly by Pearson correlation coefficient and historical volatility ratio to achieve risk quantification, source tracing, and graded early warning; and it is equipped with a multi-level collaborative deduction mechanism with three-level access control and rigid indicator constraints, supplemented by visualization comparison and strategy output capabilities. This solution can improve the accuracy of indicator calculation, enhance the fit between sand table simulation and actual power business, achieve quantitative risk grading and path tracing, unify the multi-level deduction standards of the group, and shorten the trial calculation cycle of multiple schemes.

[0100] like Figure 2 As shown, this embodiment of the invention also provides a power data processing device 20 based on sand table simulation, comprising: The first acquisition module 21 is used to acquire multi-source heterogeneous data from the power industry; The first processing module 22 is used to perform time-series alignment and standardization processing on the multi-source heterogeneous data to obtain a power time-series standard dataset; determine a subset of core factors based on the power time-series standard dataset; input the subset of core factors into a hybrid inference model for processing to obtain the initial calculation results of the indicators, wherein the hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. The second acquisition module 23 is used to acquire the simulation scenario parameters configured by the user; The second processing module 24 is used to determine the scenario game simulation results based on the simulation scenario parameters, initial indicator calculation results, and preset multi-agent game simulation rules. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Based on the scenario game simulation results, it determines risk transmission path information and risk classification judgment results. Based on the risk classification judgment results and risk transmission path information, it determines optimization strategy data and simulation report.

[0101] Optionally, the first processing module 22 is specifically used for: Time series data is aligned and deduplicated by performing time series alignment and sorting on each data source according to a uniform timestamp granularity to obtain aligned time series data. The missing values ​​in the aligned time series data are time-series filled in to obtain the filled time series data. Based on the rigid thresholds and time-series fluctuation characteristics of the power business, outlier identification and correction are performed on the completed time-series data to obtain corrected time-series data; The corrected time-series data is standardized to obtain a standard power time-series dataset.

[0102] Optionally, the first processing module 22 is also specifically used for: The standard power time series dataset is traversed and extracted to obtain multiple candidate factors and corresponding indicators; Based on the candidate factors, a multidimensional feature dataset is determined, which includes the values ​​of each candidate factor at each time point. Based on the aforementioned indicators, an indicator label set is determined, which includes the values ​​of each indicator at each time point. The multidimensional feature dataset and indicator label set are paired and divided according to time points to obtain a first time series sample set and a second time series sample set. The first time series sample set and the second time series sample set include multiple time series samples, and each time series sample includes feature data and label data corresponding to the same time point. Based on the first time series sample set and the second time series sample set, determine the basic mutual information values ​​of each candidate factor and indicator; Based on the preset time-series weighting coefficients, the basic mutual information values ​​of each candidate factor and the index are weighted and fused to obtain the weighted mutual information value of each candidate factor. According to the numerical value of the weighted mutual information, the candidate factors are sorted, and a predetermined number of candidate factors at the top of the sorting are selected as the core factor subset.

[0103] Optionally, the training process of the hybrid inference model includes: Construct a first sub-model, the structure of which consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the core factor subset, and the hidden layer uses an activation function. Construct a second sub-model, where the number of decision trees in the first sub-model is a preset value; The first calculation result output by the first sub-model and the second calculation result output by the second sub-model are weighted and fused according to the preset fusion weight to obtain the initial inference model; Obtain historical electricity datasets; The historical power dataset is divided into training set, validation set and test set according to a preset ratio; The training set is input into the initial inference model for parameter learning to obtain the intermediate inference model; Based on the validation set, the intermediate inference model is hyperparameter-tuned to obtain the optimized inference model; The generalization ability of the optimized inference model is evaluated based on the test set, and the inference model that passes the evaluation is used as the hybrid inference model after training.

[0104] Optionally, the second processing module 24 is specifically used for: Based on the scenario parameters, the initial states and behavioral strategies of multiple intelligent agents are determined, and the initial strategy set of each intelligent agent is obtained. The multiple intelligent agents include the power generation entity intelligent agent, the power grid entity intelligent agent, the user entity intelligent agent, and the regulatory entity intelligent agent. Based on the initial strategy set and the preset power unit output constraint rules, determine the output boundary value of each unit in each simulation period; Based on the output boundary value and the preset time-sharing clearing weight rule, the corrected revenue value of each unit in each simulation period is determined; Based on the corrected payout value and the initial policy set of each agent, the agents are driven to conduct multiple rounds of game iterations until an equilibrium state is reached, and the data of each agent in the equilibrium state is used as the game inference result of the scenario.

[0105] Optionally, the second processing module 24 is also specifically used for: Based on the real-time inferred values ​​of each indicator and their corresponding historical benchmark values ​​in the scenario game simulation results, the fluctuation deviation rate of each indicator is determined. The risk level of each indicator is determined based on the preset deviation range in which the volatility deviation rate of each indicator falls; Based on the risk level of each indicator, determine the target indicator pair; Based on the time-series sample values ​​of each target indicator pair in each simulation period of the game simulation results in the scenario, determine the risk transmission coefficient between each target indicator pair; Based on the risk transmission coefficient, determine the risk transmission path information between each pair of target indicators; Based on the risk transmission path information and the risk level of each indicator, the risk classification result is determined.

[0106] Optionally, the second processing module 24 is also specifically used for: Obtain the results of multiple simulation schemes; The results of the multiple simulation schemes are compared to generate multi-scheme benchmark difference data; Based on the risk classification and risk transmission path information, an optimization strategy corresponding to the current simulation condition is matched from the preset strategy rule base. Based on the comparison difference data of the multiple schemes and the optimization strategy, a simulation report is generated.

[0107] It should be noted that this device is a device corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0108] like Figure 3 As shown, this embodiment of the invention also provides a computing device 30, including a processor 31, a memory 32, and a program or instructions stored in the memory 32 and executable on the processor 31. When the program or instructions are executed by the processor 31, they implement the various processes of the above-described method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here. It should be noted that the computing device in this embodiment of the invention includes the aforementioned mobile electronic devices and non-mobile electronic devices.

[0109] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A power data processing method based on sand table simulation deduction, characterized in that, include: Acquire multi-source heterogeneous data from the power industry; The multi-source heterogeneous data is subjected to time-series alignment and standardization to obtain a standard power time-series dataset; Based on the aforementioned power time-series standard dataset, a subset of core factors is determined; The core factor subset is input into the hybrid inference model for processing to obtain the initial measurement results of the indicators. The hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. Obtain the simulation scenario parameters configured by the user; Based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules, the scenario game simulation results are determined. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Based on the game theory simulation results of the scenario, the risk transmission path information and risk classification judgment results are determined; Based on the risk classification results and risk transmission path information, optimize strategy data and simulation reports are determined.

2. The power data processing method based on sand table simulation deduction according to claim 1, characterized in that, The multi-source heterogeneous data is subjected to time-series alignment and standardization to obtain a standard power time-series dataset, including: Time series data is aligned and deduplicated by performing time series alignment and sorting on each data source according to a uniform timestamp granularity to obtain aligned time series data. The missing values ​​in the aligned time series data are time-series filled in to obtain the filled time series data. Based on the rigid thresholds and time-series fluctuation characteristics of the power business, outlier identification and correction are performed on the completed time-series data to obtain corrected time-series data; The corrected time-series data is standardized to obtain a standard power time-series dataset.

3. The sand table simulation-based deduction method for power data processing according to claim 1, characterized in that, Based on the aforementioned standard power time-series dataset, a subset of core factors is determined, including: The standard power time series dataset is traversed and extracted to obtain multiple candidate factors and corresponding indicators; Based on the candidate factors, a multidimensional feature dataset is determined, which includes the values ​​of each candidate factor at each time point. Based on the aforementioned indicators, an indicator label set is determined, which includes the values ​​of each indicator at each time point. The multidimensional feature dataset and indicator label set are paired and divided according to time points to obtain a first time series sample set and a second time series sample set. The first time series sample set and the second time series sample set include multiple time series samples, and each time series sample includes feature data and label data corresponding to the same time point. Based on the first time series sample set and the second time series sample set, determine the basic mutual information values ​​of each candidate factor and indicator; Based on the preset time-series weighting coefficients, the basic mutual information values ​​of each candidate factor and the index are weighted and fused to obtain the weighted mutual information value of each candidate factor. According to the numerical value of the weighted mutual information, the candidate factors are sorted, and a predetermined number of candidate factors at the top of the sorting are selected as the core factor subset.

4. The power data processing method based on sand table simulation according to claim 1, characterized in that, The training process of the hybrid inference model includes: Construct a first sub-model, the structure of which consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the same as the dimension of the core factor subset, and the hidden layer uses an activation function. Construct a second sub-model, where the number of decision trees in the first sub-model is a preset value; The first calculation result output by the first sub-model and the second calculation result output by the second sub-model are weighted and fused according to the preset fusion weight to obtain the initial inference model; Obtain historical electricity datasets; The historical power dataset is divided into training set, validation set and test set according to a preset ratio; The training set is input into the initial inference model for parameter learning to obtain the intermediate inference model; Based on the validation set, the intermediate inference model is hyperparameter-tuned to obtain the optimized inference model; The generalization ability of the optimized inference model is evaluated based on the test set, and the inference model that passes the evaluation is used as the hybrid inference model after training.

5. The power data processing method based on sand table simulation according to claim 1, characterized in that, Based on the scenario parameters, initial index calculation results, and preset multi-agent game simulation rules, the scenario game simulation results are determined, including: Based on the scenario parameters, the initial states and behavioral strategies of multiple intelligent agents are determined, and the initial strategy set of each intelligent agent is obtained. The multiple intelligent agents include the power generation entity intelligent agent, the power grid entity intelligent agent, the user entity intelligent agent, and the regulatory entity intelligent agent. Based on the initial strategy set and the preset power unit output constraint rules, determine the output boundary value of each unit in each simulation period; Based on the output boundary value and the preset time-sharing clearing weight rule, the corrected revenue value of each unit in each simulation period is determined; Based on the corrected payout value and the initial policy set of each agent, the agents are driven to conduct multiple rounds of game iterations until an equilibrium state is reached, and the data of each agent in the equilibrium state is used as the game inference result of the scenario.

6. The power data processing method based on sand table simulation according to claim 1, characterized in that, Based on the game theory simulation results of the scenario, the risk transmission path information and risk classification results are determined, including: Based on the real-time inferred values ​​of each indicator and their corresponding historical benchmark values ​​in the scenario game simulation results, the fluctuation deviation rate of each indicator is determined. The risk level of each indicator is determined based on the preset deviation range in which the volatility deviation rate of each indicator falls; Based on the risk level of each indicator, determine the target indicator pair; Based on the time-series sample values ​​of each target indicator pair in each simulation period of the game simulation results in the scenario, determine the risk transmission coefficient between each target indicator pair; Based on the risk transmission coefficient, determine the risk transmission path information between each pair of target indicators; Based on the risk transmission path information and the risk level of each indicator, the risk classification result is determined.

7. The power data processing method based on sand table simulation according to claim 1, characterized in that, Based on the risk classification results and risk transmission path information, determine the optimization strategy data and simulation report, including: Obtain the results of multiple simulation schemes; The results of the multiple simulation schemes are compared to generate multi-scheme benchmark difference data; Based on the risk classification and risk transmission path information, an optimization strategy corresponding to the current simulation condition is matched from the preset strategy rule base. Based on the comparison difference data of the multiple schemes and the optimization strategy, a simulation report is generated.

8. A power data processing device based on sand table simulation, characterized in that, include: The first acquisition module is used to acquire multi-source heterogeneous data from the power industry. The first processing module is used to perform time-series alignment and standardization processing on the multi-source heterogeneous data to obtain a standard power time-series dataset. Based on the aforementioned power time-series standard dataset, a subset of core factors is determined; The core factor subset is input into the hybrid inference model for processing to obtain the initial measurement results of the indicators. The hybrid inference model includes a first sub-model and a second sub-model that are fused in parallel. The second acquisition module is used to acquire the simulation scenario parameters configured by the user; The second processing module is used to determine the scenario game simulation result based on the simulation scenario parameters, the initial calculation results of the indicators and the preset multi-agent game simulation rules. The multi-agent game simulation rules include power unit output constraint rules and time-of-use clearing weight rules. Based on the game simulation results of the scenario, determine the risk transmission path information and risk classification results; based on the risk classification results and risk transmission path information, determine the optimization strategy data and simulation report.

9. A computing device, characterized in that, include: A processor, a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The system stores instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.