Virtual coach formation technology-based application method, system and equipment in power dispatching operation, and medium
By building a power dispatching model based on an LSTM neural network and reinforcement learning algorithm based on a virtual trainer, the problem of lack of a realistic simulation environment in traditional training methods is solved, efficient and safe dispatcher training is achieved, and the dispatchers' operational capabilities and system stability are improved.
Patent Information
- Application Number
- CN202510789862.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional dispatcher training methods rely on on-site guidance and theoretical explanations, and lack a highly realistic simulation environment. This makes it difficult for dispatchers to respond quickly and accurately in complex scenarios, affecting the safe and stable operation of the power system.
It adopts virtual coach formation technology, obtains historical and real-time scheduling data, builds an LSTM neural network model, combines it with reinforcement learning algorithm, provides real-time comparison and evaluation, outputs optimal operation strategies and risk warnings, sets evaluation indicators such as operation accuracy and response speed, and realizes quantitative evaluation.
It improves training efficiency and quality, enhances dispatchers' emergency response capabilities, provides real-time feedback and demonstrations, reduces actual operational risks, and saves training costs and resources.
Smart Images

Figure CN120689175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system dispatching and operation, and in particular to an application method, system, equipment and medium based on virtual trainer formation technology in power dispatching operation. Background Art
[0002] In the daily operation and management of power systems, dispatchers shoulder the heavy responsibility of managing and responding to emergencies at key facilities, such as substations. However, traditional dispatcher training methods have significant flaws. They primarily rely on on-site instruction and theoretical explanations, lacking a highly realistic simulation environment for practical operational practice. This deficiency makes it difficult for dispatchers to fully develop their intuition and skills for operating equipment during training. When faced with complex operational scenarios, especially emergencies, they may be unable to respond quickly and accurately due to a lack of practical experience, which in turn impacts the safe and stable operation of the power system.
[0003] Based on this, it has become an urgent need to develop a simulation method that can provide real-time, interactive operation experience. The application method, system, equipment and medium based on virtual trainer formation technology in power dispatching operation proposed in the present invention are developed to solve the shortcomings of the above-mentioned traditional training methods. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is how to address the significant drawbacks of traditional dispatcher training methods, which mainly rely on on-site guidance and theoretical explanations, and in which dispatchers lack a highly realistic simulation environment for actual operation practice.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for applying virtual coach formation technology in power dispatching operation, which includes obtaining historical operation data and real-time dispatching operation data, and processing the acquired data; constructing a virtual coach model for learning the relationship between power dispatching operation and system response based on the processed data, training and optimizing the virtual coach model, and outputting the optimal operation strategy; comparing the dispatcher's operation with the optimal operation strategy obtained by the virtual coach model in real time, outputting evaluation information based on the comparison results, and providing operation suggestions, risk warnings and fault handling guidance based on deviations and current system status; setting evaluation indicators according to the requirements of power dispatching operation, including operation accuracy, response speed and fault handling effect, and quantitatively evaluating the dispatcher's training results based on the evaluation indicators.
[0007] As a preferred solution of the method of applying the virtual trainer formation technology in power dispatching operation described in the present invention, the processing of the acquired data includes cleaning, labeling and classification of the acquired data for constructing a training data set.
[0008] As a preferred solution of the method for applying the virtual coach formation technology in power dispatching operation described in the present invention, the training and optimization of the virtual coach model includes inputting training data into the constructed virtual coach model, and through iterative calculation and optimization algorithm, the model learns the relationship between dispatching operations and operation results in historical data, and masters the optimal operation strategy under different scenarios. During the training process, the model is evaluated using a validation set, and the model parameters are adjusted according to the evaluation results.
[0009] As a preferred solution of the method for applying the virtual trainer formation technology in power dispatching operation described in the present invention, the evaluation information output according to the comparison results includes generating operation suggestions based on the deviation situation and the current system status when there is a deviation between the operation and the optimal strategy, performing risk warning based on the real-time system status and dispatching operation analysis, and providing operational procedures for fault location, isolation and power restoration in simulated fault scenarios.
[0010] As a preferred solution of the method for applying the virtual coach formation technology in power dispatching operation described in the present invention, the virtual coach model for learning the relationship between power dispatching operation and system response is constructed, including adopting a structure based on LSTM neural network to process time series data in power dispatching, and transferring information between different time steps of the sequence by setting cell state, forget gate, input gate and output gate, and determining the retention and update mode of information according to the current input and the state at the previous moment; the virtual coach model, as an intelligent agent in reinforcement learning, performs dispatching operations in a simulated power dispatching environment, obtains rewards or penalties according to the changes in the system state after execution, and updates the behavior strategy through the state-action-reward relationship; the parameters of the virtual coach model include neural network structure parameters, reward function parameters, learning rate and discount factor, and learns the optimal operation strategy under various dispatching scenarios through trial and error interaction and parameter adjustment process.
[0011] This preferred solution introduces a gating mechanism based on LSTM neural network into the virtual coach model, which can effectively identify the time series characteristics in power dispatching and capture the long-term dependency between operational behavior and system status. At the same time, it combines the state-action-reward feedback mechanism in reinforcement learning to enable the model to continuously optimize the dispatching strategy during the interaction process, thereby enhancing the model's learning ability and adaptability in complex dispatching environments.
[0012] As a preferred solution of the method for applying the virtual trainer formation technology in power dispatching operation described in the present invention, the real-time comparison includes obtaining the dispatcher's operation instructions, operation time and operation object information through the real-time data interface during the simulation training; comparing the obtained operation behavior with the optimal strategy, using the similarity calculation algorithm to determine the degree of matching, and judging whether the operation complies with the optimal strategy; when there is a deviation between the operation and the optimal strategy, the virtual trainer generates operation suggestions, risk warnings or fault handling guidance according to the deviation and the current system status; and quantitatively evaluates and calculates the training results of the dispatcher according to the set operation accuracy, response speed and fault handling effect evaluation indicators.
[0013] This preferred solution achieves accurate matching judgment by comparing the dispatcher's operating behavior with the optimal strategy of the virtual coach model in real time, using a similarity calculation algorithm. When there is a deviation, it triggers operation suggestions or risk warnings based on the current system status, establishing a dynamic feedback chain, thereby improving the timeliness of response to non-standard dispatch operations and the targeted nature of training.
[0014] As a preferred solution of the method for applying the virtual trainer formation technology in power dispatching operation described in the present invention, the quantitative evaluation and calculation of the training results of the dispatchers include setting evaluation indicators including operation accuracy, response speed and fault handling effect according to the requirements and goals of the power dispatching operation; during the simulation training process, the operation data and system response data of the dispatchers are collected, and combined with the evaluation indicators, the degree of compliance of the operation results with the optimal strategy, the response delay and the fault recovery process are statistically analyzed, the scores of various indicators are calculated, and the training effect is comprehensively evaluated.
[0015] This preferred solution sets evaluation indicators such as operation accuracy, response speed, and fault handling effect, and quantifies the behavioral data during the dispatching process. It can achieve an objective evaluation of the training effect of dispatchers, support sub-item analysis and comprehensive judgment of operation performance, and contribute to the continuous optimization and personalized improvement of training programs.
[0016] Another object of the present invention is to provide an application system based on virtual trainer formation technology in power dispatching operations.
[0017] In order to solve the above technical problems, the present invention provides the following technical solutions: a system for applying virtual trainer formation technology in power dispatching operation, comprising: a data acquisition and processing module, a virtual trainer model construction module, a comparison module and an evaluation module; the data acquisition and processing module is used to obtain historical operation data and real-time dispatching operation data, and process the acquired data; the virtual trainer model construction module is used to construct a virtual trainer model for learning the relationship between power dispatching operation and system response based on the processed data, train and optimize the virtual trainer model, and output the optimal operation strategy; the comparison module is used to compare the dispatcher's operation with the optimal operation strategy obtained by the virtual trainer model in real time, output evaluation information according to the comparison results, and provide operation suggestions, risk warnings and fault handling guidance based on the deviation and the current system status; the evaluation module is used to set evaluation indicators according to the requirements of power dispatching operation, including operation accuracy, response speed and fault handling effect, and quantitatively evaluate the training results of the dispatcher according to the evaluation indicators.
[0018] The present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program and is characterized in that when the processor executes the computer program, steps of a method for applying virtual trainer formation technology in power dispatching operations are implemented.
[0019] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the method for applying virtual trainer formation technology in power dispatching operation are implemented.
[0020] The beneficial effects of the present invention are as follows: the present invention acquires and cleans the historical operation data and real-time operation data of power dispatching through the data acquisition and processing module, constructs a structured training database, and provides support for the virtual coach model; the virtual coach model constructed by using the LSTM long short-term memory network combined with the reinforcement learning algorithm can learn the optimal dispatching strategy, and in the real-time guidance and evaluation module, compares the dispatching personnel's operation with the optimal strategy in real time, provides operation suggestions, risk warnings and fault handling guidance, and quantitatively evaluates the training results by setting evaluation indicators such as operation accuracy and response speed, generates a training report and gives personalized improvement suggestions; this data-driven intelligent simulation training method does not need to rely on real equipment operation. Compared with traditional training methods, it can significantly improve training efficiency and quality, realize real-time feedback and intuitive display, enhance the emergency response capabilities of dispatching personnel, provide data-driven operation evaluation and improvement suggestions, reduce actual operation risks, and save training costs and resources, bringing a new efficient and safe way to power system dispatching personnel training. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 An overall flow chart of a method for applying virtual trainer formation technology in power dispatching operations is provided as an embodiment of the present invention.
[0023] Figure 2 A module diagram of a solution for a system for applying virtual trainer formation technology in power dispatching operations, provided as an embodiment of the present invention. DETAILED DESCRIPTION
[0024] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0025] Example 1, with reference to Figure 1 , is an embodiment of the present invention, which provides a method for applying virtual trainer formation technology in power dispatching operations, including:
[0026] S1. Obtain historical operation data and real-time scheduling operation data, and process the acquired data.
[0027] S2. Based on the processed data, a virtual coach model is constructed to learn the relationship between power dispatching operations and system responses, and the virtual coach model is trained and optimized to output the optimal operation strategy.
[0028] S3. Compare the dispatcher's operations with the optimal operation strategy obtained by the virtual trainer model in real time, output evaluation information based on the comparison results, and provide operation suggestions, risk warnings and troubleshooting guidance based on the deviation situation and current system status.
[0029] S4. Set evaluation indicators according to the requirements of power dispatching operations, including operation accuracy, response speed and fault handling effect, and conduct quantitative evaluation of the training results of dispatching personnel based on the evaluation indicators.
[0030] It should be noted that in power dispatch training, traditional rule-based teaching or recorded broadcasting methods lack real-time response and personalized feedback, making it difficult to timely intervene and improve behavioral deviations in dispatch operations. The processes S1 to S4 in this embodiment not only effectively utilize historical and real-time dispatch data, but also provide intelligent comparison and precise feedback related to operations with the support of models. Ultimately, a quantifiable and traceable dispatch training evaluation mechanism is formed, thereby improving the scientific nature and training quality of the dispatch training process.
[0031] Embodiment 2 is an embodiment of the present invention, and provides a method for applying virtual trainer formation technology in power dispatching operations based on the previous embodiment, including:
[0032] In the implementation manner of the present application, historical operation data and real-time scheduling operation data are obtained in step S1, and the obtained data are processed, and the obtained data are cleaned, labeled and classified for constructing a training data set.
[0033] Specifically, the collection targets and scope include historical operation data and real-time dispatch operation data from the power dispatch system. Historical operation data includes load data (including load values for different time periods and regions), power generation data (the power generation and power generation status of each power generation device), topology data (the connection relationships and parameters of each node and line in the power network), and fault record data (fault occurrence time, location, type, and handling process). Real-time dispatch operation data includes each operation instruction of the dispatcher, the operation time, and the system status before and after the operation.
[0034] Data collection technologies and tools utilize existing power system data acquisition interfaces and protocols, as well as the data interfaces provided by SCADA (Supervisory Control and Data Acquisition) systems, to transmit data in real time to a data storage server via network communication technologies (such as TCP / IP). Historical data is extracted from the power dispatch system's historical database using database query statements.
[0035] Data cleaning involves checking the integrity and accuracy of collected data, removing duplicate data, removing data records with excessive missing values, and correcting obviously erroneous data. For abnormally large or small values in load data, statistical analysis is used to determine whether they are erroneous data and correct them accordingly.
[0036] Data labeling involves categorizing and labeling data based on its nature and purpose. Fault records are labeled by fault type (short circuit, open circuit, equipment failure, etc.), and dispatch operation data is labeled by operation purpose (load adjustment, equipment startup, fault resolution, etc.). A combination of manual labeling and machine learning-assisted labeling improves labeling efficiency and accuracy.
[0037] Data classification involves storing the processed data according to different themes and uses, dividing it into load data, power generation data, topology data, fault data, etc., and building a structured training database to facilitate subsequent model training.
[0038] In an optional embodiment, data processing includes correcting extreme values in load data by applying the median substitution method; performing structural verification and rule correction on inconsistent connection information in topological structure data; and in the labeling stage, only expert manual labeling is used to ensure the accuracy of training data.
[0039] In another optional embodiment, data processing includes classifying and labeling scheduling operation data according to different operation purposes (such as starting equipment, load adjustment, and fault handling); splitting the original data into a structured database by date and region through automated batch scripts; and adding an outlier elimination mechanism based on statistical distribution in the cleaning stage.
[0040] The implementation method of the present application ensures that the training data covers various typical operating conditions of the power system by introducing a joint cleaning strategy and a multi-dimensional classification and labeling process for multi-source data, enhances the stability of model learning, and improves the scenario adaptability of subsequent strategy outputs.
[0041] In the embodiment of the present application, in step S2, a virtual coach model for learning the relationship between power dispatching operations and system responses is constructed based on the processed data, the virtual coach model is trained and optimized, and the optimal operation strategy is output.
[0042] A structure based on LSTM neural network is used to process time series data in power dispatching. By setting cell state, forget gate, input gate and output gate, information is transmitted between different time steps of the sequence, and the retention and update mode of information is determined according to the current input and the state of the previous moment.
[0043] As an intelligent agent in reinforcement learning, the virtual coach model performs dispatching operations in a simulated power dispatching environment, obtains rewards or penalties based on the changes in system state after execution, and updates the behavioral strategy through the state-action-reward relationship.
[0044] The parameters of the virtual coach model include neural network structure parameters, reward function parameters, learning rate and discount factor, and the optimal operation strategy for various scheduling scenarios is learned through trial-and-error interaction and parameter adjustment process.
[0045] Training and optimizing the virtual coach model involves inputting training data into the constructed virtual coach model. Through iterative calculation and optimization algorithms, the model learns the relationship between scheduling operations and operating results in historical data, and masters the optimal operating strategies under different scenarios. During the training process, the model is evaluated using a validation set, and the model parameters are adjusted based on the evaluation results.
[0046] Specifically, the LSTM (Long Short-Term Memory) network is a special type of recurrent neural network (RNN) designed to address the vanishing and exploding gradient problems faced by traditional RNNs when processing long sequences of data. It is therefore well-suited for processing time series data in power dispatch, such as load fluctuations and fault development. By introducing a gating mechanism, the LSTM can selectively forget and remember information, thereby better capturing long-range dependencies. An LSTM network consists of a cell state (CellState), a forget gate (ForgetGate), an input gate (Input Gate), and an output gate (Output Gate). The cell state acts like a conveyor belt, transferring information throughout the network and allowing it to flow between different time steps in the sequence. The forget gate determines which information from the previous cell state is discarded; the input gate determines which information from the current input is added to the cell state; and the output gate generates the current output based on the cell state.
[0047] Forget gate calculation:
[0048] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0049] Among them, f t is the output of the forget gate at time t, σ is the Sigmoid activation function, and the output value is between 0 and 1, indicating the degree of forgetting; W f is the weight matrix of the forget gate; [h t-1 ,x t ] indicates the hidden state of the previous moment; h t-1 and the current input x t Splicing; b f is the bias term of the forget gate.
[0050] Input gate calculation:
[0051] i t =σ(W i , ·[h t-1 ,x t ]+b i )
[0052]
[0053] Among them, i t It is the output of the input gate at time t, which determines the degree of retention of the current input information; is the candidate cell state generated at the current moment, tanh is the hyperbolic tangent activation function, and the output value is between -1 and 1; W i 、W C are the weight matrices for input gate and candidate cell state calculations; b i 、b C are the corresponding bias terms respectively.
[0054] Cell status update:
[0055]
[0056] Among them, C t is the cell state after updating at time t, C t-1 is the cell state at the previous moment. The cell state is updated through the control of the forget gate and the input gate.
[0057] Output gate calculation:
[0058] o t =σ(W0×[h t-1 ,x t ]+b0)
[0059] h t =o t ×tanh(C t )
[0060] Among them, t is the output of the output gate at time t, which determines which information in the cell state will be output; h t It is the hidden state at time t, which is used as the output of the current moment and passed to the next moment.
[0061] Reinforcement learning algorithms are machine learning methods that enable intelligent agents to learn optimal behavioral strategies through trial and error in an environment, based on reward signals provided by the environment. In power dispatch applications, a virtual trainer model, acting as an intelligent agent, interacts with a simulated power dispatch environment, performs various dispatch operations, and receives rewards or penalties based on the resulting changes in system state. This allows the system to gradually learn the optimal dispatch strategy for various scenarios.
[0062] The core elements of reinforcement learning include an agent, an environment, a state, an action, a reward, and a policy. The agent observes the state of the environment at every moment and selects actions based on the policy. The environment updates its state after receiving the action and gives the agent a corresponding reward. The agent's goal is to maximize the long-term cumulative reward.
[0063] State-action-reward relationship, at each time step t, the agent is in state s t , perform action a t After that, transfer to the new state s t+1 , and receive a reward of r t , can be expressed as: (s t ,a t ,r t , s t+1 ).
[0064] Policy function, policy π defines the probability distribution of selecting action a under a given state s, that is, π(a|s)=P(at=a|st=s), which represents the probability of selecting action a under state s.
[0065] The value function is used to evaluate how good it is to take a specific strategy in a certain state.
[0066] V π (s) represents the expected cumulative reward that can be obtained by following the strategy π starting from state s:
[0067]
[0068] Among them, γ is the discount factor, which ranges from 0 to 1 and is used to balance the importance of current rewards and future rewards; E π represents the expectation based on the policy π. The action value function Q π (s,a) represents the expected cumulative reward that can be obtained by following strategy π after executing action a in state s:
[0069]
[0070] Q-learning is a commonly used model-free reinforcement learning algorithm that learns the optimal strategy by continuously updating the action-value function Q(s,a). Its update formula is:
[0071]
[0072] Among them, α is the learning rate, which controls the step size of each update; In the new state s t+1The maximum action value of all possible actions. By continuously iteratively updating the Q value, the optimal action value function Q is finally learned. * , thus obtaining the optimal strategy π * .
[0073] Configure the parameters of the selected algorithm, including the number of layers and nodes in the neural network, the reward function setting for reinforcement learning, the learning rate, the discount factor, etc. Through experiments and parameter optimization, the optimal parameter combination can be determined to improve the learning effect and generalization ability of the model.
[0074] Model training and optimization include the following steps:
[0075] Training data preparation: Select appropriate data from the training database as the training set, and divide the training set, validation set, and test set into a certain ratio (e.g., 7:2:1). Ensure that the training data covers various power dispatch scenarios, including normal operation, fault handling, load adjustment, and other different operating conditions.
[0076] Model training: Training data is fed into the constructed virtual trainer model. Through iterative calculations and optimization algorithms, the model learns the relationship between scheduling operations and operational results in historical data, and grasps the optimal operation strategy for different scenarios. During the training process, the model is evaluated using a validation set, and model parameters are adjusted based on the evaluation results to prevent overfitting or underfitting.
[0077] Model optimization and validation: Optimize the model by continuously adjusting algorithm parameters, improving model structure, or adding training data. Use the test set to conduct final validation of the optimized model to ensure its accuracy and reliability in practical applications.
[0078] In the implementation mode of the present application, in step S3, the dispatcher's operation is compared in real time with the optimal operation strategy obtained by the virtual trainer model, and evaluation information is output based on the comparison results. Operation suggestions, risk warnings and fault handling guidance are provided based on the deviation situation and the current system status.
[0079] When there is a deviation between the operation and the optimal strategy, operation suggestions are generated based on the deviation and the current system status. Risk warnings are issued based on real-time system status and scheduling operation analysis. In simulated fault scenarios, operational procedures for fault location, isolation, and power restoration are provided.
[0080] During the simulation training process, the dispatcher's operation instructions, operation time and operation object information are obtained through the real-time data interface.
[0081] The acquired operation behavior is compared with the optimal strategy, and the degree of matching is determined using a similarity calculation algorithm to determine whether the operation conforms to the optimal strategy.
[0082] When the operation deviates from the optimal strategy, the virtual coach generates operation suggestions, risk warnings or troubleshooting guidance based on the deviation and the current system status.
[0083] And according to the set evaluation indicators of operation accuracy, response speed and fault handling effect, the training results of the dispatchers are quantitatively evaluated and calculated.
[0084] Specifically, real-time data acquisition: During the simulation training process, the dispatcher's operational behavior data is obtained through a real-time data interface, including operation instructions, operation time, operation objects and other information.
[0085] Comparison with the Optimal Strategy: The dispatcher's actions are compared in real time with the optimal strategy learned by the virtual trainer model. Using a similarity calculation algorithm, the degree of match between the dispatcher's actions and the optimal strategy is calculated to determine whether the actions conform to the optimal strategy.
[0086] Action Suggestions: When the dispatcher's actions deviate from the optimal strategy, the virtual coach generates corresponding action suggestions based on the deviation and the current system status. For example, during load adjustment, if the dispatcher's actions result in an unreasonable load distribution, the virtual coach provides specific solutions and steps to adjust the load distribution.
[0087] Risk Warning: Based on real-time system status and dispatcher operations, combined with analysis from the virtual trainer model, early warnings are issued for potential risks. For example, if the system load approaches a critical value and the dispatcher performs an operation that may affect load balance, an overload risk warning is issued.
[0088] Fault handling guidance: In simulated fault scenarios, the virtual trainer provides dispatchers with fault handling guidance, including fault location methods, fault isolation steps, and power restoration procedures.
[0089] In the embodiment of the present application, in step S4, evaluation indicators are set according to the requirements of power dispatching operations, including operation accuracy, response speed and fault handling effect, and the training results of the dispatching personnel are quantitatively evaluated based on the evaluation indicators.
[0090] Set evaluation indicators including operation accuracy, response speed and fault handling effect according to the requirements and goals of power dispatch operations;
[0091] During the simulation training process, the dispatcher's operation data and system response data are collected, and combined with evaluation indicators, statistical analysis is conducted on the degree of compliance between the operation results and the optimal strategy, response delay and fault recovery process, the scores of various indicators are calculated, and a comprehensive evaluation of the training effect is conducted.
[0092] Specifically, according to the requirements and objectives of power dispatching operations, preset evaluation indicators are set, such as operation accuracy (the degree of conformity between the operation results and the optimal results), response speed (the time from the occurrence of a fault or the proposal of an operation requirement to the dispatcher's operation), fault handling effect (the efficiency of fault handling and the integrity of power supply restoration), etc.
[0093] The training results of dispatchers are quantitatively evaluated and calculated based on the set evaluation indicators. The relevant data during the dispatcher's operation is collected and analyzed, and the scores of various indicators are calculated to comprehensively evaluate the training effect.
[0094] Collect all operation data, system response data, evaluation index calculation results and other information of the dispatcher during the simulation training process.
[0095] Analyze and organize the collected data and write a training report. The report content includes the dispatcher's operational performance (such as case analysis of successful and failed operations), detailed data on various evaluation indicators, and comparative analysis with the optimal strategy.
[0096] Based on the problems in the dispatcher's operations analyzed in the training report and combined with the knowledge and experience of the virtual trainer model, the causes of the problems are deeply analyzed.
[0097] Generate personalized improvement suggestions based on the specific issues and characteristics of each dispatcher. For example, for dispatchers with slow operational response speeds, we recommend operational process familiarity training and rapid response training in simulated emergency situations. For dispatchers with poor troubleshooting performance, we provide more detailed troubleshooting knowledge learning materials and targeted simulation training scenarios.
[0098] In an optional embodiment, the response speed assessment is further subdivided into the response time of multiple key nodes, including instruction issuance time, system confirmation time and status response time, and statistical analysis of response deviations of different trainees is used as one of the evaluation dimensions.
[0099] In another optional embodiment, the evaluation of the fault handling effect includes scoring the dispatcher's continuity in the four stages of fault identification, location, isolation, and recovery. If there is a logical jump or confusion in the operation, the system will record it as incoherent processing and reflect it in the final evaluation.
[0100] The implementation method of this application = By constructing a multi-dimensional evaluation index system covering operational correctness, time responsiveness and fault handling capabilities, the training results of dispatchers can be presented objectively and structured, further supporting subsequent targeted training optimization and capacity improvement evaluation.
[0101] Example 3, reference Figure 2, is an embodiment of the present invention, which provides an application system based on virtual coach formation technology in power dispatching operation, including a data acquisition and processing module, a virtual coach model construction module, a comparison module and an evaluation module.
[0102] The data acquisition and processing module is used to obtain historical operation data and real-time scheduling operation data, and process the obtained data.
[0103] The virtual coach model construction module is used to build a virtual coach model for learning the relationship between power dispatching operations and system responses based on the processed data, train and optimize the virtual coach model, and output the optimal operation strategy.
[0104] The comparison module is used to compare the dispatcher's operations with the optimal operation strategy obtained by the virtual trainer model in real time, output evaluation information based on the comparison results, and provide operation suggestions, risk warnings and fault handling guidance based on the deviation situation and current system status.
[0105] The evaluation module is used to set evaluation indicators according to the requirements of power dispatching operations, including operation accuracy, response speed and fault handling effect, and to quantitatively evaluate the training results of dispatchers based on the evaluation indicators.
[0106] This embodiment also provides an electronic device, which is suitable for a method of applying virtual coach formation technology in power dispatching operation, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement a method of applying virtual coach formation technology in power dispatching operation as proposed in the above embodiment.
[0107] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for applying a virtual trainer formation technology in power dispatching operations as proposed in the above embodiment.
[0108] The storage medium proposed in this embodiment and the method for implementing a method based on virtual coach formation technology in power dispatching operation proposed in the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0109] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer's floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0110] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for applying virtual trainer formation technology in power dispatching operations, characterized by: include, Obtain historical operation data and real-time scheduling operation data, and process the acquired data; Based on the processed data, a virtual trainer model is constructed to learn the relationship between power dispatch operations and system responses. The virtual trainer model is trained and optimized to output the optimal operation strategy. Compare the dispatcher's operations with the optimal operation strategy obtained by the virtual trainer model in real time, output evaluation information based on the comparison results, and provide operation suggestions, risk warnings and troubleshooting guidance based on deviations and current system status; Evaluation indicators are set according to the requirements of power dispatching operations, including operation accuracy, response speed and fault handling effect, and the training results of dispatching personnel are quantitatively evaluated based on the evaluation indicators.
2. The method for applying the virtual trainer formation technology in power dispatching operation according to claim 1, characterized in that: The processing of the acquired data includes cleaning, labeling and classifying the acquired data for constructing a training data set.
3. The method for applying the virtual trainer formation technology in power dispatching operations according to claim 2, characterized in that: The training and optimization of the virtual coach model includes inputting training data into the constructed virtual coach model, and through iterative calculation and optimization algorithms, allowing the model to learn the relationship between scheduling operations and operation results in historical data, and master the optimal operation strategy under different scenarios. During the training process, the model is evaluated using a validation set, and the model parameters are adjusted according to the evaluation results.
4. The method for applying the virtual trainer formation technology in power dispatching operations according to claim 3, characterized in that: The output of evaluation information based on the comparison results includes generating operation suggestions based on the deviation situation and the current system status when there is a deviation between the operation and the optimal strategy, issuing risk warnings based on the real-time system status and scheduling operation analysis, and providing operational procedures for fault location, isolation and power restoration in simulated fault scenarios.
5. The method for applying the virtual trainer formation technology in power dispatching operation according to claim 4, characterized in that: The virtual trainer model for learning the relationship between power dispatching operations and system responses includes: The LSTM neural network-based structure is used to process time series data in power dispatching. By setting cell states, forget gates, input gates, and output gates, information is transferred between different time steps of the sequence, and the information retention and update methods are determined based on the current input and the previous state. The virtual coach model, as an intelligent agent in reinforcement learning, performs dispatch operations in a simulated power dispatch environment, obtains rewards or penalties based on the changes in system state after execution, and updates its behavior strategy through the state-action-reward relationship. The parameters of the virtual coach model include neural network structure parameters, reward function parameters, learning rate and discount factor, and the optimal operation strategy for various scheduling scenarios is learned through trial-and-error interaction and parameter adjustment process.
6. The method for applying the virtual trainer generation technology in power dispatching operations according to claim 4, characterized in that: The real-time comparison includes obtaining the dispatcher's operation instructions, operation time and operation object information through a real-time data interface during the simulation training; Compare the acquired operation behavior with the optimal strategy, use a similarity calculation algorithm to determine the degree of match, and judge whether the operation conforms to the optimal strategy; When the operation deviates from the optimal strategy, the virtual coach generates operational suggestions, risk warnings, or troubleshooting guidance based on the deviation and the current system status; And according to the set evaluation indicators of operation accuracy, response speed and fault handling effect, the training results of the dispatchers are quantitatively evaluated and calculated.
7. The method for applying the virtual trainer generation technology in power dispatching operations according to claim 4, characterized in that: The quantitative evaluation calculation of the training results of the dispatcher includes: Set evaluation indicators including operation accuracy, response speed and fault handling effect according to the requirements and goals of power dispatch operations; During the simulation training process, the dispatcher's operation data and system response data are collected, and combined with evaluation indicators, statistical analysis is conducted on the degree of compliance between the operation results and the optimal strategy, response delay and fault recovery process, the scores of various indicators are calculated, and a comprehensive evaluation of the training effect is conducted.
8. A system for applying virtual trainer generation technology in power dispatching operations, applying a method for applying virtual trainer generation technology in power dispatching operations according to any one of claims 1 to 7, characterized in that: include: Data acquisition and processing module, virtual coach model construction module, comparison module and evaluation module; The data acquisition and processing module is used to obtain historical operation data and real-time scheduling operation data, and process the obtained data; The virtual trainer model building module is used to build a virtual trainer model for learning the relationship between power dispatching operations and system responses based on the processed data, train and optimize the virtual trainer model, and output the optimal operation strategy; The comparison module is used to compare the dispatcher's operation with the optimal operation strategy obtained by the virtual trainer model in real time, output evaluation information based on the comparison results, and provide operation suggestions, risk warnings and fault handling guidance based on deviations and current system status; The evaluation module is used to set evaluation indicators according to the requirements of power dispatching operations, including operation accuracy, response speed and fault handling effect, and to quantitatively evaluate the training results of dispatching personnel based on the evaluation indicators.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a method for applying virtual trainer formation technology in power dispatching operations according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for applying virtual trainer formation technology in power dispatching operation according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Elevator real-time scheduling optimization system based on edge calculation
CN121573533A
An elevator real-time scheduling optimization system based on edge computing
CN121573533B