Practical operation training method, device and program product of AGC system of pumped storage power station

Through the operation and maintenance training of the AGC system of the pumped storage power station, operation data of operation and maintenance personnel are collected and optimal strategies are obtained using the cyclic pulse neural network and DQN algorithm, the problems of low efficiency and high safety risks of traditional training methods are solved, and efficient and safe operation and maintenance training is achieved.

CN120164253APending Publication Date: 2025-06-17CHINA YANGTZE POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200712.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The operation and maintenance training of traditional pumped storage power station AGC system has problems such as long training cycle, low efficiency and safety risks brought by on-site operation.

Method used

The operation data of operation and maintenance personnel is collected using preset collection frequency, and the operation sequence is obtained based on the operation action data and the cyclic pulse neural network. The operation sequence is processed using the DQN algorithm to obtain the optimal strategy, and the operation and maintenance personnel are guided to complete the operation process of practical training based on the optimal strategy.

Benefits of technology

It improves training efficiency, reduces practical risks, significantly improves the safety of practical links, and ensures a high degree of consistency between training and actual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164253A_ABST
    Figure CN120164253A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a practical operation training method and device of a pumped storage power station AGC system and a program product. The method comprises the following steps: in practical operation training of an AGC system of the pumped storage power station, acquiring operation data of operation and maintenance personnel according to a preset acquisition frequency; obtaining an operation action sequence based on the operation action data and a cyclic pulse neural network; processing the operation action sequence by adopting a DQN algorithm to obtain an optimal strategy; guiding the operation and maintenance personnel to complete the operation process of practical operation training based on the optimal strategy; wherein the operation data is continuous action data of a preset body part of the operation and maintenance personnel. According to the method, personalized training can be carried out on the operation and maintenance personnel, the training efficiency is improved, the practical operation risk is reduced, and the safety of the practical operation link is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of pumped - storage energy storage, and particularly to a practical training method, device, and program product for the AGC system of a pumped - storage power station. Background Art

[0002] A pumped - storage power station is an energy facility that efficiently utilizes the load fluctuations of the power system and is a key support for ensuring the safe and stable operation of the power grid. As the core of the pumped - storage power station, the Automatic Generation Control (AGC) system is responsible for the automatic start - stop and load distribution control of the pumped - generator units, and plays a crucial role in the stable regulation of the power grid frequency. To ensure the effective operation of the AGC system, professional training is required for the operation and maintenance personnel so that they can master the necessary professional knowledge and skills.

[0003] However, traditional AGC operation and maintenance training mainly relies on manual teaching and on - site practical operations. This method has some limitations. The training cycle is long, the efficiency is not high, and the on - site practical operation link may bring safety risks, especially in the case of high - voltage and mechanical operations. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure provide a practical training method, device, and program product for the AGC system of a pumped - storage power station, which can provide personalized training for operation and maintenance personnel, improve training efficiency, reduce practical operation risks, and significantly improve the safety of the practical operation link.

[0005] In a first aspect, the embodiments of the present disclosure provide a practical training method for the AGC system of a pumped - storage power station, adopting the following technical solutions:

[0006] In the practical training of the AGC system of a pumped - storage power station, operation data of the operation and maintenance personnel is collected at a preset sampling frequency;

[0007] Based on the operation action data and a recurrent spiking neural network, an operation action sequence is obtained;

[0008] The DQN algorithm is used to process the operation action sequence to obtain an optimal policy;

[0009] Based on the optimal policy, the operation and maintenance personnel are guided to complete the operation process of the practical training;

[0010] Wherein, the operation data is continuous action data of a preset body part of the operation and maintenance personnel.

[0011] Optionally, the obtaining of the operation action sequence based on the operation action data and a recurrent spiking neural network includes:

[0012] Normalize and pulse-code process the operation data to obtain a pulse sequence composed of multiple pulse combinations;

[0013] Input the pulse sequence into the recurrent spiking neural network, which includes recurrent units for multiple time steps. Each time step corresponds to a pulse combination. Input the pulse combination and the hidden state of the previous time step into the corresponding recurrent unit to obtain the hidden state of the corresponding time step;

[0014] Construct a human body skeleton topology graph, perform graph convolution operations on the human body skeleton topology graph to obtain spatial features;

[0015] Concatenate the hidden states of all time steps and the spatial features to obtain spatio-temporal features;

[0016] Based on the spatio-temporal features and a classifier, obtain the action recognition results for all time steps, and combine the action recognition results for all time steps into the operation action sequence.

[0017] Optionally, the normalizing and pulse-coding processing of the operation data to obtain a pulse sequence composed of multiple pulse combinations includes:

[0018] Normalize the multiple action data included in the operation data, and convert the multiple action data into multiple standard data;

[0019] Traverse the multiple standard data, and respectively determine whether each value in the standard data is greater than a preset data threshold;

[0020] If so, modify the value to 1;

[0021] If not, modify the value to 0;

[0022] Combine the modified values into a pulse combination;

[0023] The pulse combinations converted from the multiple standard data form a pulse sequence.

[0024] Optionally, the processing of the operation action sequence using the DQN algorithm to obtain an optimal policy includes:

[0025] Collect the grid frequency and pumped-storage unit load for all time steps;

[0026] Construct a state space, map the operation action sequence, the grid frequency and the pumped-storage unit load for all time steps into the state space, and obtain the first state value, the second state value and the third state value for all time steps;

[0027] Obtain the state variables for all time steps based on the first state value, second state value, and third state value for all the time steps.

[0028] Construct a value network and a target network.

[0029] Input the state variables for all the time steps into the value network respectively to obtain the Q-values of all possible actions under the state variables of the current time step, where Q is the action-value function.

[0030] Select the possible action with the largest Q-value as the current action, execute the current action, and obtain the new grid frequency and the new pumped-storage unit load.

[0031] Map the new pumped-storage unit load to the state space to obtain the new third state value for the current time step.

[0032] Based on the new grid frequency, the standard grid frequency, the third state value of the current time step, and the new third state value of the current time step, obtain the immediate reward for the current time step.

[0033] The state variables of the current time step, the current action, the immediate reward, and the state variables of the next time step form a set of experience data.

[0034] Optimize the value network and the target network based on the experience data.

[0035] Obtain the optimal policy based on the optimized value network and target network.

[0036] Optionally, the mapping of the operation action sequence, the grid frequency and the pumped-storage unit load for all time steps to the state space to obtain the first state value, second state value, and third state value for all time steps includes:

[0037] The state space includes multiple operation and maintenance actions, multiple preset frequency values, and multiple preset load values. Each operation and maintenance action corresponds to a first variable, each preset frequency value corresponds to a second variable, and each preset load value corresponds to a third variable.

[0038] Match the multiple action recognition results included in the operation action sequence with the operation and maintenance actions respectively, and record the first variable corresponding to the operation and maintenance action that matches the action recognition result as the first state value.

[0039] Among the multiple preset frequency values, select the second variable corresponding to the preset frequency value closest to the grid frequency of the current time step as the second state value.

[0040] Among the multiple preset load values, select the third variable corresponding to the preset load value closest to the pumped-storage unit load of the current time step as the third state value.

[0041] Optionally, the calculation formula of the instant reward is as follows:

[0042] R = A * R1 + B * R2 + C * R3;

[0043] Wherein, R is the instant reward; A, B, and C are preset weight values; R1 is the frequency deviation reward; R2 is the load change reward; and R3 is the start-stop operation reward.

[0044] Optionally, the practical training method for the AGC system of the pumped storage power station further includes:

[0045] Calculate the absolute value of the difference between the new grid frequency and the standard grid frequency, denoted as the first absolute value;

[0046] Determine whether the first absolute value is greater than a first threshold;

[0047] If it is greater than the first threshold, the frequency deviation reward is W1;

[0048] If it is not greater than the first threshold, the frequency deviation reward is W2.

[0049] Optionally, the practical training method for the AGC system of the pumped storage power station further includes:

[0050] Calculate the absolute value of the difference between the third state value at the current time step and the new third state value at the current time step, denoted as the second absolute value;

[0051] Determine whether the second absolute value is greater than a second threshold;

[0052] If it is greater than the second threshold, the load change reward is W3;

[0053] If it is not greater than the second threshold, the load change reward is W4.

[0054] In a second aspect, the embodiments of the present disclosure further provide a practical training system for the AGC system of a pumped storage power station, adopting the following technical solution:

[0055] An acquisition module, configured to acquire operation data of operation and maintenance personnel at a preset acquisition frequency during the practical training of the AGC system of the pumped storage power station;

[0056] An obtaining module, configured to obtain an operation action sequence based on the operation action data and a recurrent pulse neural network;

[0057] A processing module, configured to process the operation action sequence by using a DQN algorithm to obtain an optimal strategy;

[0058] A guidance module for guiding the operation and maintenance personnel to complete the operation process of practical training based on the optimal strategy;

[0059] Wherein, the operation data is the continuous action data of the preset body parts of the operation and maintenance personnel.

[0060] In a third aspect, an embodiment of the present disclosure also provides a computer device, adopting the following technical solution:

[0061] The computer device includes:

[0062] At least one processor; and,

[0063] A memory communicatively connected to the at least one processor; wherein,

[0064] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the practical training method of the pumped storage power station AGC system described in any one of the above.

[0065] In a fourth aspect, an embodiment of the present disclosure also provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the practical training method of the pumped storage power station AGC system described in any one of the above.

[0066] In a fifth aspect, an embodiment of the present disclosure also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.

[0067] The practical training method for the AGC system of a pumped - storage power station provided by the embodiments of the present disclosure efficiently and purposefully collects the operation data of operation and maintenance personnel through a preset acquisition frequency, ensuring the systematicness and comprehensiveness of the continuous action data collection of the preset body parts of operation and maintenance personnel. The cyclic pulse neural network is used to analyze the operation action data, accurately identify the operation action sequence of operation and maintenance personnel, and provide accurate input for subsequent strategy optimization. The deep Q - network (DQN) algorithm is used to process the operation action sequence, which can learn and obtain the optimal operation strategy, enhancing the intelligence level of the training method. Based on the optimal strategy, practical training guidance is provided for operation and maintenance personnel, improving the pertinence and practicability of the training and ensuring a high degree of consistency between the training and actual operation. Through the training mechanism of this method, operation and maintenance personnel can master complex operation processes faster, improving their professional skills and emergency handling capabilities. In summary, this method can conduct personalized training based on the continuous action data of the preset body parts of operation and maintenance personnel, meeting the training needs of different personnel. Combining artificial intelligence technology with practical training improves the training efficiency, provides a technological innovation for traditional training methods, and can significantly improve the operation and maintenance efficiency and safety of AGC - function - related equipment, reducing accident risks.

[0068] The above description is only an overview of the technical solutions of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the description. And in order to make the above - mentioned and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given, and in conjunction with the drawings, the details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0070] Figure 1 It is a schematic flowchart of the practical training method for the AGC system of a pumped - storage power station provided by the embodiments of the present disclosure;

[0071] Figure 2 It is a schematic flowchart of the method for obtaining the operation action sequence provided by the embodiments of the present disclosure;

[0072] Figure 3 It is a schematic flowchart of the method for obtaining the optimal strategy provided by the embodiments of the present disclosure;

[0073] Figure 4 It is a schematic block diagram of the principle of the practical training system for the AGC system of a pumped - storage power station provided by the embodiments of the present disclosure;

[0074] Figure 5 A schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Specific implementation manners

[0075] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0076] It should be clear that the embodiments of the present disclosure are described through specific specific examples below. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0077] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0078] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0079] In addition, in the following description, specific details are provided for the purpose of facilitating a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0080] Referring to Figure 1 , the present disclosure provides a practical training method for the AGC system of a pumped storage power station, including the following steps:

[0081] S1: In the practical training of the AGC system of a pumped-storage power station, collect the operation data of maintenance personnel according to a preset collection frequency;

[0082] Among them, the operation data is the continuous action data of the preset body parts of the maintenance personnel.

[0083] S2: Based on the operation action data and a recurrent pulse neural network, obtain the operation action sequence;

[0084] S3: Use the DQN algorithm to process the operation action sequence and obtain the optimal strategy;

[0085] S4: Based on the optimal strategy, guide the maintenance personnel to complete the operation process of the practical training.

[0086] The practical training method for the AGC system of the pumped-storage power station disclosed in this disclosure efficiently and purposefully collects the operation data of maintenance personnel through a preset collection frequency, ensuring the systematicness and comprehensiveness of the collection of the continuous action data of the preset body parts of the maintenance personnel. Analyze the operation action data using a recurrent pulse neural network to accurately identify the operation action sequence of the maintenance personnel, providing accurate input for subsequent strategy optimization. Using the Deep Q-Network (DQN) algorithm to process the operation action sequence can learn and obtain the optimal operation strategy, enhancing the intelligent level of the training method. Guiding the maintenance personnel in practical training based on the optimal strategy improves the pertinence and practicality of the training, ensuring a high degree of consistency between the training and actual operation. Through the training mechanism of this method, the maintenance personnel can master complex operation processes faster, enhancing their professional skills and emergency response capabilities. This real-time acquisition of operation data and strategy optimization can be timely feedback to the maintenance personnel and adjusted according to their learning progress, and the reinforcement learning algorithm enables the maintenance personnel to make higher-quality decisions based on historical data and real-time information when facing power grid regulation decisions.

[0087] In summary, this method can provide personalized training based on the continuous action data of the preset body parts of the maintenance personnel to meet the training needs of different personnel. Combining artificial intelligence technology with practical training improves the training efficiency, provides a technological innovation for traditional training methods, and can significantly improve the operation and maintenance efficiency and safety of AGC function-related equipment, reducing the accident risk.

[0088] In S1, during the practical training of the automatic generation control system (hereinafter referred to as the AGC system) of a pumped-storage power station, the operation data of the operation and maintenance personnel is collected in real time. The operation data is the continuous action data of the preset body parts of the operation and maintenance personnel. Among them, the preset body parts include at least two human key points of the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. The action data includes at least one of position, posture, posture angle, movement speed, and timestamp. Among them, the position refers to the position coordinates (x, y, z) of the human key point in three-dimensional space; the posture refers to the posture information of the human key point, including its orientation and direction, such as "upward", "forward", etc.; the posture angle refers to the pitch angle (pitch), roll angle (roll), and yaw angle (yaw) of the human key point; the movement speed refers to the linear speed and / or angular speed of the human key point; the timestamp refers to the precise time of each action data, which is used to record the time sequence of the action data. These action data are recorded in real time through sensors and motion capture systems to capture the specific actions and posture changes of the operation and maintenance personnel during the operation process, so as to provide a scientific basis for subsequent training and help improve operation efficiency and safety.

[0089] Specifically, the preset sampling frequency is 20 Hz, that is, 20 samples are taken per second. If the right wrist of the operation and maintenance personnel performs a certain action for 0.5 seconds, then 10 sampling points are generated (0.5 seconds × 20 Hz = 10 sampling points). Correspondingly, 10 groups of action data of the right wrist will be collected.

[0090] Refer to Figure 2 Referring to the flowchart of the operation action sequence acquisition method shown, in S2, based on the operation action data and the recurrent pulse neural network, the operation action sequence is obtained, including the following steps:

[0091] S21: Perform normalization processing and pulse coding processing on the operation data to obtain a pulse sequence composed of multiple pulse combinations;

[0092] S22: Input the pulse sequence into the recurrent pulse neural network. The recurrent pulse neural network includes recurrent units of multiple time steps. Each time step corresponds to a pulse combination. Input the pulse combination and the hidden state of the previous time step into the corresponding recurrent unit to obtain the hidden state of the corresponding time step;

[0093] S23: Construct a human skeleton topology graph, perform graph convolution operation on the human skeleton topology graph to obtain spatial features;

[0094] S24: Concatenate the hidden states and spatial features of all time steps to obtain spatio-temporal features;

[0095] S25: Based on the spatio-temporal features and the classifier, obtain the action recognition results of all time steps, and combine the action recognition results of all time steps into an operation action sequence.

[0096] In S21, the normalization process refers to normalizing each value of each action data in the operation data to the value range of [0, 1], and converting each action data into standard data. For example, the position of the right wrist of the operation and maintenance personnel at a certain moment is collected as (0.3, 0.5, 0.2). Each value of this action data, that is, 0.3, 0.5, and 0.2, is already within the range of [0, 1]. After the normalization process, the obtained standard data is still (0.3, 0.5, 0.2), and each value of the standard data refers to 0.3, 0.5, and 0.2.

[0097] In a specific embodiment, the pulse coding process refers to using an encoder (such as an autoencoder) to convert continuous action data, that is, operation data, into a discrete pulse sequence according to the coding rule. Specifically, it includes: traversing the standard data converted from each action data, and respectively judging whether each value in the standard data is greater than the preset data threshold; if so, modifying the value to 1; if not, modifying the value to 0; combining the modified values into a pulse combination, and each pulse combination represents the action state of a time step. For example, performing pulse coding processing on the standard data (0.3, 0.5, 0.2) obtains a pulse combination of (0, 1, 0). The pulse combinations corresponding to each action data in the operation data form a pulse sequence. Specifically, multiple pulse combinations are combined in chronological order to form a pulse sequence, that is, the pulse combinations converted from multiple standard data form a pulse sequence, which records the change of actions over time. Among them, the preset data threshold is 0.5.

[0098] In S22, the recurrent neural network (RNN) includes recurrent units for multiple time steps. Each time step corresponds to a pulse combination. The pulse combination and the hidden state of the previous time step are input into the corresponding recurrent unit to obtain the hidden state of the corresponding time step, which is the current time step. Specifically, at each time step, the corresponding recurrent unit receives two inputs: one is the pulse combination of the corresponding time step, which represents the action data at the current moment; the other is the hidden state of the previous time step, which carries the information of the previous time series. By combining these two inputs, the recurrent unit calculates and updates the hidden state of its corresponding time step, thereby capturing and maintaining the dynamic changes and temporal information in the action sequence. Among them, the previous time step refers to a time point immediately before the time step corresponding to the recurrent unit. In the time series, the action data is arranged in chronological order, and the time step is a single element in the time series, representing a specific time point in the time series. Considering the time step t in the time series, the previous time step refers to the time step t - 1, that is, the time point immediately before the time step t. The pulse sequence is a special form of manifestation in the time series, in which the action data at each time step is converted into a pulse form to represent the occurrence of a specific action or event. It can also be understood in this way: the time series contains continuous observation data of the operations and maintenance personnel's actions, such as the changes in hand position and joint angle over time; the pulse sequence is converted from the time series through sampling and normalization processing, and then into a discrete pulse form through specific coding rules. Each pulse combination represents a specific action or state.

[0099] Among them, the recurrent unit selectively retains and updates long-term memory through a gating mechanism (such as the input gate, forget gate, and output gate of LSTM). The input gate determines how much new information can enter the cell state, and the forget gate determines how much old information can be retained. For example, when the current pulse input is the "lift" action of the hand, through the gating mechanism, this action is associated with the "reach forward" action in the previous several steps to form a coherent action sequence, and then the hidden state of the corresponding time step is output, including the action information at the current moment and the historical information at the previous moment.

[0100] The recurrent neural network with pulses finally outputs the hidden states of all time steps, that is, the hidden states of consecutive multiple time steps. These states contain all the historical action information from the start of the time series to the current time step. This processing method enables the recurrent neural network with pulses to capture the long-term dependencies in the action sequence, such as associating actions like "enter the computer room" and "open the control cabinet" with the final action of "start the water pump". Among them, the action sequence refers to the set of a series of actions or action states of the operations and maintenance personnel arranged in the order of occurrence in the time series.

[0101] In S23, each human body key point is abstracted as a node of a graph, and the connection between each human body key point is abstracted as an edge of the graph, thus constructing a human body skeleton topology graph. When the preset body parts include these human body key points such as head, neck, shoulder, elbow, wrist, hip, knee, and ankle, the node transformed from the head is connected to the node transformed from the neck, the node transformed from the neck is connected to the node transformed from the shoulder, the node transformed from the shoulder is connected to the node transformed from the elbow, the node transformed from the elbow is connected to the node transformed from the wrist, the node transformed from the wrist is connected to the node transformed from the hand, the node transformed from the hip is connected to the node transformed from the knee, the node transformed from the knee is connected to the node transformed from the ankle, and the node transformed from the ankle is connected to the node transformed from the foot. The connection between each human body key point reflects the actual biomechanical characteristics, ensuring the naturalness and accuracy of action simulation.

[0102] Perform graph convolution operation on the human body skeleton topology graph to extract the local features of each node and the spatial relationship between nodes, obtaining a spatial feature representation containing human body key points. For example, the right wrist node aggregates the features of its neighbor nodes (such as the right elbow) through graph convolution to establish the association between hand actions and upper limb actions.

[0103] In S24, the spatial features output by graph convolution are concatenated with the output of the recurrent spiking neural network, which means that the human body pose features extracted by graph convolution are concatenated with the action sequence features extracted by the recurrent network. This concatenation forms a spatio-temporal unified action feature representation, completely depicting the operation and maintenance actions. The concatenated features are called spatio-temporal features.

[0104] In S25, the spatio-temporal features are input into a classifier (such as a fully connected layer followed by a Softmax activation function). The classifier includes N fully connected layers, and the N fully connected layers respectively correspond to N types of operation and maintenance actions. The Softmax activation function is applied after these fully connected layers. The Softmax function normalizes the output into a probability distribution, obtaining the action category probability distribution at each time step, and taking the action category with the highest probability as the action recognition result at the current moment. Specifically, the output layer at each time step predicts the probability that the current action belongs to each predefined operation and maintenance action category. For example, if the output layer shows that the probability that the current action belongs to "turn on the water pump" is 0.8, then this action is recognized as "turn on the water pump". This probability-based recognition method not only provides the category of the action but also reflects the confidence of the recognition. Further, the continuous output of the action recognition results is recognized to form a real-time interpretation and semantic description of the operation and maintenance actions, which means that the system can not only recognize a single action but also integrate the action recognition results of multiple consecutive time steps to form an operation action sequence. For example, the network may continuously output actions such as "enter the computer room", "open the control cabinet", and "press the start button". These actions are concatenated to form a complete semantic description of the operation and maintenance process, and the obtained operation action sequence is: "enter the computer room" -> "open the control cabinet" -> "press the start button".

[0105] The above method uses a graph convolutional network to capture the spatial topological relationship of human key points, ensuring accurate capture and understanding of complex actions. The recurrent neural network deeply analyzes time series data, effectively captures the dynamic changes and long-term dependencies of actions, enhances the coherent understanding of the action process, and forms hierarchical action understanding with features in both the time and space dimensions, which is the key to realizing intelligent recognition of operation and maintenance actions. Through the probability distribution output by the Softmax function, the system not only identifies the action category but also evaluates the confidence of the recognition, enhancing the reliability of the result. Through the probability distribution output by the Softmax function, the system not only identifies the action category but also evaluates the confidence of the recognition, enhancing the reliability of the result. The action recognition results of consecutive multiple time steps are combined into an operation action sequence, providing a data basis for obtaining the optimal strategy in the subsequent stage.

[0106] Refer to Figure 3 Referring to the flowchart of the optimal strategy acquisition method shown, in S3, the DQN algorithm is used to process the operation action sequence to obtain the optimal strategy, including the following steps:

[0107] S31: Collect the grid frequency and pumped storage unit load at all time steps;

[0108] S32: Construct a state space, map the operation action sequence, the grid frequency and the pumped storage unit load at all time steps to the state space, and obtain the first state value, the second state value and the third state value at all time steps;

[0109] S33: Based on the first state value, the second state value and the third state value at all time steps, obtain the state variables at all time steps;

[0110] S34: Construct a value network and a target network;

[0111] S35: Input the state variables at all time steps into the value network respectively to obtain the Q values of all possible actions under the state variables at the current time step, where Q is the action value function;

[0112] S36: Select the possible action with the largest Q value as the current action, execute the current action, and obtain the new grid frequency and the new pumped storage unit load;

[0113] S37: Map the new pumped storage unit load to the state space to obtain the new third state value at the current time step;

[0114] S38: Based on the new grid frequency, the standard grid frequency, the third state value at the current time step and the new third state value at the current time step, obtain the immediate reward at the current time step;

[0115] S39: The state variables at the current time step, the current action, the immediate reward, and the state variables at the next time step form a set of experience data;

[0116] S310: Optimize the value network and the target network based on the experience data;

[0117] S311: Obtain the optimal policy based on the optimized value network and target network.

[0118] In S32, the state space includes multiple operation and maintenance actions, multiple preset frequency values, and multiple preset load values. Each operation and maintenance action corresponds to a first variable, each preset frequency value corresponds to a second variable, and each preset load value corresponds to a third variable.

[0119] Match the multiple action recognition results included in the operation action sequence with the operation and maintenance actions respectively, and record the first variable corresponding to the operation and maintenance action that matches the action recognition result successfully as the first state value. For example, the state space includes 5 operation and maintenance actions, namely "start the pumped-storage unit", "stop the pumped-storage unit", "increase the load", "decrease the load", and "no operation". Among them, the first variable corresponding to "start the pumped-storage unit" is 0, the first variable corresponding to "stop the pumped-storage unit" is 1, the first variable corresponding to "increase the load" is 2, the first variable corresponding to "decrease the load" is 3, and the first variable corresponding to "no operation" is 4. If the operation action sequence includes the action recognition results of 3 time steps, namely "enter the machine room", "open the control cabinet", and "press the start button", among which, both "enter the machine room" and "open the control cabinet" match "no operation" successfully, and the corresponding first variables are both 4, then the first state values of the first 2 time steps are 4, and "press the start button" matches "start the pumped-storage unit" successfully, and the corresponding first variable is 0, then the first state value of the 3rd time step is 0.

[0120] Among multiple preset frequency values, select the second variable corresponding to the preset frequency value closest to the grid frequency at the current time step as the second state value. If the grid frequency at the current time step is equal to one of the preset frequency values, determine that this preset frequency value is the closest to the grid frequency at the current time step; if the grid frequency at the current time step is the mid-value of two preset frequency values, then determine that the grid frequency at the current time step is the closest to the smaller one of the two preset frequency values. For example, the state space contains 11 preset frequency values. The first preset frequency value is 49.5 Hz, and the corresponding second variable is 0; the second preset frequency value is 49.6 Hz, and the corresponding second variable is 1; the third preset frequency value is 49.7 Hz, and the corresponding second variable is 2; the fourth preset frequency value is 49.8 Hz, and the corresponding second variable is 3; the fifth preset frequency value is 49.9 Hz, and the corresponding second variable is 4; the sixth preset frequency value is 50 Hz, and the corresponding second variable is 5; the seventh preset frequency value is 50.1 Hz, and the corresponding second variable is 6; the eighth preset frequency value is 50.2 Hz, and the corresponding second variable is 7; the ninth preset frequency value is 50.3 Hz, and the corresponding second variable is 8; the tenth preset frequency value is 50.4 Hz, and the corresponding second variable is 9; the eleventh preset frequency value is 50.5 Hz, and the corresponding second variable is 10. If the grid frequency at a time step is 50.04 Hz, the preset frequency value closest to it is 50 Hz, and the second variable corresponding to this preset frequency value is 5, then the second state value at this time step is 5. If the grid frequency at a time step is 50.5 Hz, the preset frequency value closest to it is 50.5 Hz, and the second variable corresponding to this preset frequency value is 10, then the second state value at this time step is 10. Obtain the second state values for all time steps in the above manner.

[0121] Among multiple preset load values, select the third variable corresponding to the preset load value closest to the load of the pumped-storage unit at the current time step as the third state value. If the load of the pumped-storage unit at the current time step is equal to one of the preset load values, it is determined that this preset load value is the closest to the load of the pumped-storage unit at the current time step; if the load of the pumped-storage unit at the current time step is the intermediate value of two preset load values, it is determined that the load of the pumped-storage unit at the current time step is the closest to the smaller one of the two preset load values. For example, the state space contains 31 preset load values. The first preset load value is 0 MW, and its corresponding third variable is 0. The second preset load value is 10 MW, and its corresponding third variable is 1,..., the 30th preset load value is 290 MW, and its corresponding third variable is 29. The 31st preset load value is 300 MW, and its corresponding third variable is 30. If the load of the pumped-storage unit at a time step is 155 MW, the preset load values closest to it are 150 MW and 160 MW, and select the third variable 15 corresponding to 150 MW as the third state value for this time step. Obtain the third state values for all time steps in the above manner.

[0122] In S33, combine the first state value, the second state value, and the third state value of each time step respectively to form the state variables of all time steps. For example, if the first state value of a time step is 2, the second state value is 5, and the third state value is 15, then the state variable of this time step is (2, 5, 15).

[0123] Through the above method, map the action recognition result, the grid frequency, and the load of the pumped-storage unit to the state space, define a clear numerical representation for each variable, and make the representation of the state space clearer and more consistent. For the two continuous variables of the grid frequency and the load of the pumped-storage unit, through discretization processing, convert them into finite state values, which helps to simplify the model and reduce the computational complexity, simplifies the subsequent decision-making process of the agent, and makes it easier to learn and optimize the strategy.

[0124] In S34, the value network is a deep neural network model, and the initial target network is a replicated value network.

[0125] In S35, set an action space containing multiple possible actions, input the state variables of all time steps into the value network respectively, and for the state variable of the current time step, determine the Q values of all possible actions in the action space. For example, the state variable of the current time step is (0, 1, 5), and one of the possible actions is to start the AGC control operation to adjust the load of the pumped-storage unit. Another example, the state variable of the current time step is (2, 10, 10), and one of the possible actions is to increase the power setting value of the pumped-storage unit and increase the load of the pumped-storage unit by 20 MW.

[0126] In S36, select the possible action with the largest Q value as the current action, execute the current action in the deployed environment. After the execution is completed, the grid frequency and the load of the pumped-storage unit may change, and obtain the new grid frequency and the new load of the pumped-storage unit.

[0127] In S37, the method principle of "mapping the new load of the pumped-storage unit to the state space and obtaining the new third state value at the current time step" is the same as the method principle of obtaining the third state value in step S32, and will not be elaborated here.

[0128] In S38, the calculation formula of the immediate reward is as follows:

[0129] R = A * R1 + B * R2 + C * R3;

[0130] Among them, R is the immediate reward; A, B, and C are preset weight values; R1 is the frequency deviation reward; R2 is the load change reward; R3 is the start-stop operation reward. Among them, A is 0.5, B is 0.3, and C is 0.2.

[0131] Furthermore, the method for obtaining the frequency deviation reward includes: calculating the absolute value of the difference between the new grid frequency and the standard grid frequency, denoted as the first absolute value; determining whether the first absolute value is greater than the first threshold; if it is greater than the first threshold, the frequency deviation reward is W1, if it is not greater than the first threshold, the frequency deviation reward is W2. Among them, W1 is -1, W2 is 0, and the first threshold is 0.2 Hz.

[0132] The method for obtaining the load change reward includes: calculating the absolute value of the difference between the third state value at the current time step and the new third state value at the current time step, denoted as the second absolute value; determining whether the second absolute value is greater than the second threshold; if it is greater than the second threshold, the load change reward is W3, if it is not greater than the second threshold, the load change reward is W4. Among them, W3 is 0, W4 is 0.5, and the first threshold is 10 MW.

[0133] The method for obtaining the start-stop operation reward includes: if in the operation and maintenance actions corresponding to the action recognition results of consecutive N time steps, there are M or more unit start-stop operation and maintenance operations, the start-stop operation reward is W5, if the number of unit start-stop operation and maintenance operations is less than M, the start-stop operation reward is W6. Among them, the unit start-stop operation and maintenance operations include "start-up of the pumped-storage unit" and "shutdown of the pumped-storage unit", W5 is -1, and W6 is 0. For example, N is 3 and M is 2. In the operation and maintenance actions corresponding to the action recognition results of consecutive 3 time steps, one is "start-up of the pumped-storage unit" and the other is "shutdown of the pumped-storage unit", which means that there are 2 unit start-stop operation and maintenance operations, and the start-stop operation reward is -1.

[0134] The above method for calculating immediate rewards comprehensively considers three different factors: power grid frequency deviation, load change of pumped-storage units, and unit start-stop operation and maintenance. By setting different weights (A, B, C) for different reward components, the importance that the agent attaches to different factors can be flexibly adjusted, thereby optimizing its strategy. Among them, by setting rewards or penalties for power grid frequency deviation, the agent is encouraged to take actions to maintain the power grid frequency within the ideal range, thus improving the stability of the power grid. The calculation method of load change rewards prompts the agent to be more cautious when adjusting the load of pumped-storage units, avoiding large fluctuations in load, and contributing to the smooth management of the power grid load. By penalizing frequent start-stop operations, the wear and tear on equipment is reduced, the service life of the equipment is extended, and at the same time, the safety and reliability of power grid operation are improved.

[0135] In S39, the state variables at the current time step, the current action, the immediate reward, and the state variables at the next time step form a set of experience data, which is stored in the experience pool. Each set of experience data is a sample.

[0136] In S310, multiple sets of experience data are randomly sampled from the experience pool, and the parameters of the value network are updated using the multiple sets of experience data and the mean squared error loss function. For example, the capacity of the experience pool is 1000 sets of experience data, and 100 sets of experience data are randomly sampled from it. The mean squared error loss function is the sum of the squares of the Q-value estimation errors of these 100 sets of experience data, specifically the sum of the squares of the differences between the predicted Q-values and the target Q-values of multiple sets of experience data. The gradient descent method is used to minimize the mean squared error loss function to obtain new parameters, and the original parameters in the value network are replaced with the new parameters to complete the update of the value network. To maintain the stability of training, every preset number of steps (for example, every 100 steps of training), the parameters of the value network are copied to the target network to complete the optimization of the target network, which ensures that the expected target of Q-value estimation is relatively stable.

[0137] In S311, by continuous iteration, the Q function is optimized to enable the agent to learn to take the best actions in different states. After multiple rounds of iteration, the strategy of the agent will continuously optimize towards the direction of high cumulative rewards, and finally the training process converges to obtain the optimal strategy.

[0138] This method using the DQN algorithm (DQN, Deep Q-Network) can learn the optimal strategy from a large amount of experience data and achieve the adaptive adjustment of the AGC function. At the same time, the DQN algorithm has good generalization ability, and the trained optimized value network and target network can be applied to different power grid load conditions to obtain accurate optimal strategies.

[0139] In S4, the obtained optimal strategy is fed back to the operation and maintenance personnel, which can guide them to perform scientific operations, improve the effect of power grid frequency regulation, and enable the operation and maintenance personnel to enrich their experience and accumulate professional knowledge during the actual operation process. Among them, the optimal strategy can be presented to the operation and maintenance personnel in an easy-to-understand manner, including action sequences, decision trees, or flowcharts, etc., to ensure that the operation and maintenance personnel can understand the logic and purpose of the strategy. A simulation environment (virtual reality (VR) system, augmented reality (AR) tool, or computer simulation program) can also be prepared, which can safely simulate real operation and maintenance operations. Start the actual operation training in the simulation environment, and the operation and maintenance personnel start to execute operations according to the guidance of the optimal strategy. The training can start from simple tasks and gradually increase in difficulty. During the training process, provide real-time feedback and guidance, which can be achieved through the prompt system in the simulation environment or by the trainer guiding beside.

[0140] The present disclosure also provides 3 specific examples:

[0141] Example 1: The system monitors the power grid frequency in real time. Once a sudden drop in the power grid frequency is detected, the operation and maintenance personnel immediately adopt an emergency response mechanism. The recurrent pulse neural network identifies the start-up operations performed by the operation and maintenance personnel, including key steps such as valve opening and unit grid connection, to ensure the correctness and timeliness of these operations. Through the DQN algorithm and based on the current state and the executed "start-up" actions, corresponding rewards are given, and the start-up strategy is optimized according to these feedbacks, including adjusting the valve opening and optimizing the speed-up curve, etc., to achieve a smoother and more efficient start-up process. Using a virtual reality (VR) environment, intuitive operation prompts are provided to the operation and maintenance personnel, such as "slowly open the water inlet valve and keep the unit speed rising smoothly", etc., to ensure that they can operate according to best practices. The operation and maintenance personnel complete the start-up of the pumped-storage unit according to the guidance of the system. During this process, the system records their actions and various operation data of the pumped-storage unit, and these data will subsequently be used for further optimization of the offline strategy to continuously improve the operation and maintenance efficiency and response quality.

[0142] Example 2: The system monitors the status of standby units to determine whether the AGC load shedding operation can be performed. If so, the recurrent pulse neural network identifies the AGC load shedding operations performed by the operation and maintenance personnel, such as power regulation and switch disconnection. The DQN algorithm evaluates the impact of the AGC load shedding operation on the power grid frequency and gives the optimal strategy such as load distribution. The VR environment is used to guide the operation and maintenance personnel to standardize operations, such as "smoothly adjust the unit power and avoid sudden power changes during load shedding", etc. The operation and maintenance personnel complete the optimal strategy according to the guidance, and the system records their actions and unit data for offline strategy optimization.

[0143] Example 3: When the system detects a unit failure, such as over-limit bearing temperature, abnormal noise of the water pump, etc., it is determined that maintenance personnel need to perform emergency operations. The recurrent pulse neural network identifies the emergency responses performed by the maintenance personnel, such as emergency shutdown, fault isolation, etc. The impact of the fault on the power grid is evaluated through the DQN algorithm, and the emergency strategy is optimized, such as removing the faulty equipment, starting the standby unit, etc. The VR environment is used to guide the maintenance personnel to perform standardized responses, such as "closing the outlet valve of the faulty water pump and putting into the standby water pump", etc. The maintenance personnel complete the emergency response according to the guidance, and the system records the process data for post-event analysis and prevention.

[0144] Based on the above, the present invention proposes an innovative brain-inspired training mechanism that integrates the action recognition ability of the recurrent pulse neural network and the optimization of the reinforcement learning strategy, aiming to significantly improve the operation and maintenance training effect of automatic generation control (AGC). By combining advanced artificial intelligence technology and virtual reality (VR) environment, an immersive training experience is created for maintenance personnel, which not only accelerates the learning and mastery of operation and maintenance skills, but also promotes the effective transfer of these skills in actual operations.

[0145] In addition, the training mechanism is adaptive and can continuously adjust and optimize the training strategy according to the operation habits of maintenance personnel and the real-time dynamics of equipment. This adaptive optimization ensures that the training content always meets the actual needs and improves the pertinence and practicality of training.

[0146] In summary, the practical training method for the AGC system of the pumped storage power station provided by the present disclosure not only represents the forefront development direction of intelligent operation and maintenance technology, but also has far-reaching significance for improving the safe and stable operation of the power system. Through intelligent and immersive training methods, more professional and efficient operation and maintenance talents can be cultivated for the power industry, thus providing solid technical support and talent guarantee for the reliable operation of the power grid.

[0147] Referring to Figure 4 , the present disclosure provides a practical training system for the AGC system of a pumped storage power station, including:

[0148] An acquisition module 101, configured to collect operation data of maintenance personnel at a preset acquisition frequency during the practical training of the AGC system of the pumped storage power station;

[0149] An obtaining module 102, configured to obtain an operation action sequence based on the operation action data and the recurrent pulse neural network;

[0150] A processing module 103, configured to process the operation action sequence by using the DQN algorithm to obtain an optimal strategy;

[0151] A guiding module 104, configured to guide the maintenance personnel to complete the operation process of the practical training based on the optimal strategy;

[0152] Among them, the operation data is the continuous action data of the preset body parts of the operation and maintenance personnel.

[0153] The various change methods and specific examples in the above-provided practical training method of the AGC system of the pumped-storage power station are equally applicable to the practical training system of the AGC system of the pumped-storage power station provided in this disclosure. Through the foregoing detailed description of the practical training method of the AGC system of the pumped-storage power station, those skilled in the art can clearly know the implementation method of the practical training system of the AGC system of the pumped-storage power station. For the sake of brevity of the specification, it will not be elaborated here.

[0154] The computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0155] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the practical training method of the AGC system of the pumped-storage power station in the foregoing embodiments of the present disclosure.

[0156] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.

[0157] Such as Figure 5 It is a schematic structural diagram of a computer device provided for an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiments of the present disclosure. Figure 5 The shown computer device is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present disclosure.

[0158] Such as Figure 5As shown, a computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0159] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device, etc.; an output device including, for example, a display screen, etc.; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or wireline to exchange data. Although Figure 5 a computer device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0160] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from a storage device, or installed from the ROM. When the computer program is executed by the processor, all or part of the steps of the practical training method of the pumped storage power station AGC system according to the embodiments of the present disclosure are executed.

[0161] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0162] According to a computer-readable storage medium of an embodiment of the present disclosure, non-temporary computer-readable instructions are stored thereon. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the practical training method of the pumped storage power station AGC system according to the foregoing embodiments of the present disclosure are executed.

[0163] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as: CD-ROM and DVD), magneto-optical storage media (such as: MO), magnetic storage media (such as: magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (such as: memory card) and media with built-in ROM (such as: ROM cartridge).

[0164] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0165] The basic principles of the present disclosure have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0166] In the present disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0167] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.

[0168] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0169] Various changes, substitutions, and alterations to the technology described herein can be made without departing from the teachings defined by the appended claims. Additionally, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that are currently available or later to be developed that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0170] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0171] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A practical training method for an AGC system of a pumped storage power station, characterized in that: include: In the practical training of the AGC system of the pumped storage power station, the operation data of the operation and maintenance personnel are collected according to the preset collection frequency; Based on the operation action data and the recurrent spiking neural network, obtaining an operation action sequence; The DQN algorithm is used to process the operation action sequence to obtain the optimal strategy; Guide operation and maintenance personnel to complete the operation process of practical training based on the optimal strategy; The operation data is continuous motion data of preset body parts of the operation and maintenance personnel.

2. The practical training method for the AGC system of a pumped storage power station according to claim 1, characterized in that: The step of acquiring an operation action sequence based on the operation action data and the recurrent spiking neural network includes: Performing normalization processing and pulse coding processing on the operation data to obtain a pulse sequence composed of a plurality of pulse combinations; Inputting the pulse sequence into the recurrent spiking neural network, the recurrent spiking neural network includes recurrent units of multiple time steps, each time step corresponds to a pulse combination, inputting the pulse combination and the hidden state of the previous time step into the corresponding recurrent unit, and obtaining the hidden state of the corresponding time step; Constructing a human skeleton topology map, performing a graph convolution operation on the human skeleton topology map, and obtaining spatial features; Concatenate the hidden states of all time steps and the spatial features to obtain spatiotemporal features; Based on the spatiotemporal features and the classifier, the action recognition results of all time steps are obtained, and the action recognition results of all time steps are combined into the operation action sequence.

3. The practical training method for the AGC system of a pumped storage power station according to claim 2 is characterized in that: The normalization and pulse coding processing are performed on the operation data to obtain a pulse sequence composed of a plurality of pulse combinations, including: Normalizing the multiple action data included in the operation data to convert the multiple action data into multiple standard data; Traversing the plurality of standard data, and determining respectively whether each value in the standard data is greater than a preset data threshold; If yes, modify the value to 1; If not, modify the value to 0; Combining the modified values ​​into a pulse combination; The pulse combination converted from the plurality of standard data constitutes a pulse sequence.

4. The practical training method for the AGC system of a pumped storage power station according to claim 2 is characterized in that: The adopting of the DQN algorithm to process the operation action sequence to obtain the optimal strategy includes: Collect the grid frequency and pumped storage unit load at all time steps; Constructing a state space, mapping the operation action sequence, the grid frequency of all time steps and the pumped storage unit load to the state space, and obtaining a first state value, a second state value and a third state value of all time steps; Based on the first state value, the second state value and the third state value of all the time steps, obtaining the state variables of all the time steps; Build value networks and target networks; Inputting the state variables of all time steps into the value network respectively, obtaining the Q values ​​of all possible actions under the state variables of the current time step, where Q is the action value function; Selecting the possible action with the largest Q value as the current action, executing the current action, and obtaining a new grid frequency and a new pumped storage unit load; Map the new pumped storage unit load to the state space to obtain the new third state value of the current time step; Based on the new grid frequency, the standard grid frequency, the third state value of the current time step and the new third state value of the current time step, obtaining an instant reward for the current time step; The state variable of the current time step, the current action, the immediate reward, and the state variable of the next time step constitute a set of experience data; Optimizing the value network and the target network based on the empirical data; Obtain the optimal strategy based on the optimized value network and target network.

5. The practical training method for the AGC system of a pumped storage power station according to claim 4 is characterized in that: The step of mapping the operation action sequence, the grid frequency of all time steps and the pumped storage unit load to the state space to obtain the first state value, the second state value and the third state value of all time steps includes: The state space includes multiple operation and maintenance actions, multiple preset frequency values, and multiple preset load values, each operation and maintenance action corresponds to a first variable, each preset frequency value corresponds to a second variable, and each preset load value corresponds to a third variable; Matching the multiple action recognition results included in the operation action sequence with the operation and maintenance actions respectively, and recording the first variable corresponding to the operation and maintenance action successfully matched with the action recognition result as the first state value; Among the multiple preset frequency values, select the second variable corresponding to the preset frequency value closest to the power grid frequency at the current time step as the second state value; Among the multiple preset load values, the third variable corresponding to the preset load value closest to the load of the pumped-storage unit at the current time step is selected as the third state value.

6. The practical training method for the AGC system of a pumped storage power station according to claim 4, characterized in that: The calculation formula of the instant reward is as follows: R = A*R1+B*R2+C*R3; Among them, R is the instant reward; A, B, and C are the preset weight values; R1 is the frequency deviation reward; R2 is the load change reward; and R3 is the start-stop operation reward.

7. The practical training method for the AGC system of a pumped storage power station according to claim 6, characterized in that: Also includes: Calculating an absolute value of a difference between the new grid frequency and the standard grid frequency, and recording the difference as a first absolute value; Determining whether the first absolute value is greater than a first threshold; If it is greater than the first threshold, the frequency deviation reward is W1; If it is not greater than the first threshold, the frequency deviation reward is W2.

8. The practical training method for the AGC system of a pumped storage power station according to claim 6, characterized in that: Also includes: Calculate the absolute value of the difference between the third state value of the current time step and the new third state value of the current time step, and record it as the second absolute value; Determining whether the second absolute value is greater than a second threshold; If it is greater than the second threshold, the load change reward is W3; If it is not greater than the second threshold, the load change reward is W4.

9. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the practical training method for the AGC system of a pumped-storage power station as described in any one of claims 1-8.

10. A computer program product comprising computer instructions, characterized in that: When the computer instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.