Adaptive concentrator configuration method and system based on reinforcement learning
Through the adaptive concentrator configuration method based on reinforcement learning, the configuration parameters of the meter terminal are dynamically adjusted, which solves the problem that traditional configuration methods are difficult to respond in a timely manner when facing complex power grid environments, realizes intelligent and automated optimization of the system, and improves resource utilization and fault response capabilities.
Patent Information
- Application Number
- CN202510312763.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional concentrator configuration methods are difficult to achieve timely responses in the face of grid load fluctuations, changes in data acquisition requirements and unstable network connection conditions, resulting in data acquisition delays, network congestion and increased system energy consumption.
Adaptive concentrator configuration method based on reinforcement learning is adopted, through the policy network of the reinforcement learning module, the initial configuration action is generated according to the priority category of the meter terminal and its current state, and the configuration results are evaluated through the reward function, optimized configuration action is generated, and the configuration parameters of the relevant meter terminal are adjusted.
Intelligent and automated system optimization has been achieved, the need for manual intervention has been reduced, the system flexibility, resource utilization and fault response capabilities have been improved, and the configuration accuracy and consistency are ensured.
Smart Images

Figure CN120218532A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an adaptive concentrator configuration method and system based on reinforcement learning, belonging to the field of power technology. Background Art
[0002] A concentrator is a core device in the smart grid, undertaking important functions such as data collection, transmission, and system control. The configuration method of traditional concentrators uses static preset parameters. In the face of complex conditions such as power grid load fluctuations, changes in data collection requirements, and unstable network connection conditions, it is often difficult to respond in a timely manner, resulting in problems such as data collection delay, network congestion, and increased system energy consumption, affecting the overall performance and stability of the system.
[0003] Therefore, how to achieve the adaptive configuration of the concentrator and dynamically optimize the configuration according to the actual network conditions and collection requirements is an urgent problem to be solved. The current concentrator configuration methods mostly rely on dynamic adjustment strategies based on empirical rules. However, these strategies are usually limited to simple rule inferences, lacking the mining of deep-level relationships in data and unable to achieve refined and real-time dynamic optimization. In addition, in the face of complex and changeable power grid environments, these methods are difficult to effectively meet the dynamic requirements in multiple scenarios. On the other hand, existing methods often lack overall consideration in resource allocation and energy consumption control and are difficult to achieve an effective balance between system performance and energy consumption. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an adaptive concentrator configuration method and system based on reinforcement learning. Through the adaptive configuration method of the reinforcement learning module, it can intelligently optimize the configuration of the meter terminal, improving the system flexibility, resource utilization rate, and fault response ability.
[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0006] In the first aspect, the present invention provides an adaptive concentrator configuration method based on reinforcement learning, including:
[0007] The concentrator obtains a configuration instruction through the master control center, extracts the identity information of the meter terminal to be configured and the historical performance of its performance indicators, and determines the priority category of the meter terminal according to the identity information of the meter terminal and the historical performance of the performance indicators;
[0008] Based on the policy network of the reinforcement learning module, an initial configuration action is generated according to the priority category of the meter terminal and its current state, and the parameters corresponding to the configuration action are applied to the concentrator and the meter terminal;
[0009] After the configuration parameters are applied, the configuration results are evaluated based on the historical performance and real-time data of the electricity meter terminal performance indicators. The performance of the configuration actions is calculated through the reward function. If the evaluation results do not meet the expectations, the reinforcement learning module is used to generate optimized configuration actions to adjust the configuration parameters of the relevant electricity meter terminals;
[0010] The reinforcement learning module is trained by combining historical data and real-time feedback, continuously updating the policy network to optimize the generation of future configuration actions, generating high-priority actions for adjustment when the performance of key nodes is abnormal, and triggering an alarm and reporting the abnormal information at the same time.
[0011] Furthermore, the concentrator continuously monitors the identity information and current status of the electricity meter terminals. When the identity information or the current status of the electricity meter terminals changes, the priority category of the electricity meter terminals is dynamically adjusted.
[0012] Furthermore, the adaptive concentrator configuration method based on reinforcement learning further includes:
[0013] Configure a monitoring thread for the concentrator to monitor the configuration instructions sent by the master control center in real time;
[0014] When it is detected that the master control center sends a new configuration instruction to the concentrator, this instruction is preferentially responded to and the current configuration status of the concentrator is obtained;
[0015] When the concentrator executes the current configuration instruction, the current configuration instruction is paused and a backup mechanism is started, and the status, unfinished operations, and cached data of the current configuration instruction are saved to the backup storage area;
[0016] The backup mechanism is used to perform temporary emergency configuration on the electricity meter terminals corresponding to the new configuration instruction. After completion, the paused current configuration instruction is resumed, and configuration consistency is ensured.
[0017] Furthermore, adjusting the configuration parameters of the relevant electricity meter terminals includes:
[0018] The configuration parameters of the electricity meter terminals are refined and adjusted multiple times, the optimal configuration parameter set is determined under different network conditions, and through the terminal identity binding mechanism, the unique identifier of each electricity meter terminal is associated with its configuration parameters to ensure the consistency of the configuration adjusted each time with the corresponding electricity meter terminal.
[0019] Furthermore, the reinforcement learning module defines a state space, an action space, and a reward function. The state space includes network environment parameters, concentrator performance indicators, and the operating status of the electricity meter terminals. The action space includes configuration parameter adjustment options, and the reward function is defined according to the effect of the configuration actions.
[0020] Further, the configuration result is evaluated based on the historical performance and real-time data of the performance indicators of the electricity meter terminal, and the performance of the configuration action is calculated through a reward function, including:
[0021] Determine the historical baseline of each performance indicator according to the historical performance of the performance indicators of the electricity meter terminal to be configured, obtain the real-time data of the performance indicators of the electricity meter terminal after applying the configuration parameters, and calculate the comprehensive performance indicator by weighted averaging the real-time data of the performance indicators of the electricity meter terminal; if the comprehensive performance indicator of the electricity meter terminal is greater than the historical baseline, give a positive reward; otherwise, give a negative reward.
[0022] Further, when the performance of a key node is abnormal, generate high-priority actions for adjustment, and at the same time trigger an alarm and report the abnormal information, including:
[0023] When it is detected that the performance of a key node is abnormal, automatically issue an alarm and report the abnormal information to the master station center for timely countermeasures; at the same time, give priority to ensuring the resource allocation of the key node.
[0024] Further, the reinforcement learning module is constructed based on a deep Q-network and is used to generate future optimized configuration actions. In the training stage, the reinforcement learning module takes historical configuration data as input, through data cleaning, normalization, and standardization processing, combines the time sliding window algorithm to construct time series features, and generates a standardized training data set to train the reinforcement learning module to learn the optimal configuration strategy;
[0025] In the usage stage, taking the network environment parameters, concentrator performance indicators, and operating status of the electricity meter terminal as input, generate optimized configuration actions and apply them to the concentrator and the electricity meter terminal, and at the same time, through - The greedy strategy balances exploration and exploitation, enabling the reinforcement learning module to dynamically adjust the strategy during execution; where represents the probability of exploration;
[0026] In the optimization stage, continuously adjust the policy network in combination with historical data and real-time feedback, so that the reinforcement learning module adapts to the changing network environment, concentrator performance, and fluctuations in the status of the electricity meter terminal, ensuring the dynamic optimization of future configuration actions.
[0027] In a second aspect, the present invention provides an adaptive concentrator configuration system based on reinforcement learning, characterized by including:
[0028] The configuration management module is used to receive and parse the configuration instructions from the master control center, extract the identity information of the electricity meter terminals to be configured and the historical performance of their performance indicators, and determine the priority category of the electricity meter terminals according to the identity information of the electricity meter terminals and the historical performance of the performance indicators; Based on the policy network of the reinforcement learning module, generate an initial configuration action according to the priority category of the electricity meter terminal and its current state, and apply the parameters corresponding to the configuration action to the concentrator and the electricity meter terminal;
[0029] The parameter optimization module is used to evaluate the configuration result based on the historical performance and real-time data of the performance indicators of the electricity meter terminal, calculate the performance of the configuration action through the reward function, and if the evaluation result does not meet the expectation, use the reinforcement learning module to generate an optimized configuration action to adjust the configuration parameters of the relevant electricity meter terminal;
[0030] The reinforcement learning module is used to train the reinforcement learning module by combining historical data and real-time feedback, continuously update the policy network to optimize the generation of future configuration actions, generate high-priority actions for adjustment when the performance of key nodes is abnormal, and at the same time trigger an alarm and report the abnormal information;
[0031] Among them, the configuration management module, the parameter optimization module and the reinforcement learning module are connected through a data interaction interface to form a closed-loop control system.
[0032] Furthermore, the configuration management module also includes a monitoring thread unit, which is used to monitor the configuration instructions sent by the master control center in real time. When it detects that the master control center sends a new configuration instruction to the concentrator, it gives priority to responding to this instruction and obtains the current configuration state of the concentrator;
[0033] When the concentrator executes the current configuration instruction, pause the current configuration instruction and start the backup mechanism, and save the state, unfinished operations and cached data of the current configuration instruction to the backup storage area;
[0034] Use the backup mechanism to perform temporary emergency configuration on the electricity meter terminals corresponding to the new configuration instruction. After completion, resume the paused current configuration instruction and ensure configuration consistency.
[0035] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0036] The adaptive concentrator configuration method based on reinforcement learning of the present invention can achieve intelligent and automated system optimization by dynamically adjusting the configuration parameters of the electricity meter terminals, reducing the need for manual intervention. By dynamically adjusting configuration parameters such as the priority and real-time status of the electricity meter terminals, the system can flexibly respond to different network conditions and service requirements, ensuring the accuracy and consistency of the configuration. The priority classification and differential configuration mechanism improve the system resource utilization rate and ensure the continuity of critical services; in addition, the continuous training of the reinforcement learning module enables the system to continuously optimize the configuration strategy, improve the resource utilization rate of the electricity meter terminals, and enhance the system's response ability when abnormalities occur at critical nodes, improving the overall performance and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of a concentrator optimization configuration method based on reinforcement learning provided in Embodiment 1 of the present invention;
[0038] Figure 2 It is a module diagram of a concentrator optimization configuration system based on reinforcement learning provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.
[0040] Embodiment 1:
[0041] Figure 1 It is a flowchart of an adaptive concentrator configuration method based on reinforcement learning in Embodiment 1 of the present invention. This flowchart only shows the logical sequence of the method described in this embodiment. On the premise of not conflicting with each other, in other possible embodiments of the present invention, the steps shown or described can be completed in a different Figure 1 sequence as shown. Refer to Figure 1 , the method of this embodiment specifically includes the following steps:
[0042] The concentrator obtains configuration instructions through the master control center, extracts the identity information of the electricity meter terminals to be configured and the historical performance of their performance indicators, and determines the priority category of the electricity meter terminals according to the identity information of the electricity meter terminals and the historical performance of the performance indicators;
[0043] Based on the policy network of the reinforcement learning module, an initial configuration action is generated according to the priority category of the electricity meter terminals and their current state, and the parameters corresponding to the configuration action are applied to the concentrator and the electricity meter terminals;
[0044] After the configuration parameters are applied, evaluate the configuration results based on the historical performance and real-time data of the performance indicators of the meter terminals. Calculate the performance of the configuration actions through the reward function. If the evaluation results do not meet the expectations, use the reinforcement learning module to generate optimized configuration actions and adjust the configuration parameters of the relevant meter terminals;
[0045] Train the reinforcement learning module by combining historical data and real-time feedback, continuously update the policy network to optimize the generation of future configuration actions, generate high-priority actions for adjustment when the performance of key nodes is abnormal, and trigger an alarm and report the abnormal information at the same time.
[0046] Specifically, the concentrator configuration process mainly includes obtaining configuration instructions from the master control center and determining the set of meter terminals to be configured. Extract the identity information of the meter terminals, and determine the priority of the meter terminals according to the identity information of the meter terminals and the historical performance of the performance indicators, and classify them into high, medium, and low priority groups according to the priority.
[0047] First, the concentrator obtains configuration instructions from the master control center and determines the set of meter terminals to be configured. The master control center, as the core management platform of the power system, is responsible for issuing configuration instructions, monitoring the operation status of the system, and coordinating the communication between devices. The concentrator establishes a secure and reliable communication connection with the master control center (for example, a TCP / IP connection encrypted by SSL / TLS) and receives the configuration instructions. The concentrator parses the configuration instructions and extracts the identity information of each meter terminal. The identity information includes the meter terminal device ID (the code used to uniquely identify each meter terminal), the business type (such as residential electricity, industrial electricity, commercial electricity, etc.), the geographical location (the installation location of the meter terminal device, such as the area, street, floor, etc.), and the historical electricity consumption data (including historical electricity consumption, electricity consumption mode, etc.).
[0048] According to the identity information of the meter terminals, determine the business importance of the meter terminals and obtain the historical performance of the performance indicators of the meter terminals. After comprehensively evaluating the business importance and historical performance of the performance indicators of the meter terminals, determine the corresponding priority category of the meter terminals. The evaluation process is divided into two steps:
[0049] First, evaluate the business importance of the electricity meter terminals, and classify the electricity meter terminals into high-priority (such as hospitals, fire departments, important government agencies, etc., with stable power supply requirements and high safety requirements), medium-priority (such as large shopping malls, factories, office buildings, etc., with relatively high requirements for power supply quality), and low-priority (such as ordinary residential users, small shops, etc.). Secondly, analyze the historical performance of the performance indicators of the electricity meter terminals, including the communication success rate (the communication stability between the electricity meter terminal and the concentrator), the response time (the response speed of the electricity meter terminal to the configuration instructions), and the fault records (whether there have been frequent faults or abnormalities in history). By synthesizing the business importance of the electricity meter terminals and the historical performance of the performance indicators, and through the determined quantitative scoring criteria, assign a comprehensive score to each electricity meter terminal device, and classify it into three priority groups: high, medium, and low, to achieve the precise classification of the electricity meter terminals.
[0050] For electricity meter terminal devices with different priorities, based on the policy network of the reinforcement learning module, generate initial configuration actions according to the priority and current state of the electricity meter terminal, and apply them to the concentrator and the electricity meter terminal. The specific configuration strategy is as follows:
[0051] Set a high communication frequency (such as collecting data once a minute), preferentially allocate a larger bandwidth, and a larger data cache size for high-priority electricity meter terminals to ensure smooth communication and data integrity. Set a medium communication frequency (such as collecting data once every 5 minutes), allocate an appropriate bandwidth, and a medium cache capacity for medium-priority electricity meter terminals. Set a lower communication frequency (such as collecting data once every 15 minutes or longer), allocate a smaller bandwidth, and a basic cache capacity for low-priority electricity meter terminals. The concentrator adjusts its own operating parameters simultaneously according to the configuration requirements of electricity meter terminals with different priorities to adapt to the communication and data processing requirements of various electricity meter terminal devices, such as setting different communication queues, optimizing the priority of data processing threads, etc.
[0052] The concentrator continuously monitors the current state of the electricity meter terminals, including the power consumption load, communication status, device operating conditions, etc., combines the historical power consumption data and communication records of the electricity meter terminals, analyzes the behavior patterns of the electricity meter terminals, and predicts their future power consumption trends and communication requirements. When the business importance or the current state of the electricity meter terminal changes, dynamically adjust its priority. For example, when the power consumption load of a medium-priority industrial user suddenly increases due to an increase in production tasks, it can be temporarily upgraded to high-priority to ensure the timeliness and accuracy of data collection. The concentrator optimizes the allocation of communication resources and computing resources using resource scheduling algorithms (such as weighted round-robin, shortest remaining time first, etc.) according to the priority of the electricity meter terminals. Requests from high-priority electricity meter terminals will be processed first, and requests from low-priority electricity meter terminals will be processed when resources are idle. Through the dynamic scheduling of communication links and processor resources, avoid overloading of a single resource and improve the overall performance and stability of the system.
[0053] This embodiment preferentially guarantees the communication and data acquisition requirements of high-priority meter terminals, ensures the real-time and accuracy of key business data, and meets the requirements of the power system for power supply guarantee of key meter terminals. By classifying the priority of meter terminals and making differential configurations, the system resources are reasonably allocated, avoiding waste and inefficient use of resources, and improving the overall resource utilization rate of the system. The system can dynamically adjust the priority and configuration parameters of meter terminals according to the changes of real-time data and business requirements, adapting to the changing power grid environment and user needs. The automated acquisition of configuration instructions and priority classification reduces manual intervention, improves operation and maintenance efficiency, and reduces the operating cost of the system.
[0054] In an embodiment of the present invention, an emergency mechanism is also set up, which can respond in a timely manner to the configuration instructions sent by the master control center. By configuring a monitoring thread for the concentrator, the configuration instructions sent by the master control center are monitored in real time; when it is detected that the master control center sends a new configuration instruction to the concentrator, this instruction is preferentially responded to and the current configuration state of the concentrator is obtained; when the concentrator executes the current configuration instruction, the current configuration instruction is suspended and the backup mechanism is started.
[0055] When it is detected that the master control center sends a new configuration instruction, first, the execution of the current configuration instruction is safely suspended to ensure data consistency and task integrity. Then, the current configuration state and network environment information are obtained, and the current operating state, network load, meter terminal device state, etc. are recorded. Next, the backup mechanism is started, and the information such as the state of the current configuration instruction, unfinished operations, and cached data is saved to the backup storage area. According to the emergency configuration instruction issued by the master control center, the concentrator makes configuration adjustments to the specified meter terminal, including modifying communication parameters, updating security policies, adjusting data acquisition policies, etc. After the emergency configuration is completed, it is confirmed that the new configuration parameters have taken effect and the relevant meter terminal devices are operating normally. Finally, the backup configuration task is restored, and the unfinished configuration instruction is continued to be executed to ensure the continuity of the configuration process.
[0056] This embodiment utilizes the multi-threading technology and synchronization mechanism of the operating system to ensure the efficient operation of configuration instructions and monitoring threads. Through the backup and recovery mechanism, non-volatile memory is used to save backup data to prevent data loss caused by unexpected power outages. During the backup and recovery process, checksum or hash function is used to verify data integrity. The emergency configuration instruction is set to the highest priority to ensure that it is processed in the first place. A certain amount of system resources are reserved to cope with sudden emergency configuration requirements.
[0057] In this embodiment, through the monitoring mechanism, the system can quickly respond to the configuration instructions of the master control center in case of emergency, ensuring the continuity and reliability of critical services. Through the backup and recovery mechanism, it is ensured that the configuration instructions can continue to be executed completely after interruption, avoiding task failure or data loss. The real-time monitoring and emergency handling mechanism improves the stability of the system in the face of emergencies and reduces the risk of failures. The automated emergency handling process reduces the time of manual intervention and the probability of errors, improving the operation and maintenance efficiency.
[0058] In one embodiment of the present invention, the configuration parameters of the electricity meter terminal are refined and adjusted multiple times. The optimal configuration is determined under different network conditions, and through terminal identity binding, the consistency between the configuration adjusted each time and the corresponding electricity meter terminal is ensured.
[0059] For each electricity meter terminal, multiple rounds of fine-tuning are performed according to different network environments until the configuration parameters achieve the optimal effect under the current network environment. The goal of the refinement adjustment is to find the best configuration for each electricity meter terminal under the current network conditions, enabling the system to operate efficiently and stably. Each time the configuration is adjusted, the identity information of each electricity meter terminal and the corresponding configuration parameters are recorded to ensure the consistency and accuracy of the configuration of each electricity meter terminal.
[0060] During the configuration process, through repeated testing and modification of the configuration parameters, the best effect can be achieved under different network environments (such as different bandwidths, latencies, signal strengths, etc.). For example, when the network load is large, it may be necessary to reduce the communication frequency of some electricity meter terminals or adjust the bandwidth allocation to ensure that the overall operation of the system is not affected. When there is network congestion, it may be necessary to reduce the data acquisition frequency or adjust the bandwidth allocation to avoid system overload. When the network condition is good, the transmission frequency of the electricity meter terminal can be increased to reduce latency and ensure the real-time nature of data. Under specific network conditions, the most suitable parameters are configured for each electricity meter terminal to maximize communication stability, data transmission efficiency, and resource utilization. Through this refinement adjustment, it can be ensured that the configuration parameters of each electricity meter terminal are as well-matched as possible with its environment and requirements.
[0061] During the configuration adjustment process, different adjustments may be applied to multiple configuration schemes of the same electricity meter terminal (for example: configuring parameter A under low bandwidth conditions and parameter B under high bandwidth conditions). A connection is established between the configuration parameters and the identity of the electricity meter terminal to ensure that the configuration of each electricity meter terminal is consistent and accurate after each adjustment. The identity of the electricity meter terminal refers to the unique identifier of each electricity meter terminal. During the entire configuration process, the corresponding relationship between the configuration parameters of each electricity meter terminal and the terminal identity is recorded to ensure that each electricity meter terminal is correctly configured under different network conditions.
[0062] This association relationship ensures that even if there are multiple parameter changes during the adjustment process, the final configuration of each electricity meter terminal can accurately match its requirements and network conditions, avoiding confusion or incorrect configuration.
[0063] Based on the policy network of the reinforcement learning module, an initial configuration action is generated according to the priority category of the electricity meter terminal and its current state, and the parameters corresponding to the configuration action are applied to the concentrator and the electricity meter terminal; after the configuration parameters are applied, the concentrator collects the real-time data of each electricity meter terminal, including communication success rate, data transmission delay, terminal response time, and energy consumption. According to the collected data, a performance report of the initial configuration action is generated. The concentrator evaluates the initial configuration action in real time through the reinforcement learning module and executes necessary traffic management policies. The traffic management policy monitors and analyzes the communication traffic requirements of different electricity meter terminals, and performs dynamic network traffic management based on real-time feedback to ensure that high-priority services obtain the required resources.
[0064] The concentrator first obtains the state data of the current network through sensors. These data include indicators such as communication traffic, bandwidth utilization rate, and node response time. The reinforcement learning module is used to perform real-time analysis on these data, and a gated recurrent unit is used to process historical data and current feedback to predict future traffic requirements.
[0065] When it is detected that the traffic requirements of some nodes increase sharply, for example, when an electricity meter terminal needs to upload a large amount of data due to an emergency, the reinforcement learning module will immediately adjust the bandwidth allocation. The concentrator gives priority to processing the traffic requirements of high-priority electricity meter terminals and increases the corresponding bandwidth quota. At the same time, the reinforcement learning module will identify low-priority tasks (such as routine daily data collection tasks) and reduce their bandwidth occupancy to ensure that emergency data can be transmitted in a timely manner.
[0066] The reinforcement learning module also has the ability to learn the specific business characteristics. When an electricity meter terminal has high traffic upload requirements multiple times within a specific time period, the reinforcement learning module will reserve more bandwidth for the electricity meter terminal in the same time period in the future to prevent network congestion during high load. Through the analysis of historical data and the evaluation of real-time data, a deep understanding of the requirements of each electricity meter terminal is formed, so as to respond quickly in case of emergencies.
[0067] In this embodiment, through the real-time evaluation of the current configuration by the reinforcement learning module, the refined management and dynamic adjustment of network traffic are realized, ensuring that high-priority terminals obtain sufficient bandwidth in case of emergencies. By reducing the bandwidth occupancy in low-priority tasks, the concentrator improves the response speed to critical services and the resource utilization efficiency of the system. Real-time traffic management reduces network congestion, improves the reliability of communication and the timeliness of data transmission, and effectively enhances the user experience.
[0068] After executing the configuration instructions, evaluate the configuration result based on the historical performance and real-time data of the performance indicators of the electricity meter terminal, and calculate the performance of the configuration action through a reward function, including:
[0069] Determine the historical baseline of each performance indicator according to the historical performance of the performance indicators of the electricity meter terminal to be configured, obtain the real-time data of the performance indicators of the electricity meter terminal after applying the configuration parameters. If the real-time data of the performance indicators of the electricity meter terminal is greater than the historical baseline, give a positive reward; otherwise, give a negative reward.
[0070] Specifically, compare the real-time data of the performance indicators of the electricity meter terminal with the historical baseline to evaluate the impact of the configuration action on the system performance. Collect the operating status and performance indicators of each electricity meter terminal in real time. The performance indicators include communication success rate, data transmission delay, terminal response time, energy consumption level, etc. Calculate the comprehensive performance indicator by weighted average of the real-time data of the performance indicators of the electricity meter terminal. Then, extract the performance indicators during the historical operation period from the database and establish a historical baseline, including the average performance indicators under different time periods and different loads. Compare the comprehensive performance indicator with the historical baseline data, and calculate the difference value and change trend. Set a performance threshold, and when the change of the performance indicator exceeds the set threshold, it is determined as abnormal.
[0071] Through comparative analysis, locate the possible reasons for the performance degradation, which may involve adjustment of configuration parameters, change of network environment, or failure of the electricity meter terminal device, etc. For the identified problems, adjust the configuration parameters of the relevant electricity meter terminals, such as adjusting the communication frequency, optimizing the bandwidth allocation, modifying the data caching strategy, etc. Use professional data analysis tools to perform statistical analysis, trend prediction, and anomaly detection on the performance data. Feed back the result of the performance evaluation to the reinforcement learning module to form a closed-loop control system, generate optimized configuration actions, and automatically adjust the configuration parameters. Through continuous performance evaluation and configuration adjustment, the system can timely discover and solve performance problems, and ensure the communication quality and the reliability of data collection.
[0072] This embodiment can improve the system performance, ensure the communication quality and the reliability of data collection through continuous performance evaluation and configuration adjustment. The system can dynamically adjust the configuration strategy according to the real-time performance data to adapt to the changes of the network environment and the status of the electricity meter terminal device. The automated performance evaluation and adjustment mechanism reduces manual intervention, improves the operation and maintenance efficiency, and reduces the work intensity of the operation and maintenance personnel. By strictly monitoring and adjusting the performance indicators, the service quality provided to users is ensured, and the user satisfaction is improved.
[0073] The reinforcement learning module adopted in this embodiment is constructed based on the Deep Q-Network (DQN) algorithm, and a deep neural network is used as an approximator of the value function. The network structure of this module includes an input layer, multiple hidden layers, and an output layer. The specific steps for constructing the reinforcement learning module are as follows:
[0074] (1) Data preprocessing and partitioning: Collect historical configuration data of the data center, including communication interaction data, concentrator performance metrics, network connection quality parameters, etc. Preprocess the collected data, including data cleaning, normalization, and standardization; the standardization process is to process the historical configuration data using the time sliding window algorithm to generate a standardized data set. The standardized data set is represented in the form of a three-dimensional tensor, where the first dimension is the number of samples, the second dimension is the time window size, and the third dimension is the number of features, including features such as communication interaction, performance metrics, and network quality, so as to ensure data quality. Divide the preprocessed data into a training set, a validation set, and a test set for the training and evaluation of the reinforcement learning module.
[0075] (2) Construct an input layer, a Gated Recurrent Unit (GRU) layer, a regularization layer, a fully connected layer, and an output layer;
[0076] Input layer: Input the standardized data set processed by the time sliding window into the module.
[0077] Gated Recurrent Unit (GRU) layer: Construct a multi-layer Gated Recurrent Unit (GRU) network for processing time series data. The GRU layer can effectively capture the time dependence and sequence features of the data.
[0078] Regularization layer: Add a regularization layer (such as a Dropout layer) between the GRU layers to prevent overfitting and improve the generalization ability of the module.
[0079] Fully connected layer: Pass the output of the GRU layer to the fully connected layer to further process the extracted features.
[0080] Output layer: The output of the fully connected layer is used as the action value (Q-value) sequence of the reinforcement learning module to output the value evaluation of the configuration action.
[0081] (3) Define the reinforcement learning components: state space, action space, and reward function.
[0082] State space (S): Includes current network environment parameters (bandwidth, latency, packet loss rate), concentrator performance metrics (CPU usage, memory occupancy), and the operating status of the meter terminal (online status, data transmission rate).
[0083] Action space (A): Includes adjustment options for configuration parameters, such as adjusting the communication frequency, modifying the bandwidth allocation, changing the data acquisition strategy, etc.
[0084] Reward function (R): Design the reward function based on the changes in system performance metrics. When the system performance improves (communication success rate increases, latency decreases, energy consumption reduces), a positive reward is given; when the system performance deteriorates, a negative reward is given.
[0085] Through the above steps, the constructed reinforcement learning module is obtained. Then, train this reinforcement learning module. The specific training process is as follows:
[0086] Train in a simulation environment, interact with the environment by executing actions, and observe the new state and the obtained rewards.
[0087] Adopt an experience replay mechanism, store the experiences (state, action, reward, next state) in an experience pool, randomly extract small batches of samples for training, break data correlation, and improve training efficiency.
[0088] Adopt a method of separating the target network from the policy network, and regularly update the target network parameters to stabilize the training process.
[0089] Use the mean squared error loss function to calculate the difference between the predicted value and the target value of the Q value, and update the network parameters through backpropagation.
[0090] During the training process, continuously adjust the balance strategy between exploration and exploitation ( - greedy strategy), and gradually optimize the strategy of the reinforcement learning module.
[0091] Specifically, the system inputs the historical configuration data into the reinforcement learning module to train the reinforcement learning module and optimize the strategy. Collect historical operation data, including configuration parameters and performance metrics under different network environments and load conditions. Perform data cleaning and feature engineering, and extract key features such as network latency, packet loss rate, and terminal online status. Design the network structure of the reinforcement learning module, use GRU (Gated Recurrent Unit) layers to capture temporal features, and the output layer outputs the Q values of each configuration action.
[0092] During the training process of the reinforcement learning module, adopt the Deep Q - learning (DQN) algorithm, use the experience replay and target network mechanisms to stabilize the training process. The training process records state - action - reward - new state (SARS), uses the mean squared error (MSE) loss function to calculate the difference between the predicted Q value and the target Q value, and adopts the Adam optimizer to accelerate the module convergence. Through - greedy strategy, balance exploration and exploitation.
[0093] After training is completed, use the validation set to evaluate the performance of the reinforcement learning module, and adjust the hyperparameters of the reinforcement learning module (such as learning rate, number of network layers, number of neurons) to optimize the module effect. Use the test set to test the generalization ability of the reinforcement learning module to ensure that the reinforcement learning module has good performance on unseen data.
[0094] Deploy the trained reinforcement learning module into the system, monitor its running performance in real time, and perform online updates and optimizations according to real-time data and feedback information, so that the reinforcement learning module can adapt to the dynamic changes of the environment.
[0095] Use the deployed reinforcement learning module to analyze the current system state and output the optimal configuration strategy. According to the decisions of the module, dynamically adjust the configuration parameters of the concentrator, such as real-time adjustment of communication frequency, bandwidth allocation, data acquisition strategy and other configuration parameters, to achieve optimal configuration of the meter terminal.
[0096] The automated decision-making of the reinforcement learning module provided in this embodiment enables the system to quickly and accurately adjust the configuration, improving the configuration efficiency and effect. The module can learn and adapt to complex and changing network environments and load conditions, maintaining the efficient operation of the system. The automated policy generation and adjustment reduce the dependence on manual operations and reduce human errors. Through online learning and feedback mechanisms, the module and policies can be continuously optimized to maintain the long-term stability of the system performance.
[0097] In one embodiment of the present invention, the system performance indicators after configuration adjustment are monitored in real time to ensure the effectiveness of the optimization strategy. When performance anomalies occur at key nodes are detected, warnings are automatically issued, and the anomaly information is reported to the master station center for timely countermeasures; at the same time, resource allocation for key nodes is prioritized.
[0098] Specifically, monitor the performance indicators of key nodes in real time, including communication delay, packet loss rate, CPU usage, memory occupancy, and device status, etc. When the monitored performance indicators exceed the preset thresholds, the system triggers an anomaly event and generates a detailed anomaly report, including the time of anomaly occurrence, involved devices, current configuration parameters, exceeded threshold indicators and their values, possible reasons, etc. Through a secure and encrypted communication channel, the anomaly report is reported to the master station center to automatically notify relevant operation and maintenance personnel.
[0099] In emergency handling measures, allocate more resources to affected key nodes, such as increasing bandwidth and improving communication frequency. Temporarily reduce the resource occupancy of low-priority tasks to relieve system pressure. Take isolation measures for abnormal nodes that may affect system stability to prevent the spread of faults. Provide remote access interfaces for operation and maintenance personnel to perform fault diagnosis and handling, collect relevant log information, and assist in problem location.
[0100] In this embodiment, through the real-time monitoring and abnormal reporting of key nodes, the system can timely detect and handle potential problems, preventing system crashes or serious performance degradation. In case of abnormalities, it can prioritize the resource supply of key nodes to ensure the continuity and reliability of important services. Automated abnormal monitoring and reporting reduce the workload of operation and maintenance personnel and improve the efficiency of fault handling. Timely detection and handling of abnormal behaviors effectively prevent security threats caused by malicious attacks and equipment failures.
[0101] Embodiment 2:
[0102] As Figure 2 shown, the embodiment of the present invention also provides an adaptive concentrator configuration system based on reinforcement learning. This system consists of three main functional modules that work together to achieve the adaptive configuration function of the concentrator. The specific functions and working processes of each module are as follows:
[0103] The configuration management module is used to receive and parse the configuration instructions from the master control center, extract the identity information of the electricity meter terminals to be configured and the historical performance of their performance indicators, and determine the priority category of the electricity meter terminals according to the comprehensive score of the identity information and historical performance of the electricity meter terminals; based on the policy network of the reinforcement learning module, generate an initial configuration action according to the priority category of the electricity meter terminals and their current state, and convert the configuration action into specific parameters and apply them to the concentrator and the electricity meter terminals; the configuration management module also includes a monitoring thread that can real-time monitor new instructions from the master control center, start a backup mechanism in case of emergencies, ensure that the system can quickly respond to emergency configuration requirements, and at the same time ensure the integrity of regular configuration tasks.
[0104] The parameter optimization module is the evaluation and adjustment center of the system. Its functions mainly include collecting the real-time performance data of the electricity meter terminals after applying the configuration, obtaining historical baseline data from the database, and calculating the comprehensive performance indicators; evaluating the performance of the configuration action through the reward function, if the evaluation result does not meet the expectation, using the reinforcement learning module to generate an optimized configuration action, performing a refined adjustment of the configuration parameters, and ensuring the consistent binding with the terminal identity; maintaining the correspondence library between the configuration parameters and the network environment conditions, and adjusting the configuration parameters of the relevant electricity meter terminals; the parameter optimization module uses a multi-objective optimization algorithm to find the best balance point among communication reliability, data real-time performance, and system energy consumption to ensure that the overall performance of the system reaches the optimal.
[0105] The reinforcement learning module is the intelligent core of the system, mainly responsible for maintaining and updating the deep Q-network module based on GRU; defining and managing the state space, action space, and reward function; processing the evaluation results and reward signals from the parameter optimization module; generating initial configuration actions and optimized configuration actions; continuously training and optimizing the module by combining historical data and real-time feedback; monitoring the performance of key nodes in the system and generating high-priority actions in case of anomalies; triggering the warning mechanism and generating and reporting anomaly information at the same time. The reinforcement learning module adopts a hierarchical architecture design, with strong scalability and adaptability. It can accumulate experience continuously as the system runs, improve the accuracy and efficiency of decision-making, and gradually realize the full automation of the configuration process.
[0106] The overall operation process of the system includes: the configuration management module receives the configuration instructions from the master control center and determines the priority of the meter terminals; the reinforcement learning module generates initial configuration actions according to the current state; the configuration management module applies the configuration actions to the concentrator and meter terminals; the parameter optimization module collects real-time data and evaluates the configuration effect; the reinforcement learning module updates the module according to the evaluation results and generates optimized actions if necessary; the parameter optimization module executes the optimized actions and adjusts the configuration parameters; the whole process forms a closed loop, and the system continuously learns and optimizes.
[0107] This adaptive configuration system based on reinforcement learning has strong intelligence and adaptability, and can dynamically adjust the configuration strategy according to environmental changes and business requirements, improving the overall performance and reliability of the system.
[0108] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the flow Figure 1 of one or more processes and / or boxes
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the flow Figure 1 of one or more processes and / or boxes
[0112] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An adaptive concentrator configuration method based on reinforcement learning, characterized in that: include: The concentrator obtains the configuration instruction through the main control center, extracts the identity information of the meter terminal to be configured and the historical performance of its performance indicators, and determines the priority category of the meter terminal according to the identity information of the meter terminal and the historical performance of the performance indicators; A policy network based on a reinforcement learning module generates an initial configuration action according to the priority category and current status of the meter terminal, and applies the parameters corresponding to the configuration action to the concentrator and the meter terminal; After the configuration parameters are applied, the configuration results are evaluated based on the historical performance and real-time data of the meter terminal performance indicators. The performance of the configuration action is calculated through the reward function. If the evaluation result does not meet expectations, the reinforcement learning module is used to generate optimized configuration actions and adjust the configuration parameters of the relevant meter terminals. The reinforcement learning module is trained by combining historical data and real-time feedback, and the policy network is continuously updated to optimize the generation of future configuration actions. When the performance of key nodes is abnormal, high-priority actions are generated for adjustment, while triggering early warnings and reporting abnormal information.
2. The adaptive concentrator configuration method based on reinforcement learning according to claim 1, characterized in that: The concentrator continuously monitors the identity information and current status of the electric meter terminal, and dynamically adjusts the priority category of the electric meter terminal when the identity information or current status of the electric meter terminal changes.
3. The method for configuring an adaptive concentrator based on reinforcement learning according to claim 1, characterized in that: Also includes: Configure monitoring threads for the concentrators to monitor the configuration instructions sent by the main control center in real time; When the main control center sends a new configuration instruction to the concentrator, it responds to the instruction first and obtains the current configuration status of the concentrator; When the concentrator executes the current configuration instruction, the current configuration instruction is suspended and the backup mechanism is started to save the status of the current configuration instruction, unfinished operations and cached data to the backup storage area; The backup mechanism is used to perform temporary emergency configuration on the meter terminal corresponding to the new configuration instruction. After completion, the suspended current configuration instruction is restored and the configuration consistency is ensured.
4. The method for configuring an adaptive concentrator based on reinforcement learning according to claim 1, characterized in that: Adjust the configuration parameters of related meter terminals, including: The configuration parameters of the meter terminal are refined and adjusted multiple times to determine the optimal configuration parameter set under different network conditions. Through the terminal identity binding mechanism, a correspondence is established between the unique identifier of each meter terminal and its configuration parameters to ensure the consistency of each adjusted configuration with the corresponding meter terminal.
5. The method for configuring an adaptive concentrator based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning module defines a state space, an action space and a reward function, wherein the state space includes network environment parameters, concentrator performance indicators and the operating status of the meter terminal, the action space includes configuration parameter adjustment options, and the reward function is defined according to the effect of the configuration action.
6. The method for configuring an adaptive concentrator based on reinforcement learning according to claim 5, characterized in that: The configuration results are evaluated based on the historical performance and real-time data of the meter terminal performance indicators, and the performance of the configuration action is calculated through the reward function, including: Determine the historical baseline of each performance indicator based on the historical performance of the meter terminal performance indicator to be configured, obtain the real-time data of the meter terminal performance indicator after the configuration parameters are applied, and calculate the comprehensive performance indicator by weighted average of the real-time data of the meter terminal performance indicator; if the comprehensive performance indicator of the meter terminal is greater than the historical baseline, give a positive reward; otherwise, give a negative reward.
7. The method for configuring an adaptive concentrator based on reinforcement learning according to claim 1, characterized in that: When the performance of key nodes is abnormal, high-priority actions are generated for adjustment, and warnings are triggered and abnormal information is reported, including: When a key node is detected to have performance anomalies, an early warning is automatically issued and the abnormal information is reported to the main station center so that timely response measures can be taken; at the same time, resource allocation for key nodes is prioritized.
8. According to the method for adaptive concentrator configuration based on reinforcement learning in claim 1, the reinforcement learning module is constructed based on a deep Q network and is used to generate future optimal configuration actions. The reinforcement learning module uses historical configuration data as input during the training phase, and generates a standardized training data set by combining a time sliding window algorithm to construct time series features through data cleaning, normalization and standardization, so as to train the reinforcement learning module and enable it to learn the optimal configuration strategy; In the use phase, the network environment parameters, concentrator performance indicators and the operating status of the meter terminal are used as input to generate the optimized configuration action and apply it to the concentrator and meter terminal. - The greedy strategy balances exploration and exploitation, allowing the reinforcement learning module to dynamically adjust the strategy during execution; represents the probability of exploration; In the optimization phase, the policy network is continuously adjusted by combining historical data and real-time feedback, so that the reinforcement learning module can adapt to the changing network environment, concentrator performance and meter terminal status fluctuations, ensuring dynamic optimization of future configuration actions.
9. An adaptive concentrator configuration system based on reinforcement learning, characterized in that: include: The configuration management module is used to receive and parse the configuration instructions of the main control center, extract the identity information of the meter terminal to be configured and the historical performance of its performance indicators, and determine the priority category of the meter terminal according to the identity information of the meter terminal and the historical performance of its performance indicators; A policy network based on a reinforcement learning module generates an initial configuration action according to the priority category and current status of the meter terminal, and applies the parameters corresponding to the configuration action to the concentrator and the meter terminal; The parameter optimization module is used to evaluate the configuration results based on the historical performance and real-time data of the meter terminal performance indicators, and calculate the performance of the configuration action through the reward function. If the evaluation result does not meet the expectations, the reinforcement learning module is used to generate the optimized configuration action and adjust the configuration parameters of the relevant meter terminal; Reinforcement learning module, which is used to train the reinforcement learning module by combining historical data and real-time feedback, continuously updating the policy network to optimize the generation of future configuration actions, and generating high-priority actions for adjustment when the performance of key nodes is abnormal, while triggering early warnings and reporting abnormal information; Among them, the configuration management module, parameter optimization module and reinforcement learning module are connected through a data interaction interface to form a closed-loop control system.
10. The adaptive concentrator configuration system based on reinforcement learning according to claim 9, characterized in that: The configuration management module also includes a monitoring thread unit, which is used to monitor the configuration instructions sent by the main control center in real time. When it is detected that the main control center sends a new configuration instruction to the concentrator, it responds to the instruction first and obtains the current configuration status of the concentrator; When the concentrator executes the current configuration instruction, the current configuration instruction is suspended and the backup mechanism is started to save the status of the current configuration instruction, unfinished operations and cached data to the backup storage area; The backup mechanism is used to perform temporary emergency configuration on the meter terminal corresponding to the new configuration instruction. After completion, the suspended current configuration instruction is restored and the configuration consistency is ensured.