A method and device for uploading logs of a photovoltaic energy storage system
By using the upload strategy dynamic adjustment model in the photovoltaic energy storage system and dynamically adjusting the log upload strategy according to the equipment and network status, the problems of delay and low success rate of log upload in large-scale photovoltaic energy storage systems are solved, and more efficient and flexible log management is achieved.
Patent Information
- Application Number
- CN202510503106.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Large-scale photovoltaic energy storage systems face the problems of timely transmission and low success rates in log upload strategies, especially when the network bandwidth is limited and the number of logs is large.
The pre-trained upload strategy dynamic adjustment model is adopted to dynamically adjust the log upload strategy according to the current status parameters of each target device (including network latency, bandwidth, log generation rate, log priority and power), optimize the upload timing and path, and reduce bandwidth usage.
By adaptively adjusting the log upload strategy, the real-time and stability of log data upload is significantly improved, the overall performance of photovoltaic energy storage systems in complex network environments is improved, and a more flexible and intelligent log management solution is provided.
Smart Images

Figure CN120016697B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of new energy technologies, and more specifically, to a method and device for uploading logs of a photovoltaic energy storage system. Background Art
[0002] With the rapid development of renewable energy technologies, photovoltaic energy storage systems have gradually become a key component in the field of energy management. The stable and efficient operation of a photovoltaic energy storage system depends on the real-time monitoring and remote data uploading of the system device status and performance data. However, after the expansion of the device scale and the increase in data volume in the current photovoltaic energy storage system, a series of technical challenges have been faced, especially in the log uploading strategy.
[0003] Traditional log uploading strategies adopt a timed uploading method, usually using a fixed time interval or an event-driven uploading mechanism. Although this strategy is effective in small-scale systems, its limitations become increasingly obvious in large-scale deployments; due to limited broadband conditions, frequent uploading operations, and excessive log quantities, it is difficult for the system to ensure the timely transmission of log data, resulting in log uploading failures or delays, which in turn affects the real-time monitoring and decision-making efficiency of the system. Summary of the Invention
[0004] In view of this, the purpose of the present application is to provide a method and device for uploading logs of a photovoltaic energy storage system, which can ensure the timely transmission and success rate of log uploading in a large-scale photovoltaic energy storage system.
[0005] A method for uploading logs of a photovoltaic energy storage system provided by an embodiment of the present application is applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes:
[0006] Obtain the logs and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating conditions and network environment of the target device; the status parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths;
[0007] Input the current status parameters of each target device into a pre-trained upload strategy dynamic adjustment model, and the upload strategy dynamic adjustment model processes the current status parameters to determine a target adjustment action that matches the current status parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes multiple adjustment actions for adjusting the log uploading strategy of the target device;
[0008] Dynamically adjust the target log uploading strategy for the target device based on the target adjustment action;
[0009] Based on the target log upload policy of each target device in the target photovoltaic energy storage system, upload the log of the target device to upload the log of the photovoltaic energy storage system.
[0010] In some embodiments, in the log upload method of the photovoltaic energy storage system, the upload policy dynamic adjustment model is trained based on the following method:
[0011] Construct an upload policy dynamic adjustment model and configure an adjustment action set for the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions;
[0012] Train the upload policy dynamic adjustment model based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between the different sample state parameters of different sample devices and different adjustment actions; wherein, a loss function and a reward function are configured in the upload policy dynamic adjustment model; the reward function is used to calculate the immediate reward fed back to the sample photovoltaic energy storage system by the environment after selecting an adjustment action based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and the immediate reward fed back to the photovoltaic energy storage system by the environment.
[0013] In some embodiments, in the log upload method of the photovoltaic energy storage system, the training of the upload policy dynamic adjustment model based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between the different sample state parameters of different sample devices and different adjustment actions, includes:
[0014] At each time step, the upload policy dynamic adjustment model starts from the sample state parameters of the sample devices in the sample photovoltaic energy storage system and predicts the Q value of each adjustment action in the adjustment action set;
[0015] The upload policy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action;
[0016] The sample photovoltaic energy storage system dynamically adjusts the sample log upload policy for the sample device based on the sample adjustment action and executes the sample log upload policy to obtain the feedback state and the immediate reward fed back to the photovoltaic energy storage system by the environment; wherein, the immediate reward is calculated based on the reward function;
[0017] Based on the feedback state and the immediate reward fed back to the photovoltaic energy storage system by the environment, determine the calculation result of the loss function of the upload policy dynamic adjustment model;
[0018] Optimize the upload policy dynamic adjustment model based on the calculation result of the loss function of the model dynamically adjusted according to the upload policy until the upload policy dynamic adjustment model meets the preset training stop condition, and obtain a trained upload policy dynamic adjustment model.
[0019] In some embodiments, the method for uploading logs of the photovoltaic energy storage system further includes:
[0020] Determine multiple optimization objectives for the log upload of the photovoltaic energy storage system; the multiple optimization objectives include: the success rate of log upload, the delay time, the bandwidth occupancy, the device power, and the log priority;
[0021] Fuse the log priority and the device power to determine a combined reward function for priority and power;
[0022] Based on the success rate of the log upload, the delay time, the bandwidth occupancy, and the combined reward function for priority and power, construct a reward function for calculating the immediate reward.
[0023] In some embodiments, in the method for uploading logs of the photovoltaic energy storage system, the reward rules of the reward function include: a log upload success reward rule, a log upload failure penalty rule, a bandwidth occupancy penalty rule, and a delay penalty rule;
[0024] The log upload success reward rule includes: after the log upload is successful, obtain a reward according to the state of the sample photovoltaic energy storage system; among them, when the device power is more and / or the log priority is higher, a higher reward is obtained for the successful log upload;
[0025] The log upload failure penalty rule includes: after the log upload fails, obtain a negative reward according to the state of the sample photovoltaic energy storage system; among them, the higher the log priority, the higher the negative reward;
[0026] The bandwidth occupancy penalty rule includes: when the bandwidth occupancy of the sample photovoltaic energy storage system exceeds the preset bandwidth occupancy threshold, obtain a negative reward;
[0027] The delay penalty rule includes: when the gap between the delay time and the preset maximum allowable value is smaller, the negative reward is higher.
[0028] In some embodiments, the method for uploading logs of the photovoltaic energy storage system further includes:
[0029] Based on the immediate reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the upload policy dynamic adjustment model for the adjustment action executed in the current state of the photovoltaic energy storage system, construct a loss function.
[0030] In some embodiments, in the method for uploading logs of the photovoltaic energy storage system, the adjustment actions in the adjustment action set include: adjusting the upload frequency, selecting a network connection, and selecting a compression ratio.
[0031] In some embodiments, in the method for uploading logs of the photovoltaic energy storage system, after uploading the logs of the target devices based on the target log upload policies of each target device in the target photovoltaic energy storage system to upload the logs of the photovoltaic energy storage system, the log upload method further includes:
[0032] Recording the experience tuple for executing the target log upload policy of the target device in the target photovoltaic energy storage system and storing the experience tuple in the experience replay pool; the experience tuple includes: the current state parameters of the target device in the target photovoltaic energy storage system, the target adjustment actions taken corresponding to the current state parameters of the target device, the immediate reward, and the state parameters of the next state of the target device;
[0033] When a preset batch update condition is satisfied, randomly extracting batch data from the experience replay pool and updating the trained upload policy dynamic adjustment model based on the extracted batch data to obtain an updated upload policy dynamic adjustment model.
[0034] In some embodiments, in the method for uploading logs of the photovoltaic energy storage system, the number of the upload policy dynamic adjustment models is multiple; different upload policy dynamic adjustment models control different regions in the target photovoltaic energy storage system;
[0035] Correspondingly, inputting the current state parameters of each target device into the pre-trained upload policy dynamic adjustment model includes:
[0036] Inputting the current state parameters of each target device into the upload policy dynamic adjustment model that matches the region.
[0037] In some embodiments, there is also provided a device for uploading logs of a photovoltaic energy storage system, which is characterized in that it is applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log upload device includes:
[0038] An acquisition module, configured to acquire the logs and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters characterize the operating conditions and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths;
[0039] A determination module, configured to input the current state parameters of each of the target devices into a pre-trained upload policy dynamic adjustment model, where the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured set of adjustment actions; the set of adjustment actions includes multiple adjustment actions for adjusting the log upload policy of the target device;
[0040] An adjustment module, configured to dynamically adjust the target log upload policy for the target device based on the target adjustment action;
[0041] An upload module, configured to upload the logs of the target device based on the target log upload policy of each target device in the target photovoltaic energy storage system, so as to upload the logs of the photovoltaic energy storage system.
[0042] In an embodiment of the present application, a method and device for uploading logs of a photovoltaic energy storage system are provided, which are applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log upload method includes: obtaining the logs and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters characterize the operating conditions and network environment of the target device; the state parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, where the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured set of adjustment actions; the set of adjustment actions includes multiple adjustment actions for adjusting the log upload policy of the target device; dynamically adjusting the target log upload policy for the target device based on the target adjustment action; uploading the logs of the target device based on the target log upload policy of each target device in the target photovoltaic energy storage system, so as to upload the logs of the photovoltaic energy storage system, thereby being able to adaptively adjust the log upload policy of a single device according to factors such as device status, network bandwidth, latency, log importance, device power, log generation rate, etc. in the photovoltaic energy storage system, thereby optimizing the overall log upload policy of the system, thereby greatly improving the real-time performance and stability of log data upload, enabling the photovoltaic energy storage system to significantly improve the overall performance when dealing with complex network environments and a large number of logs, and the system provides a more flexible and intelligent log management solution, significantly improving the real-time performance, efficiency, and reliability of the system. Description of the Drawings
[0043] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0044] Figure 1 The flowchart of the method for uploading logs of the photovoltaic energy storage system according to the embodiments of the present application is shown;
[0045] Figure 2 The flowchart of the method for training the dynamic adjustment model of the upload strategy according to the embodiments of the present application is shown;
[0046] Figure 3 The schematic diagram of the framework design of the dynamic adjustment model of the upload strategy is shown;
[0047] Figure 4 The schematic diagram of the state change of the dynamic adjustment model of the upload strategy according to the embodiments of the present application is shown;
[0048] Figure 5 The flowchart of the method for uploading logs of another photovoltaic energy storage system according to the embodiments of the present application is shown;
[0049] Figure 6 The schematic diagram of the structure of the dynamic adjustment model of the upload strategy is shown;
[0050] Figure 7 The schematic diagram of the structure of the device for uploading logs of the photovoltaic energy storage system according to the embodiments of the present application is shown. Detailed implementation manners
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve for illustration and description purposes, and are not used to limit the protection scope of the present application. Additionally, it should be understood that the schematic drawings are not drawn to the physical scale. The flowcharts used in the present application show the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0052] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0053] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude adding other features.
[0054] With the rapid development of renewable energy technology, photovoltaic energy storage systems have gradually become a key component in the field of energy management. The stable and efficient operation of photovoltaic energy storage systems depends on the real-time monitoring and remote data uploading of the status and performance data of system devices. However, after the expansion of equipment scale and the increase in data volume in current photovoltaic energy storage systems, a series of technical challenges have been faced, especially in the log uploading strategy.
[0055] Traditional log uploading strategies adopt a timed uploading method, usually using a fixed time interval or an event-driven uploading mechanism. Although this strategy is effective in small-scale systems, its limitations become increasingly obvious in the case of large-scale deployment; due to limited broadband conditions, frequent uploading operations, and excessive log quantities, it is difficult for the system to ensure the timely transmission of log data, resulting in log uploading failures or delays, which in turn affect the real-time monitoring and decision-making efficiency of the system.
[0056] Specifically, first, timed uploading is prone to cause data congestion under the condition of limited network bandwidth, especially during peak load periods, and it is difficult for the system to ensure the timely transmission of log data; second, overly frequent or unnecessary uploading operations will increase the network burden and upload delay, which in turn affects the real-time monitoring and decision-making efficiency of the system; in addition, when a photovoltaic energy storage system has multiple network connection paths (such as Wi-Fi, cellular network, Bluetooth, etc.) at the same time, how to intelligently switch in different network environments and avoid unnecessary network conversions has also become an urgent problem to be solved.
[0057] Based on this, how to intelligently manage the log uploading strategy and improve the real-time performance of data uploading has become the core issue for enhancing the overall performance of photovoltaic energy storage systems.
[0058] Based on this, in the embodiments of the present application, a method and device for log uploading of a photovoltaic energy storage system are provided, which are applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes: obtaining the logs and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating conditions and network environment of the target device; the status parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; inputting the current status parameters of each target device into a pre-trained upload policy dynamic adjustment model, and the upload policy dynamic adjustment model processes the current status parameters to determine a target adjustment action that matches the current status parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes multiple adjustment actions for adjusting the log upload policy of the target device; dynamically adjusting the target log upload policy for the target device based on the target adjustment action; uploading the logs of the target device based on the target log upload policies of each target device in the target photovoltaic energy storage system to upload the logs of the photovoltaic energy storage system, so as to be able to adaptively adjust the log upload policy of a single device according to factors such as device status, network bandwidth, latency, log importance, device power, log generation rate, etc. in the photovoltaic energy storage system, thereby optimizing the overall log upload policy of the system, thereby greatly improving the real-time performance and stability of log data upload, enabling the photovoltaic energy storage system to significantly improve the overall performance when dealing with complex network environments and a large number of logs, and the system provides a more flexible and intelligent log management solution, significantly improving the real-time performance, efficiency, and reliability of the system.
[0059] Please refer to Figure 1 , Figure 1 shows a flowchart of the method for log uploading of the photovoltaic energy storage system according to the embodiments of the present application; as Figure 1 shown, the method for log uploading of the photovoltaic energy storage system is applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes the following steps S101 - S104:
[0060] S101. Obtain the logs and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating conditions and network environment of the target device; the status parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths;
[0061] S102. Input the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model. The upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured set of adjustment actions. The set of adjustment actions includes multiple adjustment actions for adjusting the log upload policy of the target device.
[0062] S103. Dynamically adjust the target log upload policy for the target device based on the target adjustment action.
[0063] S104. Based on the target log upload policies of each target device in the target photovoltaic energy storage system, upload the logs of the target device to upload the logs of the photovoltaic energy storage system.
[0064] In a large-scale photovoltaic energy storage system, a variety of key target devices will generate logs to record information such as their operating status, performance parameters, and fault alarms. Specifically, the key target devices include photovoltaic panels, battery energy storage systems, inverters, charge and discharge controllers, grid connectors, other auxiliary devices, etc. The number of some key target devices is multiple, such as photovoltaic panels, inverters, charge and discharge controllers, etc. Other auxiliary devices include environmental monitoring devices (such as temperature sensors, humidity sensors, etc.), safety protection devices (such as overcurrent and overvoltage protection devices, etc.).
[0065] Based on this, due to the large number of devices and high power involved in a large-scale photovoltaic energy storage system, the amount of data generated is also larger, and log generation is more complex. Since it is necessary to monitor the status, performance, and operation of energy storage devices in real time, a large amount of detailed parameter information and status records are included in the logs.
[0066] Small-scale photovoltaic energy storage systems are relatively simple, with a small amount of data and easier log generation. These systems usually only need to record basic device status and power parameters, and the log content is relatively simple and clear.
[0067] Therefore, in a large-scale photovoltaic energy storage system, it is necessary to optimize the log upload policy. By adaptively adjusting the log upload policy, the upload efficiency of the system can be improved, network bandwidth occupancy can be reduced, and upload latency can be reduced.
[0068] In the embodiments of the present application, a log upload decision engine is deployed in the edge server of the photovoltaic energy storage system to sense the status of the target photovoltaic energy storage system in real time, determine the log upload policy, and execute the decision.
[0069] In some embodiments, the log upload decision engine is built based on a DQN model.
[0070] In the step S101, obtain the logs and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating conditions and network environment of the target device; the status parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths.
[0071] Specifically, the log upload decision engine, through the data acquisition module, monitors the status parameters of the target device in real time, such as network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system, etc., and forms a status vector S = (s1, s2, ..., sn) as a model for dynamically adjusting the upload strategy of the model.
[0072] Among them, the log priority characterizes the importance of the log.
[0073] In some embodiments, for network latency and bandwidth, use network monitoring tools (such as Ping, Traceroute, network performance testing tools, etc.) to measure the network latency between each target device and the log processing platform, and monitor the bandwidth usage of the network interface to ensure that data transmission is not blocked due to insufficient bandwidth.
[0074] For the log generation rate, monitor the log generation rate of each device through a log collection tool or a custom script to evaluate the efficiency of log processing and potential performance bottlenecks.
[0075] For the log priority, when generating a log, set the log priority according to the importance of the log content; in some embodiments, it can be implemented through the configuration in the log management system to ensure that high-priority logs are uploaded first.
[0076] For the power of the photovoltaic energy storage system, use the monitoring interface or dedicated sensor of the energy storage system to obtain real-time power information, including remaining power, charge / discharge status, etc.
[0077] The multiple network connection paths include Bluetooth connection, WiFi connection, cellular network, etc.
[0078] First of all, timed uploads are likely to cause data congestion under the condition of limited network bandwidth, especially during peak load periods, and it is difficult for the system to ensure the timely transmission of log data. Secondly, overly frequent or unnecessary upload operations will increase the network burden and upload latency, thereby affecting the real-time monitoring and decision-making efficiency of the system. In addition, when the photovoltaic energy storage system has multiple network connection paths at the same time (such as Wi-Fi, cellular network, Bluetooth, etc.), how to intelligently switch in different network environments and avoid unnecessary network conversions has also become an urgent problem to be solved.
[0079] Therefore, collect the operating status of devices such as network latency, bandwidth, and log generation rate to intelligently adjust the log upload policy of the overall device.
[0080] In the step S102, input the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model. The upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured set of adjustment actions. The set of adjustment actions includes multiple adjustment actions for adjusting the log upload policy of the target device.
[0081] The adjustment actions in the set of adjustment actions include: adjusting the upload frequency, selecting a network connection, and selecting a compression ratio.
[0082] That is to say, in the photovoltaic energy storage system, the upload policy dynamic adjustment model can adaptively adjust the log upload policy according to factors such as device status, network bandwidth, latency, and log importance, not only optimizing the upload timing and path, but also reducing bandwidth occupancy by intelligently compressing data, thereby greatly improving the real-time performance and stability of data upload.
[0083] In some embodiments, the upload policy dynamic adjustment model is specifically implemented based on the Deep Q-Network (DQN). The Deep Q-Network (DQN) is a specific implementation of Deep Reinforcement Learning (DRL). DQN combines Q-learning in reinforcement learning with a deep neural network and can handle complex and high-dimensional state spaces, especially suitable for dynamic network environments.
[0084] Please refer to Figure 2 , in some embodiments, the upload policy dynamic adjustment model is trained based on the following steps S201-S202:
[0085] S201. Construct an upload policy dynamic adjustment model and configure the set of adjustment actions of the upload policy dynamic adjustment model. The set of adjustment actions includes multiple adjustment actions;
[0086] S202. Train the upload policy dynamic adjustment model based on the sample status parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between the different sample status parameters of different sample devices and different adjustment actions. Among them, a loss function and a reward function are configured in the upload policy dynamic adjustment model. The reward function is used to calculate the immediate reward fed back to the sample photovoltaic energy storage system by the environment after selecting an adjustment action based on the sample status parameters of the sample device. The loss function is determined based on the feedback status and immediate reward fed back to the photovoltaic energy storage system by the environment.
[0087] That is to say, the upload policy dynamic adjustment model learns the overall impact of all possible adjustment actions of different sample devices in different operating states on the photovoltaic energy storage system, so as to comprehensively consider multiple components and strategies for debugging, ensuring the performance and generalization ability of the upload policy dynamic adjustment model, and providing strong support for the log upload task of photovoltaic components.
[0088] In some embodiments, in the log upload method of the photovoltaic energy storage system, training the upload policy dynamic adjustment model based on the sample status parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between the different sample status parameters of different sample devices and different adjustment actions includes:
[0089] At each time step, the upload policy dynamic adjustment model starts from the sample status parameters of the sample devices in the sample photovoltaic energy storage system and predicts the Q value of each adjustment action in the adjustment action set.
[0090] The upload policy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action.
[0091] The sample photovoltaic energy storage system dynamically adjusts the sample log upload policy for the sample device based on the sample adjustment action, and executes the sample log upload policy to obtain the feedback status and immediate reward fed back to the photovoltaic energy storage system by the environment. Among them, the immediate reward is calculated based on the reward function.
[0092] Based on the feedback status and immediate reward fed back to the photovoltaic energy storage system by the environment, determine the calculation result of the loss function of the upload policy dynamic adjustment model.
[0093] Optimize the upload policy dynamic adjustment model based on the calculation result of the loss function of the upload policy dynamic adjustment model until the upload policy dynamic adjustment model meets the preset training stop condition, and obtain the trained upload policy dynamic adjustment model.
[0094] Specifically, please refer to Figure 3, the framework design of the above-mentioned upload policy dynamic adjustment model is as follows.
[0095] Environment: It includes the operating status of the photovoltaic energy storage system, network connection, battery power of the device, log generation rate, importance level of the log, etc. It reflects the current status and condition of the photovoltaic energy storage system and is used for the agent to learn and make decisions.
[0096] State (S): The state is a vector composed of multiple parameters, such as: current network delay (delay), available bandwidth (bandwidth), device battery level (battery_level), log generation rate (log_rate), and log priority (log_priority).
[0097] The state vector can be expressed as: s = [delay, bandwidth, battery_level, log_rate, log_priority].
[0098] Action (A): The action set includes adjusting the upload frequency, selecting a network connection (such as Wi-Fi, Bluetooth, or cellular network), and selecting the compression ratio, etc.
[0099] The action can be defined as: A = {a1, a2,... ai,... an}, where each ai represents a specific action.
[0100] Reward (R): The reward is an indicator used to evaluate the quality of the agent's behavior. For a system, the reward is usually related to factors such as upload success rate, bandwidth occupancy, delay, and power consumption. For example, when the system successfully uploads the log and reduces the bandwidth occupancy, a positive reward is given; while a negative reward is obtained for upload failure or excessive bandwidth consumption.
[0101] As Figure 3 shown, after the agent selects action A, the system executes action A, and the environment feedbacks the reward R and state S regarding action A, thereby training the agent.
[0102] Deep Q-Network (DQN) approximates the Q-value function in traditional Q-learning by introducing a deep neural network (DNN), enabling effective decision optimization in complex or high-dimensional state spaces. The specific process includes Q-value update, loss function calculation, experience replay, and the interaction between the target network and the policy network.
[0103] The following is a detailed technical description of this process.
[0104] Please refer to Figure 4 , Figure 4Shows a schematic diagram of the state change of the upload policy dynamic adjustment model described in the embodiments of the present application; as Figure 4 shown, at each time step, the system starts from the current state s(t), predicts the Q-values of all possible adjustment actions a(t) at present through a deep neural network, that is, Q(s, a). The model selects a specific adjustment action according to the Q-value, and enters the next state s(t + 1) after executing the adjustment action; starting from the state s(t + 1), execute the adjustment action a(t + 1), and enter the next state s(t + 2).
[0105] In the embodiments of the present application, the specific selection process follows the ε-greedy strategy:
[0106] a_t = arg max_a Q(s_t, a; θ);
[0107] Q(s,a) represents the expected return of executing the adjustment action a in the state s; θ is the parameter of the deep neural network, and a_t represents the optimal action in the corresponding state s_t under the maximization of the Q-value.
[0108] To avoid getting stuck in a local optimal solution in some cases, the agent will also randomly select an action (exploration) with probability ε. a optimal solution action.
[0109] Arg max a: Take the value of the corresponding variable when the following formula reaches the maximum value.
[0110] Q(s_t, a; θ): In the initial state, the initial action a will be randomly selected within a certain range. Here, a takes the maximum value within this range in the initial state, and s_t is the state corresponding to this a value. After executing the action, the environment feeds back the new state s_t+1 and the immediate reward r_t to the system; the goal of the system is to maximize the cumulative reward, so it is necessary to update the Q-value of the current state and action.
[0111] The update of the Q-value is based on the Bellman equation, which expresses the relationship between the Q-value in the current state and the maximum Q-value in the next state. Its update formula is as follows:
[0112] Q(s_t, a_t) = Q(s_t, a_t) + α * [r_t + γ * max_a' Q(s_t+1, a'; θ') -Q(s_t, a_t)];
[0113] Among them: Q(s_t, a_t) represents the Q-value of action a in the state s at time t (i.e., the current state), and a_t represents the optimal action corresponding to the current state; max_a' Q(s_t+1, a'; θ') represents the Q-value of the optimal action a' in the state s at time t+1 (i.e., the next state) until the Q-value converges to the maximum value; a' represents the adjustment action to be executed in the next step; θ' is the parameter of the target network, which is updated at a certain time interval from the parameters of the policy network; α is the learning rate, which controls the step size of each update; γ is the discount factor, which is used to balance short-term rewards and long-term rewards, and its value ranges from 0 to 1, usually set to be close to 1 (e.g., 0.99) to emphasize the importance of long-term rewards; r_t is the immediate reward obtained after executing the adjustment action a_t at time step t.
[0114] The specific implementation of the ε-greedy policy is as follows A1 - A3:
[0115] A1: Set the value of ε: ε is a value between 0 and 1, representing the probability of randomly selecting an action. The larger the value of ε, the higher the degree of exploration, but the lower the degree of exploitation; the smaller the value of ε, the lower the degree of exploration, but the higher the degree of exploitation.
[0116] A2: Generate a random number: At each decision point, generate a random number between 0 and 1.
[0117] A3: Judgment: If the random number is less than or equal to ε, randomly select an action; otherwise, select the action with the largest Q-value.
[0118] Through the ε-greedy policy, the agent selects a random action with a certain probability, even if this action may not be good. This randomness helps the agent discover new states and actions, thus finding a better strategy; at the same time, the agent selects the currently considered optimal action (i.e., the action with the largest Q-value) with a high probability. This behavior of using known knowledge helps the agent quickly obtain higher rewards.
[0119] In order to enable the DQN to be continuously updated, the DQN guides the update of network parameters by minimizing the difference between the predicted Q-value and the target Q-value.
[0120] In the method for uploading logs of the photovoltaic energy storage system described in the embodiments of the present application, the method for uploading logs further includes:
[0121] Based on the immediate reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q-value of the upload policy dynamic adjustment model for executing the adjustment action in the current state of the photovoltaic energy storage system, a loss function is constructed.
[0122] Specifically, the loss function L(θ) is defined as follows:
[0123] L(θ) = E[(r_t + γ * max_a' Q(s_t+1, a'; θ') - Q(s_t, a_t; θ))^2];
[0124] Where: r_t is the immediate reward obtained after performing action a_t at time step t; max_a' Q(s_t+1, a';θ') represents the maximum future reward that can be obtained in the next state s_t+1; γ is the discount factor used to balance short-term and long-term rewards; θ is the parameter of the current policy network; θ' is the parameter of the target network, which remains unchanged for a period of time to ensure training stability; E[ ] is the expectation operator used to calculate the average error of the network over different samples.
[0125] r_t is the immediate reward, which reflects the effect of the log upload action at this time step and is related to transmission delay, bandwidth utilization, log importance, etc.
[0126] γ * max_a' Q(s', a'; θ') represents the maximum expected return in the next state and is used to estimate the long-term benefit.
[0127] Q(s, a; θ))^2 is the predicted Q value when the neural network performs action a in state s; the square is mainly used to measure the prediction error and ensure that the network can minimize this difference through gradient descent. Specifically, the square term in the loss function is used to penalize large error values, making the model pay more attention to samples with large errors during updates, gradually adjusting the network weights, and approaching the true Q value; thus, it means trying to narrow the gap between the predicted Q value of the network and the actual target Q value at each step.
[0128] The minimization of the loss function is achieved through the Stochastic Gradient Descent (SGD) algorithm. Through backpropagation, the gradient of the loss function propagates to each layer of the neural network, gradually updating the weight parameters of the network to approximate the optimal Q value function.
[0129] In the log reporting scenario of the photovoltaic energy storage system, the reward function is a key factor in reinforcement learning for guiding the optimization decision of the DQN (DeepQ-Network). The design of the reward function directly affects how the system adjusts the upload strategy under different network conditions, device states, and log generation rates. To ensure the timeliness, effectiveness, and resource savings of log upload, the reward function must be able to reflect multi-dimensional performance metrics.
[0130] The reward function design principles include: upload success rate, bandwidth usage efficiency, log importance, energy consumption management, and latency minimization.
[0131] Upload success rate: The successful upload of logs is the main goal of the system. Therefore, a positive reward should be given for successfully uploading logs, and a negative reward for failure.
[0132] Bandwidth usage efficiency: During the upload process, if the upload task can be completed with lower bandwidth occupancy, the system should receive an additional reward; conversely, excessive bandwidth occupancy will result in a penalty.
[0133] Log importance: Different logs have different priorities, and important logs should be uploaded first. Therefore, the reward function should consider the importance weight of the logs.
[0134] Energy consumption management: The energy consumption of the device during the upload process is also an important consideration, especially when the device battery power is low. When the battery power is low, if a low-energy consumption path is selected (such as choosing a low-power network), a positive incentive should be given; conversely, high-energy consumption operations will be penalized.
[0135] Latency minimization: The timeliness of upload is also important. If the logs can be uploaded with a short latency, the system should receive a reward; if the latency is too high, a penalty should be imposed.
[0136] Based on this, in some embodiments, please refer to Figure 5 , the log upload method of the photovoltaic energy storage system further includes the following steps S501 - S503:
[0137] S501. Determine multiple optimization goals for the log upload of the photovoltaic energy storage system; the multiple optimization goals include: the success rate of log upload, latency time, bandwidth occupancy, device power, and log priority;
[0138] S502. Integrate the log priority and device power to determine a combined reward function for priority and power;
[0139] S503. Based on the success rate of log upload, latency time, bandwidth occupancy, and the combined reward function for priority and power, construct a reward function for calculating the immediate reward.
[0140] The reward rules of the reward function include: log upload success reward rule, log upload failure penalty rule, bandwidth occupancy penalty rule, and latency penalty rule;
[0141] The log upload success reward rule includes: after successfully uploading the logs, obtain a reward according to the state of the sample photovoltaic energy storage system; among them, when the device power is more and / or the log priority is higher, a higher reward is obtained for successfully uploading the logs;
[0142] The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; among them, the higher the log priority, the higher the negative reward;
[0143] The bandwidth occupancy penalty rule includes: when the bandwidth occupancy of the sample photovoltaic energy storage system exceeds the bandwidth occupancy threshold, a negative reward is obtained;
[0144] The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowable value, the higher the negative reward.
[0145] In an optional embodiment, the reward function R is specifically:
[0146] R = R_success * f(log(P_energy)) * (L_max - L) / (L_max - L_avg) - λ_1 * (B_used / B_threshold) - λ_2 * (L_delay / L_max).
[0147] R_success represents the reward for successful log upload; f(log(P_energy)) is a function of log importance and device power, used to represent the comprehensive score corresponding to different importance and power situations; for example, when the log importance is high and the power is sufficient, the system will preferentially select successful upload and give a higher reward; this function can be designed as: f(log(P_energy)) = log(1 + 1 / (1 + e^(-k*(P_energy - P_threshold)))); where k is a parameter that adjusts the steepness of the curve, P_threshold is the power threshold (set to 5% for example), P_energy is the device power; L_max represents the maximum allowable delay time; L_avg represents the average delay time of log upload; B_used represents the bandwidth used during upload; B_threshold represents the upper limit of bandwidth usage set by the system; L_delay represents the upload delay time; λ_1 is the first trade-off parameter, controlling the impact of bandwidth occupancy on the total reward; λ_2 is the second trade-off parameter, controlling the impact of network delay on the total reward.
[0148] The reward function can balance multiple objectives. The reward function comprehensively considers multiple factors such as the success rate of log upload, latency, bandwidth occupancy, and device power, enabling the system to find a balance among these objectives. It can also adapt to different scenarios: by adjusting the parameters λ_1 and λ_2, the system's sensitivity to bandwidth and latency can be adjusted according to different network environments and system requirements. It can also motivate the system to learn. The design of the reward function can encourage the system to learn the optimal upload strategy in different states, thereby improving the overall performance of the system.
[0149] A detailed explanation of the reward function is as follows.
[0150] Reward for successful log upload: After the system successfully uploads the log, it calculates the actual reward value based on the importance of the log I_log and the current device power P_energy. When the power is sufficient, a successful upload of an important log should receive a higher reward; if the power is insufficient, the system will give priority to low-power consumption networks, appropriately reducing the upload frequency or compressing the log.
[0151] Penalty for failed log upload: A failed upload not only causes a delay in uploading important log information but also may consume unnecessary network resources. Therefore, the system needs to impose a penalty according to the failure situation, especially when the log importance is high, the penalty intensity should be increased.
[0152] Penalty for bandwidth occupancy: An excessively high bandwidth occupancy rate B_used > B_threshold will affect the overall performance of the network. Therefore, when the bandwidth occupancy exceeds a certain threshold, the system will face an additional negative reward.
[0153] Penalty for latency: When the upload latency L_delay approaches the maximum value L_max allowed by the system, the system will be penalized, which helps to prompt the model to select a low-latency network or upload at an appropriate time.
[0154] To improve the efficiency and stability of learning, DQN introduces an experience replay mechanism. Each interaction between the system and the environment (state, action, reward, next state s_t+1) is stored in an experience replay pool (Replay Buffer). The data in the experience replay pool is used for training, and the specific steps are as follows:
[0155] During each training, the system randomly samples a batch (mini-batch) of data from the experience replay pool to calculate the loss and update the network parameters; this can break the temporal correlation of the data and improve the stability of training.
[0156] The randomly sampled experience data helps to avoid overfitting to certain specific state sequences, thereby enhancing the generalization ability of learning.
[0157] The replay buffer size is usually set to a fixed value, and the system will periodically discard the oldest data to make room for new experiences.
[0158] To improve the stability of Q-value learning, DQN adopts a dual-network structure of a target network and a policy network. The target network is used to calculate the target Q-value, and its parameters θ' will remain unchanged for a period of time, while the policy network is updated according to new data during each training.
[0159] The parameter update process between the target network and the policy network is as follows: θ' ← θ;
[0160] The target network synchronizes with the policy network every N steps, that is, assigns the weights θ of the policy network to the weights θ' of the target network; this mechanism reduces the frequent change of the target value, avoids instability during training, and thus improves the convergence speed of the model.
[0161] Based on this, the upload policy dynamically adjusts the model training process as follows.
[0162] First, initialize the weight parameter θ of the policy network and initialize the weight parameter θ' of the target network as θ' = θ.
[0163] Construct a replay buffer D and store the data of each interaction with the environment into the buffer.
[0164] At each time step: The system selects an adjustment action a_t using the ε-greedy policy according to the current state s_t; execute the adjustment action a_t and obtain a new state s_t+1 and an immediate reward r_t; store (s_t, a_t, r_t, s_t+1) into the replay buffer D; randomly sample a batch of data from the replay buffer D and calculate the target Q-value; minimize the loss function L(θ) and update the parameters θ of the policy network through gradient descent; every N steps, synchronize the weights θ of the policy network to the target network θ'.
[0165] Please refer to the following Figure 6 , the upload policy dynamically adjusted model includes: an input layer, a hidden layer, and an output layer.
[0166] The input layer receives the current state parameters of the target device in the system, and this vector intuitively reflects the current operating condition and network environment of the system.
[0167] For the log upload problem in the photovoltaic energy storage system, the input layer can include the following key parameters: Network Latency: Represents the time delay for data to be uploaded to the cloud or server, and is used to evaluate the network transmission quality.
[0168] Bandwidth: The currently available network bandwidth, which affects the speed and transmission efficiency of log upload. * Device Power Level: The remaining power of the device, which is used to dynamically balance log upload and power consumption to avoid frequent uploads in low power states.
[0169] Log Generation Rate: The speed at which the system generates logs, which determines the amount of logs to be uploaded.
[0170] Log Importance: The priorities of different logs, which are used to ensure that critical logs are uploaded first when network resources are scarce. These parameters can form an input vector s ∈ R^n, where n represents the state dimension.
[0171] These parameters together affect the system's decision on log upload and serve as the input to the neural network.
[0172] The hidden layer uses a fully connected network structure to map the data from the input layer to a high-dimensional space, thereby capturing the complex non-linear relationships between the original data. The hidden layer of DQN usually adopts the ReLU activation function (Rectified Linear Unit), which can enhance the non-linear expression ability of the model. The formula is as follows: h = ReLU(Ws + b); where: h is the output of the hidden layer; Ws is the weight matrix of the hidden layer, representing the connection strength between the input layer and the hidden layer; b is the bias term; The ReLU activation function is defined as ReLU(x) = max(0, x), which introduces non-linearity and avoids the model from overfitting to linear decisions.
[0173] Specifically, the function of the Ws (weight matrix) is as follows:
[0174] 1. Control the transmission strength of signals: The weight matrix determines the influence degree of each input signal (the output from the previous layer) on the output of the hidden layer. During the training process, the network will automatically adjust these weights according to the pattern of the input data, making the output closer and closer to the target value.
[0175] 2. Learn the representation of data: The hidden layer is a key part of the neural network. It is responsible for mapping the input data from the original space to a high-dimensional feature space, capturing the complex non-linear relationships between the data. Through the weight matrices of different layers, the network can extract more and more abstract features until the accurate result is obtained at the final output layer (for example, Q-value prediction).
[0176] 3. Expressive power of the model: By adjusting Ws, the network can learn complex relationships between different inputs. For example, in the scenario of log uploading, the hidden layer extracts high-order features that affect decision-making through non-linear transformation, and these features will help the network make correct decisions when facing multiple complex factors such as power, latency, and bandwidth.
[0177] The expressive power of the model can be enhanced by increasing the number of hidden layers and the number of neurons in each hidden layer.
[0178] For example, if multiple factors such as the power of the device and the latency of the network are considered to affect the priority of log uploading, the model needs to have stronger generalization ability. At this time, increasing the output of the hidden layer enables the next layer to combine multiple factors, and the weight of each factor gradually decreases to extract more abstract features.
[0179] In the log uploading scenario described in the embodiments of this application, the role of the hidden layer is to combine the device state and network conditions, and through layer-by-layer non-linear transformation, extract high-order features that affect the upload decision. For example, in the case of poor network conditions (such as high network latency but sufficient bandwidth), the system can choose to upload non-critical logs later; while in the case of both low network latency and low bandwidth, the system should give priority to uploading important logs.
[0180] The output layer generates the Q value Q(s, a) corresponding to each possible action a; in the framework of DQN, the number of neurons in the output layer is equal to the size of the action space, and each neuron corresponds to a specific adjustment action.
[0181] The Q value Q(s, a) of each action represents the long-term reward obtained by choosing this action in the current state s; at each time step, the system selects the optimal action according to the maximum Q value: a = arg max_a Q(s, a).
[0182] In this way, DQN can balance the immediate benefit and long-term benefit during log uploading. The system adaptively selects the optimal log uploading strategy under different network states.
[0183] Train the above-mentioned upload strategy dynamic adjustment model, and deploy the upload strategy dynamic adjustment model in the photovoltaic energy storage system; the key to optimizing the log reporting process lies in achieving seamless integration with the system environment and ensuring that the model can continuously adapt to changes in real-time network status and device status. To effectively improve the system performance, the deployment of the DQN model involves multiple key links, covering the entire process from data collection, model inference to policy execution.
[0184] Specifically, a data collection module, an upload policy dynamic adjustment model, and a policy execution module are deployed in the edge - side server of the target photovoltaic energy storage system. The data collection module is responsible for real - time monitoring of device operation parameters, such as network status, battery power, log generation rate, etc., and forms a state vector S=(s1, s2, ..., sn) as the input of the model. The upload policy dynamic adjustment model is the DQN model inference module, which is the core of the log policy dynamic adjustment. It runs the DQN network for policy inference. Inputting the current system state S, the model outputs the Q - value of the corresponding action, and selects the optimal adjustment action A = arg max_a Q(S, a) through the ε - greedy policy, that is, determines the next log upload action (such as stopping upload, reducing frequency, etc.). The policy execution module: executes the specific log upload policy according to the DQN inference result, such as selecting a suitable network path (Wi - Fi, cellular, etc.) or dynamically adjusting the upload time interval and compression ratio to optimize the log upload method.
[0185] Input the current state parameters of each target device into the pre - trained upload policy dynamic adjustment model, and the upload policy dynamic adjustment model processes the current state parameters to determine the target adjustment action that matches the current state parameters of the target device from the pre - configured set of adjustment actions.
[0186] Exemplarily, the following are some matching principles between the states of target devices and adjustment actions.
[0187] Network selection: When the network environment is poor or the device power is insufficient, a more reliable network such as 4G can be selected to ensure the stability of log upload.
[0188] Compression policy: When the network bandwidth is limited or the device power is insufficient, the log can be compressed to reduce the amount of data uploaded, thereby saving network costs and device power.
[0189] Priority policy: The main goal is to upload with priority, and the priority can be determined by the importance and real - time nature of the log to ensure the timely upload of high - value information.
[0190] Dynamic network adaptation: The DQN model can adaptively select the optimal network path or adjust the upload policy by continuously monitoring network parameters (such as latency, bandwidth, packet loss rate, etc.) to reduce unnecessary network switching and bandwidth waste.
[0191] Device energy efficiency management: When the device power is insufficient, the system saves energy consumption by reducing the upload frequency, lowering the data compression rate, etc., thereby extending the device's battery life.
[0192] In some embodiments, the DQN model can be deployed on edge devices or in the cloud for real-time computing to process a large amount of data and execute policies in real time.
[0193] In some embodiments, DQN model inference can utilize hardware acceleration technologies (such as GPUs and NPUs) to improve the computational efficiency of neural networks.
[0194] In some embodiments, in a large-scale photovoltaic energy storage system, the DQN model can be divided into multiple "sub-modules" to run in parallel. Each sub-module is responsible for the control of a specific area. Through the collaborative work of multiple sub-modules, it can adapt to the complexity and distributed characteristics of the entire system.
[0195] In some embodiments, in the method for uploading logs of the photovoltaic energy storage system, the number of the dynamic adjustment models of the upload policy is multiple; different dynamic adjustment models of the upload policy control different areas in the target photovoltaic energy storage system;
[0196] Correspondingly, inputting the current state parameters of each target device into the pre-trained dynamic adjustment model of the upload policy includes:
[0197] Inputting the current state parameters of each target device into the dynamic adjustment model of the upload policy that matches the area.
[0198] In step S103, based on the target adjustment action, dynamically adjust the target log upload policy for the target device.
[0199] The target log upload policy is the target log upload rule specifically executed by the target photovoltaic energy storage system.
[0200] For example, if the target adjustment action is the selection of the compression ratio, then based on the selected compression ratio, determine the compression ratio in the specific target log upload rule, and then compress the logs according to the specific compression ratio.
[0201] For example, if the target adjustment action is to adjust the upload frequency, then dynamically adjust the upload frequency in the target log upload policy for the target device.
[0202] The dynamic adjustment of the target log upload policy for the target device includes the adjustment of at least one of the upload frequency, network connection, and compression ratio in the target log upload policy.
[0203] In step S104, based on the target log upload policy of each target device in the target photovoltaic energy storage system, upload the logs of the target device to upload the logs of the photovoltaic energy storage system.
[0204] The logs of different devices in the photovoltaic energy storage system are uploaded according to their respective target log upload strategies, which may be uploaded simultaneously or not. By dynamically adjusting the target log upload strategy for the target device, the log upload strategies of different devices may already be different. For example, the upload frequency of the inverter is once every 2 minutes, but the upload frequency of the sensor is once a day, and so on.
[0205] After uploading the logs of the target device based on the target log upload strategy of each target device in the target photovoltaic energy storage system to upload the logs of the photovoltaic energy storage system, the log upload method further includes:
[0206] Recording the experience tuple for executing the target log upload strategy of the target device in the target photovoltaic energy storage system and storing the experience tuple in the experience replay pool; the experience tuple includes: the current state parameters of the target device in the target photovoltaic energy storage system, the target adjustment actions taken corresponding to the current state parameters of the target device, the immediate reward, and the state parameters of the next state of the target device;
[0207] When the preset batch update condition is satisfied, randomly extract batch data from the experience replay pool, and update the trained upload strategy dynamic adjustment model based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.
[0208] That is to say, during the operation of the photovoltaic energy storage system, the network and device states will change continuously, and the system must maintain sufficient adaptability; the batch update method can reduce the frequency of model updates, improve the update efficiency, and at the same time ensure that the model can learn diverse experiences.
[0209] Through continuous learning and optimization, the updated model is more suitable for the actual state of the target photovoltaic energy storage system in the selection and execution of the log upload strategy.
[0210] Based on the same inventive concept, an embodiment of the present application also provides a log upload device for a photovoltaic energy storage system corresponding to the log upload method of the photovoltaic energy storage system. Since the principle of solving problems by the device in the embodiment of the present application is similar to the above-mentioned log upload method of the photovoltaic energy storage system in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0211] Please refer to Figure 7 , Figure 7 which shows a schematic structural diagram of the log upload device for the photovoltaic energy storage system described in the embodiment of the present application, applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log upload device includes:
[0212] An acquisition module 701, configured to acquire logs and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating condition and network environment of the target device; the status parameters include: network latency, bandwidth, log generation rate, log priority, and the power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths;
[0213] A determination module 702, configured to input the current status parameters of each target device into a pre-trained upload policy dynamic adjustment model, and the upload policy dynamic adjustment model processes the current status parameters to determine a target adjustment action that matches the current status parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes multiple adjustment actions for adjusting the log upload policy of the target device;
[0214] An adjustment module 703, configured to dynamically adjust the target log upload policy for the target device based on the target adjustment action;
[0215] An upload module 704, configured to upload the logs of the target device based on the target log upload policy of each target device in the target photovoltaic energy storage system, so as to upload the logs of the photovoltaic energy storage system.
[0216] In some embodiments, the log upload device of the photovoltaic energy storage system further includes a training module:
[0217] The training module is configured to construct an upload policy dynamic adjustment model and configure an adjustment action set for the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions;
[0218] Train the upload policy dynamic adjustment model based on the sample status parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample status parameters of different sample devices and different adjustment actions; wherein, a loss function and a reward function are configured in the upload policy dynamic adjustment model; the reward function is used to calculate the immediate reward fed back to the sample photovoltaic energy storage system by the environment after selecting an adjustment action based on the sample status parameters of the sample device; the loss function is determined based on the feedback status and the immediate reward fed back to the photovoltaic energy storage system by the environment.
[0219] In some embodiments, in the log upload device of the photovoltaic energy storage system, when the training module trains the upload policy dynamic adjustment model based on the sample status parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample status parameters of different sample devices and different adjustment actions, it is specifically configured to:
[0220] At each time step, the upload policy dynamic adjustment model predicts the Q value of each adjustment action in the set of adjustment actions starting from the sample state parameters of the sample devices of the sample photovoltaic energy storage system;
[0221] The upload policy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action;
[0222] The sample photovoltaic energy storage system dynamically adjusts the sample log upload policy for the sample devices based on the sample adjustment action, and executes the sample log upload policy to obtain the feedback state and immediate reward fed back by the environment to the photovoltaic energy storage system; wherein, the immediate reward is calculated based on a reward function;
[0223] Based on the feedback state and immediate reward fed back by the environment to the photovoltaic energy storage system, determine the calculation result of the loss function of the upload policy dynamic adjustment model;
[0224] Optimize the upload policy dynamic adjustment model based on the calculation result of the loss function of the upload policy dynamic adjustment model until the upload policy dynamic adjustment model meets the preset training stop condition, and obtain a trained upload policy dynamic adjustment model.
[0225] In some embodiments, the training module in the log upload device of the photovoltaic energy storage system is further configured to: determine multiple optimization objectives for the log upload of the photovoltaic energy storage system; the multiple optimization objectives include: the success rate of log upload, the delay time, the bandwidth occupancy, the device power, and the log priority;
[0226] Fuse the log priority and the device power to determine a combined reward function for priority and power;
[0227] Based on the success rate of log upload, the delay time, the bandwidth occupancy, and the combined reward function for priority and power, construct a reward function for calculating the immediate reward.
[0228] In some embodiments, in the log upload device of the photovoltaic energy storage system, the reward rules of the reward function include: a log upload success reward rule, a log upload failure penalty rule, a bandwidth occupancy penalty rule, and a delay penalty rule;
[0229] The log upload success reward rule includes: after the log upload is successful, obtain a reward according to the state of the sample photovoltaic energy storage system; wherein, when the device power is more and / or the log priority is higher, a higher reward is obtained for a successful log upload;
[0230] The log upload failure penalty rule includes: after the log upload fails, obtain a negative reward according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward;
[0231] The bandwidth occupancy penalty rule includes: when the bandwidth occupancy of the sample photovoltaic energy storage system exceeds a preset bandwidth occupancy threshold, a negative reward is obtained;
[0232] The delay penalty rule includes: when the gap between the delay time and the preset maximum allowable value is smaller, the negative reward is higher.
[0233] In some embodiments, the training module in the log uploading device of the photovoltaic energy storage system is further configured to:
[0234] Based on the immediate reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action executed by the upload policy dynamic adjustment model for the current state of the photovoltaic energy storage system, a loss function is constructed.
[0235] In some embodiments, in the log uploading device of the photovoltaic energy storage system, the adjustment actions in the adjustment action set include: adjusting the upload frequency, selecting a network connection, and selecting a compression ratio.
[0236] In some embodiments, the log uploading device of the photovoltaic energy storage system further includes:
[0237] An update module, configured to upload the logs of the target device based on the target log upload policy of each target device in the target photovoltaic energy storage system. After uploading the logs of the photovoltaic energy storage system, the log uploading method further includes:
[0238] Recording the experience tuple of executing the target log upload policy of the target device in the target photovoltaic energy storage system, and storing the experience tuple in an experience replay pool; the experience tuple includes: the current state parameters of the target device in the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward, and the state parameters of the next state of the target device;
[0239] When a preset batch update condition is satisfied, batch data is randomly sampled from the experience replay pool, and the trained upload policy dynamic adjustment model is updated based on the sampled batch data to obtain an updated upload policy dynamic adjustment model.
[0240] In some embodiments, the number of the upload policy dynamic adjustment models in the log uploading device of the photovoltaic energy storage system is multiple; different upload policy dynamic adjustment models control different regions in the target photovoltaic energy storage system;
[0241] Correspondingly, when the determination module inputs the current state parameters of each target device into the pre-trained upload policy dynamic adjustment model, it is specifically configured to:
[0242] Input the current state parameters of each of the target devices into the upload policy dynamic adjustment model that matches the region.
[0243] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, which will not be elaborated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.
[0244] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0245] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0246] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.
[0247] The above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A log uploading method for a photovoltaic energy storage system, characterized in that: Applied to a target photovoltaic energy storage system, the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes: Obtaining the log and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating status and network environment of the target device; the status parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; Inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, the upload policy dynamic adjustment model processes the current state parameters, and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; Based on the target adjustment action, dynamically adjust the target log upload policy for the target device; Based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, uploading the log of the target device to upload the log of the photovoltaic energy storage system; The upload strategy dynamic adjustment model is trained based on the following method: Constructing an upload policy dynamic adjustment model, and configuring an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; The upload strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment.
2. The log uploading method of the photovoltaic energy storage system according to claim 1 is characterized in that: The uploading strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions, including: At each time step, the upload strategy dynamic adjustment model predicts the Q value of each adjustment action in the adjustment action set based on the sample state parameters of the sample equipment of the sample photovoltaic energy storage system; The upload strategy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action; The sample photovoltaic energy storage system dynamically adjusts the sample log upload strategy for the sample device based on the sample adjustment action, and executes the sample log upload strategy to obtain the feedback status and instant reward fed back to the photovoltaic energy storage system by the environment; wherein the instant reward is calculated based on the reward function; Determine the loss function calculation result of the upload strategy dynamic adjustment model based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment; The upload strategy dynamic adjustment model is optimized based on the loss function calculation result of the upload strategy dynamic adjustment model until the upload strategy dynamic adjustment model meets the preset training stop condition, thereby obtaining a trained upload strategy dynamic adjustment model.
3. The log uploading method of the photovoltaic energy storage system according to claim 1 or 2, characterized in that: The log uploading method further includes: Determine multiple optimization targets for log uploading of the photovoltaic energy storage system; the multiple optimization targets include: success rate of log uploading, delay time, bandwidth occupancy, device power, and log priority; Fusion the log priority and the device power to determine a fusion reward function regarding the priority and the power; Based on the success rate, delay time, bandwidth occupancy of the log upload and the fused reward function regarding priority and power, a reward function for calculating instant rewards is constructed.
4. The log uploading method of the photovoltaic energy storage system according to claim 3 is characterized in that: The reward rules of the reward function include: a reward rule for successful log upload, a penalty rule for failed log upload, a bandwidth occupation penalty rule, and a delay penalty rule; The reward rules for successful log upload include: after successfully uploading the log, a reward is obtained according to the status of the sample photovoltaic energy storage system; wherein, when the device has more power and / or the log priority is higher, a higher reward is obtained for successfully uploading the log; The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward; The bandwidth occupation penalty rule includes: when the bandwidth occupation of the sample photovoltaic energy storage system exceeds the preset bandwidth occupation threshold, a negative reward is obtained; The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowed value, the higher the negative reward.
5. The log uploading method of the photovoltaic energy storage system according to claim 1 or 2, characterized in that: The log uploading method further includes: Based on the instant reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action performed by the upload strategy dynamic adjustment model in the current state of the photovoltaic energy storage system, a loss function is constructed.
6. The log uploading method of the photovoltaic energy storage system according to claim 1 or 2, characterized in that: The adjustment actions in the adjustment action set include: adjusting upload frequency, selecting network connection, and selecting compression ratio.
7. The log uploading method of the photovoltaic energy storage system according to claim 1, characterized in that: After uploading the log of the target device based on the target log uploading strategy of each target device in the target photovoltaic energy storage system to upload the log of the photovoltaic energy storage system, the log uploading method further includes: Recording the experience tuple of the target log upload strategy of the target device executed in the target photovoltaic energy storage system, and storing the experience tuple in the experience playback pool; the experience tuple includes: the current state parameters of the target device of the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward and the state parameters of the next state of the target device; When the preset batch update condition is met, batch data is randomly extracted from the experience replay pool, and the trained upload strategy dynamic adjustment model is updated based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.
8. The log uploading method of the photovoltaic energy storage system according to claim 1, characterized in that: The number of the upload strategy dynamic adjustment models is multiple; different upload strategy dynamic adjustment models control different areas in the target photovoltaic energy storage system; Accordingly, the current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, including: The current state parameters of each target device are input into the region matching upload strategy dynamic adjustment model.
9. A log upload device for a photovoltaic energy storage system, characterized in that: Applied to a target photovoltaic energy storage system, the target photovoltaic energy storage system includes a plurality of target devices; the log uploading device includes: An acquisition module is used to acquire the log and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; A determination module, used for inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, wherein the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; An adjustment module, configured to dynamically adjust a target log upload policy for the target device based on the target adjustment action; An uploading module, configured to upload the log of each target device in the target photovoltaic energy storage system based on the target log upload strategy of the target device, so as to upload the log of the photovoltaic energy storage system; The upload strategy dynamic adjustment model is trained based on the following method: Constructing an upload policy dynamic adjustment model, and configuring an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; The upload strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment.
Citation Information
Patent Citations
Urban power distribution network multistage dynamic reconstruction method based on machine learning
CN114662982A
Power distribution network load state estimation method and system
CN119298076A