Log uploading method and device of photovoltaic energy storage system

By using the upload strategy dynamic adjustment model in the photovoltaic energy storage system and dynamically adjusting the log upload strategy according to the equipment and network status, the problems of delay and low success rate of log upload in large-scale photovoltaic energy storage systems are solved, and more efficient and stable log upload is achieved.

CN120016697AActive Publication Date: 2025-05-16GUANGZHOU RIMSEA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510503106.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Large-scale photovoltaic energy storage systems have problems with timely transmission and low success rates in log upload strategies, especially when broadband is limited and equipment scale is expanded.

Method used

The pre-trained upload strategy dynamic adjustment model is adopted to dynamically adjust the log upload strategy according to the current status parameters of each target device (including network latency, bandwidth, log generation rate, log priority and power), and optimize the upload timing and path.

Benefits of technology

By adaptively adjusting the log upload strategy, the real-time and stability of log data upload is significantly improved, and the overall performance of photovoltaic energy storage systems in complex network environments and large amounts of logs are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016697A_ABST
    Figure CN120016697A_ABST
Patent Text Reader

Abstract

The invention provides a log uploading method and device for a photovoltaic energy storage system, and the method is applied to a target photovoltaic energy storage system, and comprises the steps: obtaining a log and a current state parameter of each target device in the target photovoltaic energy storage system; inputting the current state parameter into a pre-trained uploading strategy dynamic adjustment model to process the current state parameter, and determining a target adjustment action matched with the current state parameter of the target equipment from a pre-configured adjustment action set; dynamically adjusting a target log uploading strategy for the target equipment based on the target adjustment action; based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded, and the log of the photovoltaic energy storage system is uploaded, so that the real-time performance and stability of log data uploading are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of new energy technology, and in particular to a log uploading method and device for a photovoltaic energy storage system. Background Art

[0002] With the rapid development of renewable energy technology, photovoltaic energy storage systems have gradually become a key component in the field of energy management. The stable and efficient operation of photovoltaic energy storage systems depends on real-time monitoring of system equipment status and performance data and remote data upload. However, the current photovoltaic energy storage system faces a series of technical challenges after the expansion of equipment scale and the increase of data volume, especially in log upload strategy.

[0003] Traditional log upload strategies use scheduled uploads, usually with fixed time intervals or event-driven upload mechanisms. Although this strategy is effective in small-scale systems, its limitations are becoming increasingly apparent in large-scale deployments. Due to limited bandwidth conditions, frequent upload operations, and an excessive number of logs, the system is unable to ensure timely transmission of log data, resulting in log upload failures or delays, which in turn affects the system's real-time monitoring and decision-making efficiency. Summary of the invention

[0004] In view of this, the purpose of the present application is to provide a log uploading method and device for a photovoltaic energy storage system, which can ensure the timely transmission and success rate of log uploading of a large-scale photovoltaic energy storage system.

[0005] A log uploading method for a photovoltaic energy storage system provided in an embodiment of the present application is applied to a target photovoltaic energy storage system, wherein the target photovoltaic energy storage system includes a plurality of target devices; the log uploading method includes: Obtaining the log and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating status and network environment of the target device; the status parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; Inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, the upload policy dynamic adjustment model processes the current state parameters, and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; Based on the target adjustment action, dynamically adjust the target log upload policy for the target device; Based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded to upload the log of the photovoltaic energy storage system.

[0006] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the upload strategy dynamic adjustment model is trained based on the following method: Constructing an upload policy dynamic adjustment model, and configuring an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; The upload strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment.

[0007] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the uploading strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions, including: At each time step, the upload strategy dynamic adjustment model predicts the Q value of each adjustment action in the adjustment action set based on the sample state parameters of the sample equipment of the sample photovoltaic energy storage system; The upload strategy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action; The sample photovoltaic energy storage system dynamically adjusts the sample log upload strategy for the sample device based on the sample adjustment action, and executes the sample log upload strategy to obtain the feedback status and instant reward fed back to the photovoltaic energy storage system by the environment; wherein the instant reward is calculated based on the reward function; Determine the loss function calculation result of the upload strategy dynamic adjustment model based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment; The upload strategy dynamic adjustment model is optimized based on the loss function calculation result of the upload strategy dynamic adjustment model until the upload strategy dynamic adjustment model meets the preset training stop condition, thereby obtaining a trained upload strategy dynamic adjustment model.

[0008] In some embodiments, the log uploading method of the photovoltaic energy storage system further includes: Determine multiple optimization targets for log uploading of the photovoltaic energy storage system; the multiple optimization targets include: success rate of log uploading, delay time, bandwidth occupancy, device power, and log priority; Fusion the log priority and the device power to determine a fusion reward function regarding the priority and the power; Based on the success rate, delay time, bandwidth occupancy of the log upload and the fused reward function regarding priority and power, a reward function for calculating instant rewards is constructed.

[0009] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the reward rules of the reward function include: a reward rule for successful log uploading, a penalty rule for failed log uploading, a bandwidth occupation penalty rule, and a delay penalty rule; The reward rules for successful log upload include: after successfully uploading the log, a reward is obtained according to the status of the sample photovoltaic energy storage system; wherein, when the device has more power and / or the log priority is higher, a higher reward is obtained for successfully uploading the log; The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward; The bandwidth occupation penalty rule includes: when the bandwidth occupation of the sample photovoltaic energy storage system exceeds the preset bandwidth occupation threshold, a negative reward is obtained; The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowed value, the higher the negative reward.

[0010] In some embodiments, the log uploading method of the photovoltaic energy storage system further includes: Based on the instant reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action performed by the upload strategy dynamic adjustment model in the current state of the photovoltaic energy storage system, a loss function is constructed.

[0011] In some embodiments, in the log uploading method for the photovoltaic energy storage system, the adjustment actions in the adjustment action set include: adjusting the upload frequency, selecting a network connection, and selecting a compression ratio.

[0012] In some embodiments, in the log uploading method of the photovoltaic energy storage system, after uploading the log of the target device based on the target log uploading strategy of each target device in the target photovoltaic energy storage system to upload the log of the photovoltaic energy storage system, the log uploading method further includes: Recording the experience tuple of the target log upload strategy of the target device executed in the target photovoltaic energy storage system, and storing the experience tuple in the experience playback pool; the experience tuple includes: the current state parameters of the target device of the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward and the state parameters of the next state of the target device; When the preset batch update condition is met, batch data is randomly extracted from the experience replay pool, and the trained upload strategy dynamic adjustment model is updated based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.

[0013] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the number of the upload strategy dynamic adjustment models is multiple; different upload strategy dynamic adjustment models control different areas in the target photovoltaic energy storage system; Accordingly, the current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, including: The current state parameters of each target device are input into the region matching upload strategy dynamic adjustment model.

[0014] In some embodiments, a log uploading device for a photovoltaic energy storage system is further provided, characterized in that it is applied to a target photovoltaic energy storage system, wherein the target photovoltaic energy storage system includes a plurality of target devices; the log uploading device includes: An acquisition module is used to acquire the log and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; A determination module, used for inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, wherein the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; An adjustment module, configured to dynamically adjust a target log upload policy for the target device based on the target adjustment action; The uploading module is used to upload the log of each target device in the target photovoltaic energy storage system based on the target log uploading strategy of the target device, so as to upload the log of the photovoltaic energy storage system.

[0015] In an embodiment of the present application, a method and device for uploading logs of a photovoltaic energy storage system are provided, which are applied to a target photovoltaic energy storage system, wherein the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes: obtaining logs and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; The current state parameters of each target device are input into a pre-trained dynamic adjustment model for upload strategy. The dynamic adjustment model for upload strategy processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set. The adjustment action set includes a plurality of adjustment actions for adjusting the log upload strategy of the target device. Based on the target adjustment action, the target log upload strategy for the target device is dynamically adjusted. Based on the target log upload strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded to upload the log of the photovoltaic energy storage system, so that the log upload strategy of a single device can be adaptively adjusted according to factors in multiple dimensions such as device status, network bandwidth, delay, log importance, device power, and log generation rate in the photovoltaic energy storage system, thereby optimizing the overall log upload strategy of the system, thereby greatly improving the real-time and stability of log data upload, so that the photovoltaic energy storage system significantly improves the overall performance when dealing with complex network environments and a large number of logs. The system provides a more flexible and intelligent log management solution, which significantly improves the real-time, efficiency and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A flow chart showing a log uploading method for a photovoltaic energy storage system according to an embodiment of the present application is shown; Figure 2 A flow chart of a method for training the upload strategy dynamic adjustment model according to an embodiment of the present application is shown; Figure 3 The schematic diagram of the framework design of the upload strategy dynamic adjustment model is shown; Figure 4 A schematic diagram showing the state changes of the upload strategy dynamic adjustment model described in an embodiment of the present application is shown; Figure 5 A flow chart showing another method for uploading logs of a photovoltaic energy storage system according to an embodiment of the present application is shown; Figure 6 A schematic diagram of the structure of the upload strategy dynamic adjustment model is shown; Figure 7 A structural schematic diagram of a log uploading device for a photovoltaic energy storage system according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0018] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.

[0019] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0020] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0021] With the rapid development of renewable energy technology, photovoltaic energy storage systems have gradually become a key component in the field of energy management. The stable and efficient operation of photovoltaic energy storage systems depends on real-time monitoring of system equipment status and performance data and remote data upload. However, the current photovoltaic energy storage system faces a series of technical challenges after the expansion of equipment scale and the increase of data volume, especially in log upload strategy.

[0022] Traditional log upload strategies use scheduled uploads, usually with fixed time intervals or event-driven upload mechanisms. Although this strategy is effective in small-scale systems, its limitations are becoming increasingly apparent in large-scale deployments. Due to limited bandwidth conditions, frequent upload operations, and an excessive number of logs, the system is unable to ensure timely transmission of log data, resulting in log upload failures or delays, which in turn affects the system's real-time monitoring and decision-making efficiency.

[0023] Specifically, first, scheduled uploading can easily cause data congestion under conditions of limited network bandwidth, especially during peak load periods, when the system can hardly ensure timely transmission of log data; second, too frequent or unnecessary uploading operations will increase the network burden and increase upload delays, which in turn will affect the system's real-time monitoring and decision-making efficiency; in addition, when the photovoltaic energy storage system has multiple network connection paths at the same time (such as Wi-Fi, cellular networks, Bluetooth, etc.), how to intelligently switch between different network environments to avoid unnecessary network conversions has also become an urgent problem to be solved.

[0024] Based on this, how to intelligently manage the log upload strategy and improve the real-time performance of data upload has become a core issue in improving the overall performance of photovoltaic energy storage systems.

[0025] Based on this, in an embodiment of the present application, a method and device for uploading logs of a photovoltaic energy storage system are provided, which are applied to a target photovoltaic energy storage system, wherein the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes: obtaining logs and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters characterize the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; The current state parameters of each target device are input into a pre-trained dynamic adjustment model for upload strategy. The dynamic adjustment model for upload strategy processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set. The adjustment action set includes a plurality of adjustment actions for adjusting the log upload strategy of the target device. Based on the target adjustment action, the target log upload strategy for the target device is dynamically adjusted. Based on the target log upload strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded to upload the log of the photovoltaic energy storage system, so that the log upload strategy of a single device can be adaptively adjusted according to factors in multiple dimensions such as device status, network bandwidth, delay, log importance, device power, and log generation rate in the photovoltaic energy storage system, thereby optimizing the overall log upload strategy of the system, thereby greatly improving the real-time and stability of log data upload, so that the photovoltaic energy storage system significantly improves the overall performance when dealing with complex network environments and a large number of logs. The system provides a more flexible and intelligent log management solution, which significantly improves the real-time, efficiency and reliability of the system.

[0026] Please refer to Figure 1 , Figure 1 A flow chart of a log uploading method for a photovoltaic energy storage system according to an embodiment of the present application is shown; Figure 1 As shown, the log uploading method of the photovoltaic energy storage system is applied to a target photovoltaic energy storage system, and the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes the following steps S101-S104: S101, obtaining the log and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; S102, inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, wherein the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; S103, dynamically adjusting the target log upload policy for the target device based on the target adjustment action; S104: Based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, upload the log of the target device to upload the log of the photovoltaic energy storage system.

[0027] In large-scale photovoltaic energy storage systems, a variety of key target devices will generate logs to record their operating status, performance parameters, fault alarms and other information; specifically, the key target devices include photovoltaic panels, battery energy storage systems, inverters, charge and discharge controllers, grid connectors, and other auxiliary equipment; some key target devices have multiple numbers, such as photovoltaic panels, inverters, charge and discharge controllers, etc.; other auxiliary equipment includes environmental monitoring equipment (such as temperature sensors, humidity sensors, etc.), safety protection equipment (such as over-current and over-voltage protection devices, etc.), etc.

[0028] Based on this, large-scale photovoltaic energy storage systems involve a large number of devices and high power, so the amount of data generated is also larger, and the log generation is more complicated. Since it is necessary to monitor the status and performance of the energy storage equipment and the operation of the power system in real time, the log contains a large amount of detailed parameter information and status records.

[0029] Small-scale photovoltaic energy storage systems are relatively simple, with smaller data volumes and easier log generation. These systems usually only need to record basic equipment status and power parameters, and the log content is relatively simple and clear.

[0030] Therefore, it is necessary to optimize the log upload strategy in large-scale photovoltaic energy storage systems. By adaptively adjusting the log upload strategy, the upload efficiency of the system can be improved, the network bandwidth usage can be reduced, and the upload delay can be reduced.

[0031] In an embodiment of the present application, a log upload decision engine is deployed in an edge-side server of a photovoltaic energy storage system to sense the status of a target photovoltaic energy storage system in real time, determine a log upload strategy, and execute a decision.

[0032] In some embodiments, the log upload decision engine is built based on a DQN model.

[0033] In step S101, the log and current state parameters of each target device in the target photovoltaic energy storage system are obtained; the current state parameters characterize the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths.

[0034] Specifically, the log upload decision engine monitors the state parameters of the target device in real time through the data acquisition module, such as network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system, and forms a state vector S = (s1, s2, ..., sn) as the upload strategy dynamic adjustment model of the model.

[0035] The log priority represents the importance of the log.

[0036] In some embodiments, with respect to network latency and bandwidth, network monitoring tools (such as Ping, Traceroute, network performance testing tools, etc.) are used to measure the network latency between each target device and the log processing platform, and to monitor the bandwidth usage of the network interface to ensure that data transmission is not hindered by insufficient bandwidth.

[0037] For the log generation rate, monitor the log generation rate of each device through log collection tools or custom scripts to evaluate the efficiency of log processing and potential performance bottlenecks.

[0038] Regarding log priority, when a log is generated, the log priority is set according to the importance of the log content; in some embodiments, this can be achieved through configuration in the log management system to ensure that high-priority logs are uploaded first.

[0039] For the power of the photovoltaic energy storage system, use the monitoring interface or dedicated sensors of the energy storage system to obtain real-time power information, including remaining power, charging / discharging status, etc.

[0040] The multiple network connection paths include Bluetooth connection, WiFi connection, cellular network, etc.

[0041] First, scheduled uploads are prone to data congestion under conditions of limited network bandwidth, especially during peak load periods, when the system cannot ensure timely transmission of log data. Second, too frequent or unnecessary uploads will increase the network burden and increase upload delays, which in turn affects the system's real-time monitoring and decision-making efficiency. In addition, when the photovoltaic energy storage system has multiple network connection paths (such as Wi-Fi, cellular networks, Bluetooth, etc.) at the same time, how to intelligently switch between different network environments to avoid unnecessary network conversions has also become an urgent problem to be solved.

[0042] Therefore, the device operating status such as network delay, bandwidth, log generation rate, etc. is collected to intelligently adjust the log upload strategy of the device.

[0043] In step S102, the current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, and the upload strategy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes multiple adjustment actions for adjusting the log upload policy of the target device.

[0044] The adjustment actions in the adjustment action set include: adjusting upload frequency, selecting network connection, and selecting compression ratio.

[0045] That is to say, in the photovoltaic energy storage system, the upload strategy dynamic adjustment model can adaptively adjust the log upload strategy according to factors such as device status, network bandwidth, delay, and log importance. It not only optimizes the upload timing and path, but also reduces bandwidth usage by intelligently compressing data, thereby greatly improving the real-time and stability of data upload.

[0046] In some embodiments, the upload strategy dynamic adjustment model is implemented based on a deep Q-network (DQN). A deep Q-network (DQN) is a specific implementation of deep reinforcement learning (DRL); DQN combines Q learning in reinforcement learning with deep neural networks, and can handle complex, high-dimensional state spaces, especially suitable for dynamic network environments.

[0047] Please refer to Figure 2 In some embodiments, the upload strategy dynamic adjustment model is trained based on the following steps S201-S202: S201, constructing an upload policy dynamic adjustment model, and configuring an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; S202. Train the upload strategy dynamic adjustment model based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and the instant reward fed back to the photovoltaic energy storage system by the environment.

[0048] That is to say, the upload strategy dynamic adjustment model learns the overall impact of all possible adjustment actions of different sample devices under different operating conditions on the photovoltaic energy storage system, thereby comprehensively considering multiple components and strategies for debugging, ensuring the performance and generalization ability of the upload strategy dynamic adjustment model, and providing strong support for the log upload task of photovoltaic components.

[0049] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the uploading strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions, including: At each time step, the upload strategy dynamic adjustment model predicts the Q value of each adjustment action in the adjustment action set based on the sample state parameters of the sample equipment of the sample photovoltaic energy storage system; The upload strategy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action; The sample photovoltaic energy storage system dynamically adjusts the sample log upload strategy for the sample device based on the sample adjustment action, and executes the sample log upload strategy to obtain the feedback status and instant reward fed back to the photovoltaic energy storage system by the environment; wherein the instant reward is calculated based on the reward function; Determine the loss function calculation result of the upload strategy dynamic adjustment model based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment; The upload strategy dynamic adjustment model is optimized based on the loss function calculation result of the upload strategy dynamic adjustment model until the upload strategy dynamic adjustment model meets the preset training stop condition, thereby obtaining a trained upload strategy dynamic adjustment model.

[0050] For details, please refer to Figure 3 ,The framework design of the upload strategy dynamic adjustment model is as follows.

[0051] Environment: includes the operating status of the photovoltaic energy storage system, network connection, battery power of the device, log generation rate, importance of logs, etc. It reflects the current status and condition of the photovoltaic energy storage system and is used for intelligent agents to learn and make decisions.

[0052] State (S): The state is a vector consisting of multiple parameters, such as current network delay (delay), available bandwidth (bandwidth), device power (battery_level), log generation rate (log_rate), and log priority (log_priority).

[0053] The state vector can be expressed as: s = [delay, bandwidth, battery_level, log_rate,log_priority].

[0054] Action (A): The action set includes adjusting the upload frequency, selecting the network connection (such as Wi-Fi, Bluetooth or cellular network), selecting the compression ratio, etc.

[0055] Actions can be defined as: A = {a1, a2, ... ai, ... an}, where each ai represents a specific action.

[0056] Reward (R): Reward is an indicator used to evaluate the behavior of an agent. For a system, rewards are usually related to factors such as upload success rate, bandwidth usage, latency, and power consumption. For example, when the system successfully uploads logs and reduces bandwidth usage, a positive reward is given; while upload failure or excessive bandwidth consumption will result in a negative reward.

[0057] like Figure 3 As shown in the figure, when the agent selects action A, the system executes action A, and the environment feeds back the reward R and state S of action A, thereby training the agent.

[0058] DQN (Deep Q-Network) approximates the Q-value function in traditional Q learning by introducing a deep neural network (DNN), enabling effective decision optimization in complex or high-dimensional state spaces. The specific process includes Q-value update, loss function calculation, experience replay, and interaction between the target network and the policy network.

[0059] The following is a detailed technical description of the process.

[0060] Please refer to Figure 4 , Figure 4 A schematic diagram showing the state changes of the upload strategy dynamic adjustment model described in the embodiment of the present application is shown; Figure 4As shown in the figure, at each time step, the system starts from the current state s(t) and predicts the Q value of all possible current adjustment actions a(t) through a deep neural network, that is, Q(s, a). The model selects a specific adjustment action based on the Q value, and enters the next state s(t+1) after executing the adjustment action; starting from state s(t+1), it executes adjustment action a(t+1) and enters the next state s(t+2).

[0061] In the embodiment of the present application, the specific selection process follows the ε-greedy strategy: a_t = arg max_a Q(s_t, a; θ); Q(s,a) represents the expected return of performing adjustment action a in state s; θ is the parameter of the deep neural network, and a_t represents the optimal action in the corresponding state s_t when the Q value is maximized.

[0062] To avoid falling into the local optimal solution in some cases, the agent will also randomly select actions (exploration) with probability ε. a The optimal solution action.

[0063] Arg max a: The value of the corresponding variable when the following formula reaches its maximum value.

[0064] Q(s_t, a; θ): In the initial state, the initial action a will be randomly selected within a certain range. Here, a takes the maximum value in the range in the initial state, and s_t is the state corresponding to the value of a. After the action is executed, the environment feeds back the new state s_t+1 and the immediate reward r_t to the system; the goal of the system is to maximize the cumulative reward, so the Q value of the current state and action needs to be updated.

[0065] The update of Q value is based on Bellman equation, which expresses the relationship between the Q value of the current state and the maximum Q value of the next state. Its update formula is as follows: Q(s_t, a_t) = Q(s_t, a_t) + α * [r_t + γ * max_a' Q(s_t+1, a'; θ') -Q(s_t, a_t)]; Among them: Q(s_t, a_t) represents the Q value of action a in state s (i.e. current state) at time t, and a_t represents the optimal action corresponding to the current state; max_a' Q(s_t+1, a'; θ') represents the Q value of the optimal action a' corresponding to state s (i.e. next state) at time t+1, until the Q value converges to the maximum value; a' represents the adjustment action performed in the next step; θ' is the parameter of the target network, which is updated at a certain time interval with the policy network parameters; α is the learning rate, which controls the step size of each update; γ is the discount factor, which is used to weigh short-term rewards and long-term rewards. Its value is between 0 and 1, and is usually set to close to 1 (e.g. 0.99) to emphasize the importance of long-term rewards; r_t is the immediate reward obtained after executing the adjustment action a_t at time step t.

[0066] The specific implementation of the ε-greedy strategy is as follows A1-A3: A1: Set the ε value: ε is a value between 0 and 1, indicating the probability of randomly selecting an action. The larger the ε value, the higher the degree of exploration, but the lower the degree of utilization; the smaller the ε value, the lower the degree of exploration, but the higher the degree of utilization.

[0067] A2: Generate random numbers: At each decision point, generate a random number between 0 and 1.

[0068] A3: Judgment: If the random number is less than or equal to ε, randomly select an action; otherwise, select the action with the largest Q value.

[0069] Through the ε-greedy strategy, the agent selects a random action with a certain probability, even if the action may not be good. This randomness helps the agent discover new states and actions, thereby finding a better strategy; at the same time, the agent selects the action that is currently considered to be the best (that is, the action with the largest Q value) with a higher probability. This behavior of using known knowledge helps the agent quickly obtain higher returns.

[0070] In order to enable DQN to continuously update, DQN guides the update of network parameters by minimizing the difference between the predicted Q value and the target Q value.

[0071] In the log uploading method of the photovoltaic energy storage system described in the embodiment of the present application, the log uploading method further includes: Based on the instant reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action performed by the upload strategy dynamic adjustment model in the current state of the photovoltaic energy storage system, a loss function is constructed.

[0072] Specifically, the loss function L(θ) is defined as follows: L(θ) = E[(r_t + γ * max_a' Q(s_t+1, a'; θ') - Q(s_t, a_t; θ))^2]; Where: r_t is the immediate reward obtained after performing action a_t at time step t; max_a' Q(s_t+1, a';θ') represents the maximum future reward that can be obtained in the next state s_t+1; γ is a discount factor used to balance short-term rewards and long-term rewards; θ is the parameter of the current policy network; θ' is the parameter of the target network, which remains unchanged for a period of time to ensure training stability; E[ ] is the expectation operator used to calculate the average error of the network on different samples.

[0073] r_t is the immediate reward, which reflects the effect of the log upload action at that time step and is related to transmission delay, bandwidth utilization, and the importance of the log.

[0074] γ * max_a' Q(s', a'; θ') represents the maximum expected return in the next state and is used to estimate the long-term benefit. Q(s, a; θ))^2 is the predicted Q value of the neural network performing action a in state s; the square is mainly to measure the prediction error and ensure that the network can minimize this difference through gradient descent. Specifically, the square term in the loss function is to punish large error values, so that the model pays more attention to samples with large errors when updating, gradually adjusts the network weights, and approaches the true Q value; therefore, it means that at each step, it tries to narrow the gap between the network's predicted Q value and the actual target Q value.

[0075] The minimization of the loss function is achieved through the Stochastic Gradient Descent (SGD) algorithm. Through back propagation, the gradient of the loss function is propagated to each layer of the neural network, gradually updating the weight parameters of the network to approach the optimal Q value function.

[0076] In the log reporting scenario of the photovoltaic energy storage system, the reward function is a key factor in reinforcement learning to guide DQN (DeepQ-Network) optimization decisions. The design of the reward function directly affects how the system adjusts the upload strategy under different network conditions, device status, and log generation rate. To ensure the timeliness, effectiveness, and resource conservation of log uploads, the reward function must be able to reflect multi-dimensional performance indicators.

[0077] The reward function design principles include: upload success rate, bandwidth utilization efficiency, log importance, energy consumption management, and delay minimization.

[0078] Upload success rate: Successful log upload is the main goal of the system, so successful log uploads should be rewarded positively, and failures should be rewarded negatively.

[0079] Bandwidth usage efficiency: During the upload process, if the upload task can be completed with lower bandwidth usage, the system should receive additional rewards; otherwise, excessive bandwidth usage will result in penalties.

[0080] Log importance: Different logs have different priorities. Important logs should be uploaded first, so the reward function should consider the importance weight of the log.

[0081] Energy management: The energy consumption of the device during the upload process is also an important consideration, especially when the device battery is low. If a low-energy path is chosen when the battery is low (such as choosing a low-power network), positive incentives should be given; conversely, high-energy operations will be punished.

[0082] Minimize latency: Timeliness of upload is also important. The system should be rewarded if logs can be uploaded within a short latency, and penalized if the latency is too high.

[0083] Based on this, in some embodiments, please refer to Figure 5 The log uploading method of the photovoltaic energy storage system further includes the following steps S501-S503: S501, determining multiple optimization targets for log uploading of the photovoltaic energy storage system; the multiple optimization targets include: success rate of log uploading, delay time, bandwidth occupancy, device power, and log priority; S502, integrating the log priority and the device power, and determining a fusion reward function for the priority and the power; S503: Based on the success rate of the log upload, the delay time, the bandwidth occupancy and the fusion reward function regarding the priority and the power, a reward function for calculating the instant reward is constructed.

[0084] The reward rules of the reward function include: a reward rule for successful log upload, a penalty rule for failed log upload, a bandwidth occupation penalty rule, and a delay penalty rule; The reward rules for successful log upload include: after successfully uploading the log, a reward is obtained according to the status of the sample photovoltaic energy storage system; wherein, when the device has more power and / or the log priority is higher, a higher reward is obtained for successfully uploading the log; The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward; The bandwidth occupation penalty rule includes: when the bandwidth occupation of the sample photovoltaic energy storage system exceeds the preset bandwidth occupation threshold, a negative reward is obtained; The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowed value, the higher the negative reward.

[0085] In an optional embodiment, the reward function R is specifically: R = R_success * f(log(P_energy)) * (L_max - L) / (L_max - L_avg) - λ_1 * (B_used / B_threshold) - λ_2 * (L_delay / L_max).

[0086] R_success represents the reward for successful log upload; f(log(P_energy)) is a function of log importance and device power, which is used to represent the corresponding comprehensive score under different importance and power conditions; for example, when the log importance is high and the power is sufficient, the system will give priority to successful upload and give a higher reward; the function can be designed as: f(log(P_energy)) = log(1 + 1 / (1 + e^(-k*(P_energy-P_threshold)))); where k is a parameter for adjusting the steepness of the curve, P_threshold is the power threshold (such as set to 5%), and P_energy is the device power; L_max represents the maximum allowed delay time; L_avg represents the average delay time for log upload; B_used represents the bandwidth used during the upload process; B_threshold represents the bandwidth usage upper limit set by the system; L_delay represents the delay time for upload; λ_1 is the first trade-off parameter, which controls the impact of bandwidth occupancy on the total reward; λ_2 is the second trade-off parameter, which controls the impact of network delay on the total reward.

[0087] The reward function can balance multiple objectives. It takes into account multiple factors such as the success rate of log upload, latency, bandwidth usage, and device power, so that the system can find a balance between these objectives. It can also adapt to different scenarios: by adjusting the parameters λ_1 and λ_2, the system's sensitivity to bandwidth and latency can be adjusted according to different network environments and system requirements. It can also encourage system learning. The design of the reward function can encourage the system to learn the optimal upload strategy under different conditions, thereby improving the overall performance of the system.

[0088] The detailed explanation of the reward function is as follows.

[0089] Reward for successful log upload: After successfully uploading a log, the system will calculate the actual reward value based on the importance of the log I_log and the current device power P_energy. When the power is sufficient, the successful upload of important logs should receive a higher reward; if the power is insufficient, the system will give priority to low-energy networks, appropriately reduce the upload frequency or compress the log.

[0090] Penalty for log upload failure: Upload failure will not only delay the upload of important log information, but also consume unnecessary network resources. Therefore, the system needs to impose penalties based on the failure situation, especially when the log is of high importance, the penalty should be increased.

[0091] Bandwidth usage penalty: A bandwidth usage rate B_usedB_threshold that is too high will affect the overall performance of the network. Therefore, when the bandwidth usage exceeds a certain threshold, the system will face additional negative rewards.

[0092] Delay penalty: When the upload delay L_delay is close to the maximum value L_max allowed by the system, the system will be penalized, which helps to encourage the model to choose a low-latency network or upload at the right time.

[0093] In order to improve the efficiency and stability of learning, DQN introduces an experience replay mechanism. Each interaction between the system and the environment (state, action, reward, next state s_t+1) will be stored in an experience replay pool (Replay Buffer). The data in the experience replay pool will be used for training. The specific steps are as follows: During each training, the system randomly extracts a mini-batch of data from the experience replay pool to calculate the loss and update the network parameters; this can break the time correlation of the data and improve the stability of training.

[0094] Randomly extracted empirical data helps avoid overfitting to certain specific state sequences, thereby improving the generalization ability of learning.

[0095] The size of the experience replay pool is usually set to a fixed value, and the system will periodically eliminate the oldest data to make room for new experience.

[0096] In order to improve the stability of Q value learning, DQN adopts a dual network structure of target network and policy network. The target network is used to calculate the target Q value, and its parameter θ' will remain unchanged for a period of time, while the policy network is updated according to new data during each training.

[0097] The parameter update process between the target network and the policy network is as follows: θ' ← θ; The target network is synchronized with the policy network every N steps, that is, the weight θ of the policy network is assigned to the weight θ' of the target network; this mechanism avoids instability during training by reducing the frequent changes of the target value, thereby improving the convergence speed of the model.

[0098] Based on this, the training process of the upload strategy dynamic adjustment model is as follows.

[0099] First, initialize the weight parameters θ of the policy network and initialize the weight parameters θ'=θ of the target network.

[0100] Construct an experience replay pool D and store the data of each interaction with the environment in the pool.

[0101] At each time step: the system uses the ε-greedy strategy to select the adjustment action a_t based on the current state s_t; executes the adjustment action a_t and obtains the new state s_t+1 and immediate reward r_t; stores (s_t, a_t, r_t, s_t+1) in the experience replay pool D; randomly extracts a batch of data from the experience replay pool D and calculates the target Q value; minimizes the loss function L(θ) and updates the parameters θ of the policy network by gradient descent; every N steps, synchronizes the weights θ of the policy network to the target network θ'.

[0102] Please refer to the following Figure 6 , the upload strategy dynamic adjustment model includes: an input layer, a hidden layer and an output layer.

[0103] The input layer receives the current status parameters of the target device in the system. This vector intuitively reflects the current system operation status and network environment.

[0104] For the log uploading problem in photovoltaic energy storage systems, the input layer can include the following key parameters: Network Latency: It indicates the time delay for uploading data to the cloud or server, which is used to evaluate the quality of network transmission. Bandwidth: The currently available network bandwidth, which affects the speed and transmission efficiency of log uploads. * Device Power Level: The remaining power of the device, which is used to dynamically balance log uploads and power consumption to avoid frequent uploads when the power is low.

[0105] Log Generation Rate: The speed at which the system generates logs determines the amount of logs that need to be uploaded. Log Importance: The priority of different logs, which is used to ensure that key logs are uploaded first when network resources are tight. These parameters can form the input vector s∈R^n, where n represents the state dimension.

[0106] These parameters jointly influence the system's decision on log upload and serve as the input of the neural network.

[0107] The hidden layer uses a fully connected network structure to map the data of the input layer to a high-dimensional space, thereby capturing the complex nonlinear relationship between the original data; the hidden layer of DQN usually uses the ReLU activation function (Rectified Linear Unit), which can enhance the nonlinear expression ability of the model. The formula is as follows: h = ReLU(Ws + b); where: h is the output of the hidden layer; Ws is the weight matrix of the hidden layer, which represents the connection strength between the input layer and the hidden layer; b is the bias term; the ReLU activation function is defined as ReLU(x) = max(0, x), which introduces nonlinear characteristics to avoid overfitting the model to linear decisions. Specifically, the function of Ws (weight matrix) is as follows: 1. Control the transmission strength of the signal: The weight matrix determines the degree of influence of each input signal (output from the previous layer) on the output of the hidden layer. During the training process, the network automatically adjusts these weights according to the pattern of the input data, making the output closer and closer to the target value.

[0108] 2. Representation of learning data: The hidden layer is a key part of the neural network, which is responsible for mapping the input data from the original space to a high-dimensional feature space and capturing the complex nonlinear relationship between the data. Through the weight matrices of different layers, the network is able to extract more and more abstract features until the final output layer (for example, Q value prediction) obtains accurate results.

[0109] 3. Model expressiveness: By adjusting Ws, the network can learn the complex relationships between different inputs. For example, in the log upload scenario, the hidden layer extracts high-order features that affect decision-making through nonlinear transformation, and these features will help the network make correct decisions when faced with complex factors such as power, latency, and bandwidth.

[0110] The expressive power of the model can be improved by increasing the number of hidden layers and the number of neurons in each hidden layer.

[0111] For example, if multiple factors such as device power and network latency are considered to affect the priority of log upload, the model needs to have stronger generalization capabilities. In this case, the output of the hidden layer is increased so that the next layer can combine multiple factors, and the weight of each factor is gradually reduced to extract more abstract features.

[0112] In the log upload scenario described in the embodiment of the present application, the role of the hidden layer is to combine the device status and network conditions, and extract high-order features that affect the upload decision through layer-by-layer nonlinear transformation. For example, in the case of poor network conditions (such as high network latency but sufficient bandwidth), the system can choose to upload non-critical logs later; while in the case of low network latency and bandwidth, the system should give priority to uploading important logs.

[0113] The output layer generates the Q value Q(s, a) corresponding to each possible action a; in the framework of DQN, the number of neurons in the output layer is equal to the size of the action space, and each neuron corresponds to a specific adjustment action.

[0114] The Q value Q(s, a) of each action represents the long-term reward of choosing this action under the current state s; at each time step, the system selects the optimal action based on the maximum Q value: a = arg max_a Q(s, a). In this way, DQN can balance the immediate benefits and long-term benefits of log uploading. The system adaptively selects the optimal log uploading strategy under different network conditions.

[0115] Train the upload strategy dynamic adjustment model and deploy it in the photovoltaic energy storage system; the key to optimizing the log reporting process is to achieve seamless integration with the system environment and ensure that the model can continuously adapt to changes in real-time network status and device status. In order to effectively improve system performance, the deployment of the DQN model involves multiple key links, covering the entire process from data collection, model reasoning to strategy execution.

[0116] Specifically, a data acquisition module, an upload strategy dynamic adjustment model and a strategy execution module are deployed in the edge server of the target photovoltaic energy storage system; the data acquisition module is responsible for real-time monitoring of equipment operating parameters, such as network status, battery power, log generation rate, etc., to form a state vector S = (s1, s2, ..., sn) as the input of the model; the upload strategy dynamic adjustment model is the DQN model reasoning module, which is the core of the dynamic adjustment of the log strategy and runs the DQN network for strategy reasoning; the current system state S is input, the model outputs the Q value of the corresponding action, and selects the optimal adjustment action A = arg max_a Q(S, a) through the ε-greedy strategy, that is, determines the next log upload action (such as stopping uploading, reducing frequency, etc.); strategy execution module: executes a specific log upload strategy according to the DQN reasoning result, such as selecting a suitable network path (Wi-Fi, cellular, etc.) or dynamically adjusting the upload time interval and compression ratio to optimize the log upload method.

[0117] The current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, and the upload strategy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set.

[0118] Exemplarily, the following are some matching principles between the states of target devices and adjustment actions.

[0119] Network selection: When the network environment is poor or the device battery is low, you can choose a more reliable network such as 4G to ensure the stability of log uploading.

[0120] Compression strategy: When network bandwidth is limited or device power is low, logs can be compressed to reduce the amount of uploaded data, thereby saving network costs and device power.

[0121] Priority strategy: The main goal is to prioritize uploading, and the priority can be determined by the importance and real-time nature of the log to ensure timely uploading of high-value information.

[0122] Dynamic network adaptation: The DQN model can adaptively select the optimal network path or adjust the upload strategy by continuously monitoring network parameters (such as latency, bandwidth, packet loss rate, etc.) to reduce unnecessary network switching and bandwidth waste.

[0123] Equipment energy efficiency management: When the device is low on power, the system saves energy by reducing the upload frequency, lowering the data compression rate, etc., thereby extending the battery life of the device.

[0124] In some embodiments, the DQN model can be deployed on edge devices or in the cloud for real-time computing to process large amounts of data and execute policies in real time.

[0125] In some embodiments, DQN model reasoning can use hardware acceleration technology (such as GPU, NPU) to speed up the computational efficiency of the neural network.

[0126] In some embodiments, in a large-scale photovoltaic energy storage system, the DQN model can be divided into multiple "sub-modules" that run in parallel, with each sub-module responsible for the control of a specific area. Through the collaborative work of multiple sub-modules, the complexity and distributed characteristics of the entire system can be adapted.

[0127] In some embodiments, in the log uploading method of the photovoltaic energy storage system, the number of the upload strategy dynamic adjustment models is multiple; different upload strategy dynamic adjustment models control different areas in the target photovoltaic energy storage system; Accordingly, the current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, including: The current state parameters of each target device are input into the region matching upload strategy dynamic adjustment model.

[0128] In the step S103, based on the target adjustment action, the target log upload strategy for the target device is dynamically adjusted.

[0129] The target log upload strategy is a target log upload rule specifically implemented by the target photovoltaic energy storage system.

[0130] For example, if the target adjustment action is compression ratio selection, the compression ratio in the specific target log upload rule is determined based on the selected compression ratio, and then the log is compressed according to the specific compression ratio.

[0131] For example, if the target adjustment action is to adjust the upload frequency, the upload frequency in the target log upload policy for the target device is dynamically adjusted.

[0132] The dynamically adjusting the target log upload strategy for the target device includes adjusting at least one of the upload frequency, network connection, and compression ratio in the target log upload strategy.

[0133] In the step S104, based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded to upload the log of the photovoltaic energy storage system.

[0134] The logs of different devices in the photovoltaic energy storage system are uploaded according to their corresponding target log upload strategies. They may be uploaded simultaneously or not. By dynamically adjusting the target log upload strategy for the target device, the log upload strategies of different devices may be different. For example, the upload frequency of the inverter is once every 2 minutes, but the upload frequency of the sensor is once a day, and so on.

[0135] After uploading the log of the target device based on the target log uploading strategy of each target device in the target photovoltaic energy storage system to upload the log of the photovoltaic energy storage system, the log uploading method further includes: Recording the experience tuple of the target log upload strategy of the target device executed in the target photovoltaic energy storage system, and storing the experience tuple in the experience playback pool; the experience tuple includes: the current state parameters of the target device of the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward and the state parameters of the next state of the target device; When the preset batch update condition is met, batch data is randomly extracted from the experience replay pool, and the trained upload strategy dynamic adjustment model is updated based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.

[0136] That is to say, when the photovoltaic energy storage system is in operation, the network and equipment status will constantly change, and the system must maintain sufficient adaptability; the batch update method can reduce the frequency of model updates, improve update efficiency, and ensure that the model can learn diverse experiences.

[0137] Through continuous learning and optimization, the updated model is more suitable for the actual status of the target PV energy storage system in terms of the selection and execution of log upload strategies.

[0138] Based on the same inventive concept, the embodiment of the present application also provides a log uploading device for a photovoltaic energy storage system corresponding to the log uploading method for a photovoltaic energy storage system. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the log uploading method for the photovoltaic energy storage system in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0139] Please refer to Figure 7 , Figure 7 A structural schematic diagram of a log uploading device for a photovoltaic energy storage system according to an embodiment of the present application is shown, which is applied to a target photovoltaic energy storage system, wherein the target photovoltaic energy storage system includes a plurality of target devices; the log uploading device includes: The acquisition module 701 is used to acquire the log and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; A determination module 702 is used to input the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, and the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes multiple adjustment actions for adjusting the log upload policy of the target device; An adjustment module 703, configured to dynamically adjust a target log upload policy for the target device based on the target adjustment action; The uploading module 704 is used to upload the log of each target device in the target photovoltaic energy storage system based on the target log uploading strategy of the target device, so as to upload the log of the photovoltaic energy storage system.

[0140] In some embodiments, the log uploading device of the photovoltaic energy storage system further includes a training module: The training module is used to construct an upload policy dynamic adjustment model and configure an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; The upload strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment.

[0141] In some embodiments, in the log uploading device of the photovoltaic energy storage system, the training module, when training the upload strategy dynamic adjustment model based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions, is specifically used to: At each time step, the upload strategy dynamic adjustment model predicts the Q value of each adjustment action in the adjustment action set based on the sample state parameters of the sample equipment of the sample photovoltaic energy storage system; The upload strategy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action; The sample photovoltaic energy storage system dynamically adjusts the sample log upload strategy for the sample device based on the sample adjustment action, and executes the sample log upload strategy to obtain the feedback status and instant reward fed back to the photovoltaic energy storage system by the environment; wherein the instant reward is calculated based on the reward function; Determine the loss function calculation result of the upload strategy dynamic adjustment model based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment; The upload strategy dynamic adjustment model is optimized based on the loss function calculation result of the upload strategy dynamic adjustment model until the upload strategy dynamic adjustment model meets the preset training stop condition, thereby obtaining a trained upload strategy dynamic adjustment model.

[0142] In some embodiments, the training module in the log uploading device of the photovoltaic energy storage system is further used to: determine multiple optimization targets for log uploading of the photovoltaic energy storage system; the multiple optimization targets include: success rate of log uploading, delay time, bandwidth occupancy, device power, and log priority; Fusion the log priority and the device power to determine a fusion reward function regarding the priority and the power; Based on the success rate, delay time, bandwidth occupancy of the log upload and the fused reward function regarding priority and power, a reward function for calculating instant rewards is constructed.

[0143] In some embodiments, in the log uploading device of the photovoltaic energy storage system, the reward rules of the reward function include: a reward rule for successful log uploading, a penalty rule for failed log uploading, a bandwidth occupation penalty rule, and a delay penalty rule; The reward rules for successful log upload include: after successfully uploading the log, a reward is obtained according to the status of the sample photovoltaic energy storage system; wherein, when the device has more power and / or the log priority is higher, a higher reward is obtained for successfully uploading the log; The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward; The bandwidth occupation penalty rule includes: when the bandwidth occupation of the sample photovoltaic energy storage system exceeds the preset bandwidth occupation threshold, a negative reward is obtained; The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowed value, the higher the negative reward.

[0144] In some embodiments, the training module in the log uploading device of the photovoltaic energy storage system is also used to: Based on the instant reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action performed by the upload strategy dynamic adjustment model in the current state of the photovoltaic energy storage system, a loss function is constructed.

[0145] In some embodiments, in the log uploading device of the photovoltaic energy storage system, the adjustment actions in the adjustment action set include: adjusting the upload frequency, selecting a network connection, and selecting a compression ratio.

[0146] In some embodiments, the log uploading device of the photovoltaic energy storage system further includes: An updating module is used to upload the log of the target device based on the target log upload strategy of each target device in the target photovoltaic energy storage system to upload the log of the photovoltaic energy storage system. The log uploading method further includes: Recording the experience tuple of the target log upload strategy of the target device executed in the target photovoltaic energy storage system, and storing the experience tuple in the experience playback pool; the experience tuple includes: the current state parameters of the target device of the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward and the state parameters of the next state of the target device; When the preset batch update condition is met, batch data is randomly extracted from the experience replay pool, and the trained upload strategy dynamic adjustment model is updated based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.

[0147] In some embodiments, in the log uploading device of the photovoltaic energy storage system, the number of the upload strategy dynamic adjustment models is multiple; different upload strategy dynamic adjustment models control different areas in the target photovoltaic energy storage system; Accordingly, when the determination module inputs the current state parameter of each target device into the pre-trained upload strategy dynamic adjustment model, it is specifically used to: The current state parameters of each target device are input into the region matching upload strategy dynamic adjustment model.

[0148] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0149] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0151] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a platform server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0152] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A log uploading method for a photovoltaic energy storage system, characterized in that: Applied to a target photovoltaic energy storage system, the target photovoltaic energy storage system includes multiple target devices; the log uploading method includes: Obtaining the log and current status parameters of each target device in the target photovoltaic energy storage system; the current status parameters characterize the operating status and network environment of the target device; the status parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; Inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, the upload policy dynamic adjustment model processes the current state parameters, and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; Based on the target adjustment action, dynamically adjust the target log upload policy for the target device; Based on the target log uploading strategy of each target device in the target photovoltaic energy storage system, the log of the target device is uploaded to upload the log of the photovoltaic energy storage system.

2. The log uploading method of the photovoltaic energy storage system according to claim 1 is characterized in that: The upload strategy dynamic adjustment model is trained based on the following method: Constructing an upload policy dynamic adjustment model, and configuring an adjustment action set of the upload policy dynamic adjustment model; the adjustment action set includes multiple adjustment actions; The upload strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions; wherein the upload strategy dynamic adjustment model is configured with a loss function and a reward function; the reward function is used to calculate the instant reward fed back to the sample photovoltaic energy storage system by the environment after the adjustment action is selected based on the sample state parameters of the sample device; the loss function is determined based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment.

3. The log uploading method of the photovoltaic energy storage system according to claim 2 is characterized in that: The uploading strategy dynamic adjustment model is trained based on the sample state parameters of different sample devices in the sample photovoltaic energy storage system, so that the sample photovoltaic energy storage system learns the relationship between different sample state parameters of different sample devices and different adjustment actions, including: At each time step, the upload strategy dynamic adjustment model predicts the Q value of each adjustment action in the adjustment action set based on the sample state parameters of the sample equipment of the sample photovoltaic energy storage system; The upload strategy dynamic adjustment model selects a sample adjustment action based on the Q value of each adjustment action; The sample photovoltaic energy storage system dynamically adjusts the sample log upload strategy for the sample device based on the sample adjustment action, and executes the sample log upload strategy to obtain the feedback status and instant reward fed back to the photovoltaic energy storage system by the environment; wherein the instant reward is calculated based on the reward function; Determine the loss function calculation result of the upload strategy dynamic adjustment model based on the feedback state and instant reward fed back to the photovoltaic energy storage system by the environment; The upload strategy dynamic adjustment model is optimized based on the loss function calculation result of the upload strategy dynamic adjustment model until the upload strategy dynamic adjustment model meets the preset training stop condition, thereby obtaining a trained upload strategy dynamic adjustment model.

4. The log uploading method of the photovoltaic energy storage system according to claim 2 or 3, characterized in that: The log uploading method further includes: Determine multiple optimization targets for log uploading of the photovoltaic energy storage system; the multiple optimization targets include: success rate of log uploading, delay time, bandwidth occupancy, device power, and log priority; Fusion the log priority and the device power to determine a fusion reward function regarding the priority and the power; Based on the success rate, delay time, bandwidth occupancy of the log upload and the fused reward function regarding priority and power, a reward function for calculating instant rewards is constructed.

5. The log uploading method of the photovoltaic energy storage system according to claim 4 is characterized in that: The reward rules of the reward function include: a reward rule for successful log upload, a penalty rule for failed log upload, a bandwidth occupation penalty rule, and a delay penalty rule; The reward rules for successful log upload include: after successfully uploading the log, a reward is obtained according to the status of the sample photovoltaic energy storage system; wherein, when the device has more power and / or the log priority is higher, a higher reward is obtained for successfully uploading the log; The log upload failure penalty rule includes: after the log upload fails, a negative reward is obtained according to the state of the sample photovoltaic energy storage system; wherein, the higher the log priority, the higher the negative reward; The bandwidth occupation penalty rule includes: when the bandwidth occupation of the sample photovoltaic energy storage system exceeds the preset bandwidth occupation threshold, a negative reward is obtained; The delay penalty rule includes: the smaller the gap between the delay time and the preset maximum allowed value, the higher the negative reward.

6. The log uploading method of the photovoltaic energy storage system according to claim 2 or 3, characterized in that: The log uploading method further includes: Based on the instant reward of the photovoltaic energy storage system, the maximum expected return of the photovoltaic energy storage system in the next state, and the predicted Q value of the adjustment action performed by the upload strategy dynamic adjustment model in the current state of the photovoltaic energy storage system, a loss function is constructed.

7. The log uploading method of the photovoltaic energy storage system according to claim 1 or 2, characterized in that: The adjustment actions in the adjustment action set include: adjusting upload frequency, selecting network connection, and selecting compression ratio.

8. The log uploading method of the photovoltaic energy storage system according to claim 1, characterized in that: After uploading the log of the target device based on the target log uploading strategy of each target device in the target photovoltaic energy storage system to upload the log of the photovoltaic energy storage system, the log uploading method further includes: Recording the experience tuple of the target log upload strategy of the target device executed in the target photovoltaic energy storage system, and storing the experience tuple in the experience playback pool; the experience tuple includes: the current state parameters of the target device of the target photovoltaic energy storage system, the target adjustment action taken corresponding to the current state parameters of the target device, the immediate reward and the state parameters of the next state of the target device; When the preset batch update condition is met, batch data is randomly extracted from the experience replay pool, and the trained upload strategy dynamic adjustment model is updated based on the extracted batch data to obtain an updated upload strategy dynamic adjustment model.

9. The log uploading method of the photovoltaic energy storage system according to claim 1, characterized in that: The number of the upload strategy dynamic adjustment models is multiple; different upload strategy dynamic adjustment models control different areas in the target photovoltaic energy storage system; Accordingly, the current state parameters of each target device are input into a pre-trained upload strategy dynamic adjustment model, including: The current state parameters of each target device are input into the region matching upload strategy dynamic adjustment model.

10. A log upload device for a photovoltaic energy storage system, characterized in that: Applied to a target photovoltaic energy storage system, the target photovoltaic energy storage system includes a plurality of target devices; the log uploading device includes: An acquisition module is used to acquire the log and current state parameters of each target device in the target photovoltaic energy storage system; the current state parameters represent the operating status and network environment of the target device; the state parameters include: network delay, bandwidth, log generation rate, log priority and power of the photovoltaic energy storage system; the target photovoltaic energy storage system corresponds to multiple network connection paths; A determination module, used for inputting the current state parameters of each target device into a pre-trained upload policy dynamic adjustment model, wherein the upload policy dynamic adjustment model processes the current state parameters and determines a target adjustment action that matches the current state parameters of the target device from a pre-configured adjustment action set; the adjustment action set includes a plurality of adjustment actions for adjusting the log upload policy of the target device; An adjustment module, configured to dynamically adjust a target log upload policy for the target device based on the target adjustment action; The uploading module is used to upload the log of each target device in the target photovoltaic energy storage system based on the target log uploading strategy of the target device, so as to upload the log of the photovoltaic energy storage system.

Citation Information

Patent Citations

  • Urban power distribution network multistage dynamic reconstruction method based on machine learning

    CN114662982A

  • Health monitoring method integrating multi-modal biological information

    CN118522438A

  • Power distribution network load state estimation method and system

    CN119298076A

  • Generating and executing context-specific neural network models based on target runtime parameters

    US20210241108A1

  • Intelligent Energy Management Systems and Methods

    US20250038530A1