Anti-countercurrent method and device of photovoltaic system and computer readable storage medium

By constructing a value table for anti-reverse current strategies through reinforcement learning, the output power of the photovoltaic system inverter is dynamically adjusted, solving the problem of high cost of reverse current prevention in existing technologies and achieving low-cost and simple reverse current prevention.

CN121333210APending Publication Date: 2026-01-13SUNGROW ICARBON TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511544031.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing photovoltaic systems, methods to prevent backflow typically involve large amounts of computation or require extensive training with historical data, resulting in high costs, inconvenient maintenance, and an inability to effectively reduce the frequency of backflow.

Method used

A value table for anti-reverse current strategy is constructed using reinforcement learning. By dynamically acquiring the inverter output power and load power of the photovoltaic system, the target power state is determined, and anti-reverse current commands are selected from the value table to adjust the inverter output power and update the strategy table in real time.

Benefits of technology

It achieves low-cost and simple backflow prevention, reduces the frequency of backflow occurrence, eliminates the need for extensive computation and historical data training, and reduces development and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121333210A_ABST
    Figure CN121333210A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-countercurrent method and device for a photovoltaic system and a computer readable storage medium, and relates to the technical field of photovoltaic systems, and the method comprises the steps: dynamically obtaining the inverter output power and the load power of the photovoltaic system at the current moment, and according to the inverter output power and the load power, in a preset power state table, carrying out the anti-countercurrent operation of the photovoltaic system; determining a target power state of the photovoltaic system; selecting an anti-reflux instruction from anti-reflux instructions corresponding to a target power state recorded in an anti-reflux strategy value table constructed by adopting reinforcement learning as a target anti-reflux instruction; the photovoltaic system adjusts the inverter output power of the photovoltaic system according to the target anti-reflux instruction; after the inverter output power of the photovoltaic system is adjusted, the new inverter output power and the new load power of the photovoltaic system are obtained, and the anti-countercurrent strategy value table is updated based on the new inverter output power and the new load power. According to the invention, the photovoltaic system can be simply and conveniently prevented from generating the countercurrent phenomenon with low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of photovoltaic system technology, and in particular to a method, device and computer-readable storage medium for preventing backflow in a photovoltaic system. Background Technology

[0002] When the power generated by a photovoltaic system exceeds the power consumption of the load, the excess power will flow back into the power grid, forming a reverse flow phenomenon.

[0003] Currently, load forecasting algorithms are commonly used to predict short-term load demand based on historical load data, allowing for advance adjustment of photovoltaic inverter output power to prevent reverse current. However, commonly used load forecasting algorithms either involve high computational demands, leading to high system costs, or rely on extensive historical data for initial training to obtain model parameters, resulting in high development costs, complex implementation, and inconvenience in operation and maintenance due to model parameter updates, leading to high maintenance costs. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, and computer-readable storage medium for preventing backflow in photovoltaic systems, which aims to prevent backflow in photovoltaic systems in a simple and low-cost manner.

[0005] This application provides a method for preventing backflow in a photovoltaic system, the method comprising: The inverter output power and load power of the photovoltaic system at the current moment are dynamically acquired, and the target power state of the photovoltaic system is determined in a preset power state table based on the inverter output power and the load power. From the anti-backflow instructions corresponding to the target power state recorded in the anti-backflow strategy value table constructed using reinforcement learning, select one anti-backflow instruction as the target anti-backflow instruction; The photovoltaic system adjusts the inverter output power according to the target anti-reverse current command; After adjusting the inverter output power of the photovoltaic system, the new inverter output power and the new load power of the photovoltaic system are obtained, and the anti-reverse current strategy value table is updated based on the new inverter output power and the new load power.

[0006] In one embodiment, the anti-reverse current strategy value table records the execution value of different anti-reverse current commands under different power states; The step of updating the anti-reverse current strategy value table based on the new inverter output power and the new load power includes: Based on the new inverter output power and the new load power, determine the target reward value corresponding to the target anti-reverse current command under the target power state; Obtain the first instruction execution value of the photovoltaic system under the new inverter output power and the new load power; Based on the target reward value and the first instruction execution value, the second instruction execution value corresponding to the target anti-backflow instruction under the target power state is adjusted to obtain a new second instruction execution value, which is then updated in the anti-backflow strategy value table.

[0007] In one embodiment, the step of determining the target reward value corresponding to the target anti-reverse current command under the target power state based on the new inverter output power and the new load power includes: If the output power of the new inverter is greater than the output power of the new load, then the preset negative reward value will be used as the target reward value. If the output power of the new inverter is less than or equal to the output power of the new load, then based on the mapping relationship between the derating ratio associated with the preset anti-reverse current command and the positive reward value, the positive reward value corresponding to the derating ratio associated with the target anti-reverse current command is obtained and used as the target reward value. Among them, the reduction ratio associated with the anti-backflow instruction is negatively correlated with the positive reward value.

[0008] In one embodiment, the step of adjusting the second instruction execution value corresponding to the target anti-backflow instruction under the target power state based on the target reward value and the first instruction execution value, to obtain a new second instruction execution value, and updating it in the anti-backflow strategy value table, includes: Calculate the product of the execution value of the first instruction and the preset discount coefficient to obtain the discounted execution value of the first instruction; The difference between the sum of the target reward value and the discounted first instruction execution value and the second instruction execution value is calculated to obtain the value difference. The value adjustment amount is obtained by multiplying the value difference by the preset learning rate. The second instruction execution value is added to the value adjustment amount to adjust the second instruction execution value, resulting in a new second instruction execution value.

[0009] In one embodiment, the method further includes: Obtain the rate of change of the inverter output power of the photovoltaic system; If the rate of change is greater than a preset rate of change threshold, then the discount coefficient is increased and the learning rate is decreased.

[0010] In one embodiment, the anti-reverse current strategy value table records the execution value of different anti-reverse current commands under different power states; The step of selecting an anti-reverse current command as the target anti-reverse current command from the anti-reverse current strategy value table constructed using reinforcement learning, corresponding to each anti-reverse current command of the target power state, includes: Based on a preset exploration probability, a target selection strategy is determined from a first preset selection strategy and a second preset selection strategy; wherein, the probability of selecting the first preset selection strategy is the exploration probability, and the probability of selecting the second preset selection strategy is the complementary probability of the exploration probability. If the target selection strategy is the first preset selection strategy, then any anti-reverse current command is randomly selected from the anti-reverse current commands corresponding to the target power state recorded in the anti-reverse current strategy value table as the target anti-reverse current command; If the target selection strategy is the second preset selection strategy, then the anti-reverse current instruction with the highest instruction execution value among the anti-reverse current instructions corresponding to the target power state recorded in the anti-reverse current strategy value table is selected as the target anti-reverse current instruction.

[0011] In one embodiment, the method further includes: Obtain the rate of change of the inverter output power of the photovoltaic system, and determine the operating state of the photovoltaic system based on the rate of change; If the working state is a low-reverse-risk state, then the exploration probability is increased; If the working state is a high risk of backflow, then the exploration probability is reduced.

[0012] In one embodiment, before the step of determining the target power state of the photovoltaic system based on the inverter output power and the load power in a preset power state table, the method further includes: Based on the rated power of the photovoltaic inverter, multiple inverter output power ranges are obtained, and based on the historical maximum load power, multiple load power ranges are obtained. Each inverter output power range is associated with each load power range to obtain multiple power states; Summarize the power states to generate the power state table.

[0013] In one embodiment, the step of constructing the anti-backflow strategy value table includes: Obtain the initial values ​​of instruction execution for all power states in the power state table under different anti-reverse current commands; The power state and anti-reverse current instruction corresponding to the initial value of each instruction execution are associated with the initial value of each instruction execution to generate the anti-reverse current strategy value table.

[0014] In one embodiment, the step of obtaining the initial values ​​of instruction execution for all power states in the power state table under different anti-reverse current commands includes: In each of the power states, obtain the maximum power difference between the inverter output power range and the load power range; For any of the power states, for each of the anti-reverse current commands, based on the preset mapping relationship between the power difference, the derating ratio associated with the anti-reverse current command, and the command execution initial value, the command execution initial value corresponding to the maximum power difference in the power state and the derating ratio associated with the anti-reverse current command is obtained, and used as the command execution initial value for the power state under the anti-reverse current command. Among them, the derating ratio associated with power difference and anti-reverse current command is positively correlated with the initial value of command execution.

[0015] In addition, to achieve the above objectives, this application also provides a photovoltaic system anti-backflow device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the photovoltaic system anti-backflow method as described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the anti-reverse current method for a photovoltaic system as described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the anti-reverse current method for a photovoltaic system as described above.

[0018] This application provides a method for preventing reverse current in a photovoltaic (PV) system. First, the inverter output power and load power of the PV system at the current moment are dynamically acquired. Based on the inverter output power and load power, the target power state of the PV system is determined from a preset power state table. Then, from the anti-reverse current commands corresponding to the target power state recorded in the anti-reverse current strategy value table constructed using reinforcement learning, one anti-reverse current command is selected as the target anti-reverse current command. Next, the PV system adjusts the inverter output power according to the target anti-reverse current command to prevent reverse current from occurring. After adjusting the inverter output power, the new inverter output power and new load power of the PV system are acquired, and the anti-reverse current strategy value table is updated based on the new inverter output power and new load power.

[0019] Therefore, the technical solution provided in this application, after determining the current power state of the photovoltaic system, i.e., after determining the target power state, will directly and quickly determine the anti-reverse current command to be executed by the photovoltaic system under the current power state by looking up a table. The anti-reverse current strategy value table is constructed using reinforcement learning and is updated directly using the real-time power data of the photovoltaic system, i.e., it is constructed and updated based on online learning. Therefore, the technical solution provided in this application, when implementing the anti-reverse current function, does not require excessive computation, nor does it require the collection of a large amount of historical data and complex model training before deployment. Furthermore, it does not require manual parameter updates and maintenance due to the accumulation of new data in new scenarios, thereby reducing implementation costs, development costs, and maintenance costs, and the implementation method is simple.

[0020] In summary, the technical solution provided in this application can easily and cost-effectively prevent reverse current phenomena in photovoltaic systems. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A schematic flowchart illustrating the anti-reverse current method for a photovoltaic system provided in the first embodiment of this application; Figure 2 This is a schematic diagram of the structure of the photovoltaic system provided in the first embodiment of this application; Figure 3 This is a schematic diagram of the structure of a photovoltaic system provided in an embodiment of this application; Figure 4 This is a schematic diagram of the hardware operating environment involved in the embodiments of this application.

[0024] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0025] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0026] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0027] When the power generated by a photovoltaic system exceeds the power consumption of the load, the excess power will flow back into the power grid, forming a reverse flow phenomenon.

[0028] Currently, load forecasting algorithms are commonly used to predict short-term load demand based on historical load data, allowing for advance adjustment of photovoltaic inverter output power to prevent reverse current. However, commonly used load forecasting algorithms either involve high computational demands, leading to high system costs, or rely on extensive historical data for initial training to obtain model parameters, resulting in high development costs, complex implementation, and inconvenience in operation and maintenance due to model parameter updates, leading to high maintenance costs.

[0029] In addition, some methods currently determine whether a reverse current phenomenon has occurred in a photovoltaic system by collecting real-time voltage, current, or power data at the grid connection point of the photovoltaic inverter. Once a reverse current phenomenon is confirmed, a derating command is sent to the photovoltaic inverter to limit its output power, which is then increased again after the system recovers. While this method has simple control logic, it cannot reduce the frequency of reverse current occurrences.

[0030] Based on this, this application provides a method for preventing reverse current in a photovoltaic system. First, the inverter output power and load power of the photovoltaic system at the current moment are dynamically acquired. Based on the inverter output power and load power, the target power state of the photovoltaic system is determined from a preset power state table. Then, from the anti-reverse current instructions corresponding to the target power state recorded in the anti-reverse current strategy value table constructed using reinforcement learning, one anti-reverse current instruction is selected as the target anti-reverse current instruction. Next, the photovoltaic system will adjust the inverter output power of the photovoltaic system according to the target anti-reverse current instruction to prevent reverse current from occurring in the photovoltaic system. After adjusting the inverter output power of the photovoltaic system, the new inverter output power and new load power of the photovoltaic system are acquired, and the anti-reverse current strategy value table is updated based on the new inverter output power and new load power.

[0031] Therefore, the technical solution provided in this application, after determining the current power state of the photovoltaic system, i.e., after determining the target power state, will directly and quickly determine the anti-reverse current command to be executed by the photovoltaic system under the current power state by looking up a table. The anti-reverse current strategy value table is constructed using reinforcement learning and is updated directly using the real-time power data of the photovoltaic system, i.e., it is constructed and updated based on online learning. Therefore, the technical solution provided in this application, when implementing the anti-reverse current function, does not require excessive computation, nor does it require the collection of a large amount of historical data and complex model training before deployment. Furthermore, it does not require manual parameter updates and maintenance due to the accumulation of new data in new scenarios, thereby reducing implementation costs, development costs, and maintenance costs, and the implementation method is simple.

[0032] In summary, the technical solution provided in this application can easily and cost-effectively prevent reverse current phenomena in photovoltaic systems.

[0033] Furthermore, the technical solution provided in this application will execute at least one anti-reverse current command regardless of the power state, so it can also prevent the photovoltaic system from experiencing reverse current in advance, thereby effectively reducing the frequency of reverse current occurrence.

[0034] The subject of the anti-backflow method for photovoltaic systems in this application can be an anti-backflow device for photovoltaic systems with data processing, network communication and program operation functions, or a control system, control circuit, etc. that can realize the above functions, or even the photovoltaic system itself. This embodiment does not specifically limit it in this regard.

[0035] The anti-reverse current device of the photovoltaic system can be integrated inside the photovoltaic inverter or exist independently; this embodiment does not make specific limitations on this.

[0036] The following description uses the anti-reverse current device of a photovoltaic system as the main implementer to illustrate the various embodiments.

[0037] This application presents a first embodiment of a photovoltaic system's anti-reverse current method; please refer to [link / reference]. Figure 1 The method for preventing backflow in a photovoltaic system may include steps S10 to S40: Step S10: Dynamically acquire the inverter output power and load power of the photovoltaic system at the current moment, and determine the target power state of the photovoltaic system based on the inverter output power and load power in the preset power state table. It should be noted that the inverter output power refers to the active power output by the photovoltaic inverter, and the load power refers to the power consumed by the loads (i.e., electrical equipment) connected to the photovoltaic system. The structure of a photovoltaic system can be exemplarily referred to [reference needed]. Figure 2 , Figure 2 P in pv That is, the inverter output power, P load That is, the load power, where ΔP is the active power at the grid connection point collected by the meter. P is calculated... pv The difference between P and ΔP can be used to obtain P. load .

[0038] Additionally, it should be noted that the power state table is used to record various power states of the photovoltaic system. The power state consists of the inverter output power range and the load power range. It is a two-dimensional state composed of the inverter output power range and the load power range. The target power state is the power state of the photovoltaic system at the current moment.

[0039] The inverter output power and load power of the photovoltaic system at the current moment are dynamically acquired. That is, the real-time inverter output power and load power of the photovoltaic system are continuously acquired. In the process of dynamic acquisition, it can be acquired in real time or periodically at a certain time interval. This embodiment does not make a specific limitation on this.

[0040] In one feasible implementation, the power state table construction step may include steps S01 to S03: Step S01: Based on the rated power of the photovoltaic inverter, multiple inverter output power ranges are obtained, and based on the historical maximum load power, multiple load power ranges are obtained. It should be noted that the maximum load power in historical statistics can be the maximum load power collected by the photovoltaic system within a certain period of time (such as 1 year) before the current moment.

[0041] When dividing the photovoltaic inverter output power range into multiple ranges based on its rated power, the first interval can be determined using the rated power of the photovoltaic inverter. Then, based on the first interval and the rated power of the photovoltaic inverter, multiple inverter output power ranges can be obtained. Alternatively, multiple inverter output power ranges can be obtained by dividing the ranges at fixed intervals. Non-uniform division can also be performed based on power distribution density, but this embodiment does not impose specific limitations on this.

[0042] Similarly, when dividing the load power range into multiple ranges based on the historical maximum load power, the second interval can be determined first using the maximum load power; then, based on the second interval and the maximum load power, multiple inverter output power ranges can be obtained. Alternatively, multiple inverter output power ranges can be obtained by using fixed intervals; non-uniform division can also be performed based on power distribution density, but this embodiment does not specifically limit this approach.

[0043] In determining the first interval using the rated power of the photovoltaic inverter, the first interval can be the product of the rated power of the photovoltaic inverter and the ratio of the first preset interval; alternatively, the ratio of the rated power of the photovoltaic inverter to the number of the first preset intervals can also be used as the first interval. This embodiment does not impose a specific limitation on this method. Similarly, in determining the second interval using the maximum load power, the second interval can be the product of the maximum load power and the ratio of the second preset interval; alternatively, the ratio of the maximum load power to the number of the second preset intervals can also be used as the second interval. This embodiment does not impose a specific limitation on this method.

[0044] In this embodiment, the first preset interval ratio, the number of first preset intervals, the second preset interval ratio, and the number of second preset intervals can all be default values, or can be flexibly set by the user according to actual conditions. This embodiment does not impose specific limitations on these settings. For example, assuming that the first preset interval ratio and the second preset interval ratio are both 5%, and the rated power of the photovoltaic inverter and the historical maximum load power are both 100kW, then the first interval and the second interval are both 5. Thus, the resulting multiple inverter output power intervals include [0,5), [5,10), [10,15), ..., [95,100], a total of 20 power intervals, and the resulting multiple load power intervals include [0,5), [5,10), [10,15), ..., [95,100], a total of 20 power intervals. Therefore, the power status table will record 400 power states.

[0045] Step S02: Associate each inverter output power range with each load power range to obtain multiple power states; Step S03: Summarize the power states to generate a power state table.

[0046] This implementation method divides the inverter output power into multiple ranges based on the rated power of the photovoltaic inverter and the load power into multiple ranges based on the historically recorded maximum load power. Then, each inverter output power range is associated with a specific load power range to obtain multiple power states. Finally, by summarizing these power states, a power state table is generated. This maps the originally continuous inverter output power and load power to a finite discrete state space, providing core support for subsequent simple and low-cost reverse current control.

[0047] Step S20: Select one anti-reverse current instruction as the target anti-reverse current instruction from the anti-reverse current strategy value table recorded by reinforcement learning, corresponding to each anti-reverse current instruction of the target power state. It should be noted that when constructing the anti-reverse current policy value table using reinforcement learning, the Sarsa reinforcement learning algorithm can be used. The anti-reverse current policy value table records the execution value of different anti-reverse current instructions under different power states. The instruction execution value is used to quantitatively evaluate the long-term expected return obtained after executing a certain anti-reverse current instruction under a power state. Taking the power states recorded in the anti-reverse current policy value table as s1~sn, the anti-reverse current instructions as a1~an, and the instruction execution value as Q(s1,a1)~Q(sn,an) as an example, the anti-reverse current policy value table can be represented as shown in Table 1 below.

[0048] Table 1:

[0049] Additionally, it should be noted that the anti-reverse current command, also known as the derating command, is a control command that adjusts the inverter's output power. It typically uses its associated derating ratio to regulate the inverter's output power in a photovoltaic system. Different anti-reverse current commands are associated with different derating ratios. For example, there can be 11 anti-reverse current commands as shown in Table 2, i.e., anti-reverse current command a∈{0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10}, each associated with a different derating ratio.

[0050] Table 2:

[0051] In one feasible implementation, step S20 may include steps S21 to S23: Step S21: Based on the preset exploration probability, determine the target selection strategy from the first preset selection strategy and the second preset selection strategy; wherein, the probability of selecting the first preset selection strategy is the exploration probability, and the probability of selecting the second preset selection strategy is the complementary probability of the exploration probability. It's important to note that the exploration probability determines the system's tendency to explore unknown or suboptimal actions (i.e., anti-backflow instructions) when making decisions, rather than blindly using actions currently considered optimal. The sum of the exploration probability and its complementary probability is 1. The first preset selection strategy is essentially an exploration strategy, meaning it doesn't rely on historical experience but randomly selects one from all possible anti-backflow instructions. The second preset selection strategy is essentially a greedy strategy, greedily selecting the anti-backflow instruction with the highest execution value recorded in the anti-backflow strategy value table under the target power state.

[0052] Step S22: If the target selection strategy is the first preset selection strategy, then arbitrarily select one anti-reverse current command from the anti-reverse current commands corresponding to the target power state recorded in the anti-reverse current strategy value table as the target anti-reverse current command. Step S23: If the target selection strategy is the second preset selection strategy, then select the anti-reverse current instruction with the highest instruction execution value among the anti-reverse current instructions corresponding to the target power state recorded in the anti-reverse current strategy value table as the target anti-reverse current instruction.

[0053] In this implementation, when selecting an anti-reverse current command as the target anti-reverse current command from among the anti-reverse current commands corresponding to the target power state recorded in the anti-reverse current strategy value table, an exploration mechanism is introduced to effectively prevent the control strategy from prematurely converging to a local optimum. Specifically, the system does not always mechanically select the anti-reverse current command with the highest execution value under current knowledge, but rather randomly explores other anti-reverse current commands with a certain probability. This randomness allows the system to step out of its currently known "comfort zone," thus having the opportunity to evaluate anti-reverse current commands that are not currently favored but may have greater potential, enabling the system to gradually approach the globally optimal control strategy.

[0054] Furthermore, to ensure that the exploration mechanism introduced in this embodiment not only does not significantly affect the anti-backflow effect but also effectively accelerates the convergence speed of the globally optimal control strategy, in a feasible embodiment, the anti-backflow method for the photovoltaic system may further include steps S201-S203: Step S201: Obtain the rate of change of the inverter output power of the photovoltaic system, and determine the operating state of the photovoltaic system based on the rate of change. It should be noted that the rate of change of inverter output power refers to the rate of change of the photovoltaic inverter's output power per unit time (e.g., 1 minute). The operating state of a photovoltaic system can include, but is not limited to, a low-reverse-risk state and a high-reverse-risk state. When a photovoltaic system is in a high-reverse-risk state, the risk of reverse current occurrence is relatively high, while when it is in a low-reverse-risk state, the risk of reverse current occurrence is relatively low.

[0055] When determining the operating state of a photovoltaic system based on the rate of change, it can be determined whether the rate of change is greater than a preset rate of change threshold. If so, the photovoltaic system is determined to be in a high reverse current risk state; otherwise, it is determined to be in a low reverse current risk state. In other embodiments, the operating state of the photovoltaic system can also be determined based on the change in the inverter output power of the photovoltaic system. This embodiment does not specifically limit this.

[0056] Step S202: If the working state is a low backflow risk state, then increase the exploration probability; Step S203: If the working state is a high risk of backflow, then reduce the exploration probability.

[0057] This implementation method, by setting the system to reduce the exploration probability when the photovoltaic system is in a high-reverse-current-risk state, prioritizes the anti-reverse-current effect, and ensures the safety of the photovoltaic system; while when the photovoltaic system is in a low-reverse-current-risk state, the exploration probability can be increased to effectively accelerate the exploration of the globally optimal control strategy.

[0058] Furthermore, to further ensure that the exploration mechanism introduced in this embodiment does not have too much impact on the anti-backflow effect, the exploration probability can be set to gradually decrease as the number of training times (i.e., the number of updates) of the anti-backflow strategy value table increases (e.g., it can be linearly decayed to 0.05). This way, in the later stages of training the anti-backflow strategy value table, the anti-backflow command will be mainly selected using a greedy strategy, thereby reducing the occurrence of backflow in the photovoltaic system due to the exploration mechanism to a certain extent.

[0059] This embodiment does not specifically limit the implementation of step S20. For example, in other feasible implementations, step S20 can also be executed by default using a greedy strategy.

[0060] Step S30: The photovoltaic system adjusts the inverter output power of the photovoltaic system according to the target anti-reverse current command; It should be noted that when the photovoltaic system adjusts the output power of the inverter according to the target anti-reverse current command, it essentially uses the derating ratio associated with the target anti-reverse current command to adjust the output power of the inverter. The specific implementation process can be expressed as the following formula 1.

[0061] P output = P pv × (1-d) Formula 1; Among them, P output P is the adjusted inverter output power. pv d represents the inverter output power before adjustment, and d is the derating ratio.

[0062] Step S40: After adjusting the inverter output power of the photovoltaic system, obtain the new inverter output power and the new load power of the photovoltaic system, and update the anti-reverse current strategy value table based on the new inverter output power and the new load power.

[0063] It should be noted that after adjusting the inverter output power of the photovoltaic system, the obtained inverter output power and load power are the new inverter output power and the new load power.

[0064] Based on the above, the technical solution provided in this embodiment, after determining the current power state of the photovoltaic system (i.e., after determining the target power state), directly and quickly determines the anti-reverse current command required by the photovoltaic system under the current power state through a table lookup method. The anti-reverse current strategy value table is constructed using reinforcement learning and is updated directly using the real-time power data of the photovoltaic system; that is, it is constructed and updated based on online learning. Therefore, the technical solution provided in this embodiment, when implementing the anti-reverse current function, not only does not involve excessive computation, but also does not require the collection of large amounts of historical data and complex model training before deployment, and it does not require manual parameter updates and maintenance due to the accumulation of new data in new scenarios. This reduces implementation costs, development costs, and maintenance costs, and the implementation method is simple.

[0065] Therefore, the technical solution provided in this embodiment can easily and cost-effectively prevent reverse current phenomena in photovoltaic systems.

[0066] Based on the first embodiment described above, a second embodiment of the anti-reverse current method for the photovoltaic system of this application is proposed. In the second embodiment, the anti-reverse current strategy value table records the execution value of different anti-reverse current commands under different power states. Step S40 may include steps S41 to S43: Step S41: Based on the new inverter output power and the new load power, determine the target reward value corresponding to the target anti-reverse current command under the target power state; It should be noted that the target reward value is the evaluation score of the immediate effect of the execution of the target anti-backflow command.

[0067] When determining the target reward value corresponding to the target anti-reverse current command under the target power state based on the new inverter output power and the new load power, if the new inverter output power is greater than the new load power, a target negative reward value (which can be a subsequent preset negative reward value) can be obtained as the target reward value; if the new inverter output power is less than or equal to the new load power, a target positive reward value (which can be a subsequent positive reward value corresponding to the derating ratio associated with the target anti-reverse current command) can be obtained as the target reward value.

[0068] In one feasible implementation, step S41 may include steps S411 to S412; Step S411: If the new inverter output power is greater than the new load power, then the preset negative reward value is used as the target reward value. Step S412: If the new inverter output power is less than or equal to the new load power, then based on the mapping relationship between the derating ratio associated with the preset anti-reverse current command and the positive reward value, obtain the positive reward value corresponding to the derating ratio associated with the target anti-reverse current command, and use it as the target reward value. Among them, the reduction ratio associated with the anti-backflow instruction is negatively correlated with the positive reward value.

[0069] It should be noted that the preset negative reward value is a negative value, which can be a default value, such as -100, or it can be flexibly set by the user according to the actual situation. This embodiment does not impose specific limitations on this. The positive reward value is a positive value. The reduction ratio associated with the anti-backflow command is negatively correlated with the positive reward value. That is, the larger the reduction ratio, the smaller the positive reward value, and vice versa.

[0070] For example, the mapping relationship between the positive reward value and the reduction ratio associated with the anti-backflow instruction can be expressed as the following formula 2.

[0071] r = 10×(1-d) Formula 2; Where r is the positive reward value and d is the reduction ratio. Based on this, in the 11 cases shown in Table 2 above, the 11 positive reward value cases can be referred to Table 3 below.

[0072] Table 3:

[0073] Furthermore, if the new inverter output power is less than or equal to the new load power, and the derating ratio associated with the target anti-reverse current command is zero, a larger positive bonus value, such as 20, can be considered.

[0074] Understandably, after the photovoltaic system executes the target anti-reverse current command under the current power state, if the output power of the new inverter in the photovoltaic system is greater than the new load power, it indicates that the target anti-reverse current command cannot achieve a good anti-reverse current effect under the target power state. In this case, a preset negative reward value can be used as the target reward value, i.e., a large negative reward is given to reduce the execution value of the target anti-reverse current command under the target power state, thereby reducing or even avoiding the execution of the target anti-reverse current command under the target power state. If the output power of the new inverter is less than or equal to the new load power, it indicates that the target anti-reverse current command has achieved a good anti-reverse current effect under the target power state. In this case, the derating ratio associated with the target anti-reverse current command can be used to evaluate the immediate effect of the execution of the target anti-reverse current command under the target power state. Specifically, based on the negative correlation between the preset derating ratio associated with the anti-reverse current command and the positive reward value, the positive reward value corresponding to the derating ratio associated with the target anti-reverse current command can be obtained as the target reward value.

[0075] This embodiment does not specifically limit the implementation of step S41. For example, in other feasible implementations, when the new inverter output power is greater than the new load power, the negative reward value corresponding to the derating ratio associated with the target anti-reverse current command can be used as the target reward value; when the new inverter output power is less than or equal to the new load power, a preset positive reward value can be used as the target reward value. The derating ratio associated with the anti-reverse current command is negatively correlated with the negative reward value, and the preset positive reward value can be a default value or can be flexibly set by the user according to actual conditions. This embodiment does not specifically limit this.

[0076] Step S42: Obtain the first instruction execution value of the photovoltaic system under the new inverter output power and the new load power; It should be noted that when obtaining the first instruction execution value of the photovoltaic system under the new inverter output power and the new load power, the new target power state and the new target anti-reverse current instruction are first determined by repeatedly executing the above steps S10 and S20; then, the instruction execution value of the new target anti-reverse current instruction under the new target power state is obtained from the anti-reverse current strategy value table, and is used as the first instruction execution value.

[0077] Step S43: Based on the target reward value and the first instruction execution value, adjust the second instruction execution value corresponding to the target anti-backflow instruction under the target power state to obtain a new second instruction execution value, and update it in the anti-backflow strategy value table.

[0078] It should be noted that the execution value of the target anti-reverse current instruction recorded in the anti-reverse current strategy value table under the target power state is called the second instruction execution value.

[0079] In one feasible implementation, step S43 may include steps S431 to S434: Step S431: Calculate the product of the execution value of the first instruction and the preset discount coefficient to obtain the discounted execution value of the first instruction; Step S432: Calculate the sum of the target reward value and the discounted first instruction execution value, and the difference between it and the second instruction execution value to obtain the value difference; Step S433: Calculate the product of the value difference and the preset learning rate to obtain the value adjustment amount; Step S434: Add the second instruction execution value to the value adjustment amount to adjust the second instruction execution value and obtain a new second instruction execution value.

[0080] It's important to note that the discount factor is generally a constant between 0 and 1, representing the degree of importance placed on future rewards. A discount factor closer to 1 indicates a greater emphasis on long-term gains (i.e., a higher emphasis on future rewards), while a factor closer to 0 indicates a greater focus on immediate returns (i.e., a lower emphasis on future rewards). The learning rate is also generally a constant between 0 and 1, controlling the magnitude of instruction execution value updates. A learning rate closer to 1 indicates a larger magnitude of instruction execution value updates, while a rate closer to 0 indicates a smaller magnitude of updates.

[0081] The implementation process of step S43 provided in this embodiment can be specifically expressed as the following formula 3.

[0082] Q'(s,a) ← Q(s,a) + α[r d + γ*Q(s',a') – Q(s,a)] Formula 3; Where Q'(s,a) is the new second instruction execution value, Q(s,a) is the second instruction execution value, α is the learning rate, and r is the learning rate. d Let Q(s',a') be the target reward value, and let Q(s',a') be the execution value of the first instruction.

[0083] Furthermore, in one feasible implementation, the anti-reverse current method for a photovoltaic system may further include steps S401-S402: Step S401: Obtain the rate of change of the inverter output power of the photovoltaic system; In step S402, if the rate of change is greater than the preset rate of change threshold, the discount coefficient is increased and the learning rate is decreased.

[0084] It should be noted that the preset rate of change threshold is used as the basis for determining whether the rate of change of the inverter output power of the photovoltaic system is large. It can be a default value, such as 20%, or it can be flexibly set by the user according to the actual situation. This embodiment does not make specific limitations on this.

[0085] In other embodiments, the discount factor can be increased and the learning rate can be decreased when the change in the inverter output power of the photovoltaic system exceeds a preset threshold. This embodiment does not specifically limit this.

[0086] This implementation increases the discount factor and decreases the learning rate when the inverter output power of the photovoltaic system changes significantly, i.e. when the light intensity fluctuates drastically. This strengthens the focus on future rewards and reduces the update range of instruction execution value, thereby avoiding risks for short-term gains and preventing sudden changes in the control strategy.

[0087] As can be seen from the above, in updating the anti-reverse current strategy value table based on the new inverter output power and the new load power, this embodiment simultaneously considers the immediate effect of the target anti-reverse current command execution under the target power state (i.e., the target reward value) and the future reward (i.e., the first command execution value). By evaluating the long-term comprehensive value of the target anti-reverse current command, the second command execution value corresponding to the target anti-reverse current command under the target power state is adjusted to obtain a new second command execution value, which is then updated in the anti-reverse current strategy value table. Therefore, the command execution values ​​recorded in the anti-reverse current strategy value table not only effectively correspond to the immediate effect of execution under the corresponding power state but also possess good foresight, thus more effectively preventing reverse current phenomena in the photovoltaic system.

[0088] In addition, in other embodiments, the anti-reverse flow strategy value table may be updated by only considering the immediate effect of the target anti-reverse flow instruction in the target power state (i.e., the target reward value) without considering future rewards (i.e., the execution value of the first instruction). This embodiment does not specifically limit the specific implementation of step S40.

[0089] Based on the first and / or second embodiments described above, a third embodiment of the anti-reverse current method for photovoltaic systems of this application is proposed. In the third embodiment, the step of constructing the anti-reverse current strategy value table may include steps S100 to S200: Step S100: Obtain the initial values ​​of instruction execution for all power states in the power state table under different anti-reverse current instructions; It should be noted that the instruction execution initial value is the initial value of the instruction execution value.

[0090] In one feasible implementation, step S100 may include steps S101 to S102: Step S101: Obtain the maximum power difference between the inverter output power range and the load power range in each power state; It should be noted that the power difference is the difference between the inverter's output power and the load power. The maximum power difference is the difference between the upper limit of the inverter's output power range and the lower limit of the load power range.

[0091] Step S102: For any power state, for each anti-reverse current command, based on the preset mapping relationship between the power difference, the derating ratio associated with the anti-reverse current command and the command execution initial value, obtain the command execution initial value corresponding to the maximum power difference and the derating ratio associated with the anti-reverse current command in the power state, and use it as the command execution initial value of the power state under the anti-reverse current command. Among them, the derating ratio associated with power difference and anti-reverse current command is positively correlated with the initial value of command execution.

[0092] It should be noted that the derating ratio associated with the power difference and anti-reverse current instruction is positively correlated with the initial value of instruction execution. That is, the larger the power difference and derating ratio, the larger the initial value of instruction execution, and the smaller the power difference and derating ratio, the smaller the initial value of instruction execution.

[0093] Understandably, in the initial construction of the anti-reverse current strategy value table, considering that the larger the difference between the inverter output power and the load power (i.e., the larger the power difference), the higher the risk of reverse current in the photovoltaic system, the more effective the anti-reverse current command should be to prevent reverse current in the photovoltaic system. Therefore, a positive correlation can be set between the power difference and the initial value of the command execution. This means that at the beginning of the construction of the anti-reverse current strategy value table, a higher initial value base is assigned to power states with high power differences, guiding the algorithm to prioritize the selection of powerful anti-reverse current actions in high-risk scenarios. Similarly, the larger the derating ratio associated with an anti-reverse current command, the better the anti-reverse current effect of the command. Therefore, a positive correlation can be set between the derating ratio associated with an anti-reverse current command and the initial value of the command execution. This ensures that at the beginning of the construction of the anti-reverse current strategy value table, the anti-reverse current command with the best anti-reverse current effect under the same power state is prioritized as the optimal choice, thus guiding the algorithm to prioritize the selection of powerful anti-reverse current actions. Therefore, the initial anti-backflow strategy can effectively prevent backflow from occurring in the photovoltaic system.

[0094] Step S200: Associate the power state and anti-reverse current command corresponding to the initial value of each command execution with the initial value of each command execution to generate an anti-reverse current strategy value table.

[0095] In this embodiment, during the construction of the anti-reverse current strategy value table, the initial values ​​of instruction execution for each power state under different anti-reverse current commands are predefined. This allows the anti-reverse current strategy value table to be updated based on the predefined initial values ​​of instruction execution, thereby effectively guiding and accelerating the convergence process of the anti-reverse current strategy value table towards the global optimal solution.

[0096] This application also provides an anti-reverse current device for a photovoltaic system. The anti-reverse current device for a photovoltaic system may include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the anti-reverse current method of the photovoltaic system in the above embodiments.

[0097] It should be noted that the anti-reverse current device of the photovoltaic system can be integrated inside the photovoltaic inverter, or it can be... Figure 3 The example shown is configured independently and communicates with the photovoltaic inverter and electricity meter in the photovoltaic system respectively. This embodiment does not impose specific limitations on this.

[0098] The following is for reference. Figure 4 It shows a structural schematic diagram of an anti-reverse current device suitable for implementing a photovoltaic system according to the embodiments of this application. Figure 4 The anti-reverse current device for the photovoltaic system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0099] like Figure 4As shown, the anti-reverse current device of the photovoltaic system may include a processing unit 101 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 102 or the program loaded from the storage device 103 into the random access memory 104. The random access memory 104 also stores various programs and data required for the operation of the anti-reverse current device of the photovoltaic system. The processing unit 101, the read-only memory 102, and the random access memory 104 are interconnected via a bus 105. An input / output interface 106 is also connected to the bus 105. Typically, the following systems can be connected to the input / output interface 106: input devices 107 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 108 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 103 including, for example, magnetic tapes, hard disks, etc.; and communication devices 109. The communication device 109 allows the anti-reverse current device of the photovoltaic system to communicate wirelessly or wiredly with other devices to exchange data. Although the diagram shows anti-reverse current devices for photovoltaic systems with various configurations, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0100] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 103, or installed from read-only memory 102. When the computer program is executed by processing device 101, it performs the functions defined in the methods of the embodiments of this application.

[0101] The anti-backflow device for photovoltaic systems provided in this application adopts the anti-backflow method for photovoltaic systems described in the above embodiments, which can easily and cost-effectively prevent backflow in photovoltaic systems. Compared with the prior art, the beneficial effects of the anti-backflow device for photovoltaic systems provided in this application are the same as those of the anti-backflow method for photovoltaic systems provided in the above embodiments, and other technical features of the anti-backflow device for photovoltaic systems are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0102] It should be understood that various parts of the embodiments of this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0103] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the above claims.

[0104] This application also provides a computer-readable storage medium storing a computer program that can run on a processor. The computer program is used to execute the anti-reverse current method of the photovoltaic system in the above embodiments.

[0105] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0106] The aforementioned computer-readable storage medium may be included in the anti-reverse current device of the photovoltaic system; or it may exist independently and not be installed in the anti-reverse current device of the photovoltaic system.

[0107] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the anti-reverse current device of the photovoltaic system, the anti-reverse current device of the photovoltaic system: dynamically acquires the inverter output power and load power of the photovoltaic system at the current moment, and determines the target power state of the photovoltaic system in a preset power state table based on the inverter output power and load power; selects an anti-reverse current instruction as the target anti-reverse current instruction from the anti-reverse current instruction corresponding to the target power state recorded in the anti-reverse current strategy value table constructed using reinforcement learning; adjusts the inverter output power of the photovoltaic system according to the target anti-reverse current instruction; after adjusting the inverter output power of the photovoltaic system, acquires the new inverter output power and new load power of the photovoltaic system, and updates the anti-reverse current strategy value table based on the new inverter output power and new load power.

[0108] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0110] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0111] The computer-readable storage medium provided in this application embodiment stores computer-readable program instructions for executing the anti-reverse current method for the photovoltaic system described above, which can easily and cost-effectively prevent reverse current phenomena in the photovoltaic system. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the anti-reverse current method for the photovoltaic system provided in the above embodiments, and will not be repeated here.

[0112] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the anti-reverse current method for a photovoltaic system as described above.

[0113] The computer program product provided in this application embodiment can easily and cost-effectively prevent reverse current phenomena in photovoltaic systems. Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as the beneficial effects of the anti-reverse current method for photovoltaic systems provided in the above embodiments, and will not be repeated here.

[0114] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A method for preventing backflow in a photovoltaic system, characterized in that, The method includes: The inverter output power and load power of the photovoltaic system at the current moment are dynamically acquired, and the target power state of the photovoltaic system is determined in a preset power state table based on the inverter output power and the load power. From the anti-backflow instructions corresponding to the target power state recorded in the anti-backflow strategy value table constructed using reinforcement learning, select one anti-backflow instruction as the target anti-backflow instruction; The photovoltaic system adjusts the inverter output power according to the target anti-reverse current command; After adjusting the inverter output power of the photovoltaic system, the new inverter output power and the new load power of the photovoltaic system are obtained, and the anti-reverse current strategy value table is updated based on the new inverter output power and the new load power.

2. The method as described in claim 1, characterized in that, The anti-reverse current strategy value table records the execution value of different anti-reverse current commands under different power states; The step of updating the anti-reverse current strategy value table based on the new inverter output power and the new load power includes: Based on the new inverter output power and the new load power, determine the target reward value corresponding to the target anti-reverse current command under the target power state; Obtain the first instruction execution value of the photovoltaic system under the new inverter output power and the new load power; Based on the target reward value and the first instruction execution value, the second instruction execution value corresponding to the target anti-backflow instruction under the target power state is adjusted to obtain a new second instruction execution value, which is then updated in the anti-backflow strategy value table.

3. The method as described in claim 2, characterized in that, The step of determining the target reward value corresponding to the target anti-reverse current command under the target power state based on the new inverter output power and the new load power includes: If the output power of the new inverter is greater than the output power of the new load, then the preset negative reward value will be used as the target reward value. If the output power of the new inverter is less than or equal to the output power of the new load, then based on the mapping relationship between the derating ratio associated with the preset anti-reverse current command and the positive reward value, the positive reward value corresponding to the derating ratio associated with the target anti-reverse current command is obtained and used as the target reward value. Among them, the reduction ratio associated with the anti-backflow instruction is negatively correlated with the positive reward value.

4. The method as described in claim 2, characterized in that, The step of adjusting the second instruction execution value corresponding to the target anti-backflow instruction under the target power state based on the target reward value and the first instruction execution value, to obtain a new second instruction execution value, and updating it in the anti-backflow strategy value table includes: Calculate the product of the execution value of the first instruction and the preset discount coefficient to obtain the discounted execution value of the first instruction; The difference between the sum of the target reward value and the discounted first instruction execution value and the second instruction execution value is calculated to obtain the value difference. The value adjustment amount is obtained by multiplying the value difference by the preset learning rate. The second instruction execution value is added to the value adjustment amount to adjust the second instruction execution value, resulting in a new second instruction execution value.

5. The method as described in claim 4, characterized in that, The method further includes: Obtain the rate of change of the inverter output power of the photovoltaic system; If the rate of change is greater than a preset rate of change threshold, then the discount coefficient is increased and the learning rate is decreased.

6. The method as described in claim 1, characterized in that, The anti-reverse current strategy value table records the execution value of different anti-reverse current commands under different power states; The step of selecting an anti-reverse current command as the target anti-reverse current command from the anti-reverse current strategy value table constructed using reinforcement learning, corresponding to each anti-reverse current command of the target power state, includes: Based on a preset exploration probability, a target selection strategy is determined from a first preset selection strategy and a second preset selection strategy; wherein, the probability of selecting the first preset selection strategy is the exploration probability, and the probability of selecting the second preset selection strategy is the complementary probability of the exploration probability. If the target selection strategy is the first preset selection strategy, then any anti-reverse current command is randomly selected from the anti-reverse current commands corresponding to the target power state recorded in the anti-reverse current strategy value table as the target anti-reverse current command; If the target selection strategy is the second preset selection strategy, then the anti-reverse current instruction with the highest instruction execution value among the anti-reverse current instructions corresponding to the target power state recorded in the anti-reverse current strategy value table is selected as the target anti-reverse current instruction.

7. The method as described in claim 6, characterized in that, The method further includes: Obtain the rate of change of the inverter output power of the photovoltaic system, and determine the operating state of the photovoltaic system based on the rate of change; If the working state is a low-reverse-risk state, then the exploration probability is increased; If the working state is a high risk of backflow, then the exploration probability is reduced.

8. The method according to any one of claims 1 to 7, characterized in that, Before the step of determining the target power state of the photovoltaic system based on the inverter output power and the load power in a preset power state table, the method further includes: Based on the rated power of the photovoltaic inverter, multiple inverter output power ranges are obtained, and based on the historical maximum load power, multiple load power ranges are obtained. Each inverter output power range is associated with each load power range to obtain multiple power states; Summarize the power states to generate the power state table.

9. The method according to any one of claims 1 to 7, characterized in that, The steps for constructing the value table of the anti-backflow strategy include: Obtain the initial values ​​of instruction execution for all power states in the power state table under different anti-reverse current commands; The power state and anti-reverse current instruction corresponding to the initial value of each instruction execution are associated with the initial value of each instruction execution to generate the anti-reverse current strategy value table.

10. The method as described in claim 9, characterized in that, The step of obtaining the initial values ​​of instruction execution for all power states in the power state table under different anti-reverse current commands includes: In each of the power states, obtain the maximum power difference between the inverter output power range and the load power range; For any of the power states, for each of the anti-reverse current commands, based on the preset mapping relationship between the power difference, the derating ratio associated with the anti-reverse current command, and the command execution initial value, the command execution initial value corresponding to the maximum power difference in the power state and the derating ratio associated with the anti-reverse current command is obtained, and used as the command execution initial value for the power state under the anti-reverse current command. Among them, the derating ratio associated with power difference and anti-reverse current command is positively correlated with the initial value of command execution.

11. A reverse current prevention device for a photovoltaic system, characterized in that, The anti-backflow device of the photovoltaic system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the anti-backflow method of the photovoltaic system as described in any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the anti-reverse current method for a photovoltaic system as described in any one of claims 1 to 10.

Citation Information

Cited By

  • Photovoltaic inverter anti-countercurrent control method and system based on power signal coupling

    CN121813512A