Photovoltaic power station and ecological restoration collaborative optimization scheduling system using reinforcement learning

CN122840550APending Publication Date: 2026-09-29CHINA UNIV OF MINING & TECH (BEIJING) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611041082.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]针对现有技术的不足,本发明提供了利用强化学习的光伏电站与生态修复协同优化调度系统,解决了现有调度方式基于固定规则或人工经验,难以对光伏出力波动、沉陷地质变化与植被生长阶段等多维动态约束进行全局协调,无法实现协同效益最大化的问题

Benefits of technology

1、本发明通过数据采集模块同步采集光伏电站运行与沉陷区生态修复多源异构监测数据,并配合数据预处理与环境状态构建模块生成标准化协同状态向量,完成光伏与生态双系统状态的统一量化表征效果,解决两类系统状态维度异构、难以纳入同一决策框架的问题,为协同调度提供精准的环境输入基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840550A_ABST
    Figure CN122840550A_ABST
Patent Text Reader

Abstract

This invention provides a photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning, belonging to the field of photovoltaic power generation and ecological restoration collaborative scheduling technology. This system includes a data acquisition module, a data preprocessing module, an environmental state construction module, a deep reinforcement learning scheduling decision module, a reward evaluation and strategy update module, and a scheduling instruction decomposition and execution module. The data acquisition module collects operational data on irradiance, temperature, and power generation of the photovoltaic power plant, as well as monitoring data on soil moisture, vegetation cover, and surface deformation in the ecological restoration area of ​​the subsidence zone. By unifying the state representation of the photovoltaic and ecological restoration systems, it outputs compliant and continuous actions under multi-dimensional constraints. Through iterative optimization of multi-objective strategies using composite rewards, it can dynamically adapt to the long-term evolution of the subsidence zone environment, ensuring the maximization of comprehensive collaborative benefits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative scheduling technology for photovoltaic power generation and ecological restoration, specifically to a collaborative optimization scheduling system for photovoltaic power plants and ecological restoration that utilizes reinforcement learning. Background Technology

[0002] The integration of comprehensive management of coal mining subsidence areas with clean energy development has become an important path for the transformation of resource-based cities. Deploying photovoltaic power stations on the water or land surface in coal mining subsidence areas can generate clean electricity from idle land and water surfaces, while also providing funding and engineering conditions for ecological restoration of the subsidence areas. However, there are significant couplings and conflicts between the efficient operation of photovoltaic power stations and ecological restoration projects in terms of resource utilization and operational timing: Operation and maintenance activities such as cleaning photovoltaic modules, tilt adjustment, and inverter control require substantial water resources and manpower, while ecological restoration projects such as vegetation irrigation, soil improvement, and terrain shaping also rely on limited water resources, equipment, and suitable climate-geological conditions. If these two processes are operated independently, problems such as competition for water and equipment, and operational interference often arise, leading to a decrease in photovoltaic power generation due to dust obstruction or improper angles, and diminished ecological restoration effects due to insufficient irrigation or inappropriate timing. Existing scheduling methods are mostly based on fixed rules or human experience, which makes it difficult to coordinate the multi-dimensional dynamic constraints such as photovoltaic power output fluctuations, subsidence geological changes, and vegetation growth stages, and thus cannot achieve the synergistic maximization of photovoltaic power generation benefits and ecological restoration effects.

[0003] In recent years, reinforcement learning algorithms have demonstrated significant advantages in sequential decision-making in complex systems, enabling them to learn optimal control strategies through interaction with the environment in high-dimensional state spaces and continuous action spaces. However, applying reinforcement learning to photovoltaic-ecological collaborative scheduling in coal mining subsidence areas requires addressing a series of challenges, including state space construction, fusion of composite constraints, definition of multi-dimensional heterogeneous action spaces, and design of multi-objective reward functions. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a photovoltaic power plant and ecological restoration collaborative optimization scheduling system that utilizes reinforcement learning. This system solves the problem that existing scheduling methods, which are based on fixed rules or human experience, are unable to globally coordinate multi-dimensional dynamic constraints such as photovoltaic output fluctuations, subsidence geological changes, and vegetation growth stages, thus failing to maximize collaborative benefits.

[0005] To achieve the above objectives, this invention provides the following technical solution: a photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning, comprising a data acquisition module, a data preprocessing module, an environmental state construction module, a deep reinforcement learning scheduling decision module, a reward evaluation and strategy update module, and a scheduling instruction decomposition and execution module. The data acquisition module collects data on the irradiance, temperature, and power generation of the photovoltaic power station, as well as data on soil moisture, vegetation coverage, and surface deformation in the ecological restoration area of ​​the subsidence zone. The data preprocessing module is connected to the output of the data acquisition module to fill in missing values ​​and normalize the various types of data collected, and extract key features that affect the photovoltaic power output level and ecological status. The environment state construction module is connected to the data preprocessing module, and generates a collaborative state vector for the current decision step based on the extracted key features. The deep reinforcement learning scheduling decision module has a built-in scheduling strategy network. Its input is connected to the output of the environmental state construction module. With the collaborative state vector as input, under the joint constraint conditions of photovoltaic power generation operation constraints and ecological restoration operation constraints, the strategy network outputs collaborative scheduling actions. The collaborative scheduling actions include two types: photovoltaic operation parameter adjustment instructions and ecological restoration operation instructions. The reward evaluation and strategy update module is connected to the deep reinforcement learning scheduling decision module. It calculates the composite reward value based on the photovoltaic power generation revenue and the improvement of ecological restoration indicators after the execution of the collaborative scheduling action, and updates the parameters of the strategy network based on the composite reward value. The scheduling instruction decomposition and execution module is connected to the action output end of the deep reinforcement learning scheduling decision module, which parses the collaborative scheduling action into control instructions that can be recognized by the photovoltaic power station control system and the ecological restoration site execution terminal, and then issues them for execution.

[0006] Preferably, the data acquisition module includes a photovoltaic power station SCADA system interface, an automatic weather station, a soil moisture sensor network, a UAV multispectral remote sensing device, and a GNSS deformation monitoring station. The UAV multispectral remote sensing device collects vegetation index and surface crack images of the subsidence area at a set period, and the GNSS deformation monitoring station acquires surface subsidence rate and horizontal displacement data in real time.

[0007] Preferably, the deep reinforcement learning scheduling decision module uses the soft actor critic algorithm to construct the scheduling logic, and the policy network outputs the mean and variance of the action space to realize the exploration decision of continuous actions. The module is also equipped with a value network to output the Q value of the state-action pair. The joint constraints include limits on the rate of change of photovoltaic active power, limits on the reactive capacity of inverters, the range of adjustment of tracking bracket angle, limits on the state of charge of energy storage, the upper limit of water intake for a single irrigation, rules for avoiding fertilization operations and rainfall events, and time windows for manual operations during the vegetation growing season.

[0008] Preferably, the reward evaluation and strategy update module calculates the composite reward according to the following formula: ; In the formula, This refers to the revenue corresponding to the actual amount of photovoltaic power generated in the current period. This refers to the quantitative values ​​for improvements in ecological restoration indicators, which include vegetation cover growth rate, relative change rate of soil organic matter content, and the rate of reduction of waterlogged area in subsidence zones. The penalties for breaching the joint constraints cover penalties for exceeding water resource extraction limits, penalties for exceeding equipment operating parameter limits, and penalties for violating ecological operation time limits. , , These are weighting coefficients, determined based on multi-objective preferences or through adaptive adjustment.

[0009] Preferably, the scheduling instruction decomposition and execution module transmits instructions via industrial Ethernet or IoT communication protocols: On the one hand, the photovoltaic operating parameter adjustment instructions are sent to the photovoltaic inverter, tracking bracket controller, module cleaning robot and energy storage converter; On the other hand, ecological restoration operation instructions are sent to irrigation solenoid valves, integrated water and fertilizer devices, replanting drones, or ground mobile operation platforms, while receiving status feedback from each execution terminal to confirm the completion of the action.

[0010] A method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning includes the following steps: S1. Real-time collection of photovoltaic power station operation data and subsidence area ecological restoration monitoring data through the data acquisition module; S2. The data preprocessing module cleans and normalizes the collected raw data, and extracts key feature vectors that characterize the photovoltaic power output status and ecological restoration status. S3. The environment state construction module generates the collaborative state vector corresponding to the current decision step based on the aforementioned key feature vectors. S4. Input the collaborative state vector into the deep reinforcement learning scheduling decision module. This module obtains the scheduling strategy through offline pre-training and online adaptive fine-tuning. Under the joint constraints of photovoltaic power generation and ecological restoration, it outputs the collaborative scheduling action at the current time step. S5. The scheduling instruction decomposition and execution module converts the collaborative scheduling actions into device-level control instructions and issues them for execution, while simultaneously collecting environmental feedback data after the instruction execution. S6. The reward evaluation and strategy update module calculates the composite reward based on the environmental feedback data, stores the experience data consisting of the current state, action, reward and the next state into the experience replay pool, and samples a small batch of data from the experience replay pool to update the strategy network parameters. S7. Return to step S1 to proceed to the next decision step and complete the rolling optimization scheduling.

[0011] Preferably, the offline pre-training process of the deep reinforcement learning scheduling decision module in step S4 is as follows: based on historical meteorological data of the subsidence area, photovoltaic operation data and ecological monitoring data, a simulation environment is built by combining the photovoltaic power output model and the eco-hydrological model. The training objective is to maximize the cumulative discount reward. The soft actor critic algorithm is used to carry out offline training. During the training process, a priority experience replay mechanism and a policy entropy regularization term are introduced. Training stops after the policy reaches a convergence state on the validation set. The online adaptive fine-tuning phase employs a low learning rate, while the scheduling strategy is updated using an exponential moving average method to ensure operational stability.

[0012] Preferably, the collaborative state vector generated in step S3 includes at least the following state quantities: photovoltaic irradiance at the current time step, module backsheet temperature, module dust accumulation rate, energy storage state of charge, grid connection point voltage, and soil volumetric moisture content, normalized vegetation index, surface subsidence rate, cumulative time since the last irrigation, current ecological restoration stage identifier, and completion rate of the area to be restored in the ecological restoration area.

[0013] Preferably, in the composite reward, the photovoltaic power generation revenue is obtained by multiplying the on-grid electricity at the current time step by the time-of-use electricity price for the corresponding time period, and the improvement in ecological restoration effect is obtained by weighted summation of the increase in vegetation coverage, the increase in soil health indicators, and the reduction in the effective water accumulation area of ​​the subsidence area. The penalties for violating the constraints include three categories: linear penalties when irrigation water intake exceeds the allowable threshold, power over-limit penalties when the photovoltaic active power change rate exceeds the grid connection standard, and time violation penalties when replanting is carried out during the vegetation dormancy period.

[0014] Preferably, the cumulative improvement value of ecological restoration indicators and the cumulative photovoltaic power generation are statistically analyzed according to the preset evaluation cycle. When the ecological restoration indicators deviate from the preset target range, the weight coefficient in the composite reward is adjusted, and the local retraining of the deep reinforcement learning strategy network is triggered to adapt to the long-term evolution of geological conditions in the subsidence area and the natural succession process of vegetation communities.

[0015] This invention provides a collaborative optimization scheduling system for photovoltaic power plants and ecological restoration utilizing reinforcement learning. It offers the following advantages: 1. This invention uses a data acquisition module to simultaneously collect multi-source heterogeneous monitoring data on photovoltaic power plant operation and ecological restoration in subsidence areas. It also uses a data preprocessing and environmental status construction module to generate a standardized collaborative state vector, thereby achieving a unified quantitative representation of the states of both photovoltaic and ecological systems. This solves the problem of heterogeneous state dimensions between the two systems and the difficulty in incorporating them into the same decision-making framework, providing a precise environmental input basis for collaborative scheduling.

[0016] 2. This invention uses a deep reinforcement learning scheduling decision module to build a dual network architecture of strategy and value by employing a soft actor critic algorithm. It also incorporates an action truncation correction mechanism in the photovoltaic and ecological joint constraint layer to achieve compliant output of continuous heterogeneous actions in high-dimensional states. At the same time, it meets multi-dimensional constraints such as power grid connection, equipment safety, and ecological operation, ensuring that scheduling actions are always within the feasible domain.

[0017] 3. This invention constructs a composite reward function that includes power generation revenue, ecological improvement, and constraint penalties through a reward evaluation and strategy update module. Combined with experience playback and soft update mechanism for strategy parameters, it achieves iterative optimization of multi-objective collaborative scheduling strategy, which can dynamically balance the economic benefits of photovoltaic power generation and the environmental benefits of ecological restoration, thereby maximizing long-term comprehensive benefits.

[0018] 4. This invention realizes the closed loop of issuing and feedback of dual-domain device-level control commands through the scheduling command decomposition and execution module. Combined with the adaptive adjustment of reward weight and the local retraining mechanism of strategy, it achieves the dynamic adaptation effect of scheduling strategy with the long-term evolution of geology and vegetation in the subsidence area, ensuring that the system maintains the best collaborative scheduling performance in long-term operation. Attached Figure Description

[0019] Figure 1 This is a system framework diagram of the present invention; Figure 2 This is a flowchart of the steps of the present invention; Figure 3 This is a line graph showing the comparative data of the present invention; Figure 4 This is an internal interaction diagram of the deep reinforcement learning decision module of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Example: like Figure 1-4As shown, this embodiment of the invention provides a photovoltaic power station and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning, including a data acquisition module, a data preprocessing module, an environmental status construction module, a deep reinforcement learning scheduling decision module, a reward evaluation and strategy update module, and a scheduling instruction decomposition and execution module. The data acquisition module collects irradiance, temperature, and power generation operation data of the photovoltaic power station, as well as soil moisture, vegetation coverage, and surface deformation monitoring data of the ecological restoration area of ​​the subsidence zone. The data acquisition module includes a photovoltaic power station SCADA system interface, an automatic weather station, a soil moisture sensor network, a UAV multispectral remote sensing device, and a GNSS deformation monitoring station. The UAV multispectral remote sensing device collects vegetation index and surface crack images of the subsidence zone at a set period, and the GNSS deformation monitoring station acquires surface subsidence rate and horizontal displacement data in real time. The data preprocessing module is connected to the output of the data acquisition module to fill in missing values ​​and normalize the various types of data collected, and to extract key features that affect the photovoltaic power output level and ecological status. The environment state construction module is connected to the data preprocessing module, and generates a collaborative state vector for the current decision step based on the extracted key features; The deep reinforcement learning scheduling decision module has a built-in scheduling strategy network. Its input is connected to the output of the environmental state construction module. Taking the collaborative state vector as input, under the joint constraints of photovoltaic power generation operation constraints and ecological restoration operation constraints, the strategy network outputs collaborative scheduling actions. The collaborative scheduling actions include two types: photovoltaic operation parameter adjustment instructions and ecological restoration operation instructions. The deep reinforcement learning scheduling decision module uses the soft actor critic algorithm to construct the scheduling logic. The strategy network outputs the mean and variance of the action space to realize the exploration decision of continuous actions. The module is also equipped with a value network to output the Q value of the state-action pair. The joint constraints include the photovoltaic active power change rate limit, the inverter reactive power capacity limit, the tracking bracket angle adjustment range, the energy storage charge state limit, the upper limit of water intake for a single irrigation, the avoidance rules for fertilization operations and rainfall events, and the time window limit for artificial operations during the vegetation growth period. The reward assessment and strategy update module is integrated with the deep reinforcement learning scheduling decision module. It calculates a composite reward value based on the photovoltaic power generation revenue and ecological restoration indicator improvement feedback after the collaborative scheduling actions are executed, and updates the parameters of the strategy network based on this composite reward value. The reward assessment and strategy update module calculates the composite reward according to the following formula: In the formula, This refers to the revenue corresponding to the actual amount of photovoltaic power generated in the current period. This refers to the quantitative values ​​for improvements in ecological restoration indicators, which include vegetation cover growth rate, relative change rate of soil organic matter content, and the rate of reduction of waterlogged area in subsidence zones. The penalties for breaching the joint constraints cover penalties for exceeding water resource extraction limits, penalties for exceeding equipment operating parameter limits, and penalties for violating ecological operation time limits. , , These are weighting coefficients, determined based on multi-objective preferences or through adaptive adjustment; The scheduling instruction decomposition and execution module is connected to the action output end of the deep reinforcement learning scheduling decision module. It parses the collaborative scheduling actions into control instructions that can be recognized by the photovoltaic power station control system and the ecological restoration field execution terminal, and issues them for execution. The scheduling instruction decomposition and execution module transmits instructions through industrial Ethernet or IoT communication protocols: on the one hand, it issues photovoltaic operating parameter adjustment instructions to photovoltaic inverters, tracking bracket controllers, module cleaning robots and energy storage converters; on the other hand, it issues ecological restoration operation instructions to irrigation solenoid valves, water and fertilizer integration devices, replanting drones or ground mobile operation platforms, and at the same time receives status feedback from each execution terminal to confirm that the action has been completed.

[0022] Example 1: When deploying this collaborative optimization scheduling system within the photovoltaic power station and ecological restoration area of ​​the coal mining subsidence zone, the data acquisition module periodically reads operating parameters such as inverter output power, module backplane temperature, combiner box current, and grid connection point voltage via the photovoltaic power station's SCADA system interface using the Modbus TCP protocol. Simultaneously, it connects to automatic weather stations arranged between the photovoltaic arrays to obtain data on total horizontal irradiance, ambient temperature, and wind speed and direction. On one side of the subsidence zone, a soil moisture sensor network is embedded in different restoration zones in a grid pattern, reporting the volumetric water content of layered soils via LoRa wireless transmission. A UAV multispectral remote sensing device conducts flight path coverage photography of the subsidence zone at a set cycle of 72 hours, and the data is processed by the ground station to generate a normalized vegetation index distribution map and surface crack identification results. A GNSS deformation monitoring station uses a BeiDou and GPS dual-mode receiver to calculate the three-dimensional displacement of monitoring points in real time, outputting the surface subsidence rate and horizontal displacement. The aforementioned multi-source heterogeneous data is aggregated through an edge gateway and uploaded to the dispatch center server via a fiber optic ring network.

[0023] Example 2: After receiving the raw data stream from the data acquisition module, the data preprocessing module first aligns the data sources according to the timestamps. For missing values ​​caused by communication interruptions or sensor malfunctions, a sliding window interpolation method based on the historical average of the same period is used to fill in the gaps. Simultaneously, the Laida criterion is used to remove abnormal jumps caused by electromagnetic interference or equipment vibration. After cleaning, the module performs maximum-minimum normalization on electrical parameters such as photovoltaic irradiance, temperature, and power, and ecological parameters such as soil moisture, vegetation index, and sedimentation rate, mapping data of different dimensions to the [0,1] interval. Subsequently, the module uses Pearson correlation analysis and recursive feature elimination to select a subset of key features significantly related to photovoltaic output level and ecological restoration status from the normalized feature set. These features include, but are not limited to, effective irradiance, estimated component dust accumulation rate, energy storage state of charge, root zone soil moisture content, normalized vegetation index increment, and sedimentation rate change rate. The feature vector is then passed to the environmental status construction module.

[0024] Example 3: After acquiring key features from the data preprocessing module, the environmental state construction module uses a fixed 15-minute decision step to concatenate the current photovoltaic power generation-side state variables and the subsidence area ecological-side state variables into a one-dimensional collaborative state vector. The photovoltaic-side state component in the vector includes the average irradiance, module backsheet temperature, estimated ash accumulation rate, remaining energy storage capacity percentage, and grid connection voltage deviation for that time step. The ecological-side state component includes the average soil volumetric moisture content of each restoration zone, the change in the latest vegetation index compared to the previous period, the root mean square value of the surface subsidence rate, the cumulative time since the last effective irrigation, and the current ecological restoration stage identifier. This collaborative state vector serves as the input to the deep reinforcement learning scheduling decision module, comprehensively depicting the real-time operating conditions of the photovoltaic power station and the ecological restoration area.

[0025] Example 4: The deep reinforcement learning scheduling decision module's embedded policy network adopts a four-layer fully connected network structure. The first two layers are shared hidden layers, each containing 256 neurons and using the ReLU activation function. The last two layers branch into photovoltaic action branches and ecological action branches, each outputting the mean and logarithmic standard deviation of the corresponding action dimension. Specific continuous collaborative scheduling actions are generated through reparameterization techniques. The value network is a three-layer fully connected network that takes the concatenation of the state vector and action vector as input and outputs the Q-value of the corresponding state-action pair, used to evaluate the long-term reward of the current policy. In the fusion constraint layer, the module encodes joint constraints such as the limit of photovoltaic active power change rate, the upper and lower bounds of the tracking support mechanical angle, the safe range of energy storage charge state, the maximum water intake for a single irrigation, the mandatory avoidance rules for fertilization and rainfall events, and the allowable operation time window corresponding to different vegetation growth stages. When sampling actions, the policy network needs to be truncated and corrected by the constraint layer to ensure that the output photovoltaic operation parameter adjustment instructions and ecological restoration operation instructions are both within the feasible domain.

[0026] Example 5: After each decision-making step is completed, the reward assessment and strategy update module calculates the composite reward value for that step based on the execution confirmation and environmental monitoring feedback returned by the scheduling instruction decomposition and execution module. The photovoltaic power generation revenue item is determined by multiplying the actual grid-connected electricity generated by the photovoltaic power station during that period by the local electricity market time-of-use price. The ecological restoration improvement item is a weighted composite of the absolute increase in vegetation coverage compared to the previous assessment period, the relative change rate of soil organic matter content in the test samples, and the reduction ratio of the effective water accumulation area in the subsidence zone, after dimensionless processing. The constraint violation penalty item includes a linearly increasing penalty for irrigation water withdrawal exceeding the allowable upper limit, a step penalty for the photovoltaic active power change rate exceeding the grid connection standard, and a one-time fixed penalty for replanting during the vegetation dormancy period. The module stores the current state, action, reward, and next state as an experience quadruple and puts them into an experience replay pool with a capacity of 100,000 records. When the amount of data in the replay pool exceeds 5,000 records, 256 experience data are randomly sampled from the pool each time. The value network parameters are updated using the mean squared error loss in the SAC algorithm. The policy network parameters are adaptively adjusted and updated using the policy gradient and temperature parameters. The target network gradually approaches the main network with a soft update method at a coefficient of 0.005.

[0027] Example 6: After receiving the collaborative scheduling actions output by the deep reinforcement learning scheduling decision module, the scheduling instruction decomposition and execution module decomposes them into a set of device-level control instructions through the action decoder. For photovoltaic operating parameter adjustment instructions, the module sends active power limit and reactive power instructions to the photovoltaic inverter via the industrial Ethernet EtherCAT protocol, sends tilt angle target values ​​to the tracking bracket controller, sends cleaning path, travel speed and water spray volume to the module cleaning robot, and sends charging and discharging power instructions to the energy storage converter. For ecological restoration operation instructions, the module sends opening duration and target flow to the irrigation solenoid valve via the IoT MQTT protocol, sends fertilizer ratio and injection rate to the water and fertilizer integration device, sends target plot coordinates and sowing density to the replanting drone, and sends trimming path and operation depth to the ground mobile operation platform. After each execution terminal completes its action, it sends a status code and execution time back to the module. The module then summarizes these information to form an action completion confirmation signal. At the same time, it packages and feeds back the actual operating parameters of each device and the changes in ecological indicators to the reward assessment and strategy update module, forming a complete closed-loop online scheduling and optimization link. This link operates continuously with a 15-minute rolling cycle, constantly improving the synergistic benefits of photovoltaic power generation and ecological restoration.

[0028] A method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning includes the following steps: S1. Real-time collection of photovoltaic power station operation data and subsidence area ecological restoration monitoring data through the data acquisition module; S2. The data preprocessing module cleans and normalizes the collected raw data, and extracts key feature vectors that characterize the photovoltaic power output status and ecological restoration status. S3. The environmental state construction module generates a collaborative state vector corresponding to the current decision step based on the aforementioned key feature vectors. The generated collaborative state vector includes at least the following state quantities: photovoltaic irradiance at the current time step, module backsheet temperature, module dust accumulation rate, energy storage state of charge, grid connection point voltage, as well as soil volumetric moisture content, normalized vegetation index, surface subsidence rate, cumulative time since the last irrigation, current ecological restoration stage identifier, and completion rate of the area to be restored in the ecological restoration area. S4. The collaborative state vector is input into the deep reinforcement learning scheduling decision module. This module obtains the scheduling strategy through offline pre-training and online adaptive fine-tuning. Under the joint constraints of photovoltaic power generation and ecological restoration, it outputs the collaborative scheduling action at the current time step. The offline pre-training process of the deep reinforcement learning scheduling decision module is as follows: Based on historical meteorological data, photovoltaic operation data and ecological monitoring data of the subsidence area, a simulation environment is built by combining the photovoltaic power output model and the eco-hydrological model. The training objective is to maximize the cumulative discount reward. The soft actor critic algorithm is used for offline training. During the training process, a priority experience replay mechanism and a policy entropy regularization term are introduced. Training stops after the policy reaches a convergent state on the validation set. In the online adaptive fine-tuning stage, a low learning rate is used, and the stability of the scheduling strategy is ensured by updating the policy parameters through an exponential moving average. S5. The scheduling instruction decomposition and execution module converts the collaborative scheduling actions into device-level control instructions and issues them for execution, while simultaneously collecting environmental feedback data after the instruction execution. S6. The reward evaluation and strategy update module calculates the composite reward based on the environmental feedback data, stores the experience data consisting of the current state, action, reward and the next state into the experience replay pool, and samples a small batch of data from the experience replay pool to update the strategy network parameters. S7. Return to step S1 to proceed to the next decision step and complete the rolling optimization scheduling.

[0029] In the composite reward system, the photovoltaic power generation revenue is obtained by multiplying the on-grid electricity volume at the current time step by the time-of-use electricity price for the corresponding time period. The improvement in ecological restoration effect is obtained by weighted summation of the increase in vegetation coverage, the increase in soil health indicators, and the reduction in the effective water accumulation area of ​​the subsidence area. The constraint violation penalties include three types: a linear penalty term when irrigation water intake exceeds the allowable threshold, a power limit violation penalty term when the photovoltaic active power change rate exceeds the grid connection standard, and a time violation penalty term when replanting is carried out during the vegetation dormancy period. The cumulative improvement value of ecological restoration indicators and the cumulative photovoltaic power generation are calculated according to the preset evaluation cycle. When the ecological restoration indicators deviate from the preset target range, the weight coefficients in the composite reward are adjusted, and the local retraining of the deep reinforcement learning strategy network is triggered to adapt to the long-term evolution of geological conditions in the subsidence area and the natural succession process of the vegetation community.

[0030] Example 7: In step S1, the data acquisition module uses the synchronous timing signal as a reference and the NTP network timing protocol to align the photovoltaic power station SCADA system, automatic weather station, soil moisture sensor network, UAV remote sensing ground station, and GNSS deformation monitoring station in time, ensuring that the data collected from different data sources have a unified timestamp. On the photovoltaic side, the SCADA system collects the inverter output power, module backsheet temperature, and grid connection point voltage at 1-second intervals, and calculates the average irradiance and power change rate every 5 minutes via a sliding window. On the ecological side, the soil moisture sensor wakes up once per hour, uploads data through the LoRa gateway after completing the measurement, the UAV performs a full coverage flight every 72 hours, transmitting visible light and multispectral images in real time during the flight, and the ground station simultaneously stitches together orthophotos and generates a vegetation index distribution map; the GNSS monitoring station continuously records at a sampling rate of 1 Hz and outputs a set of calculated three-dimensional displacement rates every 15 minutes. This heterogeneous multi-source synchronous acquisition mechanism ensures the alignment accuracy of the cooperative state vector constructed in step S3 in the time dimension.

[0031] In step S2, the data preprocessing module does not simply use mean interpolation to fill in missing values, but combines temporal proximity and spatial similarity. For data loss caused by a temporary loss of connection of a soil moisture sensor, the arithmetic mean of synchronous data from adjacent sensors within the same partition is preferentially used for filling; if the data for the entire partition is interrupted, historical data from the same period in the same vegetation growth stage of that partition are used for filling. During normalization, photovoltaic irradiance, module backsheet temperature, and power generation are normalized using maximum and minimum values, while soil volumetric water content and vegetation index are normalized using Z-score to adapt to their approximate normal distribution characteristics. The extraction of key feature vectors uses gradient boosting trees to calculate feature importance, eliminating redundant features with importance below 0.01, and finally retaining feature vectors with a dimension of less than 24 to reduce the input dimension of the subsequent scheduling decision module and accelerate inference speed.

[0032] In step S3, when generating the collaborative state vector, in addition to the state variables listed in the claims, Boolean-type identifiers are introduced based on the phased characteristics of the ecological restoration project. For example, the current ecological restoration phase identifiers are subdivided into five types: "terrain preparation phase," "soil improvement phase," "pioneer plant planting phase," "community construction phase," and "stabilization and maintenance phase," each corresponding to different permitted operation types and irrigation thresholds. The cumulative time since the last irrigation is accurately calculated by reading the historical action records of the irrigation solenoid valve. If the cumulative time exceeds the maximum drought duration recommended for this vegetation growth stage, the value of this state component will be amplified and mapped to strengthen the expression of the urgency of irrigation in the vector. The completion rate of the area to be restored is obtained by the ratio of the greened area interpreted from UAV imagery to the planned total area, and compared with the phase target progress. The difference is incorporated into the state vector, enabling the scheduling decision module to perceive the deviation in restoration progress.

[0033] The simulation environment built for offline pre-training in step S4 consists of a coupled photovoltaic power output physical model and an eco-hydrological model. The photovoltaic power output model takes horizontal irradiance, temperature, module dust accumulation rate, and tracking bracket angle as inputs. It calculates the module IV characteristics using a single-diode equivalent circuit model and then obtains the AC output power through the inverter efficiency curve. The eco-hydrological model uses an improved HYDRUS-1D numerical model to simulate soil moisture transport, taking irrigation and rainfall as inputs and outputting a sequence of soil moisture content changes at different depths. It also simulates the response of vegetation cover to moisture and accumulated temperature based on the Logistic growth equation. The simulation environment's time step is set to 15 minutes, consistent with the real system's decision-making time step. During offline training, a priority experience replay mechanism is used, giving higher sampling priority to experiences that lead to constraint violation penalties, accelerating the policy network's learning of forbidden zones. The initial coefficient of the policy entropy regularization term is 0.2, gradually annealing to 0.05 during training to transition from exploration to utilization. In the online adaptive fine-tuning stage, the learning rate of the policy network is fixed at 1×10⁻⁶. -5 Furthermore, after each update, the strategy parameters are subjected to an exponential moving average with a decay coefficient of 0.995 to suppress strategy jitter caused by online environmental noise.

[0034] In step S5, the scheduling instruction decomposition and execution module introduces an instruction conflict detection and security protection layer when converting collaborative scheduling actions into device-level instructions. Within the same decision-making time step, if there is a risk of interference between the tilt adjustment instruction of the photovoltaic tracking bracket and the walking path of the component cleaning robot, the module delays the cleaning task according to the preset safety priority, prioritizing the tilt adjustment. The cleaning instruction is activated only after the bracket completes its action and sends a feedback signal. Regarding ecological restoration instructions, when the irrigation solenoid valve opening instruction is issued, the module simultaneously reads the rainfall probability prediction from the rainfall monitoring station for the next hour. If the probability exceeds 70%, the irrigation instruction is automatically canceled and a record is generated. All issued instructions carry a checksum. The confirmation information returned by the execution terminal includes a comparison between the actual execution parameters and the instruction. When the deviation exceeds 5%, the module marks the action as an abnormal action and excludes it from the writing range of the current round of experience playback pool, preventing erroneous experience from contaminating strategy training.

[0035] Step S6, when calculating the composite reward, sets phased baselines for the increase in vegetation cover and soil health indicators in the improvement of ecological restoration effects. In the early stages of subsidence area restoration, as vegetation cover increases from near zero to 15%, its contribution to the reward is significant. Once coverage exceeds 40%, the contribution rate halves, guiding the strategy to allocate more resources to photovoltaic power generation in the later stages of restoration. The increase in soil health indicators is comprehensively assessed based on the organic matter and total nitrogen content obtained from periodic sampling and testing. Due to the long testing cycle, this indicator remains unchanged until new test results are available and is not included in real-time reward calculations. The experience replay pool adopts a circular queue structure with a storage limit of 200,000 experiences. A hybrid sampling strategy is used during sampling: 80% of batch experiences come from priority replay weighted sampling, and 20% come from uniform random sampling, to balance high-value experiences with experience diversity.

[0036] In step S7, a phased statistical analysis is conducted over a preset evaluation period of 168 hours. During this period, the moving average of the cumulative improvement value of the ecological restoration index and the cumulative photovoltaic power generation is calculated. When the vegetation coverage growth rate is lower than the lower limit of the target range for two consecutive evaluation periods, or when the rate of water receding in the subsidence area stagnates for more than four evaluation periods, the ecological restoration index is determined to have deviated from the preset target range. At this time, the weight coefficient α in the composite reward is automatically reduced by 0.05, β is automatically increased by 0.05, while γ remains unchanged, to enhance the incentive for ecological improvement. After the weight adjustment, the system triggers local retraining of the deep reinforcement learning policy network: the underlying shared hidden layer parameters of the policy network are frozen, and only the output layer parameters of the photovoltaic action branch and the ecological action branch, as well as the fully connected layer of the value network, are fine-tuned with a small sample. The data used for fine-tuning is taken from the experience within the four evaluation periods before the retraining trigger time, and the number of fine-tuning rounds is limited to within 200 rounds to avoid overfitting to historical local data. This adaptive retraining mechanism enables the scheduling strategy to follow the slowdown in the geological subsidence rate of the subsidence area or the shift in the state distribution caused by the natural succession of vegetation communities, thus maintaining a dynamic balance between the benefits of photovoltaic power generation and the effects of ecological restoration.

[0037] Table 1:

[0038] Table 2:

[0039] Table 3;

[0040] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning, comprising a data acquisition module, a data preprocessing module, an environmental state construction module, a deep reinforcement learning scheduling decision-making module, a reward evaluation and strategy update module, and a scheduling instruction decomposition and execution module, characterized in that: The data acquisition module collects data on the irradiance, temperature, and power generation of the photovoltaic power station, as well as data on soil moisture, vegetation coverage, and surface deformation in the ecological restoration area of ​​the subsidence zone. The data preprocessing module is connected to the output of the data acquisition module to fill in missing values ​​and normalize the various types of data collected, and extract key features that affect the photovoltaic power output level and ecological status. The environment state construction module is connected to the data preprocessing module, and generates a collaborative state vector for the current decision step based on the extracted key features. The deep reinforcement learning scheduling decision module has a built-in scheduling strategy network. Its input is connected to the output of the environmental state construction module. With the collaborative state vector as input, under the joint constraint conditions of photovoltaic power generation operation constraints and ecological restoration operation constraints, the strategy network outputs collaborative scheduling actions. The collaborative scheduling actions include two types: photovoltaic operation parameter adjustment instructions and ecological restoration operation instructions. The reward evaluation and strategy update module is connected to the deep reinforcement learning scheduling decision module. It calculates the composite reward value based on the photovoltaic power generation revenue and the improvement of ecological restoration indicators after the execution of the collaborative scheduling action, and updates the parameters of the strategy network based on the composite reward value. The scheduling instruction decomposition and execution module is connected to the action output end of the deep reinforcement learning scheduling decision module, which parses the collaborative scheduling action into control instructions that can be recognized by the photovoltaic power station control system and the ecological restoration site execution terminal, and then issues them for execution.

2. The photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning as described in claim 1, characterized in that: The data acquisition module includes a photovoltaic power station SCADA system interface, an automatic weather station, a soil moisture sensor network, a UAV multispectral remote sensing device, and a GNSS deformation monitoring station. The UAV multispectral remote sensing device collects vegetation index and surface crack images of the subsidence area at a set period, and the GNSS deformation monitoring station acquires surface subsidence rate and horizontal displacement data in real time.

3. The photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning as described in claim 1, characterized in that: The deep reinforcement learning scheduling decision module uses the soft actor critic algorithm to construct the scheduling logic. The policy network outputs the mean and variance of the action space to realize the exploration decision of continuous actions. The module is also equipped with a value network to output the Q value of the state-action pair. The joint constraints include limits on the rate of change of photovoltaic active power, limits on the reactive capacity of inverters, the range of adjustment of tracking bracket angle, limits on the state of charge of energy storage, the upper limit of water intake for a single irrigation, rules for avoiding fertilization operations and rainfall events, and time windows for manual operations during the vegetation growing season.

4. The photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning as described in claim 3, characterized in that: The reward evaluation and strategy update module calculates the composite reward according to the following formula: ; In the formula, This refers to the revenue corresponding to the actual amount of photovoltaic power generated in the current period. This refers to the quantitative values ​​for improvements in ecological restoration indicators, which include vegetation cover growth rate, relative change rate of soil organic matter content, and the rate of reduction of waterlogged area in subsidence zones. The penalties for breaching the joint constraints cover penalties for exceeding water resource extraction limits, penalties for exceeding equipment operating parameter limits, and penalties for violating ecological operation time limits. , , These are weighting coefficients, determined based on multi-objective preferences or through adaptive adjustment.

5. The photovoltaic power plant and ecological restoration collaborative optimization scheduling system utilizing reinforcement learning as described in claim 1, characterized in that: The scheduling instruction decomposition and execution module transmits instructions via industrial Ethernet or IoT communication protocols. On the one hand, the photovoltaic operating parameter adjustment instructions are sent to the photovoltaic inverter, tracking bracket controller, module cleaning robot and energy storage converter; On the other hand, ecological restoration operation instructions are sent to irrigation solenoid valves, integrated water and fertilizer devices, replanting drones, or ground mobile operation platforms, while receiving status feedback from each execution terminal to confirm the completion of the action.

6. A method for coordinated optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning, as described in any one of claims 1 to 5, characterized in that, Includes the following steps: S1. Real-time collection of photovoltaic power station operation data and subsidence area ecological restoration monitoring data through the data acquisition module; S2. The data preprocessing module cleans and normalizes the collected raw data, and extracts key feature vectors that characterize the photovoltaic power output status and ecological restoration status. S3. The environment state construction module generates the collaborative state vector corresponding to the current decision step based on the aforementioned key feature vectors. S4. Input the collaborative state vector into the deep reinforcement learning scheduling decision module. This module obtains the scheduling strategy through offline pre-training and online adaptive fine-tuning. Under the joint constraints of photovoltaic power generation and ecological restoration, it outputs the collaborative scheduling action at the current time step. S5. The scheduling instruction decomposition and execution module converts the collaborative scheduling actions into device-level control instructions and issues them for execution, while simultaneously collecting environmental feedback data after the instruction execution. S6. The reward evaluation and strategy update module calculates the composite reward based on the environmental feedback data, stores the experience data consisting of the current state, action, reward and the next state into the experience replay pool, and samples a small batch of data from the experience replay pool to update the strategy network parameters. S7. Return to step S1 to proceed to the next decision step and complete the rolling optimization scheduling.

7. The method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning as described in claim 6, characterized in that: The offline pre-training process of the deep reinforcement learning scheduling decision module in step S4 is as follows: Based on historical meteorological data of the subsidence area, photovoltaic operation data and ecological monitoring data, a simulation environment is built by combining photovoltaic power output model and eco-hydrological model. The training objective is to maximize the cumulative discount reward. The soft actor critic algorithm is used to carry out offline training. During the training process, a priority experience replay mechanism and a policy entropy regularization term are introduced. Training stops after the policy reaches a convergence state on the validation set. The online adaptive fine-tuning phase employs a low learning rate, while the scheduling strategy is updated using an exponential moving average method to ensure operational stability.

8. The method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning as described in claim 6, characterized in that: The collaborative state vector generated in step S3 includes at least the following state quantities: photovoltaic irradiance at the current time step, module backsheet temperature, module dust accumulation rate, energy storage state of charge, grid connection point voltage, soil volumetric moisture content, normalized vegetation index, land subsidence rate, cumulative time since the last irrigation, current ecological restoration stage identifier, and completion rate of the area to be restored.

9. The method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning as described in claim 6, characterized in that: In the aforementioned composite reward, the photovoltaic power generation revenue is obtained by multiplying the on-grid electricity volume at the current time step by the time-of-use electricity price for the corresponding time period, and the improvement in ecological restoration effect is obtained by weighted summation of the increase in vegetation coverage, the increase in soil health indicators, and the reduction in the effective water accumulation area of ​​the subsidence zone. The penalties for violating the constraints include three categories: linear penalties when irrigation water intake exceeds the allowable threshold, power over-limit penalties when the photovoltaic active power change rate exceeds the grid connection standard, and time violation penalties when replanting is carried out during the vegetation dormancy period.

10. The method for collaborative optimization scheduling of photovoltaic power plants and ecological restoration using reinforcement learning according to claim 6, characterized in that: The cumulative improvement value of ecological restoration indicators and the cumulative photovoltaic power generation are statistically analyzed according to the preset evaluation cycle. When the ecological restoration indicators deviate from the preset target range, the weight coefficient in the composite reward is adjusted, and the local retraining of the deep reinforcement learning strategy network is triggered to adapt to the long-term evolution of geological conditions in the subsidence area and the natural succession process of vegetation community.