Glacier melt water resource cooperative regulation and ecological restoration system and method
By constructing a glacial meltwater resource regulation system based on multi-source sensing data and multi-agent collaborative decision-making, the problem of insufficient water-soil-vegetation coupling regulation in traditional water resource management has been solved. This system achieves efficient water resource utilization and vegetation restoration, possesses autonomous evolution capabilities, and adapts to the non-stationarity and geological heterogeneity of the meltwater process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG UNIVERSITY
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional water resource management methods lack a systematic understanding and dynamic coordinated control of the multi-element coupling process of 'water-soil-plant', resulting in low water resource utilization efficiency, uncontrollable recharge process, unstable vegetation survival rate, sparse monitoring methods, and static control strategies, making it difficult to cope with the non-stationarity of meltwater flow and the heterogeneity of geological conditions.
A collaborative regulation and ecological restoration system for glacial meltwater resources is constructed. This system employs multi-source environmental sensing data, a multi-agent collaborative decision-making model, and an autonomous evolution mechanism to achieve high spatiotemporal resolution sensing, multi-objective collaborative decision-making, and precise execution. Through dynamic sensing via ground sensor networks and unmanned platforms, combined with a deep reinforcement learning model, hydrological prediction, recharge control, and vegetation irrigation decisions are made to form a global optimized regulation command.
It improved the efficiency of glacial meltwater resource utilization, achieved the synchronization of groundwater recharge and vegetation ecological restoration, enhanced the system's adaptability to extreme climate events, and improved groundwater recharge efficiency, vegetation survival rate, and system energy efficiency.
Smart Images

Figure CN121961808A_ABST
Abstract
Description
A system and method for coordinated regulation and ecological restoration of glacial meltwater resources Technical Field
[0001] This invention relates to the field of water conservancy engineering and ecological restoration technology, specifically to a system and method for the coordinated regulation and ecological restoration of glacial meltwater resources. Background Technology
[0002] Global climate change is accelerating glacial melting and increasing the uncertainty of seasonal meltwater runoff. At the same time, groundwater over-extraction and ecological degradation are becoming increasingly prominent problems in arid and semi-arid regions. Traditional water resource management methods often implement surface water allocation, artificial groundwater recharge, and vegetation restoration projects in a fragmented manner, lacking a systematic understanding and dynamic coordinated control of the multi-factor coupling process of "water-soil-planting". This leads to problems such as low water resource utilization efficiency, uncontrollable recharge process, and unstable vegetation survival rate.
[0003] In existing technologies, monitoring methods mostly rely on sparsely deployed fixed sensors, making it difficult to achieve high spatiotemporal resolution for full-area perception; control strategies are mostly static rules based on thresholds, which cannot adapt to the non-stationarity of meltwater flow and the heterogeneity of underlying geological conditions; in addition, the systems generally lack the ability to learn autonomously and evolve strategies from operational data, making it difficult to cope with extreme climate events and long-term environmental changes.
[0004] Therefore, there is an urgent need for an intelligent system and method that can achieve full-basin perception, multi-objective collaborative decision-making, precise execution, and autonomous evolution, in order to improve the utilization efficiency of glacial meltwater resources and simultaneously achieve groundwater replenishment and vegetation ecological restoration. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a system and method for the coordinated regulation and ecological restoration of glacial meltwater resources.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] This application provides a system for the coordinated regulation and ecological restoration of glacial meltwater resources, including:
[0008] The data sensing module is used to acquire multi-source environmental sensing data related to glacial meltwater within the target area;
[0009] The intelligent decision-making module is communicatively connected to the data perception module and is used to generate collaborative control instructions for glacial meltwater diversion control, groundwater recharge control, and vegetation irrigation control based on the multi-source environmental perception data and using a collaborative decision-making model containing multiple deep reinforcement learning agents.
[0010] The instruction execution and feedback module is communicatively connected to the intelligent decision-making module and is used to execute the coordinated control instructions and collect environmental feedback data after the instructions are executed.
[0011] The strategy evolution module is communicatively connected to the instruction execution and feedback module, and is used to continuously train and update the collaborative decision-making model based on the environmental feedback data.
[0012] Furthermore, the data sensing module includes:
[0013] A ground sensor network is fixedly deployed at key locations in the target area to periodically collect environmental data, including at least meteorological, hydrological, soil, and vegetation physiological parameters.
[0014] The dynamic perception unit based on the unmanned platform is used to respond to the instructions of the intelligent decision-making module to perform high-resolution spatial scanning of a specific area and obtain spatial distribution data of surface material composition, micro-topography and vegetation physiological state.
[0015] Furthermore, the intelligent decision-making module includes:
[0016] The hydrological prediction unit uses a recurrent neural network to predict glacial meltwater runoff in future periods based on time-series environmental sensing data.
[0017] The recharge control unit, based on soil hydrological characteristics, groundwater status and predicted runoff, uses a first deep reinforcement learning model to output control decisions on the infiltration rate of each recharge unit.
[0018] The irrigation decision unit, based on vegetation water demand characteristics, soil moisture status and weather forecasts, uses a second deep reinforcement learning model to output irrigation decisions for each vegetation unit.
[0019] The central coordinator is used to perform global multi-objective optimization on the preliminary decisions of the hydrological prediction unit, the recharge control unit, and the irrigation decision unit, and to generate the coordinated control instructions.
[0020] Furthermore, the multi-objective optimization function upon which the central coordinator is based is expressed as:
[0021]
[0022] in, This indicates the net recharge of groundwater. Indicates the amount of invalid surface runoff loss. This represents the overall score indicating the health of vegetation growth. This indicates the total energy consumption of the system. These are non-negative weighting coefficients set according to project priorities; the central coordinator outputs the globally optimal control instruction set by solving this optimization problem.
[0023] Furthermore, the instruction execution and feedback module includes:
[0024] The instruction parsing and driving unit is used to parse the coordinated control instructions into low-level control signals and drive the water diversion, recharge and irrigation execution equipment to operate.
[0025] The effect feedback unit is used to monitor the execution status of the device and collect the change data of key environmental parameters again after the instruction is executed, so as to form the environmental feedback data.
[0026] Furthermore, the strategy evolution module includes:
[0027] An experience memory is used to store historical decision and feedback data in the form of state-action-reward-new-state tuples.
[0028] An offline training engine is used to periodically sample data from the experience memory to train the deep reinforcement learning agent in the collaborative decision-making model and update its decision-making strategy.
[0029] Furthermore, the offline training engine is configured to trigger specialized training for the event when the environmental feedback data indicates that the system has encountered a preset type of extreme weather event, so as to quickly adjust the decision-making strategy.
[0030] Secondly, this application provides a method for the coordinated regulation and ecological restoration of glacial meltwater resources, including the following steps:
[0031] S1. Data acquisition steps: Continuously acquire environmental sensing data of the target area through a multi-source sensing network;
[0032] S2. Collaborative decision-making step: Based on the environmental perception data, a multi-agent deep reinforcement learning model is used to make collaborative decisions and generate a global optimization command that integrates water diversion, recharge and irrigation regulation;
[0033] S3. Instruction Execution and Feedback Steps: Execute the global optimization instruction and collect environmental feedback data after instruction execution;
[0034] S4. Policy evolution step: Based on the environmental feedback data, the multi-agent deep reinforcement learning model is trained offline to update its decision-making strategy.
[0035] Furthermore, the global optimization in the S2 collaborative decision-making step aims to maximize the net groundwater recharge and vegetation growth health, while minimizing ineffective runoff loss and system energy consumption as a comprehensive optimization objective.
[0036] Furthermore, the S4 strategy evolution step includes: when an extreme weather event is detected, using data from similar historical events to reinforce the model in order to improve the system's adaptability to sudden disturbances.
[0037] Compared with the prior art, this application has the following beneficial effects:
[0038] This invention overcomes the shortcomings of traditional monitoring methods, such as sparseness and lag, by constructing an integrated ground-based sensing system combining a fixed ground-based sensor network and dynamic aerial scanning by unmanned aerial vehicles (UAVs). It achieves high spatiotemporal resolution synchronous perception of glacial meltwater runoff, soil moisture, and vegetation status. Furthermore, it employs a multi-agent collaborative decision-making architecture based on deep reinforcement learning to decompose the fragmented water resource regulation problem into three agents: hydrological prediction, recharge control, and vegetation irrigation. These agents are then optimized locally, and a central coordinator performs global multi-objective dynamic trade-offs. This effectively addresses the non-stationarity of the meltwater process, the heterogeneity of geological conditions, and the spatiotemporal variability of ecological water demand, thus achieving… It achieves coordinated and efficient allocation of water resources in the stages of water diversion, recharge, and irrigation; through a complete control closed loop of instruction parsing, drive execution, and effect feedback, it ensures the reliable implementation and precise execution of intelligent decisions; finally, relying on a self-evolutionary learning mechanism based on experience playback and offline training, the system can continuously learn and optimize strategies from historical operating data, significantly improving groundwater recharge efficiency, vegetation survival rate, and overall system energy efficiency. At the same time, it has the ability to adapt and evolve in response to extreme climate events, fundamentally solving the systemic defects of existing technologies, such as single perception dimension, static decision logic, isolated execution units, and lack of system learning ability. Attached Figure Description
[0039] Figure 1 is a schematic diagram of the overall architecture of the glacial meltwater resource collaborative regulation and ecological restoration system in one embodiment of the present invention.
[0040] Figure 2 is a schematic diagram of the internal structure and data flow of the intelligent decision-making module in one embodiment of the present invention.
[0041] Figure 3 is a schematic diagram of the composition and collaborative operation of the data sensing module in one embodiment of the present invention.
[0042] Figure 4 is a schematic diagram of the workflow of the instruction execution and feedback module in one embodiment of the present invention.
[0043] Figure 5 is a schematic diagram of the interaction and training process between the strategy evolution module and the intelligent decision-making module in one embodiment of the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Furthermore, in this invention, an element referred to as fixed to or disposed on another element may be directly disposed on the other element, or there may be an intermediate element. When an element is considered to be connected to another element, it may be directly connected to the other element, or there may be an intermediate element present simultaneously. The terms vertical, horizontal, left, right, and similar expressions used herein are for illustrative purposes only and do not represent the only possible implementation.
[0046] Example 1
[0047] Referring to Figures 1-5, the technical solution provided in this application is applied to a typical glacial meltwater basin on the Qinghai-Tibet Plateau. This region faces multiple challenges, including accelerated glacial melting, highly non-stationary seasonal meltwater runoff, insufficient shallow aquifer recharge, and severe degradation of alpine meadow and shrub ecosystems. The system and method of this invention are deployed and implemented in this scenario to achieve efficient, intelligent, and synergistic utilization of glacial meltwater resources, simultaneously completing precise artificial recharge of groundwater and systematic ecological restoration of degraded vegetation.
[0048] Referring to Figure 1, the overall architecture of this system consists of a distributed dynamic perception network, a multi-agent collaborative decision-making module, an adaptive execution and feedback adjustment module, and a system self-evolutionary learning module, forming a complete closed loop from full-dimensional data perception to multi-agent collaborative decision-making, and then to precise physical execution and continuous policy evolution.
[0049] This embodiment elaborates on the construction and operation of the distributed dynamic sensing network:
[0050] As shown in Figure 3, the distributed dynamic sensing network is responsible for acquiring comprehensive, real-time, and multi-source heterogeneous sensing data within the watershed. This network is composed of a fixed sensing unit array and dynamically supplemented sensing units working together.
[0051] The fixed sensing unit array consists of hundreds of multi-parameter sensor nodes pre-deployed based on hydrogeological surveys and ecological zoning results. These nodes are strategically deployed in key locations: in the glacier terminus tongue region, nodes monitor near-surface air temperature and reflected radiation; at the upstream, midstream, and downstream control sections of major meltwater channels, nodes integrating ultrasonic current meters and pressure level gauges are deployed to calculate the river flow process line in real time; in the planned groundwater recharge area, nodes are deployed in a grid pattern, with each node integrating a time-domain reflectometry soil moisture sensor and a soil conductivity sensor to monitor soil volumetric water content and salinity dynamics; in the vegetation restoration area, nodes monitor soil moisture and are equipped with infrared thermometers to continuously acquire vegetation canopy temperature to calculate the crop water stress index. All sensors operate synchronously at a fixed acquisition frequency of 15 minutes, and the data is aggregated via the LoRaWAN protocol and transmitted to the edge server via satellite or NB-IoT link.
[0052] Specifically, the dynamic supplementary sensing unit consists of several hexacopter UAV platforms with long endurance and high payload capacity, along with their onboard payloads. Each UAV is equipped with a hyperspectral imager and a lightweight lidar. The hyperspectral imager covers a spectral range of 400 nm to 2500 nm, has 256 continuous spectral channels, and a spectral resolution better than 10 nm, enabling precise identification of surface mineral composition, soil organic matter content, and physiological parameters such as chlorophyll and carotenoid content in vegetation. The lidar scans at 200 Hz, acquiring hundreds of thousands of laser points per second to generate centimeter-level resolution terrain point cloud data, accurately depicting surface micro-topography, roughness, and the three-dimensional structure of vegetation. The multi-agent collaborative decision-making module generates structured sensing requirement instructions based on data anomalies in the fixed sensing network or specific decision-making task requirements. These instructions are sent to the UAV ground control station in JSON format, clearly including the latitude and longitude boundary polygon of the target area, flight altitude, flight path overlap rate, scanning mode, and the type of data product required. For example, when a fixed soil moisture sensor network detects a sharp 20% decrease in soil moisture content in the edge unit of a recharge area over two consecutive cycles, the decision module may determine that potential soil fissures or preferential flow channels have appeared in the area. It then generates a command requiring the UAV to perform a fine-grained scan of the area at centimeter resolution. Upon receiving the command, the UAV ground control station automatically plans an energy-optimal "bow"-shaped flight path, controlling the UAV to take off and reach the target area. During flight, the hyperspectral imager acquires hyperspectral images of the ground surface in pushbroom mode, while the lidar scans simultaneously. The onboard processor filters and classifies the raw point cloud data in real time, performs radiometric calibration and geometric correction on the hyperspectral images, and then performs air-to-ground joint adjustment to generate standardized digital surface models, digital elevation models, surface temperature distribution maps, and normalized difference vegetation index distribution maps, among other data products. These products are transmitted back to the edge server via the UAV's onboard 4G private network module or via a relay UAV. Thus, the distributed dynamic sensing network has constructed a three-dimensional monitoring system that integrates air and ground, and combines static and dynamic data. It has achieved all-weather, high spatiotemporal resolution monitoring of watershed meteorology and hydrology, surface runoff, soil moisture, groundwater dynamics, and vegetation physiological status, providing a solid, multi-dimensional data foundation for subsequent intelligent decision-making.
[0053] The dynamic supplementary sensing unit consists of several hexacopter UAV platforms equipped with hyperspectral imagers and lidar. The hyperspectral imager covers a spectral range of 400 nm to 2500 nm and has 256 continuous spectral channels, enabling precise identification of surface material composition and vegetation physiological parameters. The lidar has a scanning frequency of 200 Hz and can generate terrain point cloud data with centimeter-level resolution. Based on sensing demand commands (such as JSON-formatted task orders) issued by the multi-agent collaborative decision-making module, this unit performs refined scanning of hotspot areas such as areas with sudden drops in soil moisture content, acquiring spatial distribution data such as surface temperature, micro-topography, and vegetation stress index, and transmitting this data back via a dedicated 4G network.
[0054] Thus, the network has constructed a three-dimensional monitoring system that integrates air and ground, and combines static and dynamic monitoring. It has achieved all-weather, high spatiotemporal resolution monitoring of watershed meteorology and hydrology, surface runoff, soil moisture, groundwater dynamics, and vegetation physiological status. It has overcome the shortcomings of sparse and lagging monitoring by traditional fixed stations and provided a reliable data foundation for accurate decision-making.
[0055] In this embodiment, the collaborative decision-making process of the multi-agent collaborative decision-making module is described in detail:
[0056] As shown in Figure 2, the multi-agent collaborative decision-making module performs in-depth processing and collaborative optimization on the aggregated sensing data to generate globally optimal control instructions. This module includes a hydrological prediction agent, a re-irrigation control agent, a vegetation irrigation agent, and a central coordinator. All computational tasks run on a high-performance computing server deployed with GPUs.
[0057] Hydrological prediction agent: Employs a 3-layer gated recurrent unit (GRU) network structure. Inputs include the past 72-hour temperature sequence, snow albedo remote sensing data, and historical meltwater flow sequence. Outputs a probability distribution prediction (mean, variance, and quantiles) of glacial meltwater runoff for the next 24 hours to address the strong uncertainties in the meltwater process.
[0058] The recharge control agent employs a proximal policy optimization (PPO) algorithm. Its state space includes soil saturated hydraulic conductivity, groundwater level depth, predicted meltwater flow, and target total recharge volume; its action space is a continuous control vector for the infiltration rate of multiple infiltration pond units; and its reward function aims to maximize long-term recharge efficiency and minimize surface runoff loss.
[0059] Vegetation Irrigation Agent: Based on Dueling Deep Q-Network (DuelingDQN). Its state space integrates the water requirement curves of vegetation species, the current available soil moisture content, the evapotranspiration prediction for the next 24 hours, and the vegetation stress index; it outputs the optimal irrigation duration and irrigation amount decisions for different vegetation patches.
[0060] Specifically, the normalized data is input into a network structure containing three layers of gated recurrent units (GRUs). The first layer GRU contains 128 hidden neurons, whose gating mechanism effectively captures the impact of short-term temperature fluctuations on the snow-ice phase transition rate. The second layer GRU contains 64 hidden neurons, used to model the hysteresis effects and nonlinear processes of snowmelt and meltwater runoff on the glacier surface and in glacial channels. The third layer GRU contains 32 hidden neurons, used to extract the long-term trends and periodic characteristics of meltwater runoff. The network output layer uses a hybrid density network structure to output the probability distribution parameters of hourly runoff predictions for the next 24 hours, including the mean, variance, and multiple quantiles, thus providing a basis for risk quantification for downstream decision-making. The recharge control agent focuses on optimizing the core physical process of groundwater recharge. Its state space is a high-dimensional vector, specifically including: the spatial distribution matrix of soil saturated hydraulic conductivity for each recharge zone, calibrated by a digital elevation model derived from UAV lidar and combined with soil borehole measurement data; groundwater level depth data fed back in real time by pressure-type water level sensors deployed in monitoring wells around the recharge area; the expected value sequence of meltwater flow predictions for the next 24 hours provided by the hydrological prediction agent; and the monthly or quarterly target recharge volume set by the water resources management department. The agent's action space is defined as a continuous control vector of the infiltration rate of multiple independent infiltration pond units distributed within the watershed. Each action component corresponds to the target infiltration rate of an infiltration pond, and its value range is constrained to between 0 and the maximum safe infiltration capacity of the pond determined based on soil texture and structure. The agent is trained using a proximal policy optimization algorithm, with both its policy network and value network being fully connected neural networks containing two hidden layers, using the ReLU activation function. The reward function is meticulously designed, comprising four quantifiable components: the core reward is the actual monitored net groundwater recharge, calculated by comparing the changes in well water levels before and after recharge and multiplying by the aquifer specific yield and radius of influence; the main penalty is the ineffective surface runoff loss, i.e., the overflow volume monitored by flow meters deployed at the edge of the infiltration ponds; the efficiency reward is the recharge efficiency, i.e., the ratio of net recharge to total water diversion; and the stability penalty indirectly reflects the risk of soil salt leaching or accumulation that may be triggered during the recharge process by monitoring abrupt changes in soil conductivity. Through interaction with the environment, the agent learns how to intelligently allocate the infiltration load of each infiltration pond under dynamically changing soil infiltration capacity and uncertain water inflow conditions, in order to maximize long-term recharge efficiency and minimize risk. The vegetation irrigation agent is responsible for developing precise and dynamic irrigation plans for different vegetation restoration areas.Its state space integrates multi-source heterogeneous information: a pre-constructed knowledge base of vegetation species water requirement curves, containing reference values for daily average water requirements and water requirement sensitivity coefficients for different species at different phenological stages and growth phases; current effective soil water content data for the root zone of each vegetation patch provided by soil moisture sensors from a fixed sensing network; potential evapotranspiration predictions for the next 24 hours calculated based on meteorological data and the FAO Penman-Montess formula; and vegetation stress indices, such as photochemical reflectance and water stress indices, derived from UAV hyperspectral data. The agent is based on a duel-deep Q-network architecture, where the Q-network separates the state value function from the action advantage function. The action space is discrete, providing a set of predefined irrigation scheme options for each vegetation irrigation zone, each option consisting of a combination of irrigation duration and irrigation intensity. The agent's output decisions aim to accurately match the physiological water requirements of the vegetation, maximizing the conservation of precious meltwater resources and avoiding surface runoff, deep seepage, or secondary soil salinization caused by over-irrigation.
[0061] Central Coordinator: Receives preliminary action proposals from each agent and performs a global trade-off based on a pre-defined multi-objective optimization function. The function is:
[0062] .
[0063] Among them, the net replenishment of groundwater The amount of ineffective surface runoff loss is estimated by integrating the target infiltration rate of each infiltration pond over time and multiplying it by the effective replenishment coefficient obtained from regression of historical data. The simulation was performed using a one-dimensional hydrodynamic model, which considered the relationship between river water diversion capacity, dynamic changes in infiltration capacity of the infiltration pond, and surface overflow; a comprehensive score for vegetation growth health was also calculated. The total energy consumption of the system was calculated by combining the vegetation stress index, chlorophyll content inversion value, and future biomass prediction value based on the growth accumulated temperature model. The estimation is based on the rated power, operating time, and load rate of all actuators, including the gate servo motor, water pump frequency converter, and drone inspection motor. Weighting coefficients are used. The determination of the weights is a process combining scientific and engineering experience: First, experts in hydrology, ecology, energy, and economics are invited to compare the importance of the four indicators pairwise according to the core objectives of a specific engineering phase, constructing a judgment matrix. Second, the eigenvalue method is used to calculate the largest eigenvalue of the judgment matrix and its corresponding normalized eigenvector, obtaining the initial weights of each indicator, and a consistency ratio test is performed to ensure the rationality of the judgment logic. Finally, combined with historical data generated during the system's trial operation, a multi-objective Pareto front analysis is conducted to observe the distribution of system performance under different weights, fine-tuning the initial weights to ensure that the finally selected weight coefficients can guide the system to converge to a Pareto optimal solution set that meets engineering expectations and achieves balanced development in the long term. The central coordinator uses a non-dominated sorting genetic algorithm with an elitist strategy to solve this multi-objective optimization problem. This algorithm iteratively searches for the optimal solution set through selection, crossover, and mutation operations, ultimately outputting the scheme with the highest comprehensive score among the non-dominated solutions, i.e., the globally optimal control instruction set. This instruction set is a structured data package, the core contents of which include: the target flow process line for water intake from the main channel gate; the target infiltration rate time series matrix allocated to each infiltration pond unit; the daily water allocation vector and irrigation start time allocated to each vegetation irrigation zone; and the precise time window and priority for the execution of each instruction.
[0064] This architecture decomposes and coordinates the complex water-soil-plant coupling regulation problem, and performs global optimization through a central coordinator, thereby achieving efficient and dynamic allocation of water resources in the spatiotemporal dimensions.
[0065] In this embodiment, the implementation of the adaptive execution and feedback adjustment module is described in detail:
[0066] As shown in Figure 4, this module is responsible for converting decision commands into physical actions and collecting feedback.
[0067] Command parsing unit: Deployed at the industrial edge gateway, it receives the globally optimal control command set. According to the underlying device protocol, it converts the commands into specific control signals (such as PWM signals to control gate servo motors, 4-20mA signals to control regulating valves, and Modbus RTU commands to control solenoid valves).
[0068] The execution drive unit includes a programmable logic controller (PLC) and a frequency converter array. It precisely controls the opening of the head gate of the water diversion channel, adjusts the opening of the valves in the water distribution network at the bottom of the infiltration pond to match the target infiltration rate, and controls the opening and closing time of the solenoid valves in the drip irrigation network to achieve precision irrigation.
[0069] Feedback Acquisition Unit: Real-time monitoring of the actual gate opening, the actual infiltration rate of each infiltration pond, and the actual irrigation water volume. 1-2 hours after execution, it triggers a dedicated sensing network to collect environmental response data such as water level changes and soil moisture recovery, integrating these data to form environmental feedback data.
[0070] This module ensures that high-level decision-making instructions can be accurately and reliably translated into actions of underlying equipment, and obtains real-time feedback on the execution effect, forming a virtuous cycle of "perception-decision-execution-evaluation".
[0071] In this embodiment, the self-evolution process of the system's self-evolutionary learning module is described in detail.
[0072] As shown in Figure 5, this module enables the system to continuously learn from operational experience.
[0073] Experience replay buffer: in tuples ( , , , , Store complete data for each decision-execution cycle in a format that includes rewards. It is calculated by the central coordinator based on feedback data.
[0074] Policy update engine: Offline training is initiated during periods of low system load. Batch data is sampled from the buffer, and the neural network parameters of each agent are updated using the gradient descent method (e.g., the hydrological prediction agent is updated by minimizing the negative log-likelihood loss, and the recharge and irrigation agent is updated using the PPO and DQN algorithms).
[0075] Emergency Strategy Training Submodule: Automatically triggered when feedback data identifies events such as "extreme precipitation" or "sustained high temperatures." It quickly filters historical data on similar operating conditions, fine-tuning model parameters with a higher learning rate to improve the system's robustness and rapid adaptive capability in the face of sudden disturbances.
[0076] Specifically, the core of this module is the experience replay buffer and the policy update engine. The experience replay buffer is a high-performance, scalable ring-shaped storage pool deployed on a central database server. After each complete decision-execution-evaluation cycle, the system packages all relevant data from that cycle into an experience tuple and stores it in the buffer. Each tuple contains five elements: the system state at the time of the decision. This state is the union of all dimensions of the state spaces of the aforementioned agents; the actual actions performed... This refers to the set of control instructions that are ultimately output and executed by the central coordinator; the immediate rewards obtained. The new system state is calculated by the central coordinator based on a multi-objective optimization function and feedback data from the actual environment; this state is entered after the action is executed. The system consists of specific data provided by the feedback acquisition unit, and a termination flag to indicate special times such as the end of the annual meltwater season or planned system shutdown for maintenance. The policy update engine is a periodically running, automatically offline training process. It is usually started by the task scheduler at night when the system load is low or during the meltwater interval. During training, the policy update engine randomly samples a batch of historical experience tuples from the experience replay buffer using a priority experience replay algorithm, with a fixed batch size of 256. For the hydrological prediction agent, a time-series backpropagation algorithm is used, employing the actual runoff sequence as a supervision signal to calculate the negative log-likelihood loss between the predicted distribution and the actual value, and updating all weight parameters of the output layers of its gated recurrent unit network and hybrid density network using gradient descent. For the recharge control agent and the vegetation irrigation agent, their respective reinforcement learning algorithms are used for updating. Taking the reinjection control agent as an example, the engine calculates the advantage function of its actions and the temporal difference error based on generalized advantage estimation. Then, it uses a proximal policy optimization algorithm to prune the objective function and updates the parameters of its policy network and value network through gradient descent, while simultaneously updating the baseline values of each item in its reward function. The policy update engine also has a dedicated emergency policy training submodule. This submodule is automatically triggered when environmental feedback data is marked with special operating condition indicators such as "extreme precipitation events" or "sustained high temperature heat wave events." It quickly filters all historical experience data under similar extreme operating conditions from the experience replay buffer, forming a high-priority training set, and starts a reinforcement learning training loop. With a higher learning rate, it rapidly fine-tunes the policy network parameters of each agent, enabling it to learn to adopt more conservative or more aggressive control strategies under sudden disturbances, thereby significantly improving the system's robustness and rapid adaptive adjustment capability in response to abnormal climate events. Through this continuous closed loop of "perception-decision-execution-evaluation-learning," the entire system can constantly adapt to the non-stationarity of glacial meltwater patterns, the geological heterogeneity of the watershed's underlying surface, and the long-term trends brought about by climate change, achieving autonomous evolution in its intelligence level. Actual deployment data demonstrates this.
[0077] Through this continuous closed loop of perception-decision-execution-evaluation-learning, the system can continuously adapt to environmental changes and achieve autonomous evolution of its intelligence level.
[0078] In this embodiment, the method for verifying the system's operational performance is described in detail:
[0079] After being deployed in a watershed on the Qinghai-Tibet Plateau, the system demonstrated significant beneficial effects after three complete meltwater seasons (approximately 18 months) of operation and learning: compared to the baseline strategy at the initial deployment stage, the net groundwater recharge efficiency increased by approximately 35%, ineffective surface runoff loss decreased by approximately 52%, the average vegetation survival rate increased by approximately 28%, and the overall system energy consumption decreased by approximately 15%. These quantitative data validate the effectiveness of the system's self-evolutionary learning mechanism, indicating that the present invention possesses adaptive and self-evolutionary capabilities to cope with the non-stationarity of glacial meltwater, geological heterogeneity, and extreme climate events, fundamentally solving the systemic defects of the prior art, such as a single perception dimension, static decision-making logic, isolated execution units, and a lack of system learning capabilities.
[0080] Example 2: Flowchart of the corresponding method
[0081] This embodiment corresponds to the operation method of the above system, and its specific steps are as follows:
[0082] Step S110: Through the aforementioned distributed dynamic sensing network, continuously collect real-time multi-source sensing data of the glacial meltwater basin, including time-series monitoring data of fixed sensor arrays and airborne fine-scan data of UAV platforms.
[0083] Step S120: Input the data into the multi-agent collaborative decision-making module, where the hydrological prediction, recharge control, and vegetation irrigation agents each make local decisions, and then the central coordinator performs global multi-objective optimization (based on the function). This generates the globally optimal control instruction set.
[0084] Step S130: The adaptive execution and feedback adjustment module parses and executes the instruction set to drive the water diversion, recharge and irrigation equipment to perform coordinated control.
[0085] Step S140: After execution, environmental status data is collected again to form environmental feedback data.
[0086] Step S150: The system's self-evolutionary learning module stores the current experience in the experience replay buffer and starts offline training under preset conditions (such as low load periods or detection of extreme weather event indicators) to update the decision-making strategies of each agent. When an extreme weather event indicator appears, an emergency strategy training loop is started to quickly fine-tune the strategy network parameters.
[0087] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0088] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A system for the coordinated regulation and ecological restoration of glacial meltwater resources, characterized in that, include: The data sensing module is used to acquire multi-source environmental sensing data related to glacial meltwater within the target area; The intelligent decision-making module, communicatively connected to the data perception module, is used to generate coordinated control instructions for glacial meltwater diversion, groundwater recharge, and vegetation irrigation based on the multi-source environmental perception data and using a collaborative decision-making model comprising multiple deep reinforcement learning agents. The instruction execution and feedback module, communicatively connected to the intelligent decision-making module, is used to execute the coordinated control instructions and collect environmental feedback data after instruction execution. The strategy evolution module, communicatively connected to the instruction execution and feedback module, is used to continuously train and update the collaborative decision-making model based on the environmental feedback data.
2. The system according to claim 1, characterized in that, The data sensing module includes: a ground sensor network fixedly deployed at key locations in the target area, used to periodically collect environmental data including at least meteorological, hydrological, soil, and vegetation physiological parameters; and a dynamic sensing unit based on an unmanned platform, used to respond to the instructions of the intelligent decision-making module to perform high-resolution spatial scanning of a specific area and obtain spatial distribution data of surface material composition, micro-topography, and vegetation physiological state.
3. The system according to claim 1, characterized in that, The intelligent decision-making module includes: a hydrological prediction unit, which predicts glacial meltwater runoff in future periods using a recurrent neural network based on time-series environmental perception data; a recharge control unit, which outputs control decisions on the infiltration rate of each recharge unit using a first deep reinforcement learning model based on soil hydrological characteristics, groundwater status, and predicted runoff; an irrigation decision unit, which outputs irrigation decisions on each vegetation unit using a second deep reinforcement learning model based on vegetation water demand characteristics, soil moisture status, and meteorological forecasts; and a central coordinator, which performs global multi-objective optimization on the preliminary decisions of the hydrological prediction unit, the recharge control unit, and the irrigation decision unit, and generates the coordinated control instructions.
4. The system according to claim 3, characterized in that, The multi-objective optimization function upon which the central coordinator is based is expressed as: ;in, This indicates the net recharge of groundwater. Indicates the amount of invalid surface runoff loss. This represents the overall score indicating the health of vegetation growth. This indicates the total energy consumption of the system. These are non-negative weighting coefficients set according to project priorities; the central coordinator outputs the globally optimal control instruction set by solving this optimization problem.
5. The system according to claim 1, characterized in that, The instruction execution and feedback module includes: an instruction parsing and driving unit, used to parse the coordinated control instruction into low-level control signals and drive the water diversion, recharge and irrigation execution equipment to perform actions; and an effect feedback unit, used to monitor the equipment execution status and collect the change data of key environmental parameters again after the instruction is executed, forming the environmental feedback data.
6. The system according to claim 1, characterized in that, The policy evolution module includes: an experience memory for storing historical decision and feedback data in the form of tuples of state, action, reward, and new state; and an offline training engine for periodically sampling data from the experience memory to train the deep reinforcement learning agent in the collaborative decision-making model and update its decision-making policy.
7. The system according to claim 6, characterized in that, The offline training engine is configured to trigger specialized training for the event when the environmental feedback data indicates that the system has encountered a preset type of extreme weather event, so as to quickly adjust the decision-making strategy.
8. A method for synergistic regulation and ecological restoration of glacial meltwater resources, characterized in that, Includes the following steps: S1. Data acquisition steps: Continuously acquire environmental sensing data of the target area through a multi-source sensing network; S2. Collaborative Decision-Making Step: Based on the environmental perception data, a multi-agent deep reinforcement learning model is used for collaborative decision-making to generate a global optimization command that integrates water diversion, recharge, and irrigation regulation; S3. Command Execution and Feedback Step: The global optimization command is executed, and environmental feedback data is collected after the command is executed; S4. Policy evolution step: Based on the environmental feedback data, the multi-agent deep reinforcement learning model is trained offline to update its decision-making strategy.
9. The method according to claim 8, characterized in that, The global optimization in the S2 collaborative decision-making step aims to maximize the net groundwater recharge and vegetation growth health, while minimizing ineffective runoff loss and system energy consumption.
10. The method according to claim 8, characterized in that, The S4 strategy evolution steps include: when an extreme weather event is detected, using data from similar historical events to reinforce the model in order to improve the system's adaptability to sudden disturbances.