An adaptive regulation method for a water quality detection self-powered system
Patent Information
- Application Number
- CN202610877230.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本发明解决的技术问题是:野外水质监测设备在复杂环境下易断电、数据易断档、供电策略无法自动优化的问题
[0063]1、本发明通过采用马尔可夫链模型对多模态时序数据流进行工况转移概率预测,得到各候选工况状态的概率分布,并根据工况预测结果计算每个候选调控策略的净能量收益期望值与潜在能量损失期望值,将二者的比值定义为能量收益风险比,再根据能量收益风险比的高低确定各候选调控策略的选择概率,利用控制器按照所述选择概率随机抽取候选调控策略执行,将实际收益与当前工况状态存入经验回放缓冲区,并从所述经验回放缓冲区中按照预测误差大小进行优先级采样,采用参数软更新方式进行自学习,得到优化后的执行参数集对水质检测自供电系统进行自适应调控。该方案能够在光照和水流高度不确定的野外环境下,通过概率性策略探索避免陷入次优决策,并利用经验回放与参数软更新实现供电策略的持续优化,有效解决了现有技术因固定阈值规则无法适应复杂工况而导致的设备断电、数据断档以及供电策略无法自动优化的问题,提高了水质检测自供电系统在恶劣环境下的供电可靠性和自适应能力。
Smart Images

Figure CN122801196A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power supply control and energy management technology for water quality testing equipment, and in particular to an adaptive control method for a self-powered water quality testing system. Background Technology
[0002] In recent years, self-powered water quality monitoring systems have primarily adopted a multi-source complementary power supply approach combining photovoltaic (PV) and hydropower generation. These systems typically include a maximum power point tracking (MPPT) controller, energy storage batteries, and a load management unit. In conventional applications, the system switches power supply strategies based on preset fixed threshold rules. For example, when sunlight intensity exceeds a preset value, PV power generation is prioritized; when water flow velocity exceeds a preset value, hydropower generation is prioritized; and when both sources are insufficient, the energy storage battery provides power. This approach achieves a degree of self-sufficiency in outdoor environments, reducing reliance on external grid power.
[0003] However, existing self-powered water quality monitoring systems still have the following drawbacks in practical field applications: First, light intensity and water flow velocity are highly uncertain due to factors such as weather, season, and day-night cycles. Fixed threshold rules are difficult to adapt to complex and changing environmental conditions, leading to frequent power outages and shutdowns under extreme conditions such as cloudy days and slow flow, resulting in gaps in water quality monitoring data. Second, existing systems employ deterministic decision-making logic, directly determining a unique power supply strategy based on current environmental parameters, failing to explore multiple feasible strategies probabilistically. If the initial rules are not set reasonably, the system will operate in a suboptimal state for a long time. Third, existing systems lack the ability to continuously learn from historical operating experience, and the power supply strategy cannot be adaptively adjusted according to actual performance. When environmental characteristics change, parameters need to be manually reconfigured on-site, resulting in high maintenance costs and untimely response. Therefore, there is an urgent need for a self-powered control method that can adapt to uncertain environments and possesses autonomous learning and strategy optimization capabilities. Summary of the Invention
[0004] The technical problem solved by this invention is that field water quality monitoring equipment is prone to power outages, data gaps, and automatic optimization of power supply strategies in complex environments.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0006] As a preferred embodiment of the adaptive control method for the self-powered water quality detection system described in this invention, wherein:
[0007] Real-time acquisition of multimodal time-series data streams, including photovoltaic power time series, water flow velocity fluctuation variance, battery state of charge, battery temperature, and historical power consumption records of equipment;
[0008] A Markov chain model is used to predict the probability of working condition transitions in a multimodal time series data stream, and the working condition prediction results are obtained, including the probability distribution of each candidate working condition state.
[0009] Based on the operating condition prediction results, the expected net energy gain and the expected potential energy loss of each candidate control strategy are calculated respectively, and the ratio of the expected net energy gain to the expected potential energy loss is defined as the energy gain risk ratio of the candidate control strategy.
[0010] The selection probability of each candidate regulation is determined based on the energy benefit-risk ratio.
[0011] The controller randomly selects a candidate control strategy according to the selection probability and executes it to obtain the control strategy vector and actual benefit. The actual benefit and the current operating condition are stored as empirical data in the empirical playback buffer. The control strategy vector includes the generation priority switching instruction, the maximum power point tracking disturbance observation step size, the energy storage charging current limit value, and the load pre-degradation advance.
[0012] Based on the control strategy vector and the actual benefit, samples are prioritized from the experience replay buffer according to the magnitude of the prediction error, and self-learning is performed using a parameter soft update method to obtain an optimized set of execution parameters, including the operating frequency of the maximum power point tracking controller, the fine-tuning amount of the reference voltage of the multi-channel regulated output channel, the offset of the battery equalization circuit activation threshold, and the transmission power level of the wireless transmission module;
[0013] Based on the optimized set of execution parameters, the self-powered water quality detection system is adaptively controlled.
[0014] Furthermore, a Markov chain model is used to predict the operating condition transition probability of the multimodal time-series data stream, and the operating condition prediction results are obtained, including:
[0015] The photovoltaic power time series is discretized into N power levels according to the measured power values;
[0016] The variance of water flow velocity fluctuation is discretized into M turbulence levels based on the measured variance values;
[0017] The battery state of charge is discretized into K energy levels based on the measured percentage.
[0018] The battery temperature is discretized into T temperature levels based on the measured temperature values;
[0019] The device's historical power consumption records are divided into standby mode, acquisition mode, and transmission mode according to the working mode.
[0020] The power level, turbulence level, electrical quantity level, temperature level, and operating mode are combined by Cartesian product to construct a joint state space;
[0021] Based on the joint state space, the transition frequency of each joint state in historical data is statistically analyzed to construct a state transition probability matrix;
[0022] Using the actual discrete level of each joint state at the current moment and the current working mode as input, the state transition probability matrix is queried, and the probability of occurrence of each candidate joint state within a future time window is output as the working condition prediction result.
[0023] Furthermore, based on the joint state space, the transition frequency of each joint state in historical data is statistically analyzed to construct a state transition probability matrix, including:
[0024] Sample each joint state in the joint state space at a preset time interval and record the joint state at each sampling time.
[0025] The number of times the transition from the first joint state to the second joint state occurs between adjacent sampling times is counted, and this number of occurrences is taken as the transition frequency from the first joint state to the second joint state;
[0026] The total transition frequency of the first joint state is obtained by summing the transition frequencies from the first joint state to all joint states.
[0027] The transition frequency from the first joint state to the second joint state is divided by the total transition frequency of the first joint state to obtain the transition probability from the first joint state to the second joint state.
[0028] Traverse all joint states in the joint state space and construct the state transition probability matrix.
[0029] Furthermore, based on the operating condition prediction results, the expected net energy gain and expected potential energy loss of each candidate control strategy are calculated, including:
[0030] The probability of occurrence of each candidate joint state in the predicted working condition is used as the weight;
[0031] For each candidate control strategy, the power generation and power consumption under each candidate joint state are calculated. The power generation is subtracted from the power consumption, multiplied by the corresponding occurrence probability, and summed to obtain the expected net energy revenue. When the predicted power of photovoltaic power generation and the predicted power of hydropower generation are both lower than the real-time load power of the equipment under the same candidate joint state, the condition is marked as a shortage state, and a preset shortage penalty is deducted from the expected net energy revenue.
[0032] For each candidate control strategy, the probability that it will cause the battery state of charge to fall below a preset low charge threshold under each candidate joint state is calculated, as well as the probability that the rate of decrease in battery state of charge will exceed a preset rate threshold after the execution of the candidate control strategy. The two are weighted and summed and then multiplied by a preset penalty coefficient for data loss due to power failure of the device to obtain the expected value of potential energy loss. The preset rate threshold is dynamically adjusted according to the sliding average value of the device's historical power consumption records.
[0033] Furthermore, the selection probability of each candidate regulation is determined based on the energy-benefit-risk ratio, including:
[0034] Arrange the energy-benefit-risk ratios of each candidate control strategy in descending order;
[0035] The difference between two adjacent energy gain-risk ratios after sorting is used as the spacing between adjacent intervals.
[0036] The selection probability of each candidate control strategy is dynamically adjusted according to the size of the adjacent spacing, so that the probability difference between adjacent candidate control strategies with larger spacing is smaller, and the probability difference between adjacent candidate control strategies with smaller spacing is larger.
[0037] Specifically, when the energy-benefit-risk ratio of the candidate control strategy with the highest energy-benefit-risk ratio exceeds the sum of the energy-benefit-risk ratios of all other candidate control strategies, the selection probability of that candidate control strategy is forcibly set to no more than the arithmetic mean of the selection probabilities of all candidate control strategies.
[0038] Furthermore, the selection probability of each candidate control strategy is dynamically adjusted based on the size of the adjacent spacing, including:
[0039] When the adjacent spacing is greater than a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is set to be equal.
[0040] When the adjacent spacing is less than or equal to a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is allocated according to the positive correlation between energy gain and risk ratio.
[0041] Furthermore, the controller randomly selects a candidate control strategy based on the selection probability of each candidate control strategy, including:
[0042] The selection probabilities of each candidate control strategy are accumulated sequentially to construct a cumulative probability interval, wherein each candidate control strategy corresponds to a continuous cumulative probability sub-interval, and the length of the cumulative probability sub-interval is equal to the selection probability of the candidate control strategy.
[0043] Use the controller to generate a random number between 0 and 1;
[0044] The cumulative probability sub-interval into which the random number falls is determined, and the candidate control strategy corresponding to the cumulative probability sub-interval is selected as the candidate control strategy.
[0045] Further, based on the control strategy vector and the actual benefit, priority sampling is performed from the experience replay buffer according to the magnitude of the prediction error, including:
[0046] The expected return before execution is determined based on the control strategy vector, and the actual return after execution is determined based on the actual return. The absolute value of the difference between the expected return and the actual return is used as the basic prediction error.
[0047] Extract the expected return sequence of a predetermined number of consecutive empirical data under the same working condition from the experience replay buffer, and calculate the variance of the expected return sequence as the temporal instability.
[0048] The product of the basic prediction error and the temporal instability is taken as the comprehensive prediction error;
[0049] The empirical data in the empirical playback buffer are sampled according to the order of the comprehensive prediction error from largest to smallest.
[0050] Furthermore, a parameter soft update method is used for self-learning to obtain an optimized set of execution parameters, including:
[0051] The first adjustment gradient is calculated based on the empirical data obtained from priority sampling.
[0052] Extract the regret value corresponding to the current operating condition from the experience replay buffer;
[0053] A second adjustment gradient is calculated based on the regret value, and the magnitude of the second adjustment gradient is positively correlated with the regret value;
[0054] The first adjustment gradient and the second adjustment gradient are weighted and summed to obtain the comprehensive adjustment gradient;
[0055] The current set of execution parameters is moved one step along the direction of the comprehensive adjustment gradient to obtain the optimized set of execution parameters.
[0056] Furthermore, based on the optimized set of execution parameters, the self-powered water quality monitoring system is adaptively controlled, including:
[0057] Write the operating frequency of the maximum power point tracking controller from the optimized set of execution parameters into the maximum power point tracking controller;
[0058] Write the reference voltage fine-tuning amount of the multi-channel regulated output channel in the optimized execution parameter set into the load adaptation module;
[0059] Write the battery equalization circuit activation threshold offset from the optimized execution parameter set into the battery management unit.
[0060] Write the wireless transmission module transmit power level from the optimized execution parameter set into the wireless transmission unit;
[0061] Adaptive regulation is implemented for the self-powered water quality testing system.
[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0063] 1. This invention uses a Markov chain model to predict the probability of operating condition transitions in a multimodal time-series data stream, obtaining the probability distribution of each candidate operating condition state. Based on the operating condition prediction results, the expected net energy gain and the expected potential energy loss of each candidate control strategy are calculated. The ratio of these two values is defined as the energy gain-risk ratio. The selection probability of each candidate control strategy is then determined based on the energy gain-risk ratio. The controller randomly selects candidate control strategies according to the selection probability and executes them. The actual gains and the current operating condition state are stored in an experience playback buffer. Priority sampling is performed from the experience playback buffer according to the prediction error. Self-learning is performed using a soft parameter update method to obtain an optimized set of execution parameters for adaptive control of the self-powered water quality detection system. This solution can avoid suboptimal decisions through probabilistic strategy exploration in field environments with uncertain lighting and water flow, and continuously optimize the power supply strategy by using experience playback and soft parameter updates. It effectively solves the problems of equipment power outages, data gaps, and the inability to automatically optimize the power supply strategy caused by the inability of existing technologies to adapt to complex working conditions due to fixed threshold rules, and improves the power supply reliability and adaptability of the water quality detection self-powered system in harsh environments.
[0064] 2. This invention employs a Markov chain model to predict the operating condition transition probabilities of multimodal time-series data streams, obtaining the probability distribution of each candidate operating condition state. Based on the operating condition prediction results, the energy benefit-risk ratio of each candidate control strategy is calculated. The selection probability is determined according to the energy benefit-risk ratio, and the controller randomly selects and executes candidate control strategies according to the selected probabilities, thus realizing probabilistic exploration of power supply strategies. This mechanism avoids long-term suboptimal decision-making due to the mismatch between fixed threshold rules and complex environments, reducing the risk of power outages under extreme operating conditions.
[0065] 3. This invention stores the actual benefits and current operating conditions in an experience replay buffer, and samples the data from the experience replay buffer according to the magnitude of the prediction error. This allows experience data with larger prediction deviations to be used for self-learning first, which speeds up the model's adaptation to abnormal operating conditions and reduces the frequency of data gaps.
[0066] 4. This invention uses a parameter soft update method for self-learning to obtain an optimized set of execution parameters for adaptive control of the water quality detection self-powered system. This allows the power supply strategy to evolve smoothly and continuously based on actual operating results, eliminating the need for manual on-site parameter reconfiguration and reducing the operation and maintenance costs of field monitoring equipment. Attached Figure Description
[0067] Figure 1 This is a basic flowchart illustrating an adaptive control method for a self-powered water quality detection system, as provided in one embodiment of the present invention. Detailed Implementation
[0068] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0069] Example 1
[0070] like Figure 1 As shown in the figure, this embodiment introduces an adaptive control method for a self-powered water quality detection system, including:
[0071] Step 1: Real-time acquisition of multimodal time-series data streams.
[0072] In this embodiment, the multimodal time-series data stream includes photovoltaic power time series, water flow velocity fluctuation variance, battery state of charge, battery temperature, and historical power consumption records of the device.
[0073] This embodiment constructs a multimodal time-series data stream covering multiple dimensions of power generation, energy storage, and load by real-time acquisition of photovoltaic power time series, water flow velocity fluctuation variance, battery state of charge, battery temperature, and historical power consumption records of equipment. This provides a complete raw data foundation for subsequent operating condition prediction and strategy decision-making, avoiding prediction deviations caused by a single data dimension.
[0074] Step 2: Use a Markov chain model to predict the operating condition transition probability of the multimodal time series data stream and obtain the operating condition prediction results.
[0075] In this embodiment, the operating condition prediction result includes the probability distribution of each candidate operating condition state.
[0076] This embodiment uses a Markov chain model to predict the probability of working condition transitions in a multimodal time-series data stream, obtaining the probability distribution of each candidate working condition state. This quantifies the uncertainty of the field environment into a probability distribution, overcoming the shortcomings of existing technologies that cannot predict the trend of working condition changes.
[0077] Step 3: Based on the operating condition prediction results, calculate the expected net energy gain and the expected potential energy loss for each candidate control strategy, and define the ratio of the expected net energy gain to the expected potential energy loss as the energy gain risk ratio of the candidate control strategy.
[0078] This embodiment calculates the expected net energy gain and the expected potential energy loss for each candidate control strategy based on the operating condition prediction results, and defines the ratio of the two as the energy gain risk ratio. Before making a decision, the benefit potential of the strategy and the risk of power outage loss are evaluated simultaneously, so that the selection of power supply strategy takes into account both power generation efficiency and power supply security.
[0079] Step 4: Determine the selection probability of each candidate regulation based on the energy benefit-risk ratio.
[0080] This embodiment determines the selection probability of each candidate control strategy based on the energy gain-risk ratio, thereby achieving a probability allocation under the risk-reward trade-off. This allows high-reward, low-risk strategies to obtain a higher selection probability, while retaining the possibility of low-probability strategies being selected, thus avoiding strategy solidification and premature convergence.
[0081] Step 5: Execute a candidate control strategy randomly selected by the controller according to the selection probability to obtain the control strategy vector and actual benefit. Store the actual benefit and the current operating condition as empirical data in the empirical playback buffer.
[0082] In this embodiment, the control strategy vector includes a power generation priority switching command, a maximum power point tracking disturbance observation step size, an energy storage charging current limit value, and a load pre-degradation advance amount.
[0083] This embodiment utilizes the controller to randomly select a candidate control strategy according to the selection probability, obtains the control strategy vector and actual benefit, and stores the actual benefit and the current working state in the experience replay buffer. The random sampling mechanism realizes strategy exploration, and the experience data storage provides a sample source for subsequent self-learning, solving the problem that the existing technology cannot accumulate experience from historical execution results.
[0084] Step Six: Based on the control strategy vector and the actual benefit, sample from the experience replay buffer according to the prediction error size, and use the parameter soft update method for self-learning to obtain the optimized execution parameter set.
[0085] In this embodiment, the optimized set of execution parameters includes the operating frequency of the maximum power point tracking controller, the fine-tuning amount of the reference voltage of the multi-channel regulated output channel, the offset of the battery equalization circuit activation threshold, and the transmission power level of the wireless transmission module.
[0086] This embodiment prioritizes sampling from the experience replay buffer according to the magnitude of the prediction error based on the control strategy vector and the actual benefit, and uses a parameter soft update method for self-learning to obtain an optimized set of execution parameters. This allows experience data with larger prediction deviations to be used for model updates first, while maintaining smooth parameter evolution through soft updates, thereby achieving continuous optimization of the power supply strategy and avoiding system instability caused by parameter mutations.
[0087] Step 7: Based on the optimized set of execution parameters, adaptively control the self-powered water quality detection system.
[0088] This embodiment adaptively regulates the self-powered water quality detection system based on the optimized set of execution parameters. The learning results are directly applied to the operating frequency of the maximum power point tracking controller, the fine-tuning of the reference voltage of the multi-channel regulated output channel, the offset of the battery equalization circuit opening threshold, and the transmission power level of the wireless transmission module. This enables the system to automatically adjust the underlying control parameters according to the actual operating effect without the need for manual on-site intervention, thereby reducing the operation and maintenance costs of field monitoring equipment.
[0089] Example 2
[0090] This embodiment describes the implementation steps of an adaptive control method for a self-powered water quality detection system, including:
[0091] Step 1: Real-time acquisition of multimodal time-series data streams.
[0092] In this embodiment, the multimodal time-series data stream includes photovoltaic power time series, water flow velocity fluctuation variance, battery state of charge, battery temperature, and historical power consumption records of the device.
[0093] Step 2: Use a Markov chain model to predict the operating condition transition probability of the multimodal time series data stream and obtain the operating condition prediction results.
[0094] In this embodiment, the operating condition prediction result includes the probability distribution of each candidate operating condition state.
[0095] Step 2.1: Discretize the photovoltaic power time series into N power levels according to the measured power values.
[0096] This embodiment discretizes the photovoltaic power time series into N power levels according to the measured power value, transforming continuous photovoltaic power data into a finite number of discrete states. This reduces the complexity of the operating state space and provides a calculable state granularity for the subsequent probability statistics of the Markov chain model.
[0097] Step 2.2: Discretize the variance of water flow velocity fluctuation into M turbulence levels based on the measured variance values.
[0098] This embodiment discretizes the variance of water flow velocity fluctuation into M turbulence levels according to the measured variance value, quantifies the instability of water flow into discrete states, and enables the fluctuation characteristics of hydropower generation conditions to be identified and statistically analyzed by the model, avoiding the distortion of the condition description caused by using only the average flow velocity.
[0099] Step 2.3: Discretize the battery state of charge into K charge levels based on the measured percentage.
[0100] This embodiment discretizes the battery state of charge into K energy levels according to the measured percentage, and divides the continuous energy value of the energy storage system into a finite number of discrete intervals. This allows the energy state to be used as a dimension of the joint state space to participate in the transition probability statistics, providing a state basis for energy-related risk decisions.
[0101] Step 2.4: Discretize the battery temperature into T temperature levels based on the measured temperature values.
[0102] This embodiment discretizes the battery temperature into T temperature levels based on the measured temperature value, and incorporates the influence of temperature on the battery charging and discharging efficiency into the operating state space. This provides a basis for state recognition in adjusting the power supply strategy under low temperature conditions, and solves the defect of existing technologies that ignore the influence of temperature on energy storage performance.
[0103] Step 2.5: Divide the device's historical power consumption records into standby mode, acquisition mode, and transmission mode according to the working mode.
[0104] This embodiment divides the device's historical power consumption records into standby mode, acquisition mode, and transmission mode according to the working mode, so that the load characteristics under different working modes can be modeled separately, providing a basis for mode differentiation for subsequent calculation of power consumption under each candidate joint state.
[0105] Step 2.6: Combine the power level, turbulence level, electrical quantity level, temperature level and operating mode by performing a Cartesian product to construct a joint state space.
[0106] This embodiment constructs a joint state space by combining the power level, turbulence level, energy level, temperature level and operating mode through Cartesian product, thus integrating the state information of the four dimensions of power generation, energy storage, environment and load into a unified state representation, providing a complete operating condition description framework for the Markov chain model.
[0107] Step 2.7: Based on the joint state space, count the transition frequency of each joint state in historical data and construct a state transition probability matrix.
[0108] This embodiment constructs a state transition probability matrix by statistically analyzing the transition frequency of each joint state in historical data based on the joint state space. This quantifies the evolution law of historical operating conditions into probability transition relationships, enabling the system to query the probability distribution of future operating conditions based on the current state, thus providing a quantitative basis for forward-looking decision-making on power supply strategies.
[0109] Step 2.7.1: Sample each joint state in the joint state space according to a preset time interval and record the joint state at each sampling time.
[0110] This embodiment samples each joint state in the joint state space at a preset time interval, records the joint state at each sampling moment, and discretizes the continuous time process into a time-series state sequence, providing quantifiable sampling samples for subsequent statistics on transition frequencies.
[0111] Step 2.7.2: Count the number of times the transition from the first joint state to the second joint state occurs between adjacent sampling times, and use this number of occurrences as the transition frequency from the first joint state to the second joint state.
[0112] This embodiment quantifies the evolution relationship between operating states into calculable frequency data by statistically analyzing the number of times the state transitions from the first joint state to the second joint state between adjacent sampling times. This allows the state transition patterns in historical data to be quantitatively statistically analyzed.
[0113] Step 2.7.3: Sum the transition frequencies of the first joint state to all joint states to obtain the total transition frequency of the first joint state.
[0114] This embodiment obtains the total transition frequency of the first joint state by summing the transition frequencies from the first joint state to all joint states. This provides a denominator basis for the normalization calculation of subsequent transition probabilities, ensuring the mathematical completeness of the probability calculation.
[0115] Step 2.7.4: Divide the frequency of transitions from the first joint state to the second joint state by the total frequency of transitions from the first joint state to obtain the transition probability of transitioning from the first joint state to the second joint state.
[0116] This embodiment obtains the transition probability from the first joint state to the second joint state by dividing the transition frequency from the first joint state to the second joint state by the total transition frequency of the first joint state. This converts absolute frequency into relative probability, making the transition relationship between different states comparable and eliminating the influence of the difference in the number of samplings on the probability estimation.
[0117] Step 2.7.5: Traverse all joint states in the joint state space and construct the state transition probability matrix.
[0118] This embodiment constructs a state transition probability matrix by traversing all joint states in the joint state space, organizing the probability transition relationships between all joint states into a unified matrix form, and providing a data structure that can be directly called for querying the probability of future operating conditions based on the current state.
[0119] Step 2.8: Using the actual discrete level of each joint state at the current time and the current working mode as input, query the state transition probability matrix and output the probability of occurrence of each candidate joint state within a future time window as the working condition prediction result.
[0120] This embodiment uses the actual discrete level of each joint state at the current moment and the current working mode as input, queries the state transition probability matrix, and outputs the probability of occurrence of each candidate joint state within a future time window as the operating condition prediction result. This enables the system to quantitatively predict the probability of occurrence of each candidate operating condition state in the future based on the current operating condition state, providing a probabilistic operating condition prediction basis for the calculation of the energy benefit risk ratio of subsequent power supply strategies. This overcomes the shortcomings of existing technologies that rely solely on current environmental parameters for decision-making and cannot predict the trend of operating condition changes.
[0121] Step 3: Based on the operating condition prediction results, calculate the expected net energy gain and the expected potential energy loss for each candidate control strategy, and define the ratio of the expected net energy gain to the expected potential energy loss as the energy gain risk ratio of the candidate control strategy.
[0122] Step 3.1: Use the probability of occurrence of each candidate joint state in the predicted working condition as a weight.
[0123] This embodiment uses the probability of occurrence of each candidate joint state in the predicted operating conditions as a weight, so that the calculation of the expected value of net energy gain and the expected value of potential energy loss can be weighted according to the probability of occurrence of each operating condition, thus avoiding decision-making bias caused by only considering a single worst-case or best-case operating condition.
[0124] Step 3.2: For each candidate control strategy, calculate its power generation and power consumption under each candidate joint state. Subtract the power consumption from the power generation, multiply by the corresponding probability of occurrence, and sum them to obtain the expected net energy gain.
[0125] In this embodiment, when the predicted power of photovoltaic power generation and the predicted power of hydropower generation are both lower than the real-time load power of the equipment under the same candidate joint state, the condition is marked as a shortage state, and a preset shortage penalty is deducted from the net energy revenue expectation value.
[0126] This embodiment calculates the power generation and power consumption of each candidate control strategy under each candidate joint state. The power generation is subtracted from the power consumption, multiplied by the corresponding probability of occurrence, and summed to obtain the expected net energy return. The net returns under each operating state are weighted and summarized according to their probability of occurrence, ensuring that the return assessment reflects the long-term average effect rather than instantaneous extreme values. When the predicted power of photovoltaic power generation and the predicted power of hydropower generation are simultaneously lower than the real-time load power of the equipment under the same candidate joint state, this operating state is marked as a shortage state, and a shortage penalty is deducted from the expected net energy return. This provides an additional penalty for the extreme operating state where both sources are simultaneously insufficient, guiding the decision-making algorithm to prioritize avoiding such high-risk operating states.
[0127] Step 3.3: For each candidate control strategy, calculate the probability that it will cause the battery state of charge to fall below the preset low charge threshold under each candidate joint state, and the probability that the rate of decrease in battery state of charge will exceed the preset rate threshold after the execution of the candidate control strategy. Multiply the weighted sum of the two by the preset penalty coefficient for data loss due to power failure of the device to obtain the expected value of potential energy loss.
[0128] In this embodiment, the preset rate threshold is dynamically adjusted based on the sliding average of the device's historical power consumption records.
[0129] This embodiment calculates the probability that each candidate control strategy will cause the battery state of charge to fall below the low charge threshold under each candidate joint state, and the probability that the rate of decrease in battery state of charge will exceed the rate threshold after the execution of the candidate control strategy. The two are weighted and summed and then multiplied by the penalty coefficient for data loss due to power failure to obtain the expected value of potential energy loss. The power failure risk of the power supply strategy is comprehensively evaluated from two dimensions: the absolute value of charge and the rate of charge decrease. This allows dangerous trends of excessively rapid decrease in charge but not yet reaching the low charge threshold to be included in the loss assessment, thus achieving the forward-looking identification of potential power failure risks.
[0130] Step 3.4: Define the ratio of the expected net energy gain to the expected potential energy loss as the energy gain risk ratio of the candidate control strategy.
[0131] This embodiment defines the ratio of expected net energy gain to expected potential energy loss as the energy gain-risk ratio of candidate control strategies. This allows the evaluation results of each candidate control strategy to simultaneously reflect the comprehensive trade-off between its potential gains and the risk of power outages. When the expected net energy gain is the same or close, strategies with lower expected potential energy loss obtain a higher energy gain-risk ratio, thus gaining a higher probability of selection in subsequent decisions. This achieves a unified quantification of multiple objectives related to gains and risks in power supply strategy selection.
[0132] Step 4: Determine the selection probability of each candidate regulation based on the energy benefit-risk ratio.
[0133] Step 4.1: Arrange the energy-benefit-risk ratios of each candidate control strategy in descending order.
[0134] This embodiment arranges the energy-benefit-risk ratios of each candidate regulation strategy in descending order, establishing an ordered sequence basis for subsequent spacing calculations and probability adjustments between adjacent strategies. This places strategies with similar energy-benefit-risk ratios in adjacent positions within the sequence, facilitating local comparisons and differentiation processing.
[0135] Step 4.2: Use the difference between two adjacent energy gain-risk ratios after sorting as the interval between adjacent intervals.
[0136] This embodiment uses the difference between the energy gain risk ratios of two adjacent adjacent strategies after sorting as the distance between the adjacent intervals, quantifying the risk ratio differences between adjacent strategies into a calculable distance value. This allows the relative superiority or inferiority of strategies to be quantitatively evaluated, providing a quantitative basis for the dynamic adjustment of subsequent selection probabilities.
[0137] Step 4.3: Dynamically adjust the selection probability of each candidate control strategy according to the size of the adjacent spacing, so that the probability difference between adjacent candidate control strategies with larger spacing is smaller, and the probability difference between adjacent candidate control strategies with smaller spacing is larger.
[0138] When the adjacent spacing is greater than a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is set to be equal.
[0139] When the adjacent spacing is less than or equal to a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is allocated according to the positive correlation between energy gain and risk ratio.
[0140] In this embodiment, when the energy-benefit-risk ratio of the candidate control strategy with the highest energy-benefit-risk ratio exceeds the sum of the energy-benefit-risk ratios of all other candidate control strategies, the selection probability of that candidate control strategy is forcibly set to no more than the arithmetic mean of the selection probabilities of all candidate control strategies.
[0141] This embodiment dynamically adjusts the selection probability of each candidate control strategy based on the size of the adjacent spacing, so that the probability difference between adjacent candidate control strategies with larger spacing is smaller, and the probability difference between adjacent candidate control strategies with smaller spacing is larger. When the adjacent spacing is greater than a preset threshold, the selection probabilities of the two candidate control strategies corresponding to that adjacent spacing are set to be equal; when the adjacent spacing is less than or equal to the preset threshold, the selection probabilities of the two candidate control strategies corresponding to that adjacent spacing are allocated according to the positive correlation between energy-benefit-risk ratio. This adjustment mechanism assigns similar selection probabilities to adjacent strategies with large differences in energy-benefit-risk ratio, forcibly retaining the opportunity to explore low-risk strategies; at the same time, it allocates probabilities to adjacent strategies with small differences in energy-benefit-risk ratio according to the positive correlation, maintaining the ability to utilize small advantages, thus achieving an adaptive balance between exploration and utilization in strategy selection.
[0142] Step 5: Execute a candidate control strategy randomly selected by the controller according to the selection probability to obtain the control strategy vector and actual benefit. Store the actual benefit and the current operating condition as empirical data in the empirical playback buffer.
[0143] Step 5.1: Use the controller to randomly select a candidate control strategy based on the selection probability of each candidate control strategy.
[0144] Step 5.1.1: Accumulate the selection probabilities of each candidate control strategy in sequence to construct a cumulative probability interval, wherein each candidate control strategy corresponds to a continuous cumulative probability sub-interval, and the length of the cumulative probability sub-interval is equal to the selection probability of the candidate control strategy.
[0145] This embodiment constructs a cumulative probability interval by sequentially accumulating the selection probabilities of each candidate control strategy. Each candidate control strategy corresponds to a continuous cumulative probability sub-interval, and the length of the cumulative probability sub-interval is equal to the selection probability of the candidate control strategy. This transforms the discrete probability distribution into a continuous number line interval mapping, providing an indexable interval partitioning structure for subsequent random sampling operations.
[0146] Step 5.1.2: Use the controller to generate a random number between 0 and 1.
[0147] This embodiment utilizes the controller to generate a random number between 0 and 1, thus clearly defining the source of randomness in the decision-making process as the random number generator built into the controller. This makes the result of each strategy selection unpredictable and avoids the strategy rigidity caused by deterministic rules.
[0148] Step 5.1.3: Determine the cumulative probability sub-interval in which the random number falls, and select the candidate control strategy corresponding to the cumulative probability sub-interval as the candidate control strategy to be extracted.
[0149] This embodiment determines the cumulative probability sub-interval into which the random number falls, and uses the candidate control strategy corresponding to the cumulative probability sub-interval as the candidate control strategy to be extracted. This achieves random sampling according to the selection probability distribution, so that the frequency of strategy selection is consistent with its selection probability, and achieves the optimal strategy distribution in a probabilistic sense in long-term operation.
[0150] Step 5.1.4: Execute the extracted candidate control strategies to obtain actual benefits.
[0151] This embodiment obtains actual benefits by executing the extracted candidate control strategies, transforms the results of probabilistic decision-making into actual control operations on the power supply system, and forms a closed loop through the feedback of actual benefits after execution, providing real sample data for subsequent experience playback and parameter self-learning.
[0152] Step 5.2: Store the actual revenue and current operating status as empirical data in the empirical playback buffer.
[0153] In this embodiment, the control strategy vector includes a power generation priority switching command, a maximum power point tracking disturbance observation step size, an energy storage charging current limit value, and a load pre-degradation advance amount.
[0154] Step 6: Based on the control strategy vector and the actual benefit, sample from the experience replay buffer according to the priority of the prediction error, and use the parameter soft update method for self-learning to obtain the optimized execution parameter set.
[0155] Step 6.1: Based on the control strategy vector and the actual benefit, sample from the experience replay buffer according to the priority of the prediction error.
[0156] Step 6.1.1: Determine the expected return before execution based on the control strategy vector, determine the actual return after execution based on the actual return, and use the absolute value of the difference between the expected return and the actual return as the basic prediction error.
[0157] This embodiment determines the expected return before execution based on the control strategy vector, and determines the actual return after execution based on the actual return. The absolute value of the difference between the expected return and the actual return is used as the basic prediction error, and the prediction deviation is quantified into a calculable value, so that the prediction accuracy can be quantitatively evaluated and provide a ranking basis for subsequent priority sampling.
[0158] Step 6.1.2: Extract the expected return sequence of a preset number of consecutive experience data under the same working condition from the experience playback buffer, and calculate the variance of the expected return sequence as the temporal instability.
[0159] This embodiment extracts the expected return sequence of a preset number of consecutive experience data under the same working condition from the experience playback buffer, calculates the variance of the expected return sequence as the temporal instability, and quantifies the degree of fluctuation of the prediction result on the time axis as an instability index, so that the degree of inconsistency of the prediction model output under the same working condition can be quantitatively identified.
[0160] Step 6.1.3: The product of the basic prediction error and the temporal instability is taken as the comprehensive prediction error.
[0161] This embodiment uses the product of the basic prediction error and the temporal instability as the comprehensive prediction error, and integrates the errors of the two dimensions of single prediction deviation and temporal fluctuation of the prediction result into a single comprehensive index. This allows the empirical data with inaccurate and unstable predictions to obtain a higher comprehensive error value, thereby obtaining a higher sampling probability in priority sampling.
[0162] Step 6.1.4: Prioritize sampling of the empirical data in the empirical playback buffer according to the order of comprehensive prediction error from largest to smallest.
[0163] In this embodiment, the optimized set of execution parameters includes the operating frequency of the maximum power point tracking controller, the fine-tuning amount of the reference voltage of the multi-channel regulated output channel, the offset of the battery equalization circuit activation threshold, and the transmission power level of the wireless transmission module.
[0164] This embodiment prioritizes the sampling of empirical data in the empirical replay buffer according to the order of comprehensive prediction error from largest to smallest. This allows empirical data with large prediction deviations and large time-series fluctuations to be used first for model updates, thus accelerating the model's learning speed for abnormal and uncertain operating conditions.
[0165] Step 6.2: Self-learning is performed using a parameter soft update method to obtain the optimized set of execution parameters.
[0166] Step 6.2.1: Calculate the first adjustment gradient based on the empirical data obtained from priority sampling.
[0167] This embodiment calculates the first adjustment gradient based on the empirical data obtained from priority sampling, so that the high-quality empirical data obtained from sampling can guide the update direction of the execution parameter set, thus realizing the organic integration of the experience replay mechanism and parameter optimization.
[0168] Step 6.2.2: Extract the regret value corresponding to the current working condition from the experience replay buffer.
[0169] This embodiment extracts the regret value corresponding to the current working state from the experience replay buffer, and introduces the gap between the actual benefit and the theoretical best benefit in historical decisions as an additional learning signal into the parameter update process, so that the model can learn from the "missed best choice".
[0170] Step 6.2.3: Calculate the second adjustment gradient based on the regret value, wherein the magnitude of the second adjustment gradient is positively correlated with the regret value.
[0171] This embodiment calculates a second adjustment gradient based on the regret value. The magnitude of the second adjustment gradient is positively correlated with the regret value, so that the larger the decision error, the larger the adjustment gradient, thus enhancing the model's ability to correct serious erroneous decisions.
[0172] Step 6.2.4: Sum the first adjustment gradient and the second adjustment gradient with weights to obtain the comprehensive adjustment gradient.
[0173] This embodiment obtains a comprehensive adjustment gradient by weighted summing of the first adjustment gradient and the second adjustment gradient, thus merging the gradient direction from successful experience with the correction gradient direction from regret value into a unified update direction, achieving collaborative optimization of experience learning and regret minimization.
[0174] Step 6.2.5: Move the current set of execution parameters one step along the direction of the comprehensive adjustment gradient to obtain the optimized set of execution parameters.
[0175] This embodiment obtains an optimized set of execution parameters by moving the current set of execution parameters one step along the direction of the comprehensive adjustment gradient. The soft update method with single-step movement makes the parameters evolve smoothly, avoiding system instability caused by parameter mutations.
[0176] Step 7: Based on the optimized set of execution parameters, adaptively control the self-powered water quality detection system.
[0177] Step 7.1: Write the operating frequency of the maximum power point tracking controller from the optimized execution parameter set into the maximum power point tracking controller.
[0178] This embodiment writes the operating frequency of the maximum power point tracking controller into the optimized execution parameter set, enabling the maximum power point tracking response speed of the photovoltaic power generation unit to be dynamically adjusted according to the self-learning results, thereby improving the energy capture efficiency of photovoltaic power generation under different lighting conditions.
[0179] Step 7.2: Write the reference voltage fine-tuning amount of the multi-channel regulated output channel in the optimized execution parameter set into the load adapter module.
[0180] This embodiment writes the fine-tuning amount of the reference voltage of the multi-channel regulated output channel in the optimized execution parameter set into the load adaptation module, so that the output voltage of each load channel can be finely calibrated according to the self-learning results, avoiding power supply fluctuations caused by the mismatch between the fixed reference voltage and the actual load requirements.
[0181] Step 7.3: Write the battery equalization circuit activation threshold offset from the optimized execution parameter set into the battery management unit.
[0182] This embodiment writes the battery equalization circuit activation threshold offset from the optimized execution parameter set into the battery management unit, enabling the triggering conditions of the battery equalization strategy to be adaptively adjusted according to battery aging status and temperature changes, thereby extending the cycle life of the energy storage battery pack.
[0183] Step 7.4: Write the wireless transmission module transmit power level from the optimized execution parameter set into the wireless transmission unit.
[0184] This embodiment writes the wireless transmission module's transmit power level from the optimized execution parameter set into the wireless transmission unit, enabling the power consumption of wireless communication to dynamically balance data transmission reliability and energy consumption based on the self-learning results, thereby reducing communication energy consumption under unnecessary operating conditions.
[0185] Step 7.5: Adaptively regulate the self-powered water quality testing system.
[0186] This embodiment achieves closed-loop adaptive optimization of the power supply strategy of the water quality detection self-powered system by adaptively controlling the self-powered system for water quality detection and integrating all parameter writing operations in step seven into a unified control action. This enables the four modules of power generation, voltage stabilization, energy storage, and communication to operate collaboratively under the same optimized parameter set.
[0187] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0188] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An adaptive control method for a self-powered water quality detection system, characterized in that, include: Real-time acquisition of multimodal time-series data streams, including photovoltaic power time series, water flow velocity fluctuation variance, battery state of charge, battery temperature, and historical power consumption records of equipment; A Markov chain model is used to predict the probability of working condition transitions in a multimodal time series data stream, and the working condition prediction results are obtained, including the probability distribution of each candidate working condition state. Based on the operating condition prediction results, the expected net energy gain and the expected potential energy loss of each candidate control strategy are calculated respectively, and the ratio of the expected net energy gain to the expected potential energy loss is defined as the energy gain risk ratio of the candidate control strategy. The selection probability of each candidate regulation is determined based on the energy benefit-risk ratio. The controller randomly selects a candidate control strategy according to the selection probability and executes it to obtain the control strategy vector and actual benefit. The actual benefit and the current operating condition are stored as empirical data in the empirical playback buffer. The control strategy vector includes the generation priority switching instruction, the maximum power point tracking disturbance observation step size, the energy storage charging current limit value, and the load pre-degradation advance. Based on the control strategy vector and the actual benefit, samples are prioritized from the experience replay buffer according to the magnitude of the prediction error, and self-learning is performed using a parameter soft update method to obtain an optimized set of execution parameters, including the operating frequency of the maximum power point tracking controller, the fine-tuning amount of the reference voltage of the multi-channel regulated output channel, the offset of the battery equalization circuit activation threshold, and the transmission power level of the wireless transmission module; Based on the optimized set of execution parameters, the self-powered water quality detection system is adaptively controlled.
2. The adaptive control method for the self-powered water quality detection system as described in claim 1, characterized in that, A Markov chain model is used to predict the operating condition transition probability of a multimodal time-series data stream, and the operating condition prediction results are obtained, including: The photovoltaic power time series is discretized into N power levels according to the measured power values; The variance of water flow velocity fluctuation is discretized into M turbulence levels based on the measured variance values; The battery state of charge is discretized into K energy levels based on the measured percentage. The battery temperature is discretized into T temperature levels based on the measured temperature values; The device's historical power consumption records are divided into standby mode, acquisition mode, and transmission mode according to the working mode. The power level, turbulence level, electrical quantity level, temperature level, and operating mode are combined by Cartesian product to construct a joint state space; Based on the joint state space, the transition frequency of each joint state in historical data is statistically analyzed to construct a state transition probability matrix; Using the actual discrete level of each joint state at the current moment and the current working mode as input, the state transition probability matrix is queried, and the probability of occurrence of each candidate joint state within a future time window is output as the working condition prediction result.
3. The adaptive control method for the self-powered water quality detection system as described in claim 2, characterized in that, Based on the joint state space, the transition frequency of each joint state in historical data is statistically analyzed to construct a state transition probability matrix, including: Sample each joint state in the joint state space at a preset time interval and record the joint state at each sampling time. The number of times the transition from the first joint state to the second joint state occurs between adjacent sampling times is counted, and this number of occurrences is taken as the transition frequency from the first joint state to the second joint state; The total transition frequency of the first joint state is obtained by summing the transition frequencies from the first joint state to all joint states. The transition frequency from the first joint state to the second joint state is divided by the total transition frequency of the first joint state to obtain the transition probability from the first joint state to the second joint state. Traverse all joint states in the joint state space and construct the state transition probability matrix.
4. The adaptive control method for the self-powered water quality detection system as described in claim 3, characterized in that, Based on the operating condition forecast results, the expected net energy gain and expected potential energy loss of each candidate control strategy are calculated, including: The probability of occurrence of each candidate joint state in the predicted working condition is used as the weight; For each candidate control strategy, the power generation and power consumption under each candidate joint state are calculated. The power generation is subtracted from the power consumption, multiplied by the corresponding probability of occurrence, and summed to obtain the expected net energy revenue. When the predicted power of photovoltaic power generation and the predicted power of hydropower generation are both lower than the real-time load power of the equipment under the same candidate joint state, the condition is marked as a shortage state, and a preset shortage penalty is deducted from the expected net energy revenue. For each candidate control strategy, the probability that it will cause the battery state of charge to fall below a preset low charge threshold under each candidate joint state is calculated, as well as the probability that the rate of decrease in battery state of charge will exceed a preset rate threshold after the execution of the candidate control strategy. The two are weighted and summed and then multiplied by a preset penalty coefficient for data loss due to power failure of the device to obtain the expected value of potential energy loss. The preset rate threshold is dynamically adjusted according to the sliding average value of the device's historical power consumption records.
5. The adaptive control method for the self-powered water quality detection system as described in claim 4, characterized in that, The selection probability of each candidate regulation is determined based on the energy benefit-risk ratio, including: Arrange the energy-benefit-risk ratios of each candidate control strategy in descending order; The difference between two adjacent energy gain-risk ratios after sorting is used as the interval between adjacent intervals. The selection probability of each candidate control strategy is dynamically adjusted according to the size of the adjacent spacing, so that the probability difference between adjacent candidate control strategies with larger spacing is smaller, and the probability difference between adjacent candidate control strategies with smaller spacing is larger. Specifically, when the energy-benefit-risk ratio of the candidate control strategy with the highest energy-benefit-risk ratio exceeds the sum of the energy-benefit-risk ratios of all other candidate control strategies, the selection probability of that candidate control strategy is forcibly set to no more than the arithmetic mean of the selection probabilities of all candidate control strategies.
6. The adaptive control method for the self-powered water quality detection system as described in claim 5, characterized in that, The selection probability of each candidate control strategy is dynamically adjusted based on the size of the adjacent spacing, including: When the adjacent spacing is greater than a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is set to be equal. When the adjacent spacing is less than or equal to a preset threshold, the selection probability of the two candidate control strategies corresponding to the adjacent spacing is allocated according to the positive correlation between energy gain and risk ratio.
7. The adaptive control method for the self-powered water quality detection system as described in claim 6, characterized in that, The controller randomly selects a candidate control strategy based on the selection probability of each candidate control strategy, including: The selection probabilities of each candidate control strategy are accumulated sequentially to construct a cumulative probability interval, wherein each candidate control strategy corresponds to a continuous cumulative probability sub-interval, and the length of the cumulative probability sub-interval is equal to the selection probability of the candidate control strategy. Use the controller to generate a random number between 0 and 1; The cumulative probability sub-interval into which the random number falls is determined, and the candidate control strategy corresponding to the cumulative probability sub-interval is selected as the candidate control strategy.
8. The adaptive control method for the self-powered water quality detection system as described in claim 7, characterized in that, Based on the control strategy vector and the actual return, sampling is performed from the experience replay buffer according to the priority of prediction error magnitude, including: The expected return before execution is determined based on the control strategy vector, and the actual return after execution is determined based on the actual return. The absolute value of the difference between the expected return and the actual return is used as the basic prediction error. Extract the expected return sequence of a predetermined number of consecutive empirical data under the same working condition from the experience replay buffer, and calculate the variance of the expected return sequence as the temporal instability. The product of the basic prediction error and the temporal instability is taken as the comprehensive prediction error; The empirical data in the empirical playback buffer are sampled according to the order of the comprehensive prediction error from largest to smallest.
9. The adaptive control method for the self-powered water quality detection system as described in claim 8, characterized in that, Self-learning is performed using a parameter soft update method to obtain an optimized set of execution parameters, including: The first adjustment gradient is calculated based on the empirical data obtained from priority sampling. Extract the regret value corresponding to the current operating condition from the experience replay buffer; A second adjustment gradient is calculated based on the regret value, and the magnitude of the second adjustment gradient is positively correlated with the regret value; The first adjustment gradient and the second adjustment gradient are weighted and summed to obtain the comprehensive adjustment gradient; The current set of execution parameters is moved one step along the direction of the comprehensive adjustment gradient to obtain the optimized set of execution parameters.
10. The adaptive control method for the self-powered water quality detection system as described in claim 9, characterized in that, Based on the optimized set of execution parameters, the self-powered water quality detection system is adaptively controlled, including: Write the operating frequency of the maximum power point tracking controller from the optimized set of execution parameters into the maximum power point tracking controller; Write the reference voltage fine-tuning amount of the multi-channel regulated output channel in the optimized execution parameter set into the load adaptation module; Write the battery equalization circuit activation threshold offset from the optimized execution parameter set into the battery management unit. Write the wireless transmission module transmit power level from the optimized execution parameter set into the wireless transmission unit; Adaptive regulation is implemented for the self-powered water quality testing system.