Machine learning adaptive feeding device coupled with bird behavior and water quality and method of use

By using a machine learning-based adaptive feeding device to monitor bird behavior and water quality in real time and dynamically adjust the amount of food fed, the problem of insufficient or excessive feeding in existing technologies has been solved, achieving the dual goals of bird survival and water quality protection.

CN122375508APending Publication Date: 2026-07-14CHAOHU UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHAOHU UNIV
Filing Date
2026-05-22
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing feeding techniques for overwintering birds cannot dynamically adjust the amount of food according to the actual needs of the birds, resulting in insufficient or excessive food supply, water pollution, and a vicious cycle. Furthermore, there is a lack of water quality feedback and control mechanisms.

Method used

An adaptive feeding device driven by machine learning, which couples bird behavior with water quality, is used. It combines an AI camera, a water quality sensor, and machine learning algorithms to monitor bird behavior and water quality parameters in real time, and adaptively adjust the amount and strategy of feeding to form a closed-loop control.

Benefits of technology

This system enables dynamic feeding based on bird needs and water quality, avoiding food waste and water pollution, and achieving a balance between bird food supply and aquatic ecological protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122375508A_ABST
    Figure CN122375508A_ABST
Patent Text Reader

Abstract

A kind of bird behavior and water quality coupling driven machine learning adaptive feeding device and its use method.The present application relates to ecological environment engineering technical technical field, disclose a kind of bird behavior and water quality coupling driven machine learning adaptive feeding device including adaptive feeding system decision module and feeding device module;The adaptive feeding system decision module is composed of AI camera, polysilicon solar panel, wireless communication module, battery, data processor and water quality sensor, the AI camera, polysilicon solar panel, wireless communication module, battery, data processor and water quality sensor are all installed on stainless steel support, the AI camera is located at the top of stainless steel support.The present application is automatically monitored by double-layer machine learning architecture, bird population density, behavior characteristics and environmental water quality, dynamically optimizes feeding strategy, realizes the self-adapting control of feeding amount, avoids the water pollution and food waste caused by excessive feeding, and reduces the endogenous release pollution caused by bird fight and food fight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological and environmental engineering technology, and in particular to a machine learning-based adaptive feeding device driven by the coupling of bird behavior and water quality, and its usage method. Background Technology

[0002] The wintering period is a critical stage for bird survival. Affected by multiple factors such as cold climate, scarcity of natural food resources, and degradation of habitat ecology, a large number of migratory wintering birds face survival difficulties due to food shortages and insufficient energy replenishment. Artificial feeding has become an important ecological protection means to ensure the survival rate of birds and maintain biodiversity in wetlands, lakes and other bird wintering habitats. However, the current artificial feeding of wintering birds and related supporting technologies are still in a crude and passive stage, with many technical defects and ecological risks.

[0003] Traditional methods of feeding wintering birds primarily involve manual, timed, and location-based feeding or fixed-time feeding equipment. These are extensive feeding methods with fixed times and amounts, failing to dynamically adjust to the birds' actual feeding needs. On one hand, this method cannot flexibly control the feeding amount based on changes in the wintering bird population and feeding activity, easily leading to insufficient or excessive food supply, resulting in the accumulation and rotting of leftover food. On the other hand, excessive leftover food directly pollutes the overlying water, causing ammonia nitrogen and COD levels to exceed standards. Furthermore, extensive feeding can cause large-scale bird gatherings, leading to frequent bird activity, fighting, and competition for food. This physical disturbance further promotes the release of nutrients from habitat sediments, exacerbating eutrophication and other ecological problems, creating a vicious cycle of "extensive feeding - water pollution - bird survival disturbance."

[0004] Existing technologies, such as the invention patent with application number CN118104582A, achieve the identification of bird species and individual identities to determine whether feeding is needed and the type of feed, but do not involve water quality monitoring feedback and optimization of feeding strategies; the invention patent with application number CN118587736A achieves quantitative feeding of specific bird species and numbers, but it adopts fixed rules for quantitative feeding, does not introduce a machine learning self-optimization mechanism, and does not incorporate water quality feedback; the invention patent with application number CN121481176A solves the problem of unclear bidirectional coupling relationship between feeding behavior and water quality feedback, but it does not have the ability to adaptively analyze and autonomously adjust feeding parameters, and it is designed for aquaculture scenarios, which is different from the wild overwintering bird protection scenario of this invention. In the field of water environment monitoring, existing technologies mostly involve independent routine water quality testing, with monitoring points separated from feeding points. Monitoring indicators are only used for assessing the ecological quality of water bodies, failing to correlate the degree of pollution in the overlying water with feeding behavior. This lack of correlation means that feeding strategies cannot be adjusted based on water pollution feedback signals, resulting in a lack of a closed-loop management mechanism encompassing "water quality monitoring - pollution feedback - feeding adjustment." Therefore, this invention develops a machine learning-driven adaptive feeding device and its application method that couples bird behavior with water quality, balancing the dual objectives of bird food supply and aquatic ecological protection, providing scientific and efficient technical support for wetland ecological management. Summary of the Invention

[0005] To address the technical problems raised in the background section, this invention provides a machine learning-driven adaptive feeding device coupled with bird behavior and water quality, and its usage method.

[0006] This invention is achieved using the following technical solution: A machine learning-driven adaptive feeding device coupled with bird behavior and water quality includes the following steps: Adaptive feeding system decision module and feeding device module; The adaptive feeding system decision module consists of an AI camera, a polycrystalline silicon solar panel, a wireless communication module, a battery, a data processor, and a water quality sensor. The AI ​​camera, polycrystalline silicon solar panel, wireless communication module, battery, data processor, and water quality sensor are all mounted on a stainless steel bracket. The AI ​​camera is located at the top of the stainless steel bracket, the polycrystalline silicon solar panel is located above the battery, the wireless communication module is located on the battery, the data processor is located at the bottom of the battery, and the water quality sensor is located below the data processor. The feeding device module consists of an inlet, a circular plastic bucket, multiple outlets, a conveyor, a feeding inlet baffle, a feeding inlet, a high-pressure blower, a controller, and a wireless sensor. The conveyor is an inclined belt conveyor structure installed inside the circular plastic bucket. The inlet is connected to the top of the circular plastic bucket. The higher end of the conveyor is located directly below the inlet, and its lower end extends to the inner edge of the feeding bucket.

[0007] As a further improvement to the above solution, multiple discharge ports are located inside the circular plastic bucket and directly above the conveying equipment. Food particles fall into the conveying equipment through the discharge ports and are stably transported along the inclined surface towards the feeding port under the combined action of gravity and the friction of the conveyor belt on the conveying equipment. The feeding port is a rectangular opening, and its inner side forms a vertical material drop channel with the lower end of the conveying equipment. The feeding port baffle is slidably installed inside the feeding port via a slide rail. A bracket is installed on the top of the inner wall of the feeding port, and a stepper motor is installed on the bracket. The output shaft of the stepper motor is connected to a vertical lead screw through a coupling. One end of the lead screw passes through a lead screw nut fixed to the top of the feeding port baffle, and the lead screw and nut are screwed together. The controller is installed on the outside of the feeding port, and the high-pressure blower is installed at the bottom of the feeding port. The air outlet of the high-pressure blower is connected to the bottom of the feeding port.

[0008] As a further improvement to the above solution, the data processor incorporates a two-layer machine learning decision model. The data processor, in conjunction with a wireless communication module, can communicate with a cloud server and coordinate the control of multiple feeders. The data processor pre-stores a "bird species-food particle size" matching rule table, which can set a particle size of 1-2 mm for small birds and 3-5 mm for large birds based on beak size and feeding habits. This matching logic can be continuously optimized through machine learning. The matching rule table is pre-established based on the beak shape and feeding habits of different wetland birds and supports incremental learning updates and optimization via a cloud server. The water quality sensor consists of a dissolved oxygen sensor, a total dissolved solids sensor, a chemical oxygen demand sensor, a total organic carbon sensor, a total nitrogen sensor, and a total phosphorus sensor, and is signal-connected to the data processor and the wireless communication module. The inner wall of the circular plastic bucket is fitted with a semiconductor cooling chip and a desiccant box. The semiconductor cooling chip serves as a temperature control device, and the desiccant box serves as a humidity control device.

[0009] As a further improvement of the above solution, an air flow purger is provided at the feeding port and the end of the conveying device as a food residue cleaning device. The air flow purger is respectively led to the end area of the conveying plate on the inner wall of the barrel and the outer flange area of the feeding port through a Y-shaped shunt air duct. The two blowing ports are independently controlled and started regularly. The blowing port at the end of the conveying device blows downward at an angle of 45° along the inner wall of the barrel for directional purging, blowing the food debris remaining on the conveyor belt into the feeding port. The blowing port outside the feeding port blows outward along the edge of the feeding port to remove the external residues attached to the edge of the feeding port due to the wind swirl.

[0010] As a further improvement of the above solution, the double-layer machine learning decision model includes an upper-layer scene classification model, a lower-layer feeding strategy optimization model, and a scene category C. The upper-layer scene classification model is constructed using a supervised learning algorithm. The input state space feature vector S includes bird species coding, population quantity, activity intensity index, initial water quality vector, time parameter, and weather parameter, and outputs the current scene category. The scene category C includes high-density food competition scenes, low-water-temperature and low-activity scenes, exploratory feeding scenes, and pollution-sensitive scenes.

[0011] As a further improvement of the above solution, the upper-layer scene classification model uses the random forest algorithm. The specific configuration parameters are as follows: the number of decision trees is 100, the maximum depth is 15 layers, the number of randomly selected features during splitting is the integer part of the square root of the total number of features, the minimum number of samples per node is 5. The offline training data set of the upper-layer scene classification model is not less than 5,000 samples, which are artificial annotation data from different seasons and different wetland types. The annotation specifications include: high-density food competition scenes (bird density > 50 and activity intensity > 0.7), low-water-temperature and low-activity scenes (water temperature < 5°C and activity intensity < 0.3), exploratory feeding scenes (bird quantity fluctuation > 50% and feeding hesitation time > 10 seconds), pollution-sensitive scenes (0 < WQI < 50). Its model parameters are obtained through offline training and can be updated regularly through incremental learning.

[0012] As a further improvement to the above scheme, the lower-level feeding strategy optimization model is constructed using a reinforcement learning algorithm model. It uses the output of the scene classification model as input, defining a state space S, an action space A, and a reward function R. The optimal feeding strategy is learned through interaction with the environment. The reinforcement learning algorithm model employs a Deep Q-Network (DQN) algorithm. Its network structure includes: one convolutional layer with a kernel size of 3×3, a stride of 1, and 32 filters; one convolutional layer with a kernel size of 3×3, a stride of 2, and 64 filters; one fully connected layer containing 256 neurons; one fully connected layer containing 128 neurons; and one output layer matching the dimension of the action space. Except for the output layer, all layers use the ReLU activation function. Training is performed using the Adam optimizer with an initial learning rate of 0.001. The learning rate decreases to 0.95 times the original rate every 1000 training steps. The sampling strategy for the experience replay pool is priority sampling based on the absolute value of the temporal difference error.

[0013] As a further improvement to the above scheme, the reward function R is defined as follows: ; In the formula: , These are the weighting coefficients; For the feeding reward function; This is a water quality penalty function; A function to penalize food waste; ; In the formula, The percentage of birds that eat food within the area; the initial percentage of waterfowl that eat food is set to 0.75 (3 / 4), and the further away from the target, the more negative the reward. ; In the formula, , represents the changes in various water quality indicators before and after feeding; This represents the background water quality values ​​before artificial feeding. Each water quality indicator has a weight value, and the sum of all weight values ​​is 1. ; In the formula, This refers to the amount of food remaining after feeding. This represents the total amount of food fed in this round; The vector format of the action space A is: A=[Flevel, Oaperture, Rspeed, T].

[0014] As a further improvement to the above scheme, the two-layer machine learning decision model also includes the definition of a foraging-induced pollution threshold and a corresponding feeding inhibition strategy. For example, when multiple indicators trigger the threshold simultaneously, the system executes the strictest inhibition strategy. When the water quality increment does not trigger any threshold for three consecutive feeding cycles, the system gradually removes the inhibition and restores the normal feeding strategy. Based on the dynamic coupling method of the foraging-induced pollution increment weight, the real-time adjustment value of the water quality penalty weight can be dynamically calculated to obtain the following formula: ; In the formula, The basic water quality penalty weight is determined based on the scenario category. This is an adjustment factor, with a default value of 1.0; The increment of the i-th water quality indicator before and after this round of feeding; Let be the incremental threshold for the i-th water quality indicator; The contribution weight of the i-th water quality indicator in foraging-induced pollution is dynamically determined by the entropy weight method. The steps for calculating the contribution weight wi using the entropy weight method are as follows: Select water quality increment data from the most recent k feeding cycles to form a matrix, standardize the data, calculate the information entropy Ej of each indicator, and then calculate the entropy weight of each indicator. The greater the recent incremental fluctuation of a certain water quality indicator, the greater its entropy weight, and the higher its attention weight in dynamic coupling. Ultimately, the water quality penalty term is revised as follows: .

[0015] The present invention also provides a method for using a machine learning adaptive feeding device driven by the coupling of bird behavior and water quality, comprising the following steps: Step 1: System initialization and deployment. Based on the historical distribution of wetland waterbirds, the feeding device module is connected and fixed with ropes using safety buckles. The system is deployed on the shore. The adaptive feeding system decision module is deployed to the bird gathering area. The system is debugged. The sampling frequency of the AI ​​camera (1) is set to 2 times per second. The parameters of the pre-trained scene classification model and reinforcement learning model are loaded.

[0016] Step 2: Environmental perception and feature extraction. The AI ​​camera (1) collects images of birds in the feeding area in real time, identifies the species, number and activity intensity of birds, and extracts bird feature vectors. The water sample detection system simultaneously collects pollutant parameters such as dissolved oxygen (DO), total dissolved solids (TDS), chemical oxygen demand (COD), total organic carbon (TOC), total nitrogen, and total phosphorus, and constructs a state vector S=[bird species code, population size, activity intensity, water quality vector, time parameter, amount of remaining food]; Step 3: Scene recognition. Input the state vector into the scene classification model and output the current scene category C∈{high-density feeding scene, low water temperature and low activity scene, tentative feeding scene, pollution sensitive scene}. Different scene categories correspond to different exploration rates ε and model parameters to achieve adaptive adjustment. Step 4: Feeding strategy decision-making. Input the state vector S and scene category C into the reinforcement learning model. The model outputs action A = [feeding amount, feeding interval, baffle aperture opening, high-pressure fan power]. The actuator starts the feeder according to the action command, feeding gradually from small to large, while continuously monitoring environmental water quality parameters. Step 5: Real-time feedback and reward calculation; continuously monitor the proportion of birds that eat food in the area during the feeding process. When the number of birds that have eaten reaches 3 / 4 of the total bird population in the area, the total amount of food given at this time is recorded as the candidate value for the optimal feeding amount for that period. Feeding is then stopped. The remaining 1 / 4 of birds that have not eaten continue their normal activities and maintain their natural ecological behavior. After feeding, a second water sample is taken for testing, and a reward value is calculated based on the changes in water quality. ; Step 6: Online model update. Store the experience data of this decision in the experience replay pool. After every N feedings, randomly sample batch data from the experience replay pool and use the experience replay mechanism to update the Q network parameters of the reinforcement learning model. Periodically use the newly accumulated data to perform incremental learning or transfer learning on the scene classification model to improve the model's generalization ability. Step 7: Multi-feeder coordinated control. When multiple feeders are deployed in the system, the controller sequentially checks whether each feeder Fn has completed the current feeding round: if the current feeder has not completed, return to step 2; if the current feeder has completed but there are other feeders that have not completed, n=n+1, proceed to the next feeder and execute steps 2 to 6; if all feeders have completed, the current feeding round ends. Step 8: Experience recording and strategy optimization. Record complete data of each feeding process in local storage, including state sequence, action sequence, reward value sequence and water quality change curve. The system periodically uploads the data to the cloud server for offline strategy evaluation and model optimization. The optimized model parameters are then sent back to the controllers of each device. The format of the state vector: .

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a machine learning adaptive feeding device and its usage method driven by the coupling of bird behavior and water quality, and builds a two-layer scenario: the upper-layer scenario classification model identifies the type of feeding scenario, and the lower-layer reinforcement learning model adaptively optimizes the feeding strategy according to the scenario. The system has self-learning ability and can automatically optimize feeding parameters according to different wetlands, different bird species combinations, and different seasons, avoiding repeated manual parameter tuning, which is significantly better than the existing fixed rule feeding system. 2. This invention can achieve fully automatic control by constructing a three-dimensional reward function based on the deviation of the birds' satiety ratio from the target value, the degree of water quality deterioration, and the amount of food wasted. The system spontaneously learns a feeding strategy that "does not pollute the environment," thus achieving a balance between the dual goals of bird food supply and aquatic ecological protection. 3. This invention forms a complete closed-loop control chain from bird behavior monitoring → water quality testing → feeding decision-making → effect evaluation → model updating, which can solve the problem of unclear bidirectional coupling relationship between feeding behavior and water quality feedback in the prior art. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the structure of the decision-making module of the adaptive feeding system of the present invention; Figure 2 This is a schematic diagram of the feeding device module of the present invention; Figure 3 This is a schematic diagram illustrating the system deployment and collaborative operation of multiple feeders according to the present invention; Figure 4 This is a schematic diagram of the feeding strategy process of the present invention; Figure 5 This is a schematic diagram of the state space definition of the present invention; Figure 6 This is a schematic diagram illustrating the changes in water quality and weights in this invention. Figure 7 This is a vector diagram of the action space A of the present invention; Figure 8 This is a schematic diagram of the adaptive weighting coefficients for the present invention. Figure 9 This is a schematic diagram illustrating the foraging-induced contamination threshold and the corresponding feeding inhibition strategy of the present invention.

[0019] Explanation of key symbols: 1. AI camera; 2. Polycrystalline silicon solar panel; 3. Wireless communication module; 4. Battery; 5. Data processor; 6. Water quality sensor; 7. Feed inlet; 8. Round plastic bucket; 9. Discharge outlet; 11. Conveying equipment; 12. Feeding port baffle; 13. Feeding port; 14. High-pressure blower; 15. Controller; 16. Wireless sensor; 17. Stainless steel bracket. Detailed Implementation

[0020] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.

[0021] Please combine Figure 1 The bird behavior and water quality coupled-driven machine learning adaptive feeding device of this embodiment includes: an adaptive feeding system decision module and a feeding device module; The adaptive feeding system decision module consists of an AI camera (1), a polycrystalline silicon solar panel (2), a wireless communication module, a battery (4), a data processor, and a water quality sensor. The AI ​​camera (1), polycrystalline silicon solar panel (2), wireless communication module, battery (4), data processor, and water quality sensor are all mounted on a stainless steel bracket (17). The AI ​​camera (1) is located at the top of the stainless steel bracket (17), the polycrystalline silicon solar panel (2) is located above the battery (4), the wireless communication module is located on the battery (4), the data processor is located at the bottom of the battery (4), and the water quality sensor is located below the data processor. The feeding device module consists of an inlet (7), a circular plastic bucket (8), multiple outlets (9), a conveyor (11), a feeding port (13), a baffle (12), a feeding port (13), a high-pressure blower (14), a controller (15), and a wireless sensor (16). The conveyor (11) is an inclined belt conveyor structure. The conveyor (11) is installed inside the circular plastic bucket (8). The inlet (7) is connected to the top of the circular plastic bucket (8). The higher end of the conveyor (11) is located directly below the inlet (7), and its lower end extends to the inner edge of the feeding bucket.

[0022] Multiple discharge ports (9) are located inside the circular plastic bucket (8) and directly above the conveying device (11). Food particles fall onto the conveying device (11) through the discharge ports and are stably transported along the inclined surface towards the feeding port (13) under the combined action of gravity and the friction of the conveyor belt on the conveying device (11). The feeding port (13) is a rectangular opening, and its inner side forms a vertical material drop channel with the lower end of the conveying device (11). The baffle (12) of the feeding port (13) is slidably installed on the feeding port (13) via a slide rail. Inside the feeding port (13), a bracket is installed on the top of the inner wall of the feeding port (13). A stepper motor is installed on the bracket. The output shaft of the stepper motor is connected to a vertical lead screw through a coupling. One end of the lead screw passes through a lead screw nut fixed to the top of the baffle (12) of the feeding port (13). The lead screw and the nut are screwed together. The controller (15) is installed on the outside of the feeding port (13). The high-pressure blower (14) is installed at the bottom of the feeding port (13). The air outlet of the high-pressure blower (14) is connected to the bottom of the feeding port (13).

[0023] The data processor has a built-in dual-layer machine learning decision model. The data processor, together with the wireless communication module, can be used to communicate with the cloud server and coordinate the control of multiple feeders. The data processor has a pre-stored "bird species-food particle size" matching rule table. It can set the particle size to 1-2 mm for small birds and 3-5 mm for large birds according to the size of the bird's beak and feeding habits. It can also continuously optimize the matching logic through machine learning. The matching rule table is pre-established according to the beak shape and feeding habits of different wetland birds and supports incremental learning updates and optimizations through the cloud server. The water quality sensor consists of a dissolved oxygen sensor, a total dissolved solids sensor, a chemical oxygen demand sensor, a total organic carbon sensor, a total nitrogen sensor and a total phosphorus sensor, and is connected to the data processor and the wireless communication module. The inner wall of the circular plastic bucket (8) is equipped with a semiconductor cooling chip and a desiccant box. The semiconductor cooling chip serves as a temperature control device, and the desiccant box serves as a humidity control device.

[0024] An airflow sweeper is provided at the end of the feeding port (13) and the conveying device (11) as a food residue cleaning device. The airflow sweeper is led to the end area of ​​the conveying plate on the inner wall of the bucket and the outer flange area of ​​the feeding port (13) through a Y-shaped diversion air duct. The two air ports are independently controlled and start at regular intervals. The air port at the end of the conveying device (11) is tilted downward at 45° along the inner wall of the bucket to blow food scraps left on the conveyor belt into the feeding port (13). The air port on the outer side of the feeding port (13) blows outward along the edge of the feeding port (13) to remove external residues that are attached to the edge of the feeding port (13) due to the wind swirling.

[0025] The two-layer machine learning decision model includes an upper-layer scene classification model, a lower-layer feeding strategy optimization model, and scene category C. The upper-layer scene classification model is constructed using a supervised learning algorithm. The input state space feature vector S includes bird species codes, population size, activity intensity index, initial water quality vector, time parameters, and weather parameters. The output is the current scene category. The scene category C includes high-density feeding competition scene, low water temperature and low activity scene, tentative feeding scene, and pollution-sensitive scene.

[0026] The upper-layer scenario classification model uses the random forest algorithm. The specific configuration parameters are as follows: the number of decision trees is 100, the maximum depth is 15 layers, the number of randomly selected features during splitting is the integer part of the square root of the total number of features, the minimum number of samples per node is 5. The offline training dataset of the upper-layer scenario classification model has no less than 5,000 samples, which are artificial annotation data from different seasons and different wetland types. The annotation specifications include: high-density feeding competition scenario (bird density > 50 and activity intensity > 0.7), low water temperature and low activity scenario (water temperature < 5°C and activity intensity < 0.3), exploratory feeding scenario (bird number fluctuation > 50% and feeding hesitation time > 10 seconds), pollution-sensitive scenario (0 < WQI < 50). Its model parameters are obtained through offline training and can be updated by incremental learning regularly.

[0027] The lower-layer feeding strategy optimization model is constructed using the reinforcement learning algorithm model. Taking the output of the scenario classification model as the input condition, the state space S, action space A, and reward function R are defined, and the optimal feeding strategy is learned through interaction with the environment. The reinforcement learning algorithm model uses the Deep Q-Network (DQN) algorithm. Its network structure includes: a convolutional layer with a convolutional kernel size of 3×3, a stride of 1, and 32 filters; a convolutional layer with a convolutional kernel size of 3×3, a stride of 2, and 64 filters; a fully connected layer with 256 neurons; a fully connected layer with 128 neurons; and an output layer matching the dimension of the action space. Except for the output layer, the ReLU activation function is used. Its training uses the Adam optimizer, with an initial learning rate of 0.001, which decays to 0.95 times the original every 1,000 training steps. The sampling strategy of the experience replay pool is priority sampling based on the absolute value of the temporal difference error.

[0028] The formula of the reward function R is defined as follows: ; In the formula: 、 are weight coefficients; is the feeding reward function; is the water quality penalty function; is the food waste penalty function; ; In the formula, is the proportion of birds that eat food in the area; the initial value of the proportion of water birds that eat food is set to 0.75 (3 / 4), and the farther away from the target, the more negative the reward; ; In the formula, , is the change in each water quality index before and after feeding; is the water quality background value before artificial feeding; Each water quality indicator has a weight value, and the sum of all weight values ​​is 1. ; In the formula, This refers to the amount of food remaining after feeding. This represents the total amount of food fed in this round; The vector format of the action space A is: A=[Flevel, Oaperture, Rspeed, T].

[0029] The two-layer machine learning decision model also includes the definition of foraging-induced pollution thresholds and corresponding feeding inhibition strategies. For example, when multiple indicators trigger the thresholds simultaneously, the system executes the strictest inhibition strategy. When the water quality increments for three consecutive feeding cycles do not trigger any thresholds, the system gradually removes the inhibition and restores the normal feeding strategy. Based on the dynamic coupling method of foraging-induced pollution increment weights, the real-time adjustment value of the water quality penalty weights can be dynamically calculated to obtain the following formula: ; In the formula, The basic water quality penalty weight is determined based on the scenario category. This is an adjustment factor, with a default value of 1.0; The increment of the i-th water quality indicator before and after this round of feeding; Let be the incremental threshold for the i-th water quality indicator; The contribution weight of the i-th water quality indicator in foraging-induced pollution is dynamically determined by the entropy weight method. The steps for calculating the contribution weight wi using the entropy weight method are as follows: Select water quality increment data from the most recent k feeding cycles to form a matrix, standardize the data, calculate the information entropy Ej of each indicator, and then calculate the entropy weight of each indicator. The greater the recent incremental fluctuation of a certain water quality indicator, the greater its entropy weight, and the higher its attention weight in dynamic coupling. Ultimately, the water quality penalty term is revised as follows: .

[0030] The present invention also provides a method for using a machine learning adaptive feeding device driven by the coupling of bird behavior and water quality, comprising the following steps: Step 1: System initialization and deployment. Based on the historical distribution of wetland waterbirds, the feeding device module is connected and fixed with ropes using safety buckles. The system is deployed on the shore. The adaptive feeding system decision module is deployed to the bird gathering area. The system is debugged. The sampling frequency of the AI ​​camera (1) is set to 2 times per second. The parameters of the pre-trained scene classification model and reinforcement learning model are loaded.

[0031] Step 2: Environmental perception and feature extraction. The AI ​​camera (1) collects images of birds in the feeding area in real time, identifies the species, number and activity intensity of birds, and extracts bird feature vectors. The water sample detection system simultaneously collects pollutant parameters such as dissolved oxygen (DO), total dissolved solids (TDS), chemical oxygen demand (COD), total organic carbon (TOC), total nitrogen, and total phosphorus, and constructs a state vector S=[bird species code, population size, activity intensity, water quality vector, time parameter, amount of remaining food]; Step 3: Scene recognition. Input the state vector into the scene classification model and output the current scene category C∈{high-density feeding scene, low water temperature and low activity scene, tentative feeding scene, pollution sensitive scene}. Different scene categories correspond to different exploration rates ε and model parameters to achieve adaptive adjustment. Step 4: Feeding strategy decision-making. Input the state vector S and scene category C into the reinforcement learning model. The model outputs action A = [feeding amount, feeding interval, baffle aperture opening, high-pressure fan power]. The actuator starts the feeder according to the action command, feeding gradually from small to large, while continuously monitoring environmental water quality parameters. Step 5: Real-time feedback and reward calculation; continuously monitor the proportion of birds that eat food in the area during the feeding process. When the number of birds that have eaten reaches 3 / 4 of the total bird population in the area, the total amount of food given at this time is recorded as the candidate value for the optimal feeding amount for that period. Feeding is then stopped. The remaining 1 / 4 of birds that have not eaten continue their normal activities and maintain their natural ecological behavior. After feeding, a second water sample is taken for testing, and a reward value is calculated based on the changes in water quality. ; Step 6: Online model update. Store the experience data of this decision in the experience replay pool. After every N feedings, randomly sample batch data from the experience replay pool and use the experience replay mechanism to update the Q network parameters of the reinforcement learning model. Periodically use the newly accumulated data to perform incremental learning or transfer learning on the scene classification model to improve the model's generalization ability. Step 7: Multi-feeder coordinated control. When multiple feeders are deployed in the system, the controller sequentially checks whether each feeder Fn has completed the current feeding round: if the current feeder has not completed, return to step 2; if the current feeder has completed but there are other feeders that have not completed, n=n+1, proceed to the next feeder and execute steps 2 to 6; if all feeders have completed, the current feeding round ends. Step 8: Experience recording and strategy optimization. Record complete data of each feeding process in local storage, including state sequence, action sequence, reward value sequence and water quality change curve. The system periodically uploads the data to the cloud server for offline strategy evaluation and model optimization. The optimized model parameters are then sent back to the controllers of each device. The format of the state vector: .

[0032] Before system deployment, bird activity data and water quality data from different seasons, time periods, and wetland environments are collected to construct a training dataset. Each sample contains a feature vector X and a scene category label y. Feature vector X = [bird species code, population size, activity intensity (foraging frequency per unit time), COD value, TDS value, DO value, total nitrogen value, total phosphorus value, TOC value, timestamp (season / time period code), temperature, humidity]; Scene category label y is divided into four categories: high-density competition for food scene (bird density > 50 birds and activity intensity > 0.7), low water temperature and low activity scene (water temperature < 5℃ and activity intensity < 0.3), tentative feeding scene (bird number fluctuation > 50% and feeding hesitation time > 10 seconds), and pollution-sensitive scene (0 <WQI<50); To ensure the model's generalization ability and robustness in scene transfer, the total amount of labeled data is no less than 5,000 records, covering the four time periods of dawn, daytime, dusk, and night in spring, summer, autumn, and winter. The data types include actual collected samples and data-enhanced samples from different wetland ecological zones in China. The random forest algorithm was used to train the classification model. The specific configuration parameters were: 100 decision trees, 15 layers at maximum depth, sqrt strategy (i.e., the square root of the total number of features) for the number of randomly selected features during splitting, 5 minimum number of samples per node, and 5-fold cross-validation to evaluate the model performance to ensure that the classification accuracy was not less than 90%. The trained model parameters were fixed in the data processor (5).

[0033] State-action-reward table design for reinforcement learning models: The state space S includes: bird population density level (low / medium / high), activity intensity index (continuous values ​​normalized to [0,1]), food consumption ratio (real-time monitoring value), water quality comprehensive index (calculated by weighting data from each sensor), remaining food quantity (percentage), scene category code (1-4), and time parameters (seasonal code and time period code).

[0034] The action space A includes: feeding amount level (6 levels from 0 to 5, corresponding to 0 / 50g / 100g / 200g / 300g / 500g respectively), baffle hole opening (0-100% continuously adjustable), high-pressure blower power level (4 levels from 0 to 3), and feeding interval level (4 levels from 0 to 3).

[0035] The reward function is as described in the aforementioned formula. The weighting coefficients α, β, and γ are dynamically adjusted according to the scenario category: in pollution-sensitive scenarios, β is taken as a larger value to prioritize water quality protection; in high-density food competition scenarios, α is taken as a larger value to prioritize meeting the needs of birds. This embodiment uses the Deep Q-Network (DQN) algorithm for training, and the network structure is designed as follows: 1. Input layer: Receives the preprocessed state vector.

[0036] 2. First convolutional layer: The kernel size is 3×3, the stride is 1, the number of filters is 32, and the activation function is ReLU.

[0037] 3. Second convolutional layer: The kernel size is 3×3, the stride is 2, the number of filters is 64, and the activation function is ReLU.

[0038] 4. First fully connected layer: 256 neurons, activation function is ReLU.

[0039] 5. Second fully connected layer: 128 neurons, activation function is ReLU.

[0040] 6. Output layer: The number of neurons matches the total dimension of the action space, and outputs the predicted Q value of each action.

[0041] In terms of training configuration, the Adam optimizer is used, with an initial learning rate set to 0.001 and a learning rate decay scheduling is implemented: after every 1000 training steps, the learning rate decays to 0.95 times its original value. The experience replay pool has a capacity of 10,000 experience data points, and a priority-based experience replay strategy is adopted: the sampling priority is determined based on the absolute value of the temporal difference error of the samples, and the samples with larger errors have a higher probability of being extracted, thereby accelerating the model convergence speed. The target network is updated every 100 steps, and the exploration rate ε decays exponentially from 0.9 to 0.01. During training, an ε-greedy strategy is used to balance exploration and utilization.

[0042] System operation process: Five adaptive feeding systems of this invention were deployed in a wetland park, with each device spaced 50 meters apart on the water surface.

[0043] System initialization: The controller loads the pre-trained scene classification model and DQN model parameters, and sets the AI ​​camera sampling frequency to twice per second.

[0044] At 8:00 AM on a winter morning, the device detected approximately 30 mallards gathering in the area, with an activity intensity index of 0.7 (relatively high). Water quality sensor data showed: COD=15mg / L, TOC=3.2mg / L, TDS=200ppm, DO=8.5mg / L, WQI=85, indicating good overall water quality. The scene classification model output "high-density feeding competition scene." The controller input this scene information into the DQN model, which output the following actions: feeding level 3 (approximately 200g), baffle aperture opening 60%, high-pressure blower level 2. The AI ​​camera identified the mallards as medium to large-sized birds and adjusted the baffle aperture to a corresponding 3-5mm particle size according to the matching rules.

[0045] After feeding, the system continuously monitored the proportion of birds eating. After about 3 minutes, the proportion of birds eating reached 3 / 4, and the controller stopped feeding. The amount of food fed was recorded as 180g. After feeding, the water quality was tested again: COD=16mg / L (up 6.7%), TOC=3.4mg / L (up 6.25%), WQI=80 (down 5.88%). The increase in water quality indicators did not trigger the pollution threshold caused by foraging, and the water quality was within the safe range. The reward function calculated R=-0.32, and this experience data was stored in the experience playback pool.

[0046] After the feeder Fn completes the feeding cycle, the system enters the next monitoring cycle. When 100 feeding experiences are accumulated, the system triggers the online model update, updates the DQN network parameters based on priority sampling of 32 experience data from the experience playback pool, and gradually optimizes the feeding strategy. The controller starts the food residue cleaning device (20) at regular intervals to remove food residue from the feeding port and the conveyor plate.

[0047] Example of foraging-induced contamination feedback and feeding inhibition: During an autumn evening, the device detected approximately 40 wild ducks gathering to forage in the area, and the scene classification model output "high-density competition for food scene". The DQN model output the following initial actions: feed amount level 4 (approximately 300g), baffle aperture opening degree 70%, high-pressure fan level 3 (high power), and feeding interval level 2 (medium frequency). During the feeding process, approximately 3 minutes and 15 seconds later, the proportion of birds that had consumed the food reached 3 / 4, at which point the system stopped feeding for that round. The actual total amount of food fed was 278g. A second water sample was immediately taken after feeding, and the results are as follows:

[0048] The system determines that the feeding inhibition mechanism is triggered simultaneously by the three indicators COD, TP, and DO, and executes the strictest strategy (DO inhibition strategy): suspend feeding for 1 cycle.

[0049] Simultaneously, the system calculates entropy weights based on water quality increment data from the past 10 feeding cycles. Due to recent significant fluctuations in COD and TP increments (entropy weights of 0.32 and 0.28 respectively), dynamic water quality penalty weights are calculated. The empirical data was increased from the base value of 0.6 to 1.38 and stored in the experience replay pool with high priority, enabling the DQN model to learn a more conservative feeding strategy in similar scenarios.

[0050] After three cycles of inhibition and monitoring, COD recovered to 16 mg / L and DO rose to 6.5 mg / L. The system gradually de-inhibited and normal feeding resumed.

[0051] Multiple feeders working together: When multiple feeders are deployed in the system, the controller uses a polling mechanism to process the feeding tasks of each feeder in turn. The controller first starts the first feeder F1. After completing a single feeding according to the above process, it determines whether F1 has completed the feeding round. If the feeding ratio has not reached 3 / 4 and the water quality increment has not triggered the inhibition threshold, then the next feeding round of F1 is executed. If the stopping condition has been met or the feeding inhibition has been triggered, then the current feeder number n is incremented by 1, and the same process is executed for feeder F2. When all feeders have completed the feeding round, the system enters a dormant waiting state until the next monitoring cycle is triggered.

[0052] The feeding strategies of each feeder are independent of each other, while sharing the same boundary environmental information, achieving collaborative management while fully adapting to the differentiated needs of each area.

[0053] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A machine learning-driven adaptive feeding device coupled with bird behavior and water quality, characterized in that, include: Adaptive feeding system decision module and feeding device module; The adaptive feeding system decision module consists of an AI camera (1), a polycrystalline silicon solar panel (2), a wireless communication module, a battery (4), a data processor, and a water quality sensor. The AI ​​camera (1), polycrystalline silicon solar panel (2), wireless communication module, battery (4), data processor, and water quality sensor are all mounted on a stainless steel bracket (17). The AI ​​camera (1) is located at the top of the stainless steel bracket (17), the polycrystalline silicon solar panel (2) is located above the battery (4), the wireless communication module is located on the battery (4), the data processor is located at the bottom of the battery (4), and the water quality sensor is located below the data processor. The feeding device module consists of an inlet (7), a circular plastic bucket (8), multiple outlets (9), a conveyor (11), a feeding port (13), a baffle (12), a feeding port (13), a high-pressure blower (14), a controller (15), and a wireless sensor (16). The conveyor (11) is an inclined belt conveyor structure. The conveyor (11) is installed inside the circular plastic bucket (8). The inlet (7) is connected to the top of the circular plastic bucket (8). The higher end of the conveyor (11) is located directly below the inlet (7), and its lower end extends to the inner edge of the feeding bucket.

2. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 1, characterized in that, Multiple discharge ports (9) are located inside the circular plastic bucket (8) and directly above the conveying device (11). Food particles fall onto the conveying device (11) through the discharge ports and are stably transported along the inclined surface towards the feeding port (13) under the combined action of gravity and the friction of the conveyor belt on the conveying device (11). The feeding port (13) is a rectangular opening, and its inner side forms a vertical material drop channel with the lower end of the conveying device (11). The baffle (12) of the feeding port (13) is slidably installed on the feeding port (13) via a slide rail. Inside the feeding port (13), a bracket is installed on the top of the inner wall of the feeding port (13). A stepper motor is installed on the bracket. The output shaft of the stepper motor is connected to a vertical lead screw through a coupling. One end of the lead screw passes through a lead screw nut fixed to the top of the baffle (12) of the feeding port (13). The lead screw and the nut are screwed together. The controller (15) is installed on the outside of the feeding port (13). The high-pressure blower (14) is installed at the bottom of the feeding port (13). The air outlet of the high-pressure blower (14) is connected to the bottom of the feeding port (13).

3. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 1, characterized in that, The data processor is built with a two-layer machine learning decision model. The data processor, in cooperation with the wireless communication module, can be used for communicating with the cloud server and for collaborative control of multiple feeders. A matching rule table of "bird species - food particle size" is pre-stored in the data processor. According to the size of the bird's beak and feeding habits, it is set that small birds correspond to a particle size of 1 - 2 mm, and large birds correspond to a particle size of 3 - 5 mm. And the matching logic can be continuously optimized through machine learning. The matching rule table is pre-established according to the beak shapes and feeding habits of different wetland birds, and supports incremental learning update and optimization through the cloud server. The water quality sensor consists of a dissolved oxygen sensor, a total dissolved solids sensor, a chemical oxygen demand sensor, a total organic carbon sensor, a total nitrogen sensor, and a total phosphorus sensor, and is signal-connected to the data processor and the wireless communication module. A semiconductor refrigeration sheet and a drying box are attached to the inner wall of the circular plastic bucket (8). The semiconductor refrigeration sheet serves as a temperature adjustment device, and the desiccant box serves as a humidity adjustment device.

4. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 3, characterized in that, An air flow purger is provided at the end of the feeding port (13) and the conveying device (11) as a food residue cleaning device. The air flow purger is respectively led to the end area of the conveying plate on the inner wall of the bucket and the outer flange area of the feeding port (13) through a Y-shaped shunt air duct. The two blowing ports are independently controlled and started regularly. The blowing port at the end of the conveying device (11) blows downward at an angle of 45° along the inner wall of the bucket for directional purging, blowing the food debris remaining on the conveyor belt into the feeding port (13). The blowing port outside the feeding port (13) blows outward along the border of the feeding port (13) to remove the external residues adhering to the edge of the feeding port (13) due to the wind swirl.

5. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 3, characterized in that, The two-layer machine learning decision model includes an upper-layer scene classification model, a lower-layer feeding strategy optimization model, and scene category C. The upper-layer scene classification model is constructed using a supervised learning algorithm. The input state space feature vector S includes bird species encoding, population quantity, activity intensity index, initial water quality vector, time parameter, and weather parameter, and outputs the current scene category. The scene category C includes a high-density food competition scene, a low water temperature and low activity scene, a tentative feeding scene, and a pollution-sensitive scene.

6. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 5, characterized in that, The upper-layer scene classification model uses the random forest algorithm. The specific configuration parameters are: the number of decision trees is 100, the maximum depth is 15 layers, the number of randomly selected features during splitting is the integer part of the square root of the total number of features, and the minimum number of samples at a node is 5. The offline training data set of the upper-layer scene classification model is not less than 5000 samples, which are artificial annotation data from different seasons and different wetland types. The annotation specifications include: high-density food competition scene (bird density > 50 and activity intensity > 0.7), low water temperature and low activity scene (water temperature < 5°C and activity intensity < 0.3), tentative feeding scene (bird quantity fluctuation > 50% and feeding hesitation time > 10 seconds), pollution-sensitive scene (0 < WQI < 50). Its model parameters are obtained through offline training and can be updated through incremental learning regularly.

7. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 5, characterized in that, The lower-level feeding strategy optimization model is constructed using a reinforcement learning algorithm. It takes the output of the scene classification model as input, defining a state space S, an action space A, and a reward function R. The optimal feeding strategy is learned through interaction with the environment. The reinforcement learning algorithm uses a Deep Q-Network (DQN) algorithm. Its network structure includes: one convolutional layer with a kernel size of 3×3, a stride of 1, and 32 filters; one convolutional layer with a kernel size of 3×3, a stride of 2, and 64 filters; one fully connected layer with 256 neurons; one fully connected layer with 128 neurons; and one output layer matching the dimension of the action space. Except for the output layer, all layers use the ReLU activation function. Training is performed using the Adam optimizer with an initial learning rate of 0.001, which decays to 0.95 times every 1000 training steps. The sampling strategy for the experience replay pool is priority sampling based on the absolute value of the temporal difference error.

8. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 1, characterized in that, The reward function R is defined as follows: ; In the formula: , These are the weighting coefficients; For the feeding reward function; This is a water quality penalty function; A function to penalize food waste; ; In the formula, The percentage of birds that eat food within the area; the initial percentage of waterfowl that eat food is set to 0.75 (3 / 4), and the further away from the target, the more negative the reward. ; In the formula, , represents the changes in various water quality indicators before and after feeding; This represents the background water quality values ​​before artificial feeding. Each water quality indicator has a weight value, and the sum of all weight values ​​is 1. ; In the formula, This refers to the amount of food remaining after feeding. This represents the total amount of food fed in this round; The vector format of the action space A is: A=[Flevel, Oaperture, Rspeed, T].

9. The machine learning-driven adaptive feeding device based on bird behavior and water quality coupling as described in claim 1, characterized in that, The two-layer machine learning decision model also includes the definition of foraging-induced pollution thresholds and corresponding feeding inhibition strategies. For example, when multiple indicators trigger the thresholds simultaneously, the system executes the strictest inhibition strategy. When the water quality increments for three consecutive feeding cycles do not trigger any thresholds, the system gradually removes the inhibition and restores the normal feeding strategy. Based on the dynamic coupling method of foraging-induced pollution increment weights, the real-time adjustment value of the water quality penalty weights can be dynamically calculated to obtain the following formula: ; In the formula, The basic water quality penalty weight is determined based on the scenario category. This is an adjustment factor, with a default value of 1.0; The increment of the i-th water quality indicator before and after this round of feeding; Let be the incremental threshold for the i-th water quality indicator; The contribution weight of the i-th water quality indicator in foraging-induced pollution is dynamically determined by the entropy weight method. The steps for calculating the contribution weight wi using the entropy weight method are as follows: Select water quality increment data from the most recent k feeding cycles to form a matrix, standardize the data, calculate the information entropy Ej of each indicator, and then calculate the entropy weight of each indicator. The greater the recent incremental fluctuation of a certain water quality indicator, the greater its entropy weight, and the higher its attention weight in dynamic coupling. Ultimately, the water quality penalty term is revised as follows: .

10. A method for using a machine learning-driven adaptive feeding device coupled with bird behavior and water quality, characterized in that, Includes the machine learning-driven adaptive feeding device for bird behavior coupled with water quality as described in claims 1 to 9, and the following steps: Step 1: System initialization and deployment. Based on the historical distribution of wetland waterbirds, the feeding device module is connected and fixed with ropes using safety buckles. The system is deployed on the shore. The adaptive feeding system decision module is deployed to the bird gathering area. The system is debugged. The sampling frequency of the AI ​​camera (1) is set to 2 times per second. The parameters of the pre-trained scene classification model and reinforcement learning model are loaded. Step 2: Environmental perception and feature extraction. The AI ​​camera (1) collects images of birds in the feeding area in real time, identifies the species, number and activity intensity of birds, and extracts bird feature vectors. The water sample detection system simultaneously collects pollutant parameters such as dissolved oxygen (DO), total dissolved solids (TDS), chemical oxygen demand (COD), total organic carbon (TOC), total nitrogen, and total phosphorus, and constructs a state vector S=[bird species code, population size, activity intensity, water quality vector, time parameter, amount of remaining food]; Step 3: Scene recognition. Input the state vector into the scene classification model and output the current scene category C∈{high-density feeding scene, low water temperature and low activity scene, tentative feeding scene, pollution sensitive scene}. Different scene categories correspond to different exploration rates ε and model parameters to achieve adaptive adjustment. Step 4: Feeding strategy decision-making. Input the state vector S and scene category C into the reinforcement learning model. The model outputs action A = [feeding amount, feeding interval, baffle aperture opening, high-pressure fan power]. The actuator starts the feeder according to the action command, feeding gradually from small to large, while continuously monitoring environmental water quality parameters. Step 5: Real-time feedback and reward calculation; continuously monitor the proportion of birds that eat food in the area during the feeding process. When the number of birds that have eaten reaches 3 / 4 of the total bird population in the area, the total amount of food given at this time is recorded as the candidate value for the optimal feeding amount for that period. Feeding is then stopped. The remaining 1 / 4 of birds that have not eaten continue their normal activities and maintain their natural ecological behavior. After feeding, a second water sample is taken for testing, and a reward value is calculated based on the changes in water quality. ; Step 6: Online model update. Store the experience data of this decision in the experience replay pool. After every N feedings, randomly sample batch data from the experience replay pool and use the experience replay mechanism to update the Q network parameters of the reinforcement learning model. Periodically use the newly accumulated data to perform incremental learning or transfer learning on the scene classification model to improve the model's generalization ability. Step 7: Multi-feeder coordinated control. When multiple feeders are deployed in the system, the controller sequentially checks whether each feeder Fn has completed the current feeding round: if the current feeder has not completed, return to step 2; if the current feeder has completed but there are other feeders that have not completed, n=n+1, proceed to the next feeder and execute steps 2 to 6; if all feeders have completed, the current feeding round ends. Step 8: Experience recording and strategy optimization. Record complete data of each feeding process in local storage, including state sequence, action sequence, reward value sequence and water quality change curve. The system periodically uploads the data to the cloud server for offline strategy evaluation and model optimization. The optimized model parameters are then sent back to the controllers of each device. The format of the state vector: 。