Greenhouse multispectral inspection operation robot

By utilizing the dynamic self-organizing network, distributed situational consensus, and strategy self-evolution in the greenhouse multispectral inspection robot, the problem of insufficient collaborative decision-making in facility agriculture robot systems has been solved, achieving efficient distributed collaborative operation and continuous optimization, and improving the response speed and accuracy to complex agronomic problems.

CN121900483APending Publication Date: 2026-04-21CHONGQING LVZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING LVZHI TECH CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing facility agriculture robot systems, the robots fail to form organic collaboration and cannot achieve distributed perception and negotiation. This results in an inability to make quick and accurate collective decisions on complex agronomic problems, and the system has poor scalability, failing to achieve a fundamental improvement in intelligence.

Method used

This invention provides a greenhouse multispectral inspection robot that forms an organic whole with group perception, collective decision-making, experience sharing, and continuous evolution through a dynamic self-organizing network, distributed situational consensus, and strategy self-evolution among multiple autonomous operation nodes, thereby achieving efficient distributed collaborative operation.

Benefits of technology

It enables spontaneous organization and collaborative operation among multiple agricultural robots, allowing them to continuously accumulate knowledge and optimize strategies, thereby improving the speed and accuracy of response to complex agronomic problems and reducing the system's learning costs and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900483A_ABST
    Figure CN121900483A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral inspection robot for a greenhouse. The multispectral inspection robot comprises a plurality of autonomous nodes with moving, multi-modal sensing and precise operation capabilities. The autonomous nodes can spontaneously establish a temporary communication network according to task requirements to form a dynamic collaborative group; through data exchange and negotiation in the network, a collective consensus is achieved for crop stress states, and a consistent operation instruction is generated; each node records a job effect and calculates an efficiency score, and packages the efficiency score into a sharable strategy packet; and the node continuously optimizes a self decision-making model by using a local and shared strategy packet to realize intelligent evolution. According to the invention, the robot group has the capabilities of group perception, collaborative decision, experience sharing and sustainable evolution, and the greenhouse operation efficiency and the system intelligence level are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart agriculture and agricultural robot technology, specifically relating to a greenhouse multispectral inspection robot. Background Technology

[0002] With the development of precision agriculture and smart agriculture, the use of automated and intelligent equipment to replace traditional manpower in agricultural production management has become an important trend. In facility agriculture scenarios, using robots for tasks such as environmental inspection, pest and disease identification, and water and fertilizer operations can significantly improve operational efficiency, reduce labor intensity, and achieve precise resource allocation.

[0003] Existing technologies already include numerous mobile or track-based robots equipped with various sensors (such as visible light cameras, multispectral cameras, and environmental sensors). These systems typically collect data along pre-defined paths and transmit the data back to a central server or cloud platform for processing and analysis before generating operational instructions. This approach is essentially an architecture of "mobile data acquisition terminal + centralized computing brain." Its main problems are: First, all intelligence is centralized at the backend, with the robot merely acting as an execution terminal, resulting in significant response latency and an inability to respond in real-time to sudden, localized agronomic problems (such as early outbreaks of single-plant diseases); second, the perception dimensions of a single robot are limited, making it prone to misjudgments of complex agronomic conditions due to errors in a single sensor or limitations in perspective; finally, the system has poor scalability, with adding robots only expanding the coverage area without bringing about a fundamental improvement in intelligence.

[0004] In existing facility agriculture robot systems, the robots fail to form an organically collaborative intelligent group, making it impossible to make rapid and accurate collective decisions on complex agronomic problems through distributed perception and negotiation. The operational experience of each robot or each deployment site is isolated and cannot be effectively accumulated, shared, or transferred, resulting in a large amount of repetitive learning and trial-and-error costs. The system's decision-making intelligence is static and preset, lacking the ability to continuously self-optimize and evolve from actual operational feedback. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a greenhouse multispectral inspection robot. Multiple agricultural robots can spontaneously organize into an organic whole with capabilities of "group perception, collective decision-making, experience sharing, and continuous evolution." This not only enables efficient distributed collaborative operations but also allows the entire robot swarm to continuously accumulate knowledge and optimize strategies during operation.

[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a greenhouse multispectral inspection robot, comprising: Multiple autonomous operation nodes, each with the ability to move, multimodal perception, and micro-intervention execution; The dynamic self-organizing network module is used to establish temporary communication links between multiple nodes on demand, forming a dynamic topology network that serves the current intervention task; The distributed situational consensus module runs in a dynamic topology network and enables nodes to reach a consensus on the stress state of the same agronomic object through negotiation based on local perception data, and generate consistent local intervention instructions. The intervention efficacy assessment module, built into each node, is used to record the intervention actions performed by that node and the changes in the object's state before and after them, and to calculate the real-time efficacy score of the intervention based on the changes in the object's state. The strategy self-evolution module is used to encapsulate records containing actions, context, and scores into transferable policy packages, which are shared within the dynamic topology network. Based on the shared policy packages and local experience, the intervention decision model of this node is continuously optimized.

[0007] Preferably, the multiple autonomous operation nodes include at least three types of heterogeneous nodes: disease control nodes, water and fertilizer regulation nodes, and growth monitoring nodes. The three types of nodes form a collaborative sensing network through complementary sensing capabilities. Among them, the disease control node is equipped with a narrow-band multispectral filter group and an ultraviolet excitation light source, the water and fertilizer regulation node integrates a soil moisture sensor and a variable drip irrigation device, and the growth monitoring node is equipped with a chlorophyll fluorescence probe, focusing on disease identification, precise water and fertilizer supply, and growth status assessment, respectively.

[0008] Preferably, the dynamic self-organizing network module adopts the Zigbee 3.0 Mesh network protocol with on-demand wake-up. When a node detects that the crop stress index exceeds a preset threshold, it broadcasts a network request frame containing its own ID, location coordinates and stress type. Temporary communication links are established on demand among multiple autonomous operating nodes to form a dynamic topology network serving the current intervention task. Surrounding nodes decide whether to respond based on their remaining power, their own load and the Euclidean distance from the broadcasting node. The duration of the temporary communication link is positively correlated with the complexity of the intervention task. The link is automatically released after the task is completed, realizing lightweight dynamic reconstruction of the network topology.

[0009] Preferably, the distributed situation consensus module adopts an asynchronous negotiation mechanism based on the Gossip protocol. When multiple nodes simultaneously perceive the same agronomic object, each node first independently calculates the local stress index vector V_local=[H,P,W,F,T,H]ᵀ based on the data it has collected about the agronomic object, where H is the health score, P is the probability of pest and disease risk, W is the water deficit, F is the fertilizer deficit, and T is the temperature deviation. Subsequently, the nodes exchange local stress index vectors V_local in pairs and perform weighted fusion. The weight of each node's local stress index vector V_local is related to the node's sensor noise variance. After multiple iterations, the V_local of each node converges to a globally consistent stress state value V_consensus. Consensus is reached when the absolute value of the difference between V_consensus and any node's V_local is less than a preset first threshold, and all nodes generate the same intervention command based on V_consensus.

[0010] Preferably, the real-time efficacy score in the intervention efficacy assessment module The calculation method is as follows: in and These represent the stress indices of agronomic subjects before and after the intervention. This is the actual response time. For the target response time, The amount of resources consumed in this intervention includes the consumption of medicines, water, and energy. Budgeted resources for the task. , , The weighting coefficients and + + =1.

[0011] Preferably, the strategy self-evolution module encapsulates records containing actions, context, and scores into a transferable strategy package; The portable policy package is a JSON tuple containing action, context, score, and signature. It is published in the dynamic topology network via the MQTT protocol. All subscribing nodes receive it and store it in their local policy library, and old policies are replaced by the LRU algorithm.

[0012] Preferably, the continuous optimization of the intervention decision model adopts a reinforcement learning algorithm based on policy gradient. Each node performs an intervention once, randomly selects policy packages with similar contexts from the policy library for experience replay, and uses the Adam optimizer to update the model parameters, realizing online incremental learning of the intervention decision model. If the average reward improvement of the learned model reaches the preset requirement, the old model is replaced. When a strategy package in the strategy library scores higher than a preset second threshold multiple times in similar contexts, the strategy is marked as a high-quality strategy and synchronized to the same type of nodes in other greenhouses through the cloud platform, realizing cross-domain migration and symbiotic evolution of strategies.

[0013] Preferably, it also includes a cloud management platform, which is used to receive local situational data uploaded by multiple nodes and perform global fusion to generate a cross-greenhouse overall scheduling strategy; The global fusion adopts a weighted average method, and the weight coefficients are inversely proportional to the trace of the node localization covariance matrix; The cloud management platform provides three-level user permission management and supports visualization of 3D greenhouse scenes rendered by WebGL. It can replay the evolution of crop status and robot operation trajectory over the past 7 days through a timeline.

[0014] The beneficial effects of this invention are as follows: 1. Multiple agricultural robots can spontaneously organize into an organic whole with capabilities of "group perception, collective decision-making, experience sharing, and continuous evolution." This not only enables efficient distributed collaborative operations but also allows the entire robot swarm to continuously accumulate knowledge and optimize strategies during operation.

[0015] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0016] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 This is a schematic diagram of the structure of a greenhouse multispectral inspection robot according to an embodiment of the present invention. Detailed Implementation

[0017] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0018] The working principle and process of this invention application are as follows: This invention provides a greenhouse multispectral inspection robot, referring to... Figure 1 ,include: Multiple autonomous operation nodes, each with the ability to move, multimodal perception, and micro-intervention execution; The dynamic self-organizing network module is used to establish temporary communication links between multiple nodes on demand, forming a dynamic topology network that serves the current intervention task; The distributed situational consensus module runs in a dynamic topology network and enables nodes to reach a consensus on the stress state of the same agronomic object through negotiation based on local perception data, and generate consistent local intervention instructions. The intervention efficacy assessment module, built into each node, is used to record the intervention actions performed by that node and the changes in the object's state before and after them, and to calculate the real-time efficacy score of the intervention based on the changes in the object's state. The strategy self-evolution module is used to encapsulate records containing actions, context, and scores into transferable policy packages, which are shared within the dynamic topology network. Based on the shared policy packages and local experience, the intervention decision model of this node is continuously optimized.

[0019] In this embodiment, the autonomous operating node is an intelligent robot entity with autonomous mobility, deployed in a greenhouse or polytunnel environment. Each node physically integrates a mobile chassis, a sensor gimbal, an operating mechanism, a computing unit, an energy module, and a communication module. Its "multimodal perception" capability refers to the node's ability to simultaneously collect various physical, chemical, and biological information about the target object through different types of sensors. For example, it can acquire visual images through a camera, obtain reflectance spectra in specific bands through a multispectral sensor, obtain temperature, humidity, and light intensity through environmental sensors, and obtain three-dimensional spatial information through lidar. This multi-source information fusion perception, compared to a single sensor, can more comprehensively and accurately assess the crop status.

[0020] In this embodiment, "micro-intervention execution capability" refers to the node's ability to perform precise physical or chemical operations on localized, small-scale agronomic problems based on decision-making instructions. Examples include targeted pesticide application, variable-rate drip irrigation, physical straightening, and localized supplemental lighting for individual plants or small areas of crop. This "micro-intervention" differs from traditional large-scale, unified operations, emphasizing precision and minimizing interference.

[0021] In this embodiment, the dynamic self-organizing network module is the core of the system's distributed collaboration. Its "on-demand creation" characteristic means that the network topology is not pre-fixed or persistent, but rather triggered by specific tasks. A node initiates the networking process only when it detects an abnormal event requiring collaborative handling. The resulting "dynamic topology network" is a temporary, task-oriented communication and collaboration group. After the task is completed, the network automatically disbands, and nodes either return to independence or wait for the next networking iteration. This design significantly saves communication resources and node energy consumption.

[0022] In this embodiment, the distributed situational consensus module aims to address the potential inconsistency in the perception of the same agronomic object's state among multiple nodes. Due to differences in node location, sensor accuracy, and observation angle, their "locally perceived data" may vary. This module operates a pre-defined set of negotiation rules (such as the Gossip protocol and voting mechanisms) to enable relevant nodes within the network to exchange data, compare and judge, and through multiple iterations, ultimately reach a consensus result acceptable to all participants regarding the object's health, disease risk, water and fertilizer requirements, and other "stress states." The "local intervention instructions" generated based on this consensus possess consistency and authority, ensuring unified subsequent collaborative actions.

[0023] In this embodiment, the intervention effectiveness evaluation module achieves closed-loop feedback for a single operation. After performing an intervention action (such as spraying pesticides), the node does not immediately end the task but instead re-perceives the same object after a certain time interval, recording the "object status change." For example, if the disease index was 0.8 before spraying and 0.3 two hours later, the index is retested. By comparing the status data before and after the intervention, and combining it with response time, resource consumption, etc., the "real-time effectiveness score of this intervention" can be calculated. This score is a key indicator for quantitatively evaluating whether the operation is effective and efficient.

[0024] In this embodiment, the strategy self-evolution module embodies the system's learning capability. It encapsulates a complete intervention event (including the environmental "context" at the time of triggering, the executed "action," and the resulting effect "score") into a structured "transferable strategy package." This strategy package is a readable and executable "experience case." By "sharing" within the dynamic network, other nodes can learn this experience. Based on the shared strategy package and local experience, nodes use machine learning algorithms (such as reinforcement learning) to adjust and optimize the parameters of their internal "intervention decision model." This allows the node's decision-making ability to continuously improve with the increase in the number of tasks performed, i.e., achieving "self-evolution." The entire system can thus learn from practice and become increasingly intelligent with use.

[0025] Based on the first preferred embodiment, the multiple autonomous operation nodes include at least three types of heterogeneous nodes: disease control nodes, water and fertilizer regulation nodes, and growth monitoring nodes. These three types of nodes form a collaborative sensing network through complementary sensing capabilities. Specifically, the disease control node is equipped with a narrow-band multispectral filter group and an ultraviolet excitation light source to identify early symptoms of diseases; the water and fertilizer regulation node integrates a soil moisture sensor and a variable-rate drip irrigation device to achieve precise water and fertilizer management; and the growth monitoring node is equipped with a chlorophyll fluorescence probe and a 3D scanner to assess plant growth vitality and morphology.

[0026] In this embodiment, "heterogeneous nodes" refer to nodes of different types with specialized designs in hardware configuration and core functions, each with its own strengths. Compared to homogeneous general-purpose nodes, this design can provide more professional and accurate sensing and operational capabilities at the same cost. The three types of nodes constitute a functionally complementary "cooperative sensing network." For example, if a disease node detects abnormal leaves, it can summon a water and fertilizer node to detect soil conditions to rule out physiological diseases, and at the same time summon a growth node to assess the overall health of the plant, thereby making a more comprehensive diagnosis.

[0027] In this embodiment, the "narrow-band multispectral filter group" at the disease control node typically consists of filters with center wavelengths in characteristic bands such as 450nm (blue), 550nm (green), 660nm (red), 730nm (red edge), and 850nm (near-infrared). By calculating the combination of reflectance in these bands (such as Normalized Difference Vegetation Index (NDVI) and Photochemical Reflectance Index (PRI), physiological information such as chlorophyll content and water stress can be effectively retrieved, enabling early and non-destructive detection of diseases. The "ultraviolet excitation light source" is typically a UV-LED array with a wavelength of 365nm. When certain pathogens (such as powdery mildew and downy mildew) are irradiated, they are excited to produce specific autofluorescence. By capturing this fluorescence through matching filters and sensors, highly specific identification of specific diseases can be achieved, effectively distinguishing diseases from interference such as pesticide spots and dust.

[0028] In this embodiment, the "soil moisture sensor" at the water and fertilizer regulation node is a probe that can be inserted into the substrate to measure parameters such as volumetric water content (VWC, unit: %), electrical conductivity (EC, unit: mS / cm, reflecting salt or nutrient concentration), and temperature in the root zone in real time. The "variable drip irrigation device" includes a water pump, solenoid valve, Venturi fertilizer applicator, EC / pH online monitoring instrument, and controller. It can dynamically adjust the irrigation duration, flow rate, and A / B mother liquor absorption ratio in real time based on the VWC and EC values ​​fed back by the sensor, achieving on-demand supply and precise control of the water and fertilizer environment in the root zone.

[0029] In this embodiment, the "chlorophyll fluorescence probe" at the growth monitoring node is an active modulation fluorometer. It emits a beam of measurement light of a specific intensity toward the leaves, and by detecting the fluorescence signal emitted after chlorophyll molecules are excited, it can calculate parameters such as maximum photochemical efficiency (Fv / Fm). This parameter is a sensitive indicator reflecting the health status of the plant's photosynthetic system and whether it is suffering from environmental stress. The "3D scanner" typically uses LiDAR or structured light technology. By acquiring dense point cloud data of the plant canopy, it can reconstruct three-dimensional morphological parameters such as plant height, canopy diameter, leaf area index (LAI), and stem thickness, and quantitatively assess growth and plant structure.

[0030] In a preferred embodiment, the dynamic self-organizing network module adopts the on-demand wake-up Zigbee 3.0 Mesh network protocol. When a node detects that the crop stress index exceeds a preset threshold, it broadcasts a network request frame containing its own ID, location coordinates, and stress type. Temporary communication links are established on demand among multiple autonomous operating nodes to form a dynamic topology network serving the current intervention task. Surrounding nodes decide whether to respond based on their remaining power, their own load, and the Euclidean distance from the broadcasting node. The duration of the temporary communication link is positively correlated with the complexity of the intervention task. The link is automatically released after the task is completed, realizing lightweight dynamic reconfiguration of the network topology.

[0031] In this embodiment, "on-demand wake-up" means that the network communication module is normally in a low-power listening or sleep state, and is only activated to full-power operation when networking or communication tasks are required. This can significantly reduce the standby power consumption of nodes and extend the operation time of a single charge. Zigbee 3.0 is a low-speed, low-power, and highly reliable wireless personal area network protocol based on the IEEE 802.15.4 standard. Its "mesh" network topology allows multi-hop relay communication between nodes and has self-organizing and self-healing characteristics, making it very suitable for building flexible temporary networks in complex greenhouse environments (with obstructions).

[0032] In this embodiment, the "crop stress index" is a comprehensive quantitative indicator. It can be a scalar value or a vector as in the first embodiment. It is the output of the node's internal algorithm after processing and fusing perceived data (such as multispectral indices, fluorescence intensity, temperature, and humidity), used to characterize the severity of crop stress (disease, drought, nutrient deficiency, etc.). The "preset threshold" is preset by the administrator based on crop type, growth stage, and tolerance; for example, the disease risk probability threshold is set to 0.6. When the detected value exceeds this threshold, it indicates that the problem has reached a level requiring attention and action.

[0033] In this embodiment, the "network request frame" is a data packet following a specific format. In addition to the sending node's unique identifier (ID) and current precise location (e.g., obtained via RTK-GPS or UWB positioning), it must also include the "stress type" (e.g., "fungal disease," "water stress"), which helps potential responding nodes determine whether they have the capability to handle such tasks (functionality matching). After broadcasting, all nodes within the communication range will receive the request.

[0034] In this embodiment, the "response decision" of surrounding nodes is a multi-factor trade-off process. The node's built-in decision logic calculates a response priority score. For example: Score_response = α*(E_remain / E_total) + β*(1-CPU_load) - γ*Distance. Where E_remain is the remaining battery power, CPU_load is the current computing load, Distance is the Euclidean distance to the requesting node, and α, β, and γ are weighting coefficients. Nodes with higher remaining battery power, less computing time, and closer proximity receive higher scores. Only nodes with scores exceeding an internal threshold send a response confirmation, agreeing to join the temporary network. This ensures that the network consists of nodes in good condition, improving the task success rate.

[0035] In this embodiment, the "temporary communication link duration T_link" is not a fixed value, but is dynamically set based on an estimate of the "intervention task complexity." Complexity can be estimated based on factors such as the target area size, the number of required node types, and the expected number of operation steps. For example, a simple single-point spraying task might have T_link set to 5 minutes; a complex multi-node collaborative trimming and irrigation task might have T_link set to 30 minutes. After the task is completed, all nodes automatically execute the link disconnection protocol, releasing communication resources, and the network topology disappears, achieving "lightweight dynamic reconfiguration."

[0036] In a preferred embodiment, the distributed situation consensus module adopts an asynchronous negotiation mechanism based on the Gossip protocol. When multiple nodes simultaneously perceive the same agronomic object, each node first independently calculates the local stress index vector V_local=[H,P,W,F,T,H]ᵀ based on the data it has collected about the agronomic object, where H is the health score, P is the probability of pest and disease risk, W is the water deficit, F is the fertilizer deficit, and T is the temperature deviation. Subsequently, the nodes exchange local stress index vectors V_local in pairs and perform weighted fusion. The weight of each node's local stress index vector V_local is related to the node's sensor noise variance. After multiple iterations, the V_local of each node converges to a globally consistent stress state value V_consensus. Consensus is reached when the absolute value of the difference between V_consensus and any node's V_local is less than a preset first threshold, and all nodes generate the same intervention command based on V_consensus.

[0037] In this embodiment, the "Gossip protocol" is a commonly used information propagation and consensus algorithm in distributed systems, simulating the spread of epidemics or rumors in social networks. Its core is that nodes randomly or according to certain rules select other nodes in the network to exchange information. In this system, it functions as an "asynchronous negotiation mechanism," meaning that nodes do not require strict synchronization clocks and can send and receive messages at their own convenient times. It possesses strong fault tolerance and scalability, making it suitable for dynamically changing wireless ad hoc network environments.

[0038] In this embodiment, the "local stress index vector V_local" is a multidimensional feature vector that quantifies various aspects of crop status into a computable and comparable mathematical form. For example: H (health score) can be obtained by normalizing the normalized vegetation index NDVI, ranging from [0,1], with higher values ​​indicating better health; P (probability of disease and pest risk) can be calculated from the output confidence of the disease identification model or a specific spectral index, ranging from [0,1]; W (water deficit) can be derived from the difference between canopy temperature and air temperature (CWSI) or soil moisture content, ranging from [0,1], with higher values ​​indicating greater water deficiency; F (fertilizer deficit) can be derived from leaf nitrogen content or SPAD value, ranging from [0,1]; T (temperature deviation) can be represented by the normalized absolute value of the difference between the current temperature and the optimal growth temperature of the crop, ranging from [0,1].

[0039] In this embodiment, "weighted fusion" is a key step in reaching consensus. Each node, upon receiving V_local data from its neighbors, does not simply average the data; instead, it assigns different weights based on the reliability of the information source. Here, the "noise variance σ_i of the sensor" is used as a reliability metric. The variance σ_i can be obtained by periodically measuring and calibrating a standard reference object; it reflects the stability and accuracy of the sensor's measurements. The smaller the variance, the more reliable the sensor, and the larger its data weight w_i. The formula w_i=(1-σ_i) / ∑(1-σ_j) ensures that the sum of all weights is 1, and that more reliable nodes have greater influence.

[0040] In this embodiment, the consensus process is "multi-round iteration." Assume nodes A, B, and C perceive the same goal. In the first round: A and B exchange vectors, updating their respective vectors according to the weights mentioned above; A and C exchange and update; B and C exchange and update. After one round, each node's V_local incorporates some information from other nodes. Then, the second and third rounds of exchange and updates are performed. As the number of iterations increases, because all nodes are constantly absorbing information from a common "opinion pool" and adjusting themselves, their V_locals gradually converge, mathematically manifested as a continuous decrease in the Euclidean distance between the vectors.

[0041] In this embodiment, "converging to a globally consistent V_consensus" is a dynamic determination process. It can be set that when the changes in V_local of all nodes (such as the sum of the absolute values ​​of changes in each dimension) are less than a minimum value δ in two consecutive iterations, the vector is considered stable, and the weighted average of the current V_local values ​​of all nodes is taken as V_consensus. Then, a final verification is performed: checking whether the absolute value (or Euclidean distance) of the difference between V_consensus and the V_local of each node in the last round is less than a "preset first threshold ε". If both are satisfied, then "consensus is reached". This threshold ε defines the tightness of the consensus, for example, set to 0.05. Based on V_consensus, all nodes independently run the same instruction generation algorithm, naturally producing "the same intervention instructions", such as: "Administer medication to the target, with the dosage coefficient taken as the value of P_consensus".

[0042] In a preferred embodiment, the real-time efficacy score in the intervention efficacy assessment module The calculation method is as follows: in and These represent the stress indices of agronomic subjects before and after the intervention. This is the actual response time. For the target response time, The amount of resources consumed in this intervention includes the consumption of medicines, water, and energy. Budgeted resources for the task. , , The weighting coefficients and + + =1.

[0043] In this embodiment, real-time performance scoring It is a comprehensive performance indicator designed to evaluate the quality of an intervention action from three dimensions: effectiveness, timeliness, and cost. Its value theoretically ranges between (-∞, +∞), with a positive value indicating that the overall performance is better than the benchmark, and a negative value indicating that it is worse than expected.

[0044] In this embodiment, the first item Evaluate the effectiveness of the intervention. and It is usually a comprehensive stress index (such as the main indicator in V_consensus mentioned above, such as P), or a specific indicator for this intervention target (such as the disease risk probability P for disease intervention). This represents the absolute change in the indicator, divided by... Normalization has been performed to represent the relative improvement rate. For example, before spraying... =0.8, after spraying If the stress index is 0.2, then the improvement rate is (0.2-0.8) / 0.8 = -0.75. Since we want the stress index to decrease, a negative improvement rate indicates effectiveness, and the larger the absolute value, the better the effect. In the formula, this term is usually expected to be negative, and the larger the absolute value, the greater the positive contribution to the total score (depending on...). The symbol processing in this formula may involve inverting this term in actual calculations to ensure that "the better the effect, the higher the score".

[0045] In this embodiment, the second item Assess response timeliness. It is an ideal response time preset based on the degree of urgency, for example, for highly contagious diseases. It could be set to 30 minutes; for mild nutritional deficiencies, It can be set to 24 hours. Indicates the time saved (positive if the response is faster). Divide by Normalize the result. This item encourages rapid response. For example, if the target time is 30 minutes and the actual time is 15 minutes, then this item is (30-15) / 30=0.5.

[0046] In this embodiment, the third item Assess resource efficiency. The actual consumption can be quantified as: reagent consumption (ml) * unit cost coefficient + water consumption (L) * water cost coefficient + electricity consumption (Wh) * electricity cost coefficient, ultimately yielding a comprehensive cost value. It is the budgeted cost allowed to complete this type of task. This represents the budget execution rate; the closer it is to 1, the closer the consumption is to the budget, and less than 1 indicates savings. The negative sign before the formula means that the less consumption (the smaller the ratio of this item), the smaller the negative drag on the overall score, i.e., the higher the overall score. This encourages resource conservation.

[0047] In this embodiment, the weighting coefficient , , This allows managers to adjust based on their management strategies. For example, in the early stages of a pest or disease outbreak, more emphasis might be placed on effectiveness and timeliness, and specific settings might be implemented. =0.5, =0.4, =0.1; In routine management, more emphasis may be placed on cost and resource conservation, and settings may be adjusted accordingly. =0.3, =0.3, =0.4. The sum of all weights is 1, ensuring that the total relative importance of the three dimensions remains constant.

[0048] In a preferred embodiment, the policy self-evolution module encapsulates records containing actions, context, and scores into a transferable policy package; The portable policy package is a JSON tuple containing action, context, score, and signature. It is published in the dynamic topology network via the MQTT protocol. All subscribing nodes receive it and store it in their local policy library, and old policies are replaced by the LRU algorithm.

[0049] In this embodiment, the "portable strategy package" serves as the fundamental carrier of system knowledge. Its JSON format ensures the data's structure, readability, and cross-platform compatibility.

[0050] In this embodiment, the "Context" field records in detail the status during policy execution, used for future policy matching. It goes beyond a simple stress index, including crop type, growth stage, and environmental conditions, making the applicability assessment of the policy more accurate. The "Action" field must describe the specific operations and their parameters in detail to ensure reproducibility. The "Score" field stores the calculated performance score. The "Signature" field is generated by the node that generated the policy using its private key to encrypt the core content of the policy package (such as policy ID, context, and action hash), used to verify the integrity and authenticity of the policy package and prevent tampering.

[0051] In this embodiment, MQTT (Message Queuing Telemetry Transport) is a lightweight publish / subscribe messaging protocol, well-suited for asynchronous communication between IoT devices. Nodes "publish" policy packets to agreed-upon topics, and other nodes "subscribed" to those topics automatically receive them. This approach decouples message senders and receivers; the sender doesn't need to know who the receivers are, facilitating information dissemination in dynamic networks.

[0052] In this embodiment, the "local policy repository" is a database or collection of files on each node, used to store policy packages learned from its own experience and shared with other nodes. Due to limited storage space, an eviction mechanism is needed. The "Least Recently Used" (LRU) algorithm is a cache eviction policy that assumes recently used policies are more likely to be used again in the future. In implementation, a last access timestamp can be maintained for each policy package. When the policy repository is full and new policies need to be added, the policy package with the earliest last access time (i.e., the "least recently used") is identified and deleted. This ensures that the policy repository stores relatively active and useful experience.

[0053] In a preferred embodiment, the continuous optimization of the intervention decision model employs a policy gradient-based reinforcement learning algorithm. Each node, after each intervention, randomly selects a policy package from its local policy library that shares a "similar context" with the current situation for experience replay. The Adam optimizer is used to update the parameters θ of the local decision model. The learning rate η adopts a decay strategy, set to η = 0.001 × decay^{episode}, where decay = 0.99 and episode is the total number of training episodes. This enables online incremental learning of the intervention decision model. If the learned model achieves a preset requirement (e.g., an improvement of more than 15%) in the next 5 consecutive decision episodes, the old model is replaced with the new model. When a policy package in the policy library is executed multiple times consecutively under similar context conditions and the score is always greater than a preset second threshold (e.g., 0.85), the policy is marked as a high-quality policy and synchronized to similar nodes in other greenhouses via a cloud platform, achieving cross-domain policy migration and system-level symbiotic evolution.

[0054] In this embodiment, "policy gradient-based reinforcement learning" is a paradigm of machine learning. In this system, a node can be considered an agent, whose observation (State) is the currently perceived "context" (such as a stress vector or environmental data), whose action (Action) is the chosen intervention "action" (such as what pesticide to spray and how much), and whose reward (Reward) is the "efficiency score" obtained after execution. The decision model (such as a neural network) is the policy function π(as), which outputs the probability distribution of action a based on state s. The goal of the policy gradient algorithm is to maximize the expected cumulative reward obtained by the agent after performing a series of actions by directly optimizing the parameters θ of the policy function.

[0055] In this embodiment, "experience replay" is a key technique for stable training in reinforcement learning. Nodes store each complete interaction as an experience tuple in a buffer. During training, instead of directly using newly generated experiences, a batch of historical experiences is randomly drawn from the buffer to update the model. This breaks the temporal correlation between data, improves sample utilization, and makes the learning process more stable. Here, "randomly drawing policy packs with similar contexts from the policy library" is a variant of experience replay that combines case-based reasoning; it directly utilizes encapsulated success / failure cases (policy packs) for learning.

[0056] In this embodiment, the "Adam optimizer" is a stochastic gradient descent optimization algorithm with an adaptive learning rate. It combines the advantages of momentum and adaptive learning rate adjustment, and has fast convergence speed and stable performance in deep learning training, making it very suitable for online learning scenarios.

[0057] In this embodiment, the decay strategy of the learning rate η, η=0.001×0.99^{episode}, means that the initial learning rate is 0.001, and the learning rate is multiplied by 0.99 after each training episode. As training progresses, the learning rate gradually decreases, allowing the model to learn quickly in the early stages and fine-tune in the later stages, which helps to converge to a better solution.

[0058] In this embodiment, "online incremental learning" means that the model is not trained offline after collecting a large amount of data all at once, but rather its parameters are updated on a small scale every time a new experience is generated or a new policy package is received during the continuous operation of the nodes. This allows the model to continuously adapt to environmental changes and new knowledge.

[0059] In this embodiment, the validation condition for model updates and replacements—an average reward increase of more than 15% over 5 episodes—is a rigorous A / B testing approach. The new model must consistently demonstrate significantly better performance than the old model in real-world applications to be adopted, thus avoiding model degradation caused by a single, accidental performance error.

[0060] In this embodiment, "cross-domain transfer of high-quality strategies" is key to the system's knowledge enhancement. When a strategy is proven to be highly effective (consistently high scores) in multiple similar scenarios, it transcends the individual experience of a single node and rises to become a generalizable "best practice." Through synchronization via a "cloud platform," nodes in other similar greenhouses can directly load these high-quality strategies when initializing or encountering new problems, significantly shortening the learning curve and achieving collective intelligent evolution where "learning in one place benefits everywhere."

[0061] In a preferred embodiment, a cloud management platform is also included. This platform receives local situational data (such as V_local, location, and status of each node) uploaded by multiple nodes and performs global fusion to generate a cross-greenhouse overall scheduling strategy. The global fusion employs a weighted average method, where the weight coefficient w_i is inversely proportional to the trace (Σ_i) of the node's localization covariance matrix, i.e., w_i ∝ 1 / Trace (Σ_i). The cloud management platform provides three levels of user access control: administrator, agronomist, and operator, and supports visualization of 3D greenhouse scenes rendered using WebGL technology. Users can use interactive controls on the timeline to replay the evolution of crop status and robot operation trajectories over any past time period (e.g., within 7 days).

[0062] In this embodiment, the "cloud management platform" serves as the system's "brain" and command center, responsible for macro-level monitoring, data analysis, strategy management, and human-computer interaction. It is typically deployed on a remote server or cloud server, communicating with the gateways of each edge greenhouse via a wide area network or directly with the nodes.

[0063] In this embodiment, "global fusion" aims to grasp the operational status of the entire agricultural facility from a higher dimension. The "local situational data" uploaded by each node may have issues such as incomplete spatial coverage, temporal asynchrony, and accuracy differences. Weighted averaging is a common fusion method, and the key lies in determining the weights. Here, the weights are inversely proportional to the "trace(Σ_i) of the node positioning covariance matrix".

[0064] In this embodiment, the "location covariance matrix Σ_i" describes the uncertainty in the node's own position estimation. In SLAM or filtered localization algorithms, in addition to providing the position coordinates, a covariance matrix is ​​also estimated. The smaller the trace of this matrix (i.e., the sum of the variances on the diagonal), the more accurately the node understands its own position, and the higher the reliability of its reported, location-bound sensing data (such as "the 5th plant in area A is diseased"). Therefore, assigning higher weight to its data is reasonable. This reflects the consideration of data source quality during fusion.

[0065] In this embodiment, the "three-level user access control" ensures system security and separation of duties. The administrator has the highest level of access, and can configure the system, manage users, and view all data; the agronomist, as a domain expert, can view all agricultural data, review or adjust agronomic strategies generated by the system, and set various operational thresholds; the operator is mainly responsible for daily monitoring, executing confirmed tasks, and handling simple alarms, and their view and operation functions are limited.

[0066] In this embodiment, "WebGL 3D visualization" utilizes the graphics acceleration capabilities of the browser to render a realistic 3D greenhouse scene in real time on a webpage. Digital crops, robots, and equipment are modeled and placed in virtual space. This provides a more intuitive and comprehensive situational awareness than traditional 2D charts or camera footage.

[0067] In this embodiment, the "timeline playback" function is a powerful tool for data analysis. Users can select a historical time period, and the system will dynamically display, like playing a video, how crop color changes (reflecting health), how diseased areas spread, and how the robot moves and performs its tasks within that time period. This helps in post-event analysis, tracing the root causes of problems, evaluating operational effectiveness, and optimizing management strategies. The playback data is based on previously continuously stored global fusion results and robot logs.

[0068] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A greenhouse multispectral inspection robot, characterized in that, include: Multiple autonomous operation nodes, each with the ability to move, multimodal perception, and micro-intervention execution; The dynamic self-organizing network module is used to establish temporary communication links between multiple nodes on demand, forming a dynamic topology network that serves the current intervention task; The distributed situational consensus module runs in a dynamic topology network and enables nodes to reach a consensus on the stress state of the same agronomic object through negotiation based on local perception data, and generate consistent local intervention instructions. The intervention efficacy assessment module, built into each node, is used to record the intervention actions performed by that node and the changes in the object's state before and after them, and to calculate the real-time efficacy score of the intervention based on the changes in the object's state. The strategy self-evolution module is used to encapsulate records containing actions, context, and scores into transferable policy packages, which are shared within the dynamic topology network. Based on the shared policy packages and local experience, the intervention decision model of this node is continuously optimized.

2. The greenhouse multispectral inspection robot according to claim 1, characterized in that, Multiple autonomous operation nodes include at least three types of heterogeneous nodes: disease control nodes, water and fertilizer regulation nodes, and growth monitoring nodes. These three types of nodes form a collaborative sensing network through complementary sensing capabilities. Among them, the disease control node is equipped with a narrow-band multispectral filter group and an ultraviolet excitation light source, the water and fertilizer regulation node integrates a soil moisture sensor and a variable drip irrigation device, and the growth monitoring node is equipped with a chlorophyll fluorescence probe, focusing on disease identification, precise water and fertilizer supply, and growth status assessment, respectively.

3. The greenhouse multispectral inspection robot according to claim 1, characterized in that, The dynamic self-organizing network module adopts the Zigbee 3.0 Mesh network protocol with on-demand wake-up. When a node detects that the crop stress index exceeds a preset threshold, it broadcasts a network request frame containing its own ID, location coordinates and stress type. Temporary communication links are established on demand among multiple autonomous operating nodes to form a dynamic topology network serving the current intervention task. Surrounding nodes decide whether to respond based on their remaining power, their own load and the Euclidean distance from the broadcasting node. The duration of the temporary communication link is positively correlated with the complexity of the intervention task. The link is automatically released after the task is completed, realizing lightweight dynamic reconfiguration of the network topology.

4. The greenhouse multispectral inspection robot according to claim 1, characterized in that, The distributed situation consensus module adopts an asynchronous negotiation mechanism based on the Gossip protocol. When multiple nodes simultaneously perceive the same agronomic object, each node first independently calculates the local stress index vector V_local=[H,P,W,F,T,H]ᵀ based on the data it has collected about the agronomic object, where H is the health score, P is the probability of pest and disease risk, W is the water deficit, F is the fertilizer deficit, and T is the temperature deviation. Subsequently, the nodes exchange local stress index vectors V_local in pairs and perform weighted fusion. The weight of each node's local stress index vector V_local is related to the node's sensor noise variance. After multiple iterations, the V_local of each node converges to a globally consistent stress state value V_consensus. Consensus is reached when the absolute value of the difference between V_consensus and any node's V_local is less than a preset first threshold, and all nodes generate the same intervention command based on V_consensus.

5. The greenhouse multispectral inspection robot according to claim 1, characterized in that, Real-time efficacy score in the intervention efficacy assessment module The calculation method is as follows:

6. Among them and These represent the stress indices of agronomic subjects before and after the intervention. This is the actual response time. For the target response time, The amount of resources consumed in this intervention includes the consumption of medicines, water, and energy. Budgeted resources for the task. , , The weighting coefficients and + + =1.

7. The greenhouse multispectral inspection robot according to claim 1, characterized in that, The strategy self-evolution module encapsulates records containing actions, context, and scores into a transferable strategy package; The portable policy package is a JSON tuple containing action, context, score, and signature. It is published in the dynamic topology network via the MQTT protocol. All subscribing nodes receive it and store it in their local policy library, and old policies are replaced by the LRU algorithm.

8. A greenhouse multispectral inspection robot according to claim 5, characterized in that, The continuous optimization of the intervention decision model adopts a policy gradient-based reinforcement learning algorithm. Each node performs an intervention once, randomly selects policy packages with similar contexts from the policy library for experience replay, and uses the Adam optimizer to update the model parameters, realizing online incremental learning of the intervention decision model. If the average reward improvement of the learned model reaches the preset requirement, the old model is replaced. When a strategy package in the strategy library scores higher than a preset second threshold multiple times in similar contexts, the strategy package is marked as a high-quality strategy and synchronized to the same type of nodes in other greenhouses through the cloud platform, realizing cross-domain migration and symbiotic evolution of strategies.

9. A greenhouse multispectral inspection robot according to claim 1, characterized in that, It also includes a cloud management platform, which receives local situational data uploaded by multiple nodes and performs global fusion to generate cross-greenhouse overall scheduling strategies; The global fusion adopts a weighted average method, and the weight coefficients are inversely proportional to the trace of the node localization covariance matrix; The cloud management platform provides three-level user permission management and supports visualization of 3D greenhouse scenes rendered by WebGL. It can replay the evolution of crop status and robot operation trajectory over the past 7 days through a timeline.