Weeding robot operation fine control system based on reinforcement learning

By constructing a unified state modeling mechanism driven by multimodal perception and an Actor-Critic reinforcement learning model, combined with the conditional risk value method, the problem of multi-source information modeling and tail risk perception in complex environments of existing intelligent weeding technologies is solved, and the efficient, stable and safe fine control of the weeding robot is realized.

CN122195018BActive Publication Date: 2026-07-24AUBO (BEIJING) ROBOTICS TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AUBO (BEIJING) ROBOTICS TECH CO LTD
Filing Date
2026-05-18
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing intelligent weeding technologies suffer from insufficient multi-source information modeling capabilities in complex environments, inadequate reward distribution characterization, lack of tail risk perception capabilities, and low efficiency in multi-device collaborative optimization, resulting in insufficient stability and safety in complex environments.

Method used

A unified state modeling mechanism driven by multimodal perception is constructed. An Actor-Critic reinforcement learning model based on parameterized quantile representation is introduced. The reward distribution is explicitly characterized by combining the conditional value at risk method. The distributed optimization of risk structure consistency is achieved through the reward distribution alignment mechanism of Wasserstein distance and the cross-node federated commentator collaborative update mechanism.

Benefits of technology

It enhances the risk perception and decision-making safety of weeding robots in complex environments, improves the convergence efficiency and robustness of the strategy optimization process, and enhances the generalization ability and collaborative learning efficiency of the model under multiple field and environmental conditions, thus ensuring the refined and stable control of weeding operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195018B_ABST
    Figure CN122195018B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of decision control, and discloses a weeding robot operation fine control system based on reinforcement learning. The system performs space-time alignment and structured processing on images, point clouds and multi-source sensing data through multi-modal environment sensing and unified state modeling to construct an operation state vector. On this basis, an Actor-Critic model based on parameterized quantile representation is introduced to perform distributed modeling on returns, and combined with conditional risk value, the crop injury, border crossing and collision risks are constrained. At the same time, through cognitive uncertainty guidance strategy exploration, combined with return distribution alignment and federal collaborative updating mechanism, stable optimization and generalization improvement in a multi-device environment are realized, so that a joint control strategy considering operation efficiency and risk control is generated, and fine and stable control of weeding operation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of decision control technology, and discloses a fine control system for weeding robots based on reinforcement learning. Background Technology

[0002] With the continuous improvement of the intelligence level of agricultural equipment, weeding robot systems based on multimodal perception and automatic control are gradually being applied to precision field operations. Existing intelligent weeding systems typically identify and locate crops and weeds by fusing image, point cloud, and multi-sensor data, and generate operation instructions by combining path planning or rule-based control strategies, thus achieving automation of the weeding process to a certain extent. However, existing technologies mostly rely on static rules or single-step decision-making mechanisms, which are insufficient for dynamic evolution modeling of multi-source heterogeneous information in complex field environments, making it difficult to form a stable and high-precision unified state representation. At the same time, some schemes that introduce reinforcement learning still focus on expected reward optimization, lacking the ability to characterize the reward distribution structure, and are difficult to effectively represent the coupling relationship between multiple objectives such as crop damage, missed removal, repetitive operations, and energy consumption. In terms of risk control, existing methods lack explicit modeling of tail risks, making it difficult to constrain extreme adverse operation conditions, resulting in insufficient stability and security of strategies in complex scenarios. In addition, in multi-device collaborative scenarios, centralized or independent training methods are generally adopted, lacking efficient federated collaborative optimization mechanisms, making it difficult to achieve knowledge sharing and generalization improvement across field environments. Therefore, existing intelligent weeding technologies still have shortcomings in terms of adaptability to complex environments, risk perception capabilities, and distributed collaborative optimization capabilities, and need further improvement. Summary of the Invention

[0003] This invention addresses the shortcomings of existing intelligent weeding technologies, such as insufficient multi-source information modeling capabilities, inadequate characterization of reward distribution, lack of tail risk perception, and low efficiency of multi-device collaborative optimization. It proposes a fine-grained control system for weeding robots based on reinforcement learning. This system constructs a unified state modeling mechanism driven by multimodal perception, performing time synchronization, spatial alignment, and structured representation processing on image data, point cloud data, and multi-source sensor data to form an operational state vector that continuously characterizes the evolution of the field operation environment. Based on this, an Actor-Critic reinforcement learning model based on parameterized quantile representation is introduced to perform distributed modeling of future cumulative rewards. Furthermore, the conditional value-at-risk method is combined to explicitly characterize the lower tail of the reward distribution, thereby enabling the assessment of risks such as crop damage, boundary violations, and collisions. The system employs quantitative constraints for extreme adverse scenarios. Simultaneously, a cognitive uncertainty estimation mechanism is constructed based on the dispersion of the reward distribution to guide the strategy to perform adaptive perturbation exploration in high-uncertainty regions, thereby improving strategy search efficiency and robustness. Furthermore, by introducing a reward distribution alignment mechanism based on Wasserstein distance and a cross-node federated commentator collaborative update mechanism, distributed optimization with consistent risk structure is achieved in a multi-device environment. Asymmetric trust domain constraints are combined to perform shrinkage and compression processing on the aggregation results to suppress distribution drift and improve model stability, thus enhancing the model's cross-scenario generalization ability while ensuring local environmental adaptability. Based on these mechanisms, the system can generate joint control strategies that balance operational efficiency and risk control in complex field environments, achieving refined and stable control of the weeding process.

[0004] This invention provides a fine control system for a weeding robot based on reinforcement learning. The system includes: a mobile operation platform, a multimodal environment perception module, a state modeling and encoding module, a reinforcement learning decision-making module, and an execution control module.

[0005] The mobile operation platform, installed on the chassis of the weeding robot, includes a wheeled mobile chassis, a steering drive mechanism, an on-board power supply unit, an inertial measurement unit, a wheel speed encoder, a satellite positioning unit, an on-board computing controller, and a communication module. It is used to perform autonomous movement, positioning, speed adjustment, heading correction, and operation path tracking within the target field.

[0006] A multimodal environment perception module is used to construct precise field operation status data.

[0007] The state modeling and encoding module, deployed in the onboard computing controller, is used to acquire the robot's own pose and velocity information, the working status of the actuators, and receive the field precision operation status data output by the multimodal environment perception module. It performs unified time axis alignment, spatial coordinate mapping, and structured encoding processing on the aforementioned multi-source heterogeneous data to construct a unified state space representation containing environmental semantic information, geometric structure information, and operation constraint information, and generates the corresponding operation state vector. Based on the current operation state vector, the action space is defined as a joint action space composed of chassis motion control actions and weeding execution actions.

[0008] A reinforcement learning decision-making module is deployed in the onboard computing controller of each weeding robot client to construct a risk-aware distributed federated Actor-Critic model. It receives the job state vector output by the state modeling and encoding module and, based on the risk-aware distributed federated Actor-Critic model, sequentially executes uncertainty-guided strategy generation, distributed value assessment, tail risk modulation, local critic update, federated critic collaborative aggregation, and constrained backfeeding of aggregation parameters within the current control cycle. This outputs a target joint control strategy adapted to the heterogeneous environment of the current field. The risk-aware distributed federated Actor-Critic model includes a local policy network (Actor), a parametric quantile critic network (Critic) federated collaborative training module, and a client trust domain constraint module.

[0009] The execution control module, deployed in the chassis control unit and the work execution mechanism of the weeding robot, receives the target joint control strategy and performs instruction parsing, constraint mapping, timing scheduling and closed-loop control processing of the target joint control strategy to generate chassis motion control instructions and weeding execution instructions, thereby driving the chassis control unit and the work execution mechanism to work together to complete the fine weeding operation.

[0010] Furthermore, the output process of the target joint control strategy specifically includes the following steps:

[0011] Step S1: Decision Input Construction: Perform splicing, normalization mapping, and structured encoding on the environmental semantic features, geometric structure features, target weed features, robot motion state features, actuator state features, and safety constraint features in the current operation state vector to construct a decision input representation corresponding to the current control cycle; construct a historical reward distribution cache;

[0012] Step S2: Uncertainty-Guided Strategy Generation: The decision input representation is input into the local policy network Actor, forward policy reasoning is performed, and an initial joint control strategy consisting of chassis motion control actions and weeding execution actions is output. The chassis motion control actions include chassis linear velocity control, chassis angular velocity control, and lateral correction. The weeding execution actions include actuator lifting displacement, nozzle opening mode, spray flow rate, spray duration, cutter head speed, cutter depth, and emergency braking action. Cognitive uncertainty estimation parameters matching the corresponding state-action pair are read from the historical reward distribution cache and used as the basis for initializing policy particles within the current control cycle. An exploration mechanism based on cognitive uncertainty is introduced to perform adaptive perturbation processing on the initial joint control strategy to obtain candidate control actions at the current moment.

[0013] Step S3: Distributed Value Assessment: Combine the current job state vector with the candidate control action at the current moment to form a state-action pair, and input it into the parameterized quantile critic network Critic; the parameterized quantile critic network Critic outputs a distributed quantile representation of the future cumulative return for the state-action pair; calculate the cognitive uncertainty estimation parameters based on the distributed quantile representation, the cognitive uncertainty estimation parameters are used to characterize the quantile dispersion of the corresponding state-action pair, and the cognitive uncertainty estimation parameters are written into the historical return distribution cache;

[0014] Step S4: Tail Risk Modulation: Extract the lower tail distribution based on distributed quantile representation, determine the tail interval range according to the preset tail truncation ratio, and perform weighted processing on each tail quantile in the lower tail distribution according to the position distribution of each tail quantile within the tail interval; perform lower conditional risk CVaR value calculation on the weighted lower tail distribution to obtain the tail risk measurement result corresponding to the current state-action pair; the tail risk measurement result is used to characterize the potential risk level of crop damage, omission expansion, repeated application, dangerous collision, boundary crossing and abnormal energy consumption caused by the current control action under the most unfavorable operation situation; perform expectation calculation on each quantile based on distributed quantile representation to obtain the distribution mean return; jointly map the tail risk measurement result and the distribution mean return to construct the risk-sensitive value characterization result;

[0015] Step S5: Local Commentator Update: Construct a distributed Bellman target distribution based on the current state transition sample. Calculate the quantile regression error between the distributed Bellman target distribution and the distributed quantile representation, and define the calculation result as the quantile regression error term. Write the distributed quantile representation corresponding to the current round into the historical return distribution cache. Calculate the lower conditional value at risk (CVaR) statistic for each historical return distribution in the historical return distribution cache, and generate risk weights based on the CVaR statistic corresponding to each historical return distribution. Based on the risk weights, perform Wasserstein-1 centroid solving on multiple historical return distributions to construct a CVaR-weighted Wasserstein-1 centroid distribution, which serves as the risk perception reference distribution during the commentator network update process in the current round. Calculate the distribution distance between the current return distribution represented by the distributed quantile representation and the risk perception reference distribution to form a return distribution consistency constraint term. By jointly minimizing the quantile regression error term and the return distribution consistency constraint term, perform parameter updates on the parameterized quantile commentator network to obtain the local updated commentator parameters corresponding to the current weeding robot client.

[0016] Step S6: Federated Commentator Collaborative Aggregation: The federated collaborative training module performs cross-node collaborative processing on the locally updated commentator parameters of each weeding robot client; the federated collaborative training module is implemented through hierarchical collaboration between the field edge collaborative station and the central federated management server; among them, the field edge collaborative station is responsible for parameter aggregation and forwarding, and the central federated management server is responsible for weighted aggregation to generate global aggregated commentator parameters, and broadcasts them back to each weeding robot client through the field edge collaborative station;

[0017] Step S7: Restricted Feedback of Aggregate Parameters: The client trust domain constraint module receives the global aggregate commentator parameters broadcast back, and performs trust domain constraint update on the output distribution corresponding to the global aggregate commentator parameters with the risk perception reference distribution as the anchor point, generating restricted aggregate commentator parameters applicable to the current client;

[0018] Step S8: Strategy Optimization Output: Based on the risk-sensitive value representation results, the mean advantage term and the tail risk advantage term are calculated respectively, and weighted fusion is performed according to the preset risk weights to form a risk-enhanced advantage estimate; then the risk-enhanced advantage estimate is input into the strategy gradient update process to optimize the parameters of the local strategy network Actor, so that the local strategy network Actor gradually improves the effectiveness of weed removal, the reliability of crop protection, the boundary compliance ability, and the energy consumption control ability in continuous training iterations; after completing the Actor update for the current control cycle, the updated local strategy network Actor, combined with the constrained aggregated critic parameters, performs strategy optimization on the initial joint control strategy and re-outputs the target joint control strategy for the next control time.

[0019] Furthermore, an exploration mechanism based on cognitive uncertainty is introduced to perform adaptive perturbation processing on the initial joint control strategy to obtain the candidate control action at the current moment. The process specifically includes the following steps:

[0020] Step B1: Based on the initial joint control strategy, construct a set of strategy particles corresponding to the current control cycle, and use the initial joint control strategy as the initial strategy center of the particle swarm; configure the cumulative benefit recording parameters and cognitive uncertainty estimation parameters for each strategy particle, and construct a state hierarchy tree with the current operation state as the root node; the state hierarchy tree is used to record the state transition path of each strategy particle in the rolling exploration process and its parent-child dependency relationship, thereby forming a backtrackable multi-path search structure;

[0021] Step B2: Perform rolling exploration on each strategy particle in the strategy particle set. Under the joint modulation of the perturbation action distribution and the cognitive uncertainty estimation parameters, each strategy particle performs multi-step evolution along the joint action space. During the rolling cycle, joint action sampling is performed with the current state as input to generate joint control actions of chassis motion control actions and weeding execution actions. Based on the joint control actions, the state transition results are obtained, and the cognitive uncertainty estimation parameters are dynamically updated. When the cumulative gain of any strategy particle increases, the current rolling process of that strategy particle is immediately terminated, and its cumulative gain record parameters and corresponding state nodes are retained to obtain the set of surviving strategy particles and the set of failed strategy particles.

[0022] Step B3: Sort and filter the cumulative returns of each strategy particle in the surviving strategy particle set to determine the candidate particle corresponding to the maximum cumulative return; when there are multiple candidate particles, perform a secondary sorting based on the quantitative index corresponding to the cognitive uncertainty estimation parameter, and determine the strategy particle with the largest quantitative index as the winner particle according to the lexicographical order sorting rule of "cumulative return first, uncertainty second best"; based on the winner particle, perform state reset on each failed strategy particle in the failed strategy particle set, so that each failed strategy particle is repositioned to the state node corresponding to the winner particle and inherits its state context and search frontier, and the remaining surviving strategy particles are retained to continue to perform rolling exploration along their respective paths, thereby realizing the centralized allocation of search resources to high-value and high-uncertainty areas, and obtaining the allocated strategy particle set;

[0023] Step B4: Perform parent node backtracking along the state phylogenetic tree for the allocated policy particle set. Select the corresponding ancestor state node for each policy particle within the preset ancestor generation range, and roll back each policy particle to the corresponding ancestor state node and restart joint action sampling to obtain the policy particle set after rollback and recovery. The ancestor generation range is used to control the backtracking depth so as to avoid getting trapped in local optimum or unrecoverable state while maintaining the continuity of local search.

[0024] Step B5: After completing the preset exploration rounds, perform joint sorting on the cumulative gains and cognitive uncertainty estimation parameters of each strategy particle in the rolled-back and restored strategy particle set to determine the group winner particle; and uniformly reset all strategy particles in the rolled-back and restored strategy particle set to the state node corresponding to the group winner particle to achieve group convergence; extract the joint action parameters from the state node corresponding to the group winner particle, and use the joint action parameters as the candidate control action output at the current moment.

[0025] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:

[0026] This invention constructs a unified state modeling mechanism driven by multimodal perception, which fuses and expresses image data, point cloud data, and multi-source sensor data under a unified time axis and spatial coordinate system. This achieves high-precision structured representation of complex field environment information, improves the perception accuracy and environmental understanding ability of weeding robots regarding crop row structure, weed distribution, and operational constraints, and solves the problems of information fragmentation and incomplete state expression in the process of multi-source heterogeneous data fusion in existing intelligent systems. Thus, it provides a stable, continuous, and highly reliable decision input foundation for the subsequent generation of refined control strategies.

[0027] This invention introduces an Actor-Critic reinforcement learning model based on parameterized quantile representation and combines it with conditional value at risk to model the lower tail of the reward distribution. This enables explicit quantitative constraints on extreme adverse situations such as crop damage, out-of-bounds operation, collision risk, and abnormal energy consumption, improving the risk perception capability and decision-making safety of weeding robots in complex field environments. It solves the problem that existing methods rely solely on expected reward optimization, making it difficult to balance efficiency and safety. At the same time, through a policy perturbation mechanism driven by cognitive uncertainty, it achieves adaptive exploration of high uncertainty regions, enhancing the convergence efficiency and robustness of the policy optimization process, enabling the system to stably output high-quality control strategies under different operating scenarios.

[0028] This invention achieves distributed optimization of risk structure consistency among multiple devices by constructing a reward distribution alignment mechanism based on Wasserstein distance and a cross-node federated reviewer collaborative update mechanism. This improves the model's generalization ability and collaborative learning efficiency under multiple field and environmental conditions. Furthermore, by introducing asymmetric trust domain constraints, it effectively suppresses the offset and distortion of reward distribution during federated aggregation, enhancing the model's stability and reliability during continuous iteration. This enables the weeding robot to continuously generate joint control strategies that balance operational efficiency and risk control in complex dynamic environments, significantly improving the overall quality of precision weeding operations and system performance. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of a module of a fine control system for a weeding robot based on reinforcement learning proposed in this invention.

[0030] Figure 2 This is a schematic diagram illustrating the relationship between the current return distribution, historical return distribution, and risk perception reference distribution proposed in Example 3. Detailed Implementation

[0031] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0032] Example 1, according to Figure 1 This invention provides a fine control system for weeding robots based on reinforcement learning, which can be deployed in the field autonomous operation environment of farmland, orchard, tea garden, breeding experimental field, facility agriculture planting belt or inter-row passage. The system includes: a mobile operation platform, a multimodal environment perception module, a state modeling and coding module, a reinforcement learning decision module and an execution control module.

[0033] This embodiment is deployed in the inter-row passages of an open-field vegetable planting plot to perform fine weeding operations. The plot area is 120m × 48m, with 16 planting ridges. The center-to-center distance between adjacent planting ridges is 0.75m, and the average passage width between ridges is 0.42m. The crop is lettuce seedlings, with an average plant spacing of 0.18m and a seedling height of 9cm to 14cm. The plot contains areas with mixed growth of gramineous and broadleaf weeds, and there are complex operating conditions such as local stones, exposed drip irrigation tape, turned-up mulch film, and uneven furrow and ridge edges. The system operates in an ambient temperature of 22.6℃, relative humidity of 61%, wind speed of 1.8m / s, and soil volumetric moisture content of 23.4%.

[0034] The mobile operation platform, installed on the chassis of the weeding robot, includes a wheeled mobile chassis, a steering drive mechanism, an on-board power supply unit, an inertial measurement unit, a wheel speed encoder, a satellite positioning unit, an on-board computing controller, and a communication module. It is used to perform autonomous movement, positioning, speed adjustment, heading correction, and operation path tracking within the target field.

[0035] The multimodal environmental perception module, deployed on the front mast, top of the robot, and near the work execution end, connects to an RGB camera, depth camera, LiDAR, near-ground ranging sensor, soil moisture sensor, crop row detection sensor, non-contact plant spacing detector, and wind speed, temperature, and humidity sensors. It collects forward-facing and top-facing field-of-view images, depth point cloud data, laser scan contour data, and near-ground ranging information, and performs time synchronization, extrinsic parameter calibration, and coordinate alignment processing on the multi-source perception data. Based on crop row geometry, plant spacing patterns, leaf morphology and texture features, and ground elevation differences, it identifies crop plants, target weeds, non-target weeds, bare soil areas, stones, mulch film boundaries, and drip irrigation tape in the field. The system segments, detects, and identifies edges of furrows and ditches, obstacles, high-risk restricted areas, personnel, and livestock. The identification results are projected onto the robot coordinate system and the field map coordinate system to generate a set of operational objects containing the center location of weed targets, object category identifiers, spatial size parameters, occlusion relationships, crop protection radius, executable windows, non-executable areas, and passage costs. Based on multi-source sensing data processed with time synchronization and extrinsic parameter calibration, the set of operational objects is structured and uniformly represented to construct detailed field operation status data. This detailed field operation status data includes forward field-of-view images, top-of-view images, depth point cloud data, laser scan contour data, near-ground ranging data, and information on the set of operational objects.

[0036] The state modeling and coding module, deployed in the onboard computing controller, is used to acquire the robot's own pose and velocity information, the working status of the actuators, and receive the field precision operation status data output by the multimodal environment perception module. It performs unified time axis alignment, spatial coordinate mapping, and structured coding processing on the aforementioned multi-source heterogeneous data to construct a unified state space representation containing environmental semantic information, geometric structure information, and operation constraint information, and generates corresponding operation state vectors. The operation state vectors include: the robot's current position, attitude angle, velocity, angular velocity, remaining battery power, current operating mode of the actuators, crop row center deviation, target weed location, target weed type, target weed size, target weed density, distance to neighboring crops, distance to obstacles, passage cost, surface slope, surface adhesion conditions, historical missed removal markers, and historical repeated operation markers. Based on the current operation state vector, the action space is defined as a joint action space composed of chassis motion control actions and weeding execution actions.

[0037] A reinforcement learning decision-making module, deployed in the onboard computing controller of each weeding robot client, is based on the Actor-Critic model. It introduces a distributed quantile value representation mechanism, a tail risk modeling mechanism based on conditional risk value, a policy perturbation mechanism driven by cognitive uncertainty, a reward distribution alignment mechanism based on Wasserstein-1 centroid, and a cross-node federated critic collaborative update mechanism to jointly optimize the value representation method, risk modeling capability, and cross-environment generalization capability of the Actor-Critic model, thereby constructing a risk-aware distributed federated Actor-Critic model. The module receives the job state vector output by the state modeling encoding module and, based on the risk-aware distributed federated Actor-Critic model, sequentially executes uncertainty-guided policy generation, distributed value assessment, tail risk modulation, local critic update, federated critic collaborative aggregation, and constrained backfeeding of aggregation parameters within the current control cycle, thereby outputting a target joint control strategy adapted to the heterogeneous environment of the current field. The risk-aware distributed federated Actor-Critic model includes a local policy network (Actor), a parametric quantile critic network (Critic) federated collaborative training module, and a client trust domain constraint module.

[0038] In this embodiment, the output target joint control strategy is as follows: chassis linear velocity control amount: 0.41m / s; chassis angular velocity control amount: 0.06rad / s; lateral correction amount: -0.021m; lifting displacement amount: 0.030m; nozzle opening mode: right nozzle; spraying flow rate: 14ml / min; spraying duration: 0.22s; cutter head rotation speed: 840rpm; cutter penetration depth: 9mm; emergency braking action: 0.

[0039] The execution control module, deployed in the chassis control unit and work execution mechanism of the weeding robot, receives the target joint control strategy and performs instruction parsing, constraint mapping, timing scheduling, and closed-loop control processing on the target joint control strategy to generate chassis motion control commands and weeding execution commands, thereby driving the chassis control unit and work execution mechanism to collaboratively complete the precision weeding operation. The chassis control unit, installed on the chassis of the weeding robot, connects the wheeled mobile chassis, steering drive mechanism, drive motor, braking mechanism, inertial measurement unit, wheel speed encoder, and positioning unit. It receives chassis motion control commands and performs closed-loop control processing on chassis linear velocity control, chassis angular velocity control, lateral correction, and emergency braking actions, driving the weeding robot to complete autonomous movement, speed adjustment, heading correction, path tracking, crop row alignment, and obstacle avoidance. It also combines inertial measurement information, wheel speed feedback information, and positioning information to make real-time corrections to the chassis motion state, improving the robot's stability and road handling in complex field environments. The robot prioritizes path tracking accuracy and operational safety. Its operational execution mechanism, mounted at the rear of the chassis, includes a mechanical cutting component, a rotary hoe component, a directional spraying component, a solenoid valve assembly, a nozzle array, a lifting adjustment component, a lateral movement mechanism, and a posture fine-tuning component. This mechanism receives weeding commands and performs cutting, uprooting, crushing, targeted spraying, or combined weed suppression treatments on target weeds. The mechanical cutting and rotary hoe components perform physical weeding operations, the directional spraying component uses the solenoid valve assembly and nozzle array for precise pesticide spraying control, the lifting adjustment component adjusts the execution depth, the lateral movement mechanism compensates for lateral position, and the posture fine-tuning component corrects the attitude at the execution end. Driven by control commands, each execution unit adjusts the execution depth, lateral offset, spraying flow rate, spraying duration, blade rotation speed, and execution window in real time. Combined with the chassis's movement, this achieves dynamic matching and precise control during operation, thereby improving weed removal efficiency and reducing repetitive work and energy consumption while ensuring crop protection.

[0040] In this embodiment, after receiving the above-mentioned target joint control strategy, the execution control module parses its execution instructions, maps constraints, schedules timing, and processes closed-loop control to form chassis motion control instructions and weeding execution instructions; the chassis control unit executes linear velocity control, angular velocity control, and lateral correction control at a frequency of 50Hz; the work execution mechanism executes lifting, lateral movement, spraying, and cutter head attitude adjustment control at a frequency of 20Hz.

[0041] During the subsequent 8.6 seconds of operation following this control cycle, the robot processed 14 target weeds within the local sensing window, including: mechanical cutting of 6 weeds; directional spraying of 5 weeds; and combined weed suppression of 3 weeds.

[0042] The actual operation results are as follows: target weed removal success rate: 92.86%; crop damage rate: 3.57%; missed removal rate: 7.14%; repetitive operation rate: 4.29%; number of local boundary crossings: 0; number of obstacle collisions: 0; energy consumption per unit area: 0.117 kWh / 100 m².

[0043] Example 2 differs from Example 1 in that: the reinforcement learning decision module constructs a risk-aware distributed federated Actor-Critic model; it receives the job state vector output by the state modeling and encoding module, and outputs a target joint control strategy based on the risk-aware distributed federated Actor-Critic model; the reinforcement learning decision module differs from Example 1 in that it specifically includes the following: constructing an Actor-Critic model; receiving the job state vector output by the state modeling and encoding module, and based on the Actor-Critic model, inputting the job state vector into the Actor network within the current control cycle to generate a joint control strategy consisting of chassis motion control actions and weeding execution actions, while simultaneously inputting it into the Critic network to perform state value evaluation on the joint control strategy, and iteratively updating the parameters of the Actor network and Critic network based on the value evaluation results, thereby outputting a target joint control strategy adapted to the current heterogeneous field environment.

[0044] Example 3, according to Figure 2 This embodiment is based on Embodiment 1. In this embodiment, the output process of the target joint control strategy specifically includes the following steps:

[0045] Step S1: Decision Input Construction: Perform splicing, normalization mapping, and structured encoding on the environmental semantic features, geometric structure features, target weed features, robot motion state features, actuator state features, and safety constraint features in the current operation state vector to construct a decision input representation corresponding to the current control cycle; construct a historical reward distribution cache;

[0046] Step S2: Uncertainty-Guided Strategy Generation: The decision input representation is input into the local policy network Actor, forward policy reasoning is performed, and an initial joint control strategy consisting of chassis motion control actions and weeding execution actions is output. The chassis motion control actions include chassis linear velocity control, chassis angular velocity control, and lateral correction. The weeding execution actions include actuator lifting displacement, nozzle opening mode, spray flow rate, spray duration, cutter head speed, cutter depth, and emergency braking action. Cognitive uncertainty estimation parameters matching the corresponding state-action pair are read from the historical reward distribution cache and used as the basis for initializing policy particles within the current control cycle. An exploration mechanism based on cognitive uncertainty is introduced to perform adaptive perturbation processing on the initial joint control strategy to obtain candidate control actions at the current moment.

[0047] Step S3: Distributed Value Assessment: Combine the current job state vector with the candidate control action at the current moment to form a state-action pair, and input it into the parameterized quantile critic network (Critic). The Critic network outputs a distributed quantile representation of the future cumulative returns for the state-action pair to characterize the long-term return distribution structure of the current control action over multiple job cycles. The return distribution structure reflects the combined impact of weed removal returns, crop damage costs, omission costs, duplicate operation costs, boundary crossing costs, collision costs, and energy consumption costs. Based on the distributed quantile representation, cognitive uncertainty estimation parameters are calculated. These parameters characterize the quantile dispersion of the corresponding state-action pair and are written into the historical return distribution cache.

[0048] In this embodiment, the cognitive uncertainty estimation parameter is calculated based on the distributed quantile representation output by the parameterized quantile commentator network. Specifically, for the current state-action pair, the parameterized quantile commentator network outputs predicted reward values ​​corresponding to multiple quantile levels. These multiple predicted reward values ​​constitute a distributed quantile representation, which serves as a discrete representation of the reward distribution. Based on the distributed quantile representation, the dispersion between the reward values ​​of each quantile is quantitatively analyzed. By statistically analyzing the deviation of each quantile reward value from its overall mean, the cognitive uncertainty estimation parameter used to characterize the discrete characteristics of the reward distribution of the state-action pair is obtained. When the reward values ​​of each quantile in the distributed quantile representation are concentrated, it indicates that the reward structure corresponding to the current state-action pair is stable, and its cognitive uncertainty estimation parameter has a low value. When the reward values ​​of each quantile show significant dispersion or an increased distribution span, it indicates that the reward structure corresponding to the current state-action pair has strong fluctuations, and its cognitive uncertainty estimation parameter has a high value, thereby effectively characterizing the cognitive uncertainty level of the model for the state-action pair.

[0049] Step S4: Tail Risk Modulation: Extract the lower tail distribution based on distributed quantile representation, determine the tail interval range according to the preset tail truncation ratio, and perform weighted processing on each tail quantile in the lower tail distribution according to the position distribution of each tail quantile within the tail interval; perform lower conditional risk CVaR value calculation on the weighted lower tail distribution to obtain the tail risk measurement result corresponding to the current state-action pair; the tail risk measurement result is used to characterize the potential risk level of crop damage, omission expansion, repeated application, dangerous collision, boundary crossing and abnormal energy consumption caused by the current control action under the most unfavorable operation situation; perform expectation calculation on each quantile based on distributed quantile representation to obtain the distribution mean return; jointly map the tail risk measurement result and the distribution mean return to construct the risk-sensitive value characterization result;

[0050] The distribution of the mean return satisfies:

[0051] ;

[0052] in, Indicates time The current job status, Indicates time In state The following strategy The current control action output. Indicating in strategy Under its influence, at all times Corresponding state With action The mean return of the state-action pair is used to characterize the average return level of the cumulative return distribution corresponding to the state-action pair over multiple future control periods. Indicating in strategy Under the influence of the state-action pair The corresponding distribution of future cumulative returns is used to represent the action to be performed starting from the current moment. Then, the strategy is executed continuously at multiple subsequent time points. The cumulative returns that may be generated at that time are randomly distributed; Represents the mathematical expectation operation;

[0053] The conditional risk CVaR value is satisfied:

[0054] ;

[0055] in, Indicating in strategy Under its influence, at all times status With action The corresponding underside conditional risk CVaR value is used to characterize the average return level when the future cumulative return distribution is in the tail interval; The conditional value at risk operator is used to truncate the value at a preset tail ratio. Under constraints, perform conditional expectation calculations on the lower tail distribution of the future cumulative return distribution; This represents the total number of tail fractions involved in the tail risk calculation. Indicates the index of the tail portion location; This indicates the current state-action pair. Next, the The weight coefficients corresponding to each tail segment site are used to characterize the contribution of that tail segment site in the tail risk aggregation process. Indicates the first A quantile level parameter that falls within the tail interval; Indicates the relationship with the first Individual tails rank The corresponding predicted return value, i.e., the parameterized quantile commentator network for state-action pairs Output quantile return estimates;

[0056] Step S5: Local Commentator Update: Construct a distributed Bellman target distribution based on the current state transition sample. Calculate the quantile regression error between the distributed Bellman target distribution and the distributed quantile representation, and define the calculation result as the quantile regression error term. Write the distributed quantile representation corresponding to the current round into the historical return distribution cache. Calculate the lower conditional value at risk (CVaR) statistic for each historical return distribution in the historical return distribution cache, and generate risk weights based on the CVaR statistic corresponding to each historical return distribution. Based on the risk weights, perform Wasserstein-1 centroid solving on multiple historical return distributions to construct a CVaR-weighted Wasserstein-1 centroid distribution, which serves as the risk perception reference distribution during the commentator network update process in the current round. Calculate the distribution distance between the current return distribution represented by the distributed quantile representation and the risk perception reference distribution to form a return distribution consistency constraint term. By jointly minimizing the quantile regression error term and the return distribution consistency constraint term, perform parameter updates on the parameterized quantile commentator network to obtain the local updated commentator parameters corresponding to the current weeding robot client.

[0057] In this embodiment, the relationship between the return distribution modeling process, the construction method of the historical return distribution set, and the generation of the risk perception reference distribution is as follows: Figure 2 As shown; Figure 2 This diagram illustrates the relationship between the current return distribution, historical return distribution, and risk perception reference distribution; for example... Figure 2As shown, the horizontal axis represents the cumulative return value, and the vertical axis represents the distribution quality. The dashed lines of different colors in the figure represent different distribution instances in the historical return distribution set. Among them, the blue dashed line represents historical return distribution 1, the orange dashed line represents historical return distribution 2, the green dashed line represents historical return distribution 3, and the red dashed line represents historical return distribution 4. The purple solid line in the figure represents the current return distribution. The brown solid line in the figure represents the risk perception reference distribution.

[0058] Among them, the risk perception reference distribution is a reference distribution obtained by weighting the lower conditional value at risk (CVaR) statistic and solving the Wasserstein-1 centroid based on the historical return distribution set. It is used to characterize the overall return distribution structure after considering tail risk. The difference between the current return distribution and the risk perception reference distribution is used to calculate the distribution distance in the subsequent process to construct a return distribution consistency constraint, thereby constraining and optimizing the update process of the parameterized quantile commentator network.

[0059] Quantile regression error term:

[0060] ;

[0061] in, This represents the quantile regression error term. This represents the total number of quantiles. This represents the index of the current critic network output quantile. Index representing the quantile of the target distribution; Indicates the first quantile level parameters Indicates quantile level The corresponding quantile regression loss function; This indicates that the target return distribution is at the quantile level. The target quantile value at that location. This indicates the current parameterized quantile critic network at the quantile level. The quantile return estimate at the output;

[0062] Consistency constraint of reward distribution:

[0063] ;

[0064] in, This represents the consistency constraint term for the return distribution. This represents the quantile index. Indicates the first quantile level parameters This indicates the current parameterized quantile critic network in the state-action pair. Below, at the quantile level The estimated return value output at the location; This indicates that the risk perception reference distribution is in the state-action pairs. Below, at the quantile level The corresponding reference return value is indicated by the superscript. This indicates that the quantile value originates from the Wasserstein-1 barycentric distribution;

[0065] Joint optimization objective:

[0066] ;

[0067] in, This represents the joint optimization objective, also known as the commentator network's total loss function; This represents the weighting coefficients for the consistency constraint of the return distribution;

[0068] Step S6: Federated Commentator Collaborative Aggregation: The federated collaborative training module performs cross-node collaborative processing on the locally updated commentator parameters of each weeding robot client; the federated collaborative training module is implemented through hierarchical collaboration between the field edge collaborative station and the central federated management server; among them, the field edge collaborative station is responsible for parameter aggregation and forwarding, and the central federated management server is responsible for weighted aggregation to generate global aggregated commentator parameters, and broadcasts them back to each weeding robot client through the field edge collaborative station;

[0069] Step S7: Restricted Feedback of Aggregate Parameters: The client trust domain constraint module receives the global aggregate commentator parameters broadcast back, and performs trust domain constraint update on the output distribution corresponding to the global aggregate commentator parameters with the risk perception reference distribution as the anchor point, generating restricted aggregate commentator parameters applicable to the current client;

[0070] Step S8: Strategy Optimization Output: Based on the risk-sensitive value representation results, the mean advantage term and the tail risk advantage term are calculated respectively, and weighted fusion is performed according to the preset risk weights to form a risk-enhanced advantage estimate; then the risk-enhanced advantage estimate is input into the strategy gradient update process to optimize the parameters of the local strategy network Actor, so that the local strategy network Actor gradually improves the effectiveness of weed removal, the reliability of crop protection, the boundary compliance ability, and the energy consumption control ability in continuous training iterations; after completing the Actor update for the current control cycle, the updated local strategy network Actor, combined with the constrained aggregated critic parameters, performs strategy optimization on the initial joint control strategy and re-outputs the target joint control strategy for the next control time.

[0071] Example 4, based on Example 3, introduces an exploration mechanism based on cognitive uncertainty to perform adaptive disturbance processing on the initial joint control strategy to obtain the candidate control action at the current moment. The specific steps include:

[0072] Step B1: Based on the initial joint control strategy, construct a set of strategy particles corresponding to the current control cycle, and use the initial joint control strategy as the initial strategy center of the particle swarm; configure the cumulative benefit recording parameters and cognitive uncertainty estimation parameters for each strategy particle, and construct a state hierarchy tree with the current operation state as the root node; the state hierarchy tree is used to record the state transition path of each strategy particle in the rolling exploration process and its parent-child dependency relationship, thereby forming a backtrackable multi-path search structure;

[0073] Step B2: Perform rolling exploration on each strategy particle in the strategy particle set. Under the joint modulation of the perturbation action distribution and the cognitive uncertainty estimation parameters, each strategy particle performs multi-step evolution along the joint action space. During the rolling cycle, joint action sampling is performed with the current state as input to generate joint control actions of chassis motion control actions and weeding execution actions. Based on the joint control actions, the state transition results are obtained, and the cognitive uncertainty estimation parameters are dynamically updated. When the cumulative gain of any strategy particle increases, the current rolling process of that strategy particle is immediately terminated, and its cumulative gain record parameters and corresponding state nodes are retained to obtain the set of surviving strategy particles (particles with increased cumulative gains) and the set of failed strategy particles (particles that have not obtained effective gains or have entered infeasible states).

[0074] pseudocode:

[0075] enter:

[0076] P / / Set of policy particles;

[0077] s_t / / Current job status;

[0078] D / / Distribution of disturbance actions;

[0079] U / / Set of parameters for estimating cognitive uncertainty;

[0080] Output:

[0081] P_alive / / ​​Set of particles for the survival strategy;

[0082] P_fail / / Set of failure strategy particles;

[0083] 1: P_alive ← ∅ / / Initialize the survival strategy particle set;

[0084] 2: P_fail ← ∅ / / Initialize the set of particles for the failure strategy;

[0085] 3: for each p_i in P do / / Iterate through each strategy particle;

[0086] 4: s_i ← s_t / / Use the current job state as the initial state of the particle;

[0087] 5: C_i ← GetReturnRecord(p_i) / / Reads the cumulative profit record parameter corresponding to the particle;

[0088] 6: U_i ← GetUncertainty(p_i, U) / / Reads the cognitive uncertainty estimation parameters corresponding to the particle;

[0089] 7: while not StopRollout(p_i) do / / If the preset rollout termination condition is not met, continue exploring;

[0090] 8: a_i ← SampleJointAction(s_i, D, U_i) / / Perform joint action sampling based on the current state, perturbation action distribution, and uncertainty parameters;

[0091] 9: (s_i', r_i, flag_i) ← StepEnv(s_i, a_i) / / Execute the joint control action to obtain the next state, benefit feedback, and state flag;

[0092] 10: U_i ← UpdateUncertainty(U_i, s_i, a_i, s_i') / / Dynamically update the cognitive uncertainty estimation parameters based on the state transition results;

[0093] 11: C_i' ← UpdateReturnRecord(C_i, r_i) / / Update the cumulative profit record parameter;

[0094] 12: if C_i' > C_i then / / If the cumulative profit increases;

[0095] 13: SaveStateNode(p_i, s_i', C_i', U_i) / / Saves the updated cumulative profit record parameters and corresponding state nodes;

[0096] 14: P_alive ← P_alive ∪ {p_i} / / Add this particle to the survival strategy particle set;

[0097] 15: break / / Immediately terminates the current scrolling process of this particle;

[0098] 16: end if

[0099] 17: if flag_i == infeasible or ReachMaxStep(p_i) then / / If the infeasible state is entered or the preset maximum number of scroll steps is reached;

[0100] 18: P_fail ← P_fail ∪ {p_i} / / Add this particle to the set of failed strategy particles;

[0101] 19: break / / Terminates the current rolling process of this particle;

[0102] 20: end if

[0103] 21: s_i ← s_i' / / Update the current state and proceed to the next scroll step;

[0104] 22: C_i ← C_i' / / Update the cumulative profit record parameter;

[0105] 23: end while

[0106] 24: end for

[0107] 25: return P_alive, P_fail / / Output the set of particles for the survival strategy and the set of particles for the failure strategy.

[0108] Step B3: Sort and filter the cumulative returns of each strategy particle in the surviving strategy particle set to determine the candidate particle corresponding to the maximum cumulative return; when there are multiple candidate particles, perform a secondary sorting based on the quantitative index corresponding to the cognitive uncertainty estimation parameter, and determine the strategy particle with the largest quantitative index as the winner particle according to the lexicographical order sorting rule of "cumulative return first, uncertainty second best"; based on the winner particle, perform state reset on each failed strategy particle in the failed strategy particle set, so that each failed strategy particle is repositioned to the state node corresponding to the winner particle and inherits its state context and search frontier, and the remaining surviving strategy particles are retained to continue to perform rolling exploration along their respective paths, thereby realizing the centralized allocation of search resources to high-value and high-uncertainty areas, and obtaining the allocated strategy particle set;

[0109] Partial pseudocode:

[0110] enter:

[0111] P_alive / / ​​Set of particles for the survival strategy;

[0112] P_fail / / Set of failure strategy particles;

[0113] Output:

[0114] p_win / / Winner particle;

[0115] P_alloc / / Set of policy particles after allocation;

[0116] 1: CandidateSet ← ∅ / / Initializes the candidate particle set;

[0117] 2: maxReturn ← max { GetCumulativeReturn(p_i) | p_i ∈ P_alive} / / Get the maximum cumulative return value in the survival strategy particle set;

[0118] 3: CandidateSet ← { p_i | p_i ∈ P_alive and GetCumulativeReturn(p_i) = maxReturn} / / Filter candidate particles whose cumulative return equals the maximum value;

[0119] 4: if |CandidateSet| = 1 then / / Determine if the candidate particle is unique;

[0120] 5: p_win ← the only particle in CandidateSet / / If it is unique, then it is directly determined as the winning particle;

[0121] 6: else

[0122] 7: p_win ← LexicographicArgMax(CandidateSet, ReturnFirst,UncertaintySecond) / / Perform a secondary sort according to the lexicographical sorting rule of "cumulative return first, uncertainty second";

[0123] 8: end if

[0124] 9: v_win ← GetStateNode(p_win) / / Read the state node corresponding to the winning particle;

[0125] 10: for each p_j ∈ P_fail do / / Iterate through each failed strategy particle;

[0126] 11: ResetState(p_j, v_win) / / Resets the failed policy particle to the state node corresponding to the winner particle;

[0127] 12: InheritContext(p_j, p_win) / / Makes the failed policy particle inherit the state context and search frontier of the winner particle;

[0128] 13: end for

[0129] 14: P_alloc ← P_alive ∪ P_fail / / Merges the failed policy particles after state reset with the remaining surviving policy particles to form a set of allocated policy particles;

[0130] 15: return p_win, P_alloc / / Output the winning particle and the set of policy particles after allocation.

[0131] Step B4: Perform parent node backtracking along the state phylogenetic tree for the allocated policy particle set. Select the corresponding ancestor state node for each policy particle within the preset ancestor generation range, and roll back each policy particle to the corresponding ancestor state node and restart joint action sampling to obtain the policy particle set after rollback and recovery. The ancestor generation range is used to control the backtracking depth so as to avoid getting trapped in local optimum or unrecoverable state while maintaining the continuity of local search.

[0132] Step B5: After completing the preset exploration rounds, perform joint sorting on the cumulative gains and cognitive uncertainty estimation parameters of each strategy particle in the rolled-back and restored strategy particle set to determine the group winner particle; and uniformly reset all strategy particles in the rolled-back and restored strategy particle set to the state node corresponding to the group winner particle to achieve group convergence; extract the joint action parameters from the state node corresponding to the group winner particle, and use the joint action parameters as the candidate control action output at the current moment.

[0133] Example 5, this example is based on Example 4. In this example, step S6 specifically includes the following steps:

[0134] Step S61: Edge-side aggregation and preprocessing: The field edge collaborative station deployed in the farm's machine room establishes communication connections with multiple weeding robot clients, receiving locally updated commentator parameters uploaded by each client; it performs node identity binding, timestamp alignment, parameter integrity verification, and abnormal communication data removal on the locally updated commentator parameters, and performs batch caching and ordered arrangement of commentator parameters from different clients according to a preset communication scheduling strategy to form an edge-side parameter set; subsequently, the edge-side parameter set is packaged in a unified format and uploaded to the central federated management server;

[0135] Step S62: Central-side weighted aggregation processing: The central federated management server, deployed in the cloud server, receives the edge-side parameter set uploaded from each field edge collaboration station, performs unified decoding and version consistency verification on the parameters of each local update commentator in the set; based on the sample size, task completion quality, historical risk stability index, and communication effectiveness index of each client, constructs multi-factor weighting coefficients, performs weighted aggregation operation on each commentator network parameter, and generates global aggregated commentator parameters; wherein, the central federated management server only performs aggregation processing on the local update commentator parameters, and does not aggregate the local policy network Actor parameters of each client, thereby maintaining the adaptive ability and stability of the policy model in the heterogeneous environment of each field to local environmental characteristics, risk distribution structure, and action execution constraints;

[0136] Step S63: Edge-side distribution and feedback: The central federated management server distributes the generated global aggregated commentator parameters to each field edge collaboration station; after receiving the global aggregated commentator parameters, each field edge collaboration station performs parameter version marking, node mapping matching and distribution scheduling processing, and broadcasts the global aggregated commentator parameters back to the corresponding weeding robot clients, so that each client can perform subsequent trust domain constraint updates and local policy optimization processing.

[0137] Example 6, this example is based on Example 5. In this example, step S7 specifically includes the following steps:

[0138] Step S71: Asymmetric Trust Domain Boundary Construction: Establish upper and lower allowable offset radii for each quantile of the risk perception reference distribution, thereby constructing an asymmetric distribution offset constraint interval at each quantile.

[0139] Step S72: Shrinking pre-update processing: The distributed quantile output corresponding to the global aggregated commentator parameters is linearly shrunk and mapped to the risk perception reference distribution according to the preset shrinkage coefficient to reduce the distribution offset introduced by cross-round federated aggregation and obtain the intermediate distribution update result;

[0140] Step S73: Asymmetric compression constraint mapping: Input the intermediate distribution update result into the asymmetric compression operator constructed with the hyperbolic tangent function as the core, and perform direction-sensitive compression mapping processing on the offset of each quantile. This makes the offset above the reference distribution limited by the upper allowable offset radius, and the offset below the reference distribution limited by the lower allowable offset radius, thereby restricting the distribution update result within the asymmetric distribution offset constraint interval and obtaining the asymmetric compression constraint quantile output result.

[0141] Step S74: Restricted Commentator Parameter Reconstruction: Based on the asymmetric compression constraint quantile output results, perform backfitting on the parameter representation of the parameterized quantile commentator network to obtain the commentator network parameters that satisfy the trust domain constraint, forming restricted aggregated commentator parameters suitable for the current client, so as to suppress the cross-wheel drift, abnormal oscillation and worst-case gain estimation distortion of the commentator output in the tail distribution after federated aggregation.

[0142] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. A precision control system for a weeding robot based on reinforcement learning, characterized in that, The system includes: A multimodal environment perception module is used to construct precise field operation status data. The state modeling and coding module acquires the robot's own data and combines it with the field precision operation state data processing to generate an operation state vector; The reinforcement learning decision-making module constructs a risk-aware distributed federated Actor-Critic model; it receives the current job state vector and, based on the risk-aware distributed federated Actor-Critic model, sequentially executes uncertainty-guided strategy generation, distributed value assessment, tail risk modulation, local commentator update, federated commentator collaborative aggregation, and aggregate parameter restricted backfeeding processing within the current control cycle, and outputs the target joint control strategy. The execution control module processes the target joint control strategy, generates chassis motion control commands and weeding execution commands, and completes the fine weeding operation. The risk-aware distributed federated Actor-Critic model includes a local policy network, a parameterized quantile critic network, and a federated collaborative training module. The output process of the target joint control strategy specifically includes the following steps: Step S1: Process the features in the current job state vector to construct a decision input representation; construct a historical reward distribution cache; Step S2: Input the decision input representation into the local policy network, perform forward policy reasoning, and output an initial joint control policy consisting of chassis motion control action and weeding execution action; read the cognitive uncertainty estimation parameters from the historical reward distribution cache, and introduce an exploration mechanism based on cognitive uncertainty to perform adaptive perturbation processing on the initial joint control policy to obtain the candidate control action at the current time. Step S3: Combine the current job state vector with the candidate control action at the current time to form a state-action pair, and input it into the parameterized quantile critic network to output a distributed quantile representation; Step S4: Extract the lower tail distribution based on the distributed quantile representation, and perform weighting on each tail quantile in the lower tail distribution; calculate the lower conditional risk (CVaR) value on the weighted lower tail distribution to obtain the tail risk measurement result; calculate the expected value of each quantile based on the distributed quantile representation to obtain the distribution mean return; jointly map the tail risk measurement result with the distribution mean return to construct the risk-sensitive value representation result. Step S5: Construct a distributed Bellman target distribution, calculate the quantile regression error by combining the distributed Bellman target distribution with the distributed quantile representation, and define the calculation result as the quantile regression error term; calculate the CVaR statistic for multiple historical return distributions in the historical return distribution cache, and generate risk weights based on the CVaR statistic corresponding to each historical return distribution; perform Wasserstein-1 centroid solving on multiple historical return distributions based on the risk weights, and construct a CVaR-weighted Wasserstein-1 centroid distribution as the risk perception reference distribution; calculate the distribution distance between the current return distribution represented by the distributed quantile representation and the risk perception reference distribution to form a return distribution consistency constraint term; by jointly minimizing the quantile regression error term and the return distribution consistency constraint term, perform parameter updates on the parameterized quantile commentator network to obtain the local updated commentator parameters corresponding to the current weeding robot client; Step S6: Perform cross-node collaborative processing on the local update commentator parameters of each weeding robot client through the federated collaborative training module to generate global aggregated commentator parameters; Step S7: Using the risk perception reference distribution as the anchor point, perform trust domain constraint update on the output distribution corresponding to the global aggregated commentator parameters to generate restricted aggregated commentator parameters; Step S8: Based on the risk-sensitive value representation results, form a risk-enhancing advantage estimate and perform parameter optimization on the local policy network; after the update is completed, the updated local policy network, combined with the constrained aggregate critic parameters, performs policy optimization on the initial joint control policy and re-outputs the target joint control policy for the next control time step.

2. The precision control system for a weeding robot based on reinforcement learning according to claim 1, characterized in that: The chassis motion control actions include chassis linear velocity control, chassis angular velocity control, and lateral correction. The weeding execution actions include actuator lifting displacement, nozzle opening mode, spraying flow rate, spraying duration, cutter head speed, cutter depth, and emergency braking action.

3. The precision control system for a weeding robot based on reinforcement learning according to claim 1, characterized in that: The process of introducing an exploration mechanism based on cognitive uncertainty to perform adaptive perturbation processing on the initial joint control strategy and obtain candidate control actions at the current time includes the following steps: Step B1: Based on the initial joint control strategy, construct a set of strategy particles; configure the cumulative revenue recording parameters and cognitive uncertainty estimation parameters for each strategy particle, and construct a state hierarchy tree with the current operation state as the root node; Step B2: Perform rolling exploration on each strategy particle in the strategy particle set, perform joint action sampling within the rolling cycle, generate joint control actions, obtain state transition results based on joint control actions, dynamically update the cognitive uncertainty estimation parameters, and obtain the surviving strategy particle set and the failed strategy particle set. Step B3: Perform sorting, filtering, secondary sorting, and state reset on the cumulative gains of each strategy particle in the surviving strategy particle set to obtain the allocated strategy particle set; Step B4: Perform parent node backtracking along the state hierarchy tree for the allocated policy particle set to obtain the rolled-back and restored policy particle set; Step B5: Perform joint sorting on the cumulative returns and cognitive uncertainty estimation parameters of each strategy particle in the strategy particle set after rollback recovery to determine the group winner particle; and uniformly reset to the state node corresponding to the group winner particle to achieve group convergence; extract joint action parameters from the state node corresponding to the group winner particle, and use the joint action parameters as candidate control actions at the current moment.

4. The precision control system for a weeding robot based on reinforcement learning according to claim 3, characterized in that: Step B3 specifically includes: sorting and filtering the cumulative returns of each strategy particle in the surviving strategy particle set to determine the candidate particle corresponding to the maximum cumulative return; when there are multiple candidate particles, performing a secondary sorting based on the quantitative index corresponding to the cognitive uncertainty estimation parameter to determine the strategy particle with the largest quantitative index as the winner particle; and resetting the state of each failed strategy particle in the failed strategy particle set based on the winner particle to obtain the allocated strategy particle set.

5. The precision control system for a weeding robot based on reinforcement learning according to claim 1, characterized in that: The federated collaborative training module is implemented through a hierarchical collaboration between field edge collaborative stations and a central federated management server. The field edge collaborative stations are responsible for parameter aggregation and forwarding, while the central federated management server is responsible for weighted aggregation to generate global aggregated commentator parameters, and broadcasting them back to each weeding robot client via the field edge collaborative stations.