Low altitude pollination control method based on closed loop multi-objective reinforcement optimization algorithm
By using multi-source sensor data and reinforcement learning algorithms, canopy geometric channels and aerodynamic bias attention were constructed, solving the problem of sedimentation estimation and multi-objective constraint coordination in low-altitude pollination by UAVs, and achieving stable real-time control and uniform deposition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN UNIV OF SCI & ENG
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-15
AI Technical Summary
Existing UAV low-altitude pollination technology struggles to achieve real-time deposition estimation and feedback under dynamic wind fields, complex canopy geometry, and flower position distribution. It also faces challenges in coordinating multiple objectives and safety constraints, and in generating actions that fail to meet feasible domain and on-site constraints, leading to problems such as insufficient deposition uniformity and non-target drift.
A canopy geometric channel is constructed using multi-source sensor data. The local aerodynamic and particle migration fields are solved using Fourier neural operators. Deposition and uncertainty estimation are performed by combining a spatiotemporal attention network. Safety constraints and multi-objective rewards are constructed based on differentiable barrier functions. Adaptive trade-offs are achieved through world model reinforcement learning. Candidate actions are refined and closed-loop feedback is implemented using a diffusion model.
It improves deposition uniformity, reduces non-target drift, suppresses mechanical impact and energy consumption of flowers, meets minimum safety distance and no-spray zone constraints, and achieves stable real-time control.
Smart Images

Figure CN121411265B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude pollination by unmanned aerial vehicles (UAVs), and more particularly to a low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm. Background Technology
[0002] In recent years, low-altitude drones have been widely used in orchard pollination and plant protection operations, and research on powdering particle size control, rotor downwash airflow modeling, trajectory planning and spray volume scheduling has been continuously advancing.
[0003] Existing technologies mostly rely on preset routes and empirical rules for open-loop operations. Some operations use sensors such as wind speed and direction, and powder spraying flow rate to perceive the environment, and use empirical formulas, simplified CFD or data-driven models to estimate deposition and drift. There are also attempts to optimize paths and parameters based on reinforcement learning or model predictive control.
[0004] However, under the combined influence of dynamic wind fields, complex canopy geometry, and flower position distribution, real-time and reliable sedimentation feedback is still lacking. The main shortcomings of existing technologies are:
[0005] 1. Deposition volume is difficult to estimate and quantify uncertainty in real time: Estimations based on offline simulation or empirical models are difficult to adapt to rapid changes in local downwash and environmental wind, lack explicit utilization of canopy normal, porosity and flower position spatial distribution, and usually do not provide uncertainty quantification that can be used for safety decisions.
[0006] 2. Insufficient coupling between multi-objectives and safety constraints: Optimization is mostly based on a single coverage or efficiency index, and lacks overall planning with non-target drift, flower mechanical impact, energy consumption and operation time, etc. Constraints such as no-spraying areas, maximum allowable wind speed, and minimum safety distance are mostly handled with hard rules, which are difficult to adaptively balance under complex conditions.
[0007] 3. Lack of closed-loop action refinement oriented towards feasible domain: Even if candidate parameters or trajectories are generated, there is often a lack of post-processing mechanisms that combine aerodynamic field and constraint gradient, which can easily lead to violation of constraints or insufficient deposition uniformity in actual execution, making it difficult to form stable closed-loop control.
[0008] Therefore, a method for low-altitude pollination by drones that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0009] One objective of this invention is to propose a low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm. Addressing the problems in existing technologies such as difficulty in real-time estimation and feedback of pollination deposition, difficulty in coordinating multi-objective and safety constraints, and difficulty in meeting feasible domain and field constraints in action generation, this invention proposes a technical solution that utilizes multi-source sensing and geometric enhancement to construct canopy geometric channels, uses Fourier neural operators to solve local aerodynamic and particle migration fields and constructs aerodynamic bias attention, employs a spatiotemporal attention network for deposition and uncertainty estimation, constructs safety constraints and multi-objective rewards based on differentiable barrier functions and adaptively balances them in world model reinforcement learning, and then refines candidate actions and provides closed-loop feedback using a diffusion model guided by aerodynamic fields and constraints. This invention achieves technical effects such as improving deposition uniformity, reducing non-target drift, suppressing mechanical impact and energy consumption of flowers, meeting minimum safety distance and no-spraying zones, and realizing stable real-time control under uncertain conditions.
[0010] A low-altitude pollination control method based on a closed-loop multi-objective enhancement optimization algorithm according to an embodiment of the present invention is characterized by comprising the following steps:
[0011] S1. Collect multi-source sensor data and form an input data set. Then, perform geometric enhancement processing on the input data set to obtain the canopy geometric channel.
[0012] S2. Input the input data set and the canopy geometry channel into the Fourier neural operator model, and output the local aerodynamic and particle migration fields;
[0013] S3. The local aerodynamic and particle migration fields are mapped into an attention bias matrix through the bias construction module, and the input data set, canopy geometric channels, and attention bias matrix are combined to form the deposition estimation input;
[0014] S4. Input the deposition estimate into the spatiotemporal attention network and output the deposition estimate output and uncertainty index.
[0015] S5. Calculate the multi-objective reward vector by combining the sedimentation estimation output, uncertainty index, and target weight vector in the input dataset through the multi-objective construction module. Construct a constraint set by combining the sedimentation estimation output, uncertainty index, canopy geometric channel, field boundaries, and no-spraying zone map in the input dataset through the constraint generation module.
[0016] S6. Combine the input data set, sedimentation estimation output and uncertainty index to form a job state vector. Input the job state vector, multi-objective reward vector and constraint set into the world model multi-objective reinforcement learning module to output candidate action sequence.
[0017] S7. Input the candidate action sequence and constraint set into the diffusion action refinement module. The diffusion action refinement module guides diffusion sampling with local aerodynamics and particle migration field, and performs constraint guidance and noise reduction with the differentiable barrier function corresponding to the constraint set, and outputs the refined action sequence.
[0018] S8. Input the refined action sequence into the flight control execution module for execution, collect the multi-source sensor data after execution to obtain data set two, and use data set two as the input data set for the next control cycle.
[0019] Optionally, step S1 specifically includes:
[0020] The system collects camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle, field boundary and no-spraying zone maps, and target weight vectors. It also calibrates, synchronizes, and spatially registers each channel to a unified coordinate system.
[0021] Distortion correction and denoising are performed on camera images, and canopy depth representation is established by combining the relative distance to the canopy and flight attitude and position information;
[0022] The canopy normal and canopy porosity are calculated based on canopy depth representation and camera images, and the flower position probability is obtained based on flower organ detection and region statistics from camera images.
[0023] The input data set is formed by summarizing camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle, field boundary and no-spraying zone map, and target weight vector. The canopy geometric channel is constructed by combining the canopy normal, canopy porosity and flower position probability.
[0024] Optionally, step S2 specifically includes:
[0025] Under a unified coordinate system, the local computational domain is determined based on the relative distance between the flight attitude and position information in the input dataset and the canopy, combined with the canopy geometric channel.
[0026] The input data set and the canopy geometric channel are used to construct the conditional input of the Fourier neural operator model through feature encoding. The conditional input includes boundary conditions and source terms. The boundary conditions include at least the far-field boundary of the ambient wind speed and direction, the near-field boundary of the rotor-induced airflow, and the canopy boundary characteristics determined by the canopy normal and canopy porosity. The source terms include at least the particle injection parameters determined by the powder flow rate and nozzle angle.
[0027] Within the current control cycle, the spatial coordinates within the local computational domain are mapped by Fourier features and input together with the conditional input into the Fourier neural operator model to obtain the local aerodynamic and particle migration fields, and output the rinsing velocity field and particle concentration field with a preset spatial resolution.
[0028] The Fourier neural operator model employs supervised learning with measured wind speed and particle concentration data during the training phase, and applies continuity constraints and mass conservation constraints.
[0029] Optionally, step S3 specifically includes:
[0030] The downwash velocity field and particle concentration field in the local aerodynamic and particle migration field are processed by the bias construction module. The streamline direction field is determined based on the downwash velocity field, and the weights along the streamline direction and the canopy normal and canopy porosity in the canopy geometric channel are combined.
[0031] Based on the particle concentration field and the weights, the attention bias value is calculated for the spatiotemporal index pairs of the input data set and the canopy geometric channel. An attention bias matrix with the same size as the attention score matrix of the self-attention mechanism in the spatiotemporal attention network is generated. The attention bias matrix is then normalized and truncated to limit the range of bias values and ensure numerical stability.
[0032] The input data set, canopy geometry channel, and attention bias matrix are combined in a time-synchronized and spatially registered manner to form the deposition estimation input. This allows the attention bias matrix to implement aerodynamic bias attention in the self-attention mechanism of the spatiotemporal attention network, with weighted focusing along the streamline direction and the canopy normal.
[0033] Optionally, step S4 specifically includes:
[0034] Under a unified coordinate system, the deposition estimation input is aligned temporally and spatially. A spatiotemporal attention network is used to encode the deposition estimation input temporally and spatially to form a spatiotemporal label sequence. During the calculation of attention score in the self-attention mechanism, an attention bias matrix is introduced to implement aerodynamic bias attention. The spatiotemporal label sequence is weighted and aggregated and then subjected to nonlinear transformation and normalization by a feedforward network. The deposition estimation output and uncertainty index are output. The deposition estimation output includes a deposition index to characterize the amount of pollen that can fall per unit area, a deposition uniformity index to characterize the uniformity of cover, and a drift risk index to characterize the risk of drifting to non-target areas. The uncertainty index is obtained by calculating the variance or confidence interval of the output distribution.
[0035] During the training phase, a loss function is constructed and parameters are learned by considering the mass conservation constraints between the total amount of powder sprayed, the amount of deposition, the amount of non-target drift, and the suspension loss.
[0036] Optionally, step S5 specifically includes:
[0037] The sedimentation index, sedimentation homogeneity index, drift risk index, and uncertainty index in the sedimentation estimation output are unified and normalized in terms of dimensions, and sedimentation homogeneity reward and non-target drift penalty are obtained according to the preset monotonic mapping rule.
[0038] The energy consumption per unit area is estimated and the energy consumption penalty is obtained based on the powder spraying flow rate, flight attitude and position information in the input data set. The operation time penalty is calculated based on the timestamp and flight path progress in the input data set. The mechanical impact penalty of the flower is constructed based on the relative distance to the canopy and the uncertainty index.
[0039] The deposition uniformity reward, non-target drift penalty, energy consumption penalty, operation time penalty and flower mechanical impact penalty are weighted and combined according to the target weight vector to form a multi-target reward vector, and the weight of the safety-related component is increased when the uncertainty index exceeds the threshold.
[0040] Meanwhile, based on the sedimentation estimation output and uncertainty index, the canopy geometric channel and the field boundary and no-spraying area map in the input data set, and combined with the wind speed, wind direction and dust spraying flow rate in the input data set, minimum safe distance constraint, maximum allowable wind speed constraint, no-spraying area constraint and dust spraying flow rate upper limit constraint are constructed respectively. Each constraint is expressed in the form of a differentiable barrier function, and the constraint weights are adaptively adjusted according to the uncertainty index to obtain the constraint set.
[0041] Optionally, step S6 specifically includes:
[0042] The input data set, sedimentation estimation output, and uncertainty index are synchronized in time and spatially registered and summarized into an operation state vector. The operation state vector includes at least flight attitude and position information, relative distance to the canopy, wind speed and direction, dust spraying flow rate and nozzle angle, field boundary and no-spraying area map, sedimentation index, sedimentation uniformity index, drift risk index, and uncertainty index.
[0043] The job state vector, multi-objective reward vector, and constraint set are input together into the world model multi-objective reinforcement learning module. The world model multi-objective reinforcement learning module adopts the DreamerV3 structure. It performs rolling prediction based on the job state vector to evaluate the state transition and the expectation of the multi-objective reward vector within a preset time window. In the strategy optimization, the constraint terms are constructed using the differentiable obstacle function corresponding to the constraint set, and the constraint weights are adaptively adjusted using the Lagrange multiplier method. This ensures that the strategy maximizes the multi-objective reward weighted by the objective weight vector while satisfying the minimum safe distance, maximum allowable wind speed, no-spray zone, and upper limit of powder spraying flow rate.
[0044] Output candidate action sequences, which include flight altitude, flight speed, roll angle, yaw angle, powder flow rate, and nozzle angle.
[0045] Optionally, step S7 specifically includes:
[0046] The candidate action sequence, constraint set, and local aerodynamic and particle migration field are input into the diffusion action refinement module. The candidate action sequence is used as the initial trajectory for the diffusion denoising process for noise injection and denoising iteration.
[0047] In each round of denoising iteration, the action guidance term is calculated based on the local aerodynamic and particle migration field and superimposed on the denoising update. At the same time, the gradient of the differentiable barrier function corresponding to the constraint set with respect to the current action sequence is superimposed on the denoising update as a constraint guidance term, so that the generated action sequence converges toward the feasible region that satisfies the constraint set.
[0048] The iteration stops when the constraint set is satisfied or the preset number of iterations is reached, and a set of action samples that satisfy the constraint set is obtained. Then, one action sequence is selected as the refined action sequence according to the criterion of minimizing the constraint cost calculated by the differentiable obstacle function.
[0049] Optionally, step S8 specifically includes:
[0050] The refined action sequence is input into the flight control execution module in the order of timestamps. The flight control execution module converts the flight altitude, flight speed, roll angle, yaw angle, powder flow rate and nozzle angle in the refined action sequence into the corresponding actuator set values and issues control commands to drive the UAV attitude control, trajectory control and powder spraying system to execute.
[0051] During execution, real-time operational data is recorded within a unified coordinate system, and measurements of camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle are collected. Time synchronization, spatial registration and necessary calibration are completed, and the data are summarized to form Data Set Two.
[0052] Data set two is used as the input data set for the next control cycle to start the next control cycle, thereby achieving closed-loop control.
[0053] The beneficial effects of this invention are:
[0054] 1. Real-time sedimentation estimation and uncertainty feedback: The local aerodynamic and particle migration fields are obtained through Fourier neural operators, and an aerodynamic biased attention-guided spatiotemporal attention network is constructed to output the sedimentation index, sedimentation uniformity, drift risk and its uncertainty. The estimation accuracy and robustness are improved by combining mass conservation and continuity constraints. The local computational domain is used to reduce the computational burden and meet the online feedback requirements within the control cycle.
[0055] 2. Adaptive trade-offs in multi-objective safety optimization: Deposition uniformity, non-target drift, energy consumption, operation time and mechanical impact of flowers are incorporated into the multi-objective reward. The minimum safe distance, maximum allowable wind speed, no-spray zone and powder spraying limit are expressed by a differentiable obstacle function. In the DreamerV3 world model, Lagrange multipliers are used for adaptive weighting to achieve uniformity improvement and drift suppression under the premise of safety, while also taking efficiency into account.
[0056] 3. Improved action feasibility and closed-loop stability: The diffusion action refinement is adopted, which uses the guiding terms of local aerodynamics and particle migration field and constraint gradients to guide the candidate action to converge to the feasible region and minimize the constraint cost. Combined with the rolling prediction of the world model, the execution deviation and trajectory jitter are reduced, thereby improving the stability of closed-loop control and field adaptability. Attached Figure Description
[0057] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0058] Figure 1 This is a flowchart of a low-altitude pollination control method based on a closed-loop multi-objective enhancement optimization algorithm proposed in this invention. Detailed Implementation
[0059] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0060] refer to Figure 1 A low-altitude pollination control method based on a closed-loop multi-objective enhancement optimization algorithm, characterized by the following steps:
[0061] S1. Collect multi-source sensor data and form an input data set. Then, perform geometric enhancement processing on the input data set to obtain the canopy geometric channel.
[0062] S2. Input the input data set and the canopy geometry channel into the Fourier neural operator model, and output the local aerodynamic and particle migration fields;
[0063] S3. The local aerodynamic and particle migration fields are mapped into an attention bias matrix through the bias construction module, and the input data set, canopy geometric channels, and attention bias matrix are combined to form the deposition estimation input;
[0064] S4. Input the deposition estimate into the spatiotemporal attention network and output the deposition estimate output and uncertainty index.
[0065] S5. Calculate the multi-objective reward vector by combining the sedimentation estimation output, uncertainty index, and target weight vector in the input dataset through the multi-objective construction module. Construct a constraint set by combining the sedimentation estimation output, uncertainty index, canopy geometric channel, field boundaries, and no-spraying zone map in the input dataset through the constraint generation module.
[0066] S6. Combine the input data set, sedimentation estimation output and uncertainty index to form a job state vector. Input the job state vector, multi-objective reward vector and constraint set into the world model multi-objective reinforcement learning module to output candidate action sequence.
[0067] S7. Input the candidate action sequence and constraint set into the diffusion action refinement module. The diffusion action refinement module guides diffusion sampling with local aerodynamics and particle migration field, and performs constraint guidance and noise reduction with the differentiable barrier function corresponding to the constraint set, and outputs the refined action sequence.
[0068] S8. Input the refined action sequence into the flight control execution module for execution, collect the multi-source sensor data after execution to obtain data set two, and use data set two as the input data set for the next control cycle.
[0069] In this specific embodiment, S1 specifically refers to:
[0070] The system first uses a unified coordinate system The process of acquiring, calibrating, and aligning multi-source data is completed, specifically including camera images. Optical scattering signal electrostatic charge signal Ambient wind speed and direction Relative distance from the canopy Flight attitude and position information ( Powder spraying flow rate With nozzle angle Map of field boundaries and no-spray zones and target weight vector The system includes multiple channels, and performs internal and external parameter calibration, time synchronization, and spatial registration on all channels to ensure data consistency in time and space. Time synchronization uses a master clock reference for alignment, ensuring the sensor... The original timestamp is Fixed offset is The unified time is midnight. Then the unified time after alignment is:
[0071] ;
[0072] in Indicates the unified time after alignment. Indicates the first Sampling time of each sensor, This indicates a fixed time offset of the sensor relative to the master clock. Indicates the zero point of the unified timeline. For sensor indexing;
[0073] Spatial registration transforms local measurements to a unified coordinate system based on the extrinsic parameters of each sensor, making the three-dimensional points in the sensor coordinate system... Rotation to a unified coordinate system is Translation The coordinates after registration are:
[0074] ;
[0075] in Representing a three-dimensional point in a unified coordinate system Represents the rotation matrix from the sensor coordinate system to the unified coordinate system. This represents the corresponding translation vector;
[0076] Visual channel pair Distortion correction and denoising are performed to obtain stable pixel-level observations, and in Combined ranging and position Establish a depth representation of the canopy and a dense 3D point cloud, making the camera intrinsic parameters as follows: The extrinsic parameters of the camera to the unified coordinate system are The homogeneous coordinates of the pixels are The depth along the pixel ray is Then pixel The corresponding three-dimensional points of the canopy are:
[0077] ;
[0078] in Representing the three-dimensional coordinates of the canopy points, Represents the camera intrinsic parameter matrix, Represents its inverse matrix, Represents the rotation matrix from the camera coordinate system to the unified coordinate system. Indicates the corresponding translation vector, This represents the depth value estimated jointly by the ranging prior and the local plane assumption. Homogeneous coordinates of pixels;
[0079] Based on this 3D reconstruction, the canopy normal field is obtained by local difference and normal vector normalization of the depth map to characterize the receptor orientation affected by wind direction and downwash. Canopy porosity is estimated through image segmentation and depth consistency to reflect permeability, allowing the region to... The internal occupation indication is The number of pixels in the region is The porosity is:
[0080] ;
[0081] in Indicates the porosity of the region, This indicates that the pixel belongs to the canopy entity. Indicates gaps or background;
[0082] Pixel-level flower position probability is obtained through a flower detection network. This data, combined with regional statistics, forms the spatial distribution of flower positions for powder spraying alignment.
[0083] Finally, the calibrated, aligned and registered data will be... Summarized into the input data set The canopy geometric channels are constructed by combining the canopy normal, porosity, and flower position probability. For use in subsequent steps.
[0084] In this specific embodiment, S2 specifically refers to:
[0085] The system in a unified coordinate system Based on flight attitude and position information and the relative distance from the canopy And combined with the canopy normal in the canopy geometry channel With porosity To determine the local computational domain for online solving in order to reduce inference burden and maintain physical consistency, the local computational domain is represented as a machine-aligned cuboid in a unified coordinate system as follows:
[0086] ;
[0087] in Represents the local computational domain, Representing a unified coordinate system Spatial points in Indicates that drones are in The position in the middle, Indicates from machine system to rotation matrix, Represents local offset coordinates in the machine system. and Indicates horizontal and vertical dimensions, and Indicates the vertical range below and above, symbol Indicates set membership; the symbol | indicates conditional constraints.
[0088] Subsequently, the conditional input of the Fourier neural operator FNO is constructed, encoding the boundary conditions and source terms into location-related features and solving them in a unified coordinate system within the current control cycle. The far-field boundary is represented by constant constraints on ambient wind speed and direction as follows:
[0089] ;
[0090] in Represents the gas velocity field, Represents the far-field boundary of the local computational domain, Vectors representing ambient wind speed and direction, and their symbols. Represents boundary operators;
[0091] The rotor near-field boundary is given by the induced airflow distribution as follows:
[0092] ;
[0093] in Indicates the near-field boundary of the rotor, Indicates the location The rotor-induced velocity field generation function at the location, Represents the set of rotor parameters;
[0094] Crown boundary characteristics are and The combined inputs, which reflect permeability and directionality, are used as conditional features in the operator. The source term, defined by the powder injection flow rate and nozzle angle, expresses the particle injection density as follows:
[0095] ;
[0096] in Indicates position Particle injection source density at the location, Indicates powder spraying flow rate, This indicates a support indicator function relative to nozzle geometry and orientation. Indicates the nozzle angle;
[0097] To enhance the ability to represent flow patterns at different scales, Fourier feature mapping is performed on the local spatial coordinates as follows:
[0098] ;
[0099] in Indicates position Fourier eigenvectors, Indicates the first Frequency vectors Indicates the number of frequency vectors used. and Representing the cosine and sine functions respectively, Representing pi, symbol Indicates vector transpose;
[0100] The above conditions and Fourier features are fed into FNO to solve for the rinsing velocity field and particle concentration field at the preset spatial resolution, and expressed in operator output form as follows:
[0101] ;
[0102] in Indicates position The downwash velocity vector at the location, Indicates position Particle concentration at the location, Represents Fourier neural operators, This represents the set of conditional inputs that includes boundary conditions and source terms;
[0103] During the training phase, supervised learning based on measured wind speed and particle concentration data is employed, and physical consistency constraints are applied. The continuity constraint is expressed as an incompressible condition:
[0104] ;
[0105] in The symbol • represents the vector differential operator, and the symbol • represents the divergence operation. The velocity field is represented by 0, and the zero scalar is represented by 0. The mass conservation constraint is constructed by balancing the total amount of powder sprayed, the amount of sediment, and the non-target drift and suspension loss to construct a loss term for the stabilizing operator.
[0106] In this specific embodiment, S3 specifically refers to:
[0107] The system is based on the rinsing velocity field and particle concentration field, combined with the normal and porosity in the canopy geometric channels. After completing spatiotemporal synchronization and spatial registration, it constructs aerodynamic bias attention and generates deposition estimation input. First, the local streamline direction is determined based on the rinsing velocity to reflect the flow-dominated migration characteristics. Let the position be... The velocity vector at the point is Its Euclidean norm is Then the direction of the unit streamline is:
[0108] ;
[0109] in Representing spatial position in a unified coordinate system This represents the downwash velocity vector at that location. This indicates taking the Euclidean norm of a vector. Indicates the direction of the unit streamline;
[0110] Subsequently, a weighted coefficient along the streamline-normal direction was combined with the canopy normal and porosity to characterize the combined effect of receptor orientation and permeability, with the canopy normal being... The porosity of the canopy is The symbol • represents the dot product of vectors. If we denote the absolute value, then the weight is defined as:
[0111] ;
[0112] in This indicates the weighting coefficient used for attention bias at that location. Represents the unit normal of the canopy surface, Indicates the porosity of the canopy;
[0113] Next, attention bias matrix elements are calculated on the spatiotemporal marker pairs to guide self-attention to focus along high-concentration and effective channels, letting the first... The spatial coordinates of the markers are The timestamp is , No. The spatial coordinates of the markers are The timestamp is ,Location The particle concentration is Spatial scale hyperparameter is The time-scale hyperparameter is Scaling factor is ,symbol If the function is exponential, then the attention bias is defined as:
[0114] ;
[0115] in Indicates the first With the The bias between the markers and Representing the spatial location of the two spatiotemporal markers respectively, and Indicates the corresponding timestamp, Indicates the location particle concentration, and Controlling the attenuation magnitude in space and time, Control the overall bias size;
[0116] The bias matrix is then normalized and truncated to ensure numerical stability and match the scale of the attention score, with the mean of the bias matrix set to be... Standard deviation is The stability constant is The lower limit and the upper limit are respectively and ,symbol If a function is used to crop the input to a specified interval, then the normalized bias is:
[0117] ;
[0118] in Represents the normalized and truncated bias value. and These represent the mean and standard deviation of the elements of the bias matrix, respectively. Used to avoid a denominator of zero. and Limit the bias range;
[0119] To match the size of the attention score matrix in the self-attention mechanism, let the bias matrix be... ,in The number of spatiotemporal markers is represented. Finally, the input dataset, canopy geometric channels, and normalized bias matrix are combined according to temporal synchronization and spatial registration to form the deposition estimation input. Let the input dataset be... Canopy geometric channels are ,symbol If the cascading operation is performed along a predetermined dimension, then the deposition estimation input is:
[0120] ;
[0121] in This represents the depositional estimation input used by the spatiotemporal attention network. This represents the multi-source input data set formed in step S1. Represents the geometric channels of the canopy, This represents the normalized and truncated attention bias matrix.
[0122] In this specific embodiment, S4 specifically refers to:
[0123] The system performs temporal and spatial alignment and spatiotemporal encoding on the sedimentation estimation input within a unified coordinate system. Elements of the bias matrix are injected into the self-attention score to form aerodynamic bias attention, thereby extracting spatiotemporal features more sensitive to sedimentation and drift discrimination while maintaining physical consistency. Specifically, let the... The and the first The query, key, and value vectors of each spatiotemporal marker are as follows: ;
[0124] in For attention scale, Let the value vector dimension be the element of the normalized attention bias matrix from step S3. Let the transpose operator be The square root operator is Then the attention score and weight are:
[0125] ;
[0126] in Indicates the first With the Marked score Indicates in Normalized weights of dimensions Indicates along the index The exponentially normalized mapping, let the total number of labels be... Let the summation operator be The weighted aggregation is then:
[0127] ;
[0128] in For the first Aggregation features of individual tags;
[0129] The aggregation results are then processed by a feedforward network and normalized before being converged in the spatiotemporal dimensions to obtain the deposition estimation output. Let the output vector be... ;
[0130] in The deposition index is the amount of powder that can fall per unit area. To cover uniformity indicators, This is an indicator of drift risk in non-target areas;
[0131] To quantify uncertainty, Monte Carlo sampling is used to obtain multiple output variances. Let the number of samplings be... , No. The next output is The variance operator is Then the uncertainty is:
[0132] ;
[0133] in Corresponding to Variance estimation;
[0134] During the training phase, robustness and generalization are improved by combining the objectives of supervised loss and quality conservation constraints, resulting in a total powder spray volume of [missing information]. The amount of sediment is Non-target drift amount is Suspension loss is The loss due to mass conservation is:
[0135] Supervision loss is Weights are The total loss is ;
[0136] in For absolute value operators, To train and optimize the objective function, The mean square error of the measured deposition or concentration label can be used without affecting the process of this step.
[0137] In this specific embodiment, S5 specifically includes:
[0138] Based on the system deposition index D, depositional uniformity index U, drift risk index R, and uncertainty components U_D, U_U, and U_R, and combined with the target weight vector and map, dimensional unification, normalization, and reward / penalty construction are performed. The normalization uses linear dimensionless normalization for subsequent weighting, allowing the original indices to be... The upper and lower bounds are referenced as follows: The stability constant is The normalization result is:
[0139] ;
[0140] in Indicates the dimensionless index, The original index to be normalized, Indicates the corresponding upper and lower reference bounds, This is used to avoid instability caused by an excessively small denominator;
[0141] Subsequently, a monotonic rule was used to obtain a reward for deposition uniformity and a penalty for non-target drift, making ,in Indicates a reward for uniform deposition. Indicates non-target drift penalty, and They represent respectively to and The normalized result, to estimate the energy consumption penalty per unit area, considers the dominant influence of flight speed and powder flow rate, and sets the UAV velocity vector as follows: Powder spraying flow rate is The coefficient is ,definition:
[0142] ;
[0143] in Indicates energy consumption penalty, Represents the Euclidean norm, Represents the velocity vector of the drone, Indicates powder spraying flow rate, This represents the weighting coefficient for speed and flow rate;
[0144] The task duration penalty is characterized by time efficiency, and the control cycle duration is set to... Reference duration is ,definition:
[0145] ;
[0146] in Indicates penalty for homework duration. Indicates the duration of the current period, To represent the normalized baseline, and to enhance conservatism regarding the relative distance and uncertainty between the flower's mechanical impact penalty and the canopy, the relative distance between the flower and the canopy is set to... The minimum safe distance is Uncertainty aggregation amount is The magnification factor is The soft penalty function is ,definition ;
[0147] in Indicates mechanical impact punishment on flowers, Represents the natural logarithm, Represents the natural constant, The independent variable representing the soft penalty function, Indicates the safety threshold, Indicates the current distance, The average value representing uncertainty Indicates the degree of uncertainty amplification;
[0148] In multi-objective combination, an adaptive amplification is introduced to increase the weight of safety-related components under high uncertainty, and the objective weight vector is set as follows: Element-wise product is Safety amplification factor is Sigmoid is Define the weighted vector as And let the multi-objective vector be The weighted multi-objective reward vector is obtained. ;
[0149] in This represents the multi-objective reward after adaptive amplification of weights and uncertainties. Indicates a large amount, Indicates slope, Indicates the uncertainty threshold, The independent variable representing the sigmoid function is... Indicates element-wise product, sign Indicates the arrangement of column vectors;
[0150] Furthermore, based on the sedimentation estimation output and uncertainty, canopy geometry and map, and combined with wind speed, direction, and dust injection rate, a set of differentiable safety constraints is constructed for subsequent optimization. The minimum safety distance constraint is represented using soft barriers. The maximum permissible wind speed constraint makes the environmental wind vector as follows: The upper limit is ,definition The no-spray zone constraint is based on a signed distance function of the map, whereby the current position or nozzle projection is... The signed distance from the no-spray zone to the point is (Positive outside, negative inside), definition The upper limit constraint on powder spraying flow rate is set to be [value]. ,definition To make the constraints more stringent under high uncertainty, the obstacle weights are adaptively amplified, and the baseline weights are set as follows:
[0151] and define Finally, the constraint set is obtained.
[0152] In this specific embodiment, S6 specifically refers to:
[0153] The system synchronizes and registers multi-source measurements and estimates in a unified coordinate system, and encodes them into job state vectors for input into the world model's multi-objective reinforcement learning module. Let the job state vector be... Indicates the current moment The status includes attitude rotation, position, relative distance to the canopy, ambient wind speed and direction, dust spraying flow rate and nozzle angle, map coding, and information such as deposition index, depositional homogeneity, drift risk and its uncertainty components, and symbols. express 3D real space, Representing state dimension, Representing discrete-time indexes, and employing compact motion parameterization consistent with the flight control interface, the motion vector is... Indicates at time The control inputs include components and symbols such as flight altitude setting, speed setting, roll angle, yaw angle, powder flow rate, and nozzle angle. express 3D real space, Representing the action dimension, in policy optimization, the system uses a world model with a DreamerV3 structure to make rolling predictions of the future and introduces constraints in the form of differentiable obstacle functions. The multi-objective reward vectors are weighted and converged to obtain a scalar reward, which is then set as follows: Indicates at time The overall benefits make the world model dynamics... Indicates parameters The state transition approximation is given by the policy as follows: Indicates parameters Given the conditional distribution, let the discount factor be... To represent time preference, let the predicted time domain length be... Let the constraint number be , order the The differentiable barrier of each constraint is Let the violation cost of a state-action pair be represented by an adaptive weight. Denotes the tightened constraint coefficients under high uncertainty, and let the Lagrange multipliers be... Indicates the first Let the penalty coefficients of each constraint be and let the expectation operator be . Let the summation operator be the expectation on the rolling trajectory induced by the world model and policy. The summation over time steps is represented by the Lagrange objective:
[0154] ;
[0155] in Indicates the initial state The expected optimization goal is as follows: Indicates the step index of the rolling prediction. Indicates the first Scalar reward of step Indicates the first Step 1 The consequences of violating these constraints;
[0156] Based on this objective, candidate action sequences are solved in each control cycle for flight control feedforward. Let the action sequence be... Indicates from time arrive Let the set of actions be such that the optimal sequence is... Let represent the optimized output candidate action sequence, then:
[0157] ;
[0158] in This represents an operator that maximizes the objective function;
[0159] To achieve adaptive control over the constraint strictness and maintain convergence of the feasible region, the Lagrange multipliers are updated using projected gradients, with the learning rate set to [value missing]. To indicate the update step size, let the projection operator be... If we project the input onto the non-negative region, then we have ;
[0160] in Indicates parameter update, This indicates the constraint cost sampling on the rolling trajectory. and They respectively represent general states and actions;
[0161] The final output includes a candidate sequence of actions, including flight altitude, speed, roll angle, yaw angle, powder flow rate, and nozzle angle, which is used for subsequent steps to refine diffusion actions and execute flight control.
[0162] In this specific embodiment, S7 specifically refers to:
[0163] The system uses candidate action sequences as initial trajectories and combines them with constraint sets, rinsing velocity fields, and particle concentration fields. Under a unified coordinate system, it employs diffusion action refinement through noise injection and denoising iterations to converge towards a better physically feasible region while adhering to differentiability constraints. The core denoising update uses a combined form of score-guided, physics-guided, and constraint-guided approaches, allowing the... The trajectory of the step is The previous trajectory was The length of the time horizon is Action dimension is Step size is The scoring network is (Parameters are) Returns the approximate logarithmic density gradient at the diffusion time index; the diffusion time index is... The physical guiding coefficient is The constraint guidance coefficient is Noise intensity is Gaussian noise is (With zero mean and unit covariance), the noise reduction update is as follows:
[0164] ;
[0165] in Represents the gradient over the entire action sequence. Representing constraint sets, symbols Indicates scaling and addition of the items in parentheses, and the symbol. Although it does not appear in this formula, it indicates the meaning of parameter update;
[0166] The physical guidance target focuses on high concentrations along the rolling trajectory, aligns with the canopy normal, and also considers low-porosity regions, making the trajectory as follows: Scrolling state (Originated from differentiable kinematics propagation), with weights of Particle concentration field is The downward washing velocity vector is The unit streamline direction is The canopy normal is The porosity of the canopy is The dot product symbol is Euclidean norm is Absolute value Then we have:
[0167] and ;
[0168] in Represents the weighted convergence of physically guided targets, This represents the weighting coefficient determined by both streamline alignment and permeability.
[0169] The constraint cost is given in the form of a time-aggregated differentiable barrier, let the number of constraints be... , No. The barrier function of each constraint is (Including minimum safety distance, maximum wind speed, no-spray zones, and powder spraying limits, etc.), its adaptive weight is The scrolling state that is consistent with the action is ,definition:
[0170] ;
[0171] in This represents the total cost of constraint violations across the entire sequence;
[0172] During the iteration process, early stopping is implemented using a dual criterion of tolerance and number of steps, and repeated sampling is used to improve the coverage of feasible solutions. Finally, the refined result is selected based on the criterion of minimizing constraint cost, with the number of samples set to [value missing]. , No. The samples after denoising are The optimal sample is The minimization operator is Then there is ;
[0173] in This indicates the output of a refined sequence of actions for flight control execution. Indicates sample index, Indicates the number of sample records, This represents the operator that minimizes the objective function.
[0174] In this specific embodiment, S8 specifically refers to:
[0175] The system implements "gating-tracking-deployment" on refined actions and constraint sets in a unified coordinate system. First, it performs constraint gating on the refined actions to obtain safe and executable actions, and automatically reverts when soft obstacles are triggered, making the refined actions after diffusion... Safe alternative actions are The gate control indicator is Indicator functions are The current status is , No. The differentiable barrier of each constraint is The number of constraints is Then the gating and substitution are:
[0176] and ;
[0177] in Indicates the action after gating, This represents the maximum value among all constraint indices;
[0178] The desired output is then generated from the gated action and superimposed with the measurement-constructed tracking error by feedforward-proportional-derivative control to form a joint command, with the discrete-time index set to... The desired output vector is (Depend on Mapped (obtained), measurement status is Output mapping is The error is Feedforward mapping is , proportional gain is The differential gain is The discrete derivative of the error is The error of the previous cycle is The control cycle is Joint command is The following is given:
[0179] and ;
[0180] in This represents the selection mapping from state to the tracked channel (height, velocity, attitude, powder flow rate, etc.). and Gain matrix grouped by channel, This is the joint command to be sent to the flight controller;
[0181] To ensure that powder coating quality and compliance are adaptively maintained within a closed loop, the powder coating channel employs online scheduling with a deposition-uniformity objective, allowing the refining flow rate to be... The flow command is The sedimentation target is The goal of uniformity is Current sedimentary estimates are The current homogeneity estimate is The feedback coefficient is The maximum traffic limit is The clipping function is Indicates scalar Clip to range Then we have:
[0182] ;
[0183] in Replace The amount of powder sprayed in the solution must meet the target and upper limit constraints.
[0184] Apart from the key calculations mentioned above, the remaining steps are completed in text form, including limiting and saturating the rate of each channel to match the actuator, performing redundancy checks and readback confirmations at the communication layer, linking height increase and deceleration and nozzle return to safety actions when the gate is triggered, and summarizing and sending back the wind speed and direction, pose error, flow command, deposition and uniformity estimation and constraint trigger records within the sliding window for subsequent model updates and parameter self-tuning.
[0185] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0186] This application establishes a closed-loop link from online deposition and drift estimation to constraint compliance action generation and execution feedback through a combination of "multi-source sensing and geometric enhancement - Fourier neural operator local aerodynamics and particle migration field - spatiotemporal attention network with aerodynamic bias attention - multi-objective reward and differentiable obstacle constraint - world model reinforcement learning - diffusion action refinement - flight control execution and closed-loop data acquisition". It outputs deposition index, deposition homogeneity, drift risk and its uncertainty in real time under a unified coordinate system, so that safety-related decisions have a reliable quantitative basis. In the strategy optimization stage, target weights and Lagrange multipliers are used to adaptively coordinate deposition homogeneity, non-target drift, energy consumption, operation time and flower mechanical impact. In the generation stage, local aerodynamics and particle migration field and differentiable obstacle function are jointly used to guide diffusion denoising, so that candidate actions converge to the feasible region and minimize constraint costs. Combined with gating-tracking-distribution and data acquisition, real-time stable control is achieved in the face of dynamic wind fields and complex canopy geometry. This achieves the technical effects of improving deposition uniformity, reducing non-target drift, suppressing mechanical impact and energy consumption of the flowers, and maintaining closed-loop operation compliance under the constraints of minimum safe distance, maximum allowable wind speed, no-spraying areas, and upper limit of powder spraying flow rate.
[0187] To address the technical issues, this case makes several targeted improvements to the algorithm structure: Fourier neural operators are used within the local computational domain, and continuity and mass conservation constraints are applied during the training phase to ensure physical consistency and online usability of aerodynamic and particle field estimations; "Aerodynamic bias attention" is introduced at the deposition estimation end, constructing an attention bias matrix using the downwash streamline direction, canopy normal, porosity, and particle concentration, enabling spatiotemporal feature extraction to focus along the real migration channel, improving the accuracy and robustness of deposition and drift discrimination; At the optimization end, differentiable barrier functions are used to transform the minimum safe distance, maximum wind speed, no-spray zone, and dust spraying upper limit into continuously optimizable constraints, which are adaptively tightened according to uncertainty, combined with DreamerV3 rolling prediction and online Lagrange multiplier weighting, to dynamically balance the strategy between multi-objective benefits and constraint compliance; At the generation end, diffusion action refinement is adopted, directly injecting the physical guiding terms and constraint gradients of the local aerodynamic and particle migration fields into the denoising update, supplemented by early stopping and minimum constraint cost selection, significantly improving the feasibility and field adaptability of the actions.
[0188] The aforementioned structural improvements, combined with the closed-loop design, enable this project to achieve the expected technical effects more efficiently and robustly under conditions of uncertainty and strong coupling.
Claims
1. A low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm, characterized in that, Includes the following steps: S1. Collect multi-source sensor data and form an input data set. Then, perform geometric enhancement processing on the input data set to obtain the canopy geometric channel. S2. Input the input data set and the canopy geometry channel into the Fourier neural operator model, and output the local aerodynamic and particle migration fields; S3. The local aerodynamic and particle migration fields are mapped into an attention bias matrix through the bias construction module, and the input data set, canopy geometric channels, and attention bias matrix are combined to form the deposition estimation input; S4. Input the deposition estimate into the spatiotemporal attention network and output the deposition estimate output and uncertainty index. S5. Calculate the multi-objective reward vector by combining the sedimentation estimation output, uncertainty index, and target weight vector in the input dataset through the multi-objective construction module. Construct a constraint set by combining the sedimentation estimation output, uncertainty index, canopy geometric channel, field boundaries, and no-spraying zone map in the input dataset through the constraint generation module. S6. Combine the input data set, sedimentation estimation output and uncertainty index to form a job state vector. Input the job state vector, multi-objective reward vector and constraint set into the world model multi-objective reinforcement learning module to output candidate action sequence. S7. Input the candidate action sequence and constraint set into the diffusion action refinement module. The diffusion action refinement module guides diffusion sampling with local aerodynamics and particle migration field, and performs constraint guidance and noise reduction with the differentiable barrier function corresponding to the constraint set, and outputs the refined action sequence. S8. Input the refined action sequence into the flight control execution module for execution, collect the multi-source sensor data after execution to obtain data set two, and use data set two as the input data set for the next control cycle; S1 includes: The system collects camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle, field boundary and no-spraying zone maps, and target weight vectors. It also calibrates, synchronizes, and spatially registers each channel to a unified coordinate system. Distortion correction and denoising are performed on camera images, and canopy depth representation is established by combining the relative distance to the canopy and flight attitude and position information; The canopy normal and canopy porosity are calculated based on canopy depth representation and camera images, and the flower position probability is obtained based on flower organ detection and region statistics from camera images. The input data set is formed by summarizing camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle, field boundary and no-spraying zone map, and target weight vector, and the canopy geometric channel is constructed by canopy normal, canopy porosity and flower position probability. S2 includes: Under a unified coordinate system, the local computational domain is determined based on the relative distance between the flight attitude and position information in the input dataset and the canopy, combined with the canopy geometric channel. The input data set and the canopy geometric channel are used to construct the conditional input of the Fourier neural operator model through feature encoding. The conditional input includes boundary conditions and source terms. The boundary conditions include at least the far-field boundary of the ambient wind speed and direction, the near-field boundary of the rotor-induced airflow, and the canopy boundary characteristics determined by the canopy normal and canopy porosity. The source terms include at least the particle injection parameters determined by the powder flow rate and nozzle angle. Within the current control cycle, the spatial coordinates within the local computational domain are mapped by Fourier features and input together with the conditional input into the Fourier neural operator model to obtain the local aerodynamic and particle migration fields, and output the rinsing velocity field and particle concentration field with a preset spatial resolution. The Fourier neural operator model employs supervised learning with measured wind speed and particle concentration data during the training phase, and applies continuity constraints and mass conservation constraints.
2. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S3 specifically refers to: The downwash velocity field and particle concentration field in the local aerodynamic and particle migration field are processed by the bias construction module. The streamline direction field is determined based on the downwash velocity field, and the weights along the streamline direction and the canopy normal and canopy porosity in the canopy geometric channel are combined. Based on the particle concentration field and the weights, the attention bias value is calculated for the spatiotemporal index pairs of the input data set and the canopy geometric channel. An attention bias matrix with the same size as the attention score matrix of the self-attention mechanism in the spatiotemporal attention network is generated. The attention bias matrix is then normalized and truncated to limit the range of bias values and ensure numerical stability. The input data set, canopy geometry channel, and attention bias matrix are combined in a time-synchronized and spatially registered manner to form the deposition estimation input. This allows the attention bias matrix to implement aerodynamic bias attention in the self-attention mechanism of the spatiotemporal attention network, with weighted focusing along the streamline direction and the canopy normal.
3. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S4 specifically refers to: Under a unified coordinate system, the deposition estimation input is aligned temporally and spatially. A spatiotemporal attention network is used to encode the deposition estimation input temporally and spatially to form a spatiotemporal label sequence. During the calculation of attention score in the self-attention mechanism, an attention bias matrix is introduced to implement aerodynamic bias attention. The spatiotemporal label sequence is weighted and aggregated and then subjected to nonlinear transformation and normalization by a feedforward network. The deposition estimation output and uncertainty index are output. The deposition estimation output includes a deposition index to characterize the amount of pollen that can fall per unit area, a deposition uniformity index to characterize the uniformity of cover, and a drift risk index to characterize the risk of drifting to non-target areas. The uncertainty index is obtained by calculating the variance or confidence interval of the output distribution. During the training phase, a loss function is constructed and parameters are learned by considering the mass conservation constraints between the total amount of powder sprayed, the amount of deposition, the amount of non-target drift, and the suspension loss.
4. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S5 specifically refers to: The sedimentation index, sedimentation homogeneity index, drift risk index, and uncertainty index in the sedimentation estimation output are unified and normalized in terms of dimensions, and sedimentation homogeneity reward and non-target drift penalty are obtained according to the preset monotonic mapping rule. The energy consumption per unit area is estimated and the energy consumption penalty is obtained based on the powder spraying flow rate, flight attitude and position information in the input data set. The operation time penalty is calculated based on the timestamp and flight path progress in the input data set. The mechanical impact penalty of the flower is constructed based on the relative distance to the canopy and the uncertainty index. The deposition uniformity reward, non-target drift penalty, energy consumption penalty, operation time penalty and flower mechanical impact penalty are weighted and combined according to the target weight vector to form a multi-target reward vector, and the weight of safety-related components is increased when the uncertainty index exceeds the threshold. Meanwhile, based on the sedimentation estimation output and uncertainty index, the canopy geometric channel and the field boundary and no-spraying area map in the input data set, and combined with the wind speed, wind direction and dust spraying flow rate in the input data set, minimum safe distance constraint, maximum allowable wind speed constraint, no-spraying area constraint and dust spraying flow rate upper limit constraint are constructed respectively. Each constraint is expressed in the form of a differentiable barrier function, and the constraint weights are adaptively adjusted according to the uncertainty index to obtain the constraint set.
5. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S6 specifically refers to: The input data set, sedimentation estimation output, and uncertainty index are synchronized in time and spatially registered and summarized into an operation state vector. The operation state vector includes at least flight attitude and position information, relative distance to the canopy, wind speed and direction, dust spraying flow rate and nozzle angle, field boundary and no-spraying area map, sedimentation index, sedimentation uniformity index, drift risk index, and uncertainty index. The job state vector, multi-objective reward vector, and constraint set are input together into the world model multi-objective reinforcement learning module. The world model multi-objective reinforcement learning module adopts the DreamerV3 structure. It performs rolling prediction based on the job state vector to evaluate the state transition and the expectation of the multi-objective reward vector within a preset time window. In the strategy optimization, the constraint terms are constructed using the differentiable obstacle function corresponding to the constraint set, and the constraint weights are adaptively adjusted using the Lagrange multiplier method. This ensures that the strategy maximizes the multi-objective reward weighted by the objective weight vector while satisfying the minimum safe distance, maximum allowable wind speed, no-spray zone, and upper limit of powder spraying flow rate. Output candidate action sequences, which include flight altitude, flight speed, roll angle, yaw angle, powder flow rate, and nozzle angle.
6. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S7 specifically refers to: The candidate action sequence, constraint set, and local aerodynamic and particle migration field are input into the diffusion action refinement module. The candidate action sequence is used as the initial trajectory for the diffusion denoising process for noise injection and denoising iteration. In each round of denoising iteration, the action guidance term is calculated based on the local aerodynamic and particle migration field and superimposed on the denoising update. At the same time, the gradient of the differentiable barrier function corresponding to the constraint set with respect to the current action sequence is superimposed on the denoising update as a constraint guidance term, so that the generated action sequence converges toward the feasible region that satisfies the constraint set. The iteration stops when the constraint set is satisfied or the preset number of iterations is reached, and a set of action samples that satisfy the constraint set is obtained. Then, one action sequence is selected as the refined action sequence according to the criterion of minimizing the constraint cost calculated by the differentiable obstacle function.
7. The low-altitude pollination control method based on a closed-loop multi-objective reinforcement optimization algorithm according to claim 1, characterized in that, S8 specifically refers to: The refined action sequence is input into the flight control execution module in the order of timestamps. The flight control execution module converts the flight altitude, flight speed, roll angle, yaw angle, powder flow rate and nozzle angle in the refined action sequence into the corresponding actuator set values and issues control commands to drive the UAV attitude control, trajectory control and powder spraying system to execute. During execution, real-time operational data is recorded within a unified coordinate system, and measurements of camera images, optical scattering signals, electrostatic charge signals, wind speed and direction, relative distance to the canopy, flight attitude and position information, powder spraying flow rate and nozzle angle are collected. Time synchronization, spatial registration and necessary calibration are completed, and the data are summarized to form Data Set Two. Data set two is used as the input data set for the next control cycle to start the next control cycle, thereby achieving closed-loop control.