Agricultural unmanned aerial vehicle terrain following method and device based on vision
By combining the adversarial training GAN network and depth estimation model with genetic algorithm to optimize trajectory planning, the security threat caused by uneven terrain of agricultural drone in hilly and mountainous areas is solved, and efficient and low-cost agricultural drone terrain follow-up flight is achieved.
Patent Information
- Application Number
- CN202510374162.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
AI Technical Summary
In the hilly and mountainous environment, agricultural drones face flight safety threats and degraded operating quality during operation. The existing technology has high cost, high weight, high energy consumption, and insufficient terrain data, which affects the drone's response speed and operating efficiency.
The agricultural drone is used to obtain infrared and visible light images for offline confrontation training of GAN network, the generator is integrated to enhance terrain information, and the crop canopy distance estimation model is used to estimate crop canopy distances, and the trajectory planning is optimized through genetic algorithms to construct an agricultural drone terrain following method based on airborne monocular vision.
The agricultural drone terrain is achieved in hilly and mountainous areas, improving operation quality and safety, reducing costs and weight, and improving reaction speed and operation efficiency.
Smart Images

Figure CN120255536A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent agriculture, and is a method and device for realizing terrain following technology of agricultural drones based on vision, which can be applied to agricultural fields such as crop monitoring, pesticide spraying, fertilization, etc., and is particularly suitable for agricultural operations in complex terrains such as hills and mountains. Background Art
[0002] As the development direction of modern agriculture, smart agriculture has gradually emerged globally and shown great potential in improving agricultural production efficiency, saving resources, reducing environmental impacts, etc. Promoting the research and development and application of intelligent agricultural machinery and equipment such as drones is an important technical support for realizing smart agriculture. At present, agricultural drones have been widely used in agricultural operations such as land and soil analysis, aerial seeding, spraying operations, crop monitoring, agricultural irrigation, and crop health assessment. However, the agricultural environment in China is special, and about 70% of the country's territory is hilly and mountainous areas. In the hilly and mountainous environment, due to the differences in crop growth conditions and the uneven terrain, agricultural drones face major challenges during operation. An inappropriate flight height may damage farmland or cause unnecessary losses, resulting in a decline in the quality of agricultural operations and posing a threat to the flight safety of agricultural drones. At present, many scholars have devoted a lot of research to realizing terrain perception through devices such as lidar (LiDAR) to achieve the stable flight of agricultural drones, but the corresponding cost is higher, the weight is heavier, and the energy consumption is greater. At the same time, researchers have also developed a variety of path planning algorithms, such as three-dimensional path planning based on terrain data, adaptive path optimization, etc., enabling drones to safely and effectively complete tasks in an environment with drastic terrain changes. However, in the case of inaccurate terrain data, the effect of the algorithm will be greatly reduced, and the high computational load may increase the processing time, affecting the reaction speed and operation efficiency of the drone. In contrast, using a monocular vision sensor and combining deep learning technology to accurately estimate terrain height information from a monocular image and taking into account the crop canopy during the depth estimation process not only has a lower cost, lighter weight, but also more accurate data, which helps the informatization and intelligent development of agricultural machinery and equipment in China. Summary of the Invention
[0003] The present invention provides a vision-based terrain following method and device for agricultural drones. First, a large number of infrared and visible light images obtained by the agricultural drone are used for offline adversarial training of the GAN network. Secondly, the pre-trained GAN generator is used to fuse the infrared and visible light images obtained by the agricultural drone in real time to enhance the terrain information. Then, a method for estimating the distance of the crop canopy using a depth estimation model for monocular vision is provided. By taking pictures with a single camera and using the enhanced terrain information image fused by the pre-trained GAN generator, the estimation of the distance of the crop canopy is realized. The crop canopy depth map D obtained based on onboard monocular vision p As a constraint of the objective function, considering other constraint conditions during the flight of the agricultural drone, an objective function is constructed to optimize the trajectory planning to cope with multiple constraints. The genetic algorithm in the optimization algorithm is selected to optimize the objective function J, and the optimized curve is used as the flight path of the drone to achieve terrain following flight.
[0004] To achieve the above object, the technical solutions provided by the embodiments of the present invention are as follows:
[0005] A vision-based terrain following method for agricultural drones, the method comprising:
[0006] S1. Using a large number of infrared and visible light images obtained by the agricultural drone for offline adversarial training of the GAN network;
[0007] S2. Using the pre-trained GAN generator to fuse the infrared and visible light images obtained by the agricultural drone in real time to enhance the terrain information;
[0008] S3. Providing a method for estimating the distance of the crop canopy using a depth estimation model for monocular vision, and realizing the estimation of the distance of the crop canopy by taking pictures with a single camera and using the enhanced terrain information image fused by the pre-trained GAN generator;
[0009] S4. Using the crop canopy depth map D obtained based on onboard monocular vision p As a constraint of the objective function, considering other constraint conditions during the flight of the agricultural drone, an objective function is constructed to optimize the trajectory planning to cope with multiple constraints;
[0010] S5. Selecting the genetic algorithm in the optimization algorithm to optimize the objective function H, and using the optimized curve as the flight path of the drone to achieve terrain following flight;
[0011] Further, the step S1 is specifically:
[0012] S11. Use the visual sensor carried by an agricultural drone to obtain visible light images and infrared images of hilly and mountainous areas. Conduct multiple flights under different lighting conditions and time periods to ensure the diversity and comprehensiveness of the data, and define the infrared image I ir ., and the visible light image I vis ;
[0013] S12. Normalize the obtained infrared images and visible light images to make them suitable for subsequent deep learning processing. Specifically, it is formula (1)
[0014]
[0015] I is the original image, and I norm is the normalized image, and I min ., I max are the minimum and maximum pixel values of the image respectively;
[0016] S13. Input the processed infrared images and visible light images into the generator G, use the convolutional neural network (CNN) architecture, and extract features layer by layer;
[0017] S14. Insert the Squeeze-and-Excitation (SE) module after each convolutional layer to enhance the attention to important features;
[0018] S15. Perform the Squeeze operation in the SE module, conduct global average pooling on the features output by each convolutional layer to generate the global feature vector Z. Specifically, it is formula (2):
[0019]
[0020] where Z C is the global feature vector of channel C, Z represents the global feature vector obtained through the global average pooling operation, with a dimension of C, and C represents the number of channels of the feature map; x i,j,c is the value of the input feature map at position (i, j) and channel C, x represents the input feature map, with a dimension of H×W×C, i and j represent the spatial positions of the feature map, and H and W respectively represent the height and width of the feature map;
[0021] Perform the Excitation operation in the SE module. Through two fully connected layers and activation functions (ReLU and Sigmoid), generate the feature weight S. First, pass through the first fully connected layer and the Relu activation function to generate the output vector e of the first fully connected layer, which is formula (3):
[0022] e = ζ(W1·Z) (3)
[0023] Then, it passes through the second fully-connected layer and the Sigmoid activation function to generate the output vector S of the second fully-connected layer, which is specifically represented by Equation (4):
[0024] S = σ(W2 · e) (4)
[0025] where W1 and W2 respectively represent the weight matrices of the first and second fully-connected layers, ζ is the Relu activation function, and σ is the Sigmoid activation function;
[0026] Finally, the weight s of the Excitation operation is reapplied to the original feature map for recalibration of the feature map, which is specifically represented by Equation (5):
[0027]
[0028] is the reweighted feature map, represents the recalibrated feature map, S c is the weight of channel C;
[0029] S16. After the SE module operation inserted in each convolutional layer, a fused image with enhanced crop canopy information is generated; the fused image and the real image are input into the discriminator D, and multiple convolutional layers are used for layer-by-layer feature extraction, and binary classification is performed through the fully-connected layer to determine whether it is real or generated, and the authenticity of the image is reflected according to the real score (between 0 and 1) of the output image;
[0030] S17. Use the processed infrared image and visible light image datasets as training data to train the generative adversarial network (GAN). During the training process, the cross-entropy loss function is used to measure the performance of the generator and the discriminator. The loss function of the generator is defined as Equation (6):
[0031]
[0032] To enhance the enhancement of crop canopy information in the generated image, a feature enhancement loss can be additionally added to the generator loss. Suppose there is a feature extraction network F, and its feature loss intensity is defined as Equation (7):
[0033]
[0034] where, x ∼ p data (x) represents that the sample x is drawn from the real data distribution p data (x); z ∼ p z (z) represents that the random noise vector is drawn from the noise distribution p z (z);
[0035] Combined with the cross - entropy loss and the feature enhancement loss, the total loss function of the generator is given by Equation (8):
[0036]
[0037] Using the cross - entropy loss function to measure the discriminator's ability to distinguish between real images and generated images, the loss function of the discriminator can be defined as Equation (9):
[0038]
[0039] Where \(G\) is the generator that generates images according to the input noise \(z\); \(D\) is the discriminator that discriminates whether the input image is real or generated; \(E\) represents the expectation; \(x\) is the real image; \(G(z)\) is the generated image, \(F\) is the feature extraction network that outputs the feature representation of the image; represents the L2 norm; \(\lambda\) is the balance weight parameter; \(p\) z (z) represents the noise distribution; \(p\) data (x) represents the real data distribution;
[0040] S18. During the training process, fix the parameters of the discriminator, train the generator, then fix the parameters of the generator and train the discriminator, and alternate the training. Repeat the above steps until the generated images are both real and contain enhanced terrain information;
[0041] S19. After completing the training, save the weights of the generator for later use in online processing;
[0042] Further explanation: The specific steps of step S2 are as follows:
[0043] S21. Use the visual sensors carried by the agricultural drone to obtain infrared and visible light images in real - time, and define the infrared image as \(I\) ir and the visible light image as \(I\) vis ;
[0044] S22. Use the pre - trained GAN generator to fuse the real - time obtained infrared and visible light images to generate an image \(I\) F containing enhanced crop canopy information;
[0045] Further explanation: The specific steps of step S3 are as follows:
[0046] S31. Use the agricultural drone equipped with visual sensors to collect infrared images and visible light images in different scenarios such as farmland, hills, and plains. At the same time, use the depth sensor LiDAR to synchronously collect depth information in different scenarios to form a labeled data set, denoted as where \(x\) i is the \(i\) - th image in the labeled data set, and \(d\) iis the corresponding depth annotation, and M is the total number of samples in the annotation dataset; at the same time, a large number of unannotated agricultural UAV image data are collected, denoted as and N is the total number of samples in the unannotated dataset.
[0047] S32. Use the annotation dataset D l to train the initial monocular depth estimation model (teacher model T′), and use the affine invariant loss function for training to handle the problems of depth scale and offset between different datasets; the affine invariant loss function is shown in formula (10):
[0048]
[0049] where H and W are the height and width of the on-board monocular image respectively, d i are the true crop canopy depth and the predicted crop canopy depth of the i-th pixel respectively, is the loss function between the predicted crop canopy depth value and the true crop canopy depth value, are the normalized true crop canopy depth value and the predicted crop canopy depth value respectively, median(d) represents the median of the depth value d, t(d) is the median of the crop canopy depth map d, and s(d) is the average of the absolute deviations of the crop canopy depth map d;
[0050] S33. Use the trained teacher model T′ to predict the crop canopy depth of the unannotated dataset D u to generate a pseudo-annotation dataset, denoted as T′(u i ) represents the crop canopy depth prediction map of the teacher model T′ for the unannotated on-board monocular image u i ;
[0051] S34. Combine the annotation dataset D l and the pseudo-annotation dataset to form a comprehensive training set, and train the student model S on the comprehensive training set. During the training process, strong perturbations, including color distortion and spatial distortion, such as CutMix, are injected into the unannotated on-board monocular images, so that the student model S actively seeks additional visual knowledge and obtains invariant representations from the unlabeled on-board monocular images; first, insert a pair of random unlabeled on-board monocular images u a and u b spatially, specifically represented by formula (11):
[0052] u ab = u a ⊙M + u b ⊙(1 - M) (11)
[0053] where ua , u b are a pair of randomly unlabeled airborne monocular images, M is a binary mask with the rectangle area set to 1, and u ab represents the fused image:
[0054] Unlabeled loss is obtained by calculating the affine invariant loss in the effective regions defined by M and 1 - M, specifically represented by formula (12):
[0055]
[0056] where ⊙ represents element - wise multiplication, that is, the corresponding elements of two matrices are multiplied;
[0057] The two losses are added through weighted average to obtain the unlabeled loss specifically represented by formula (13)
[0058]
[0059] where S is the student model, ρ is the loss function, ∑M and ∑(1 - M) are the sums of the masks;
[0060] S35. In the depth estimation task, use the semantic segmentation model DINOv2 as auxiliary supervision, adopt the feature alignment loss to ensure that the student model S retains more semantic information in the pre - trained model; set a tolerance margin to allow flexibility in feature alignment, focusing on the most relevant semantic features; the final loss function is the affine invariant loss function Unlabeled loss Feature alignment loss
[0061] The feature alignment loss is formula (14):
[0062]
[0063] where f i and f′ i are the features of the student model and the DINOv2 model respectively;
[0064] The final loss function is represented by formula (15):
[0065]
[0066] S36. Input the feature map generated by the DINOv2 encoder into the DPT decoder to generate an airborne monocular crop canopy depth map. The DPT decoder first reorganizes the feature map into a multi-scale image representation and then gradually fuses these representations to generate a dense depth prediction. The feature reorganization operation includes three steps: the Read operation, the Concatenate operation, and the Resample operation.
[0067] The Read operation processes the feature map containing the read-out tokens, converting it into a form suitable for subsequent processing, including ignoring the read-out tokens, directly discarding the read-out tokens, and only retaining the tokens. A token refers to a small patch extracted from the image, and each patch is converted into a feature vector.
[0068]
[0069] Readi gnore Read(t) represents the reading result obtained by ignoring certain information from the input t under the Read operation.
[0070] Add the read-out token, adding the value of the read-out token to each image token.
[0071]
[0072] Read add Read(t) represents the set of results generated by adding each time step of the read operation to another specific value t0.
[0073] Project the read-out token: Concatenate the read-out token with each image token and then project it through a multi-layer perceptron (MLP).
[0074]
[0075] Read proj Read(t) represents processing and projecting each time step of the input and generating the processed feature representation.
[0076] Among them, t0 is the read-out token, t1 to are other tokens, MLP represents the multi-layer perceptron, and cat represents the concatenation operation.
[0077] Reorganize the tokens into a two-dimensional feature map according to their positions in the original image, and t reshaped represents the feature map arranged by spatial position. Use the Resample operation to perform spatial resampling on the reorganized feature map and use convolution operations to adjust the resolution of the feature map to match the resolution required for subsequent processing.
[0078] Resample s (t) = Conv3x3(Conv1x1(t))
[0079] Resample s (t) represents a resampling process that includes two convolution operations; Conv1x1(t) represents applying a 1x1 convolution (Conv1x1) operation to the input feature map t.
[0080] Use a residual convolution unit to perform a convolution operation on the feature map and maintain the resolution of the feature map; gradually increase the resolution of the feature map through an upsampling operation, so that the finally generated feature map gradually approaches the resolution of the input image;
[0081] The output head receives the high-resolution feature map from the feature fusion module of the DPT decoder, further through multiple convolutional layers and upsampling operations, the resolution is increased to the same as that of the input image; finally, a 1x1 convolutional layer is used to generate the final absolute depth map D of the crop canopy based on airborne monocular vision p , depth map D p Each value d(t) in represents the distance from the agricultural UAV to the crop canopy at time t;
[0082] x conv = ConvLayer(x input )
[0083] x output = Upsample(x conv )
[0084] D p = Conv1x1(x output )
[0085] Among them, x conv represents the feature map processed by the convolutional layer; ConvLayer(x input ) represents using the convolutional layer ConvLayer on the input feature map x input ; x output represents the upsampled feature map, and Upsample(x conv ) represents performing the upsampling operation Upsample on the convolutional output feature map x conv ; D p represents the final output feature map, and Conv1x1(x output ) represents applying the 1x1 convolution Conv1x1 to the upsampled feature map x output .
[0086] Furthermore, the specific step S4 is as follows:
[0087] S41. Considering the mobility of the UAV, the growth of crops, and the terrain undulation comprehensively, establish the constraint conditions;
[0088] For the mobility constraint, consider the speed constraint and acceleration constraint of the UAV. The acceleration constraint and speed constraint of the UAV are represented by formulas (16) and (17) respectively:
[0089] ||a(t)|| ≤ a max (16)
[0090] ||v(t)|| ≤ v max (17)
[0091] That is, at any time t, the magnitude of the acceleration and speed of the UAV cannot exceed the predetermined maximum acceleration a max and the predetermined maximum speed v max ; where ||a(t)|| represents the magnitude of the acceleration of the UAV at time t, that is, the norm of the acceleration vector a(t); ||v(t)|| represents the magnitude of the speed of the UAV at time t, that is, the norm of the velocity vector v(t).
[0092] Based on the crop canopy depth map D obtained by the on-board monocular vision p Each value d(t) in it represents the distance from the agricultural UAV to the crop canopy at time t; the flight altitude of the UAV is z u (t), then the crop height is represented by formula (18):
[0093] z crop (t) = z u (t) - d(t) (18)
[0094] It is necessary to ensure that the flight altitude z u (t) of the UAV is higher than the crop height z crop (t) plus the minimum safety distance h min , so:
[0095] z u (t) ≥ z crop (t) + h min
[0096] z u (t) ≥ z u (t) - d(t) + h min
[0097] d(t) ≥ h min
[0098] where h min is the minimum safety distance between the UAV and the crop;
[0099] In a mountainous environment, the terrain has a large undulation. Therefore, the flight trajectory of the UAV needs to consider the constraints brought by the terrain undulation, which is specifically expressed by formula (19):
[0100]
[0101] Among them, is the magnitude of the gradient vector, representing the magnitude of the terrain slope; θ max is a pre-set maximum slope threshold, representing the maximum terrain slope at which the UAV can fly safely;
[0102] S42. Analyze the coupling characteristics of multiple constraints and construct a penalty function to handle relatively complex and non-linear constraints such as terrain undulation and crop canopy height:
[0103] The speed constraint is constraint condition ①. When the speed exceeds the maximum allowable speed v max , a speed penalty function C1(r(t)) is generated:
[0104] C1(r(t)) = max(0, ||v(t)|| - v max ) ①
[0105] If the speed ||v(t)|| exceeds v max , the value of the penalty function is ||v(t)|| - v max , otherwise it is 0; the same applies to other constraints;
[0106] The acceleration constraint is constraint condition ②. When the acceleration exceeds the maximum allowable acceleration a max , an acceleration penalty function C2(r(t)) is generated:
[0107] C2(r(t)) = max(0, ||a(t)|| - a max ) ②
[0108] The crop canopy height constraint is constraint condition ③. When the UAV height is lower than the sum of the crop canopy height and the minimum safety distance z crop (t) + h min , a crop canopy height penalty function C3(r(t)) is generated:
[0109] C3(r(t)) = max(0, z crop + h min - z u (t)) ③
[0110] The terrain undulation constraint is constraint condition ④. When the magnitude of the terrain slope exceeds the pre-set maximum slope threshold θ max , a terrain undulation penalty function C4(r(t)) is generated:
[0111]
[0112] Combining the above penalty terms together, a comprehensive penalty function Φ(r(t)) is formed and represented by formula (20):
[0113] Φ(r(t)) = λ1C1(r(t)) + λ2C2(r(t)) + λ3C3(r(t)) + λ4C4(r(t))
[0114] That is:
[0115] C i (r(t)) represents the penalty function of the i-th constraint condition, and λ i is the penalty coefficient of the i-th constraint condition;
[0116] S43. By introducing a cubic Bézier curve, the flight trajectory of an agricultural drone is established. The flight requirements and environmental constraints of the drone are fully considered, and an objective function is constructed to optimize the trajectory planning;
[0117] Based on the mission requirements of the agricultural drone, four control points of the cubic Bézier curve are set, namely P0, P1, P2, and P3; P0 is the take-off location of the drone, and P1 is a point slightly deviating from the crop row, which can be above the crop canopy to ensure that the drone has sufficient height during operation. P2 can be set between the crop rows to adjust the curvature and make the flight path of the drone smoother. P3 is selected as the target location of the mission to ensure that the drone operates in this area. For the trajectory planning of the drone, the cubic Bézier curve can represent the position information of the drone at different time points;
[0118] The equation of the cubic Bézier curve is represented by formula (21):
[0119] B(t ′ ) = (1 - t ′ ) 3 p0 + 3t ′ (1 - t ′ ) 2 P1 + 3t ′2 (1 - t ′ )P2 + t ′3 P3 (21)
[0120] Among them, P0 and P3 are the starting position and the final target position of the drone respectively; P1 and P2 are the initially selected intermediate control points; the parameter t ′ is used to represent the position of points on the curve and is a parameter from 0 to 1;
[0121] By discretizing t ′ , a series of trajectory points B(t′ );Take the trajectory point B(t ′ ) as the UAV flight trajectory r(t); Considering the smoothness, energy consumption and various constraints of the UAV flight trajectory, a multi-constraint trajectory objective optimization function is constructed, which is specifically represented by formula (22):
[0122] J = ∫(w1||r″(t)|| 2 + w2||r′(t)|| 2 + w3Φ(r(t)))dt (22)
[0123] That is:
[0124] J = ∫0 1 (w1||B″(t′)|| 2 + w2||B′(t′)|| 2 + w3Φ(B(t′)))dt′
[0125] where t′ is a parameter from 0 to 1; w1, w2, and w3 are weight coefficients;
[0126] The constraint condition is formula (23):
[0127]
[0128] Further explanation, the specific steps of step S5 are as follows:
[0129] Encoding: Encode the positions of the control points P1, P2, and the end point P3 into chromosomes;
[0130] Randomly generate an initial population: Randomly generate a group of initial populations, and each individual represents a different combination of control points and end points;
[0131] Fitness function: Calculate the fitness value of each chromosome, that is, calculate the value of the objective function J. The lower the fitness value, the better the path of the individual (i.e., the combination of control points and end points);
[0132] Perform selection operation: Use the roulette wheel selection method to select better individuals for mating according to the fitness value;
[0133] Perform crossover operation: Perform a crossover operation on the selected individuals to generate new individuals, that is, new combinations of control points and end points;
[0134] Perform mutation operation; Apply the mutation operation to the newly generated individuals to introduce new individuals and prevent the algorithm from falling into local optimum;
[0135] Replacement: Add the newly generated individuals to the population and replace the individuals with poor fitness to maintain the population size;
[0136] Termination condition: Stop the iteration when the maximum number of iterations of the algorithm runs or the value of the fitness function converges to a certain threshold;
[0137] Optimal solution: Generate a final curve based on the optimized control points and end points, and this curve serves as the flight path for the UAV to achieve terrain following.
[0138] Based on the above method, the present invention also proposes an agricultural UAV device. The UAV device is equipped with a vision sensor. The lidar is used to obtain a point cloud image containing depth information, and the vision sensor is used to obtain an infrared image and a visible light image; the UAV device includes the trained GAN network model in the above S1; the UAV device can use this GAN network model to implement the content of step S2; the UAV device realizes terrain following flight according to the content of steps S3, S4, and S5.
[0139] Advantages of the present invention:
[0140] 1. The present invention enhances terrain information based on infrared and visible light fused images to generate an image containing enhanced crop canopy information;
[0141] 2. The present invention can realize the terrain following ability of agricultural UAVs in hilly and mountainous areas based on agricultural UAV vision;
[0142] 3. The present invention can be applied to agricultural operations such as crop monitoring, pesticide spraying, fertilization, etc., and is especially suitable for agricultural operations in complex terrains such as hills and mountains. Description of the drawings
[0143] Figure 1 Overall process schematic diagram of the present invention;
[0144] Figure 2 Flowchart of the depth estimation method of the present invention; Detailed implementation manners
[0145] The present invention will be further described below with reference to the drawings.
[0146] As Figure 1 shown, the present invention is a method for realizing a vision-based agricultural UAV terrain following technology, including the following steps:
[0147] S1. Use a large number of infrared and visible light images obtained by an agricultural UAV to perform offline adversarial training on the GAN network;
[0148] S2. Use the pre-trained GAN generator to fuse the infrared and visible light images obtained by the agricultural UAV in real time to enhance terrain information;
[0149] S3. Provide a method for estimating the distance of a monocular vision crop canopy using a depth estimation model. Estimate the distance of the crop canopy by capturing an image of the enhanced terrain information fused by a pre-trained GAN generator with a single camera.
[0150] S4. Use the crop canopy depth map D obtained based on airborne monocular vision p as a constraint of the objective function. Considering other beam conditions during the flight of the agricultural drone comprehensively, construct an objective function to optimize the trajectory planning to cope with multiple constraints.
[0151] S5. Select the genetic algorithm in the optimization algorithm to optimize the objective function J, and use the optimized curve as the flight path of the drone to achieve terrain-following flight.
[0152] Further explanation: The specific step S1 is as follows:
[0153] Perform offline adversarial training on the GAN network using a large number of infrared and visible light images obtained by the agricultural drone.
[0154] S11. Use the visual sensor carried by the agricultural drone to obtain visible light images and infrared images of hilly and mountainous areas, fly multiple times under different lighting conditions and time periods to ensure the diversity and comprehensiveness of the data, and define the infrared image I ir and the visible light image I vis .
[0155] S12. Normalize the obtained infrared and visible light images to be suitable for subsequent deep learning processing. Specifically, it is formula (1)
[0156]
[0157] I is the original image, I norm is the normalized image, I min , I max are the minimum and maximum pixel values of the image respectively.
[0158] S13. Input the processed infrared and visible light images into the generator G, use the convolutional neural network (CNN) architecture, and extract features layer by layer.
[0159] S14. Insert a Squeeze-and-Excitation (SE) module after each convolutional layer to enhance the attention to important features.
[0160] S15. Perform the Squeeze operation in the SE module, perform global average pooling on the features output by each convolutional layer to generate a global feature vector Z. Specifically, it is formula (2):
[0161]
[0162] where Z C is the global feature of channel C. Z represents the global feature vector obtained through the global average pooling operation, with a dimension of C, where C represents the number of channels of the feature map; x i,j,c is the value of the input feature map at position (i, j) and channel C. x represents the input feature map, with a dimension of H×W×C, where i and j represent the spatial positions of the feature map, and H and W respectively represent the height and width of the feature map;
[0163] Perform the Excitation operation in the SE module. Through two fully connected layers and activation functions (ReLU and Sigmoid), generate the feature weight S. First, pass through the first fully connected layer and the Relu activation function to generate the output vector e of the first fully connected layer, as shown in Equation (3):
[0164] e = ζ(W1·Z) (3)
[0165] Then, pass through the second fully connected layer and the Sigmoid activation function to generate the output vector S of the second fully connected layer, which is specifically represented by Equation (4):
[0166] S = σ(W2·e) (4)
[0167] where W1 and W2 respectively represent the weight matrices of the first and second fully connected layers, ζ is the Relu activation function, and σ is the Sigmoid activation function;
[0168] Finally, reapply the weight s of the Excitation operation to the original feature map for recalibration of the feature map, which is specifically represented by Equation (5):
[0169]
[0170] is the reweighted feature map, represents the recalibrated feature map, S c is the weight of channel c;
[0171] S16. After performing the operation of the SE module inserted in each convolutional layer, generate a fused image with enhanced crop canopy information; input the fused image and the real image into the discriminator D, use multiple convolutional layers for layer-by-layer feature extraction, and perform binary classification through the fully connected layer to determine whether it is real or generated, and reflect the authenticity of the image according to the real score (between 0 and 1) of the output image;
[0172] S17. Use the processed infrared and visible light image datasets as training data to train a generative adversarial network (GAN). During the training process, use the cross-entropy loss function to measure the performance of the generator and the discriminator. The loss function of the generator is defined as Equation (6):
[0173]
[0174] To enhance the enhancement of crop canopy information in the generated images, a feature enhancement loss can be additionally added to the generator loss. Suppose there is a feature extraction network F, and its feature loss intensity is defined as Equation (7):
[0175]
[0176] Combining the cross-entropy loss and the feature enhancement loss, the total loss function of the generator is Equation (8):
[0177]
[0178] Use the cross-entropy loss function to measure the discriminator's ability to distinguish between real images and generated images. The loss function of the discriminator can be defined as Equation (9):
[0179]
[0180] Where G is the generator, which generates an image from the input noise z; D is the discriminator, which outputs the probability that the input image is real; E represents the expectation; x is the real image; G(z) is the generated image, F is the feature extraction network, which outputs the feature representation of the image; represents the L2 norm; λ is the balancing weight parameter; p z represents the noise distribution; p data represents the real data distribution;
[0181] S18. During the training process, fix the parameters of the discriminator and train the generator. Then fix the parameters of the generator and train the discriminator, and alternate the training. Repeat the above steps until the generated images are both real and contain enhanced terrain information;
[0182] S19. After completing the training, save the weights of the generator for later use during online processing;
[0183] Further explanation, the specific content of step S2 is as follows:
[0184] Use the pre-trained GAN generator to fuse the infrared and visible light images obtained in real time by the agricultural drone to enhance the terrain information;
[0185] S21. Use the visual sensors carried by the agricultural drone to obtain infrared and visible light images in real time, and define the infrared image as Iir The visible light image is I vis ;
[0186] S22. Use the pre-trained GAN generator to fuse the real-time acquired infrared and visible light images to generate an image I containing enhanced crop canopy information F ;
[0187] Further explanation, the specific step S3 is as follows:
[0188] As Figure 2 shown, provide a method for monocular vision crop canopy distance estimation using a depth estimation model, and estimate the crop canopy distance by the image of enhanced terrain information captured by a single camera and fused by the pre-trained GAN generator;
[0189] S31. Use an agricultural drone equipped with a vision sensor to collect infrared images and visible light images in different scenarios of farmland, hills, and plains, and at the same time use the depth sensor LiDAR to synchronously collect the depth information in different scenarios to form a labeled data set, denoted as where x i is the i-th image in the labeled data set, d i is the corresponding depth annotation, and M is the total number of samples in the labeled data set; at the same time, collect a large amount of unlabeled agricultural drone image data, denoted as
[0190] S32. Use the labeled data set D l to train the initial monocular depth estimation model (teacher model T′), and use the affine invariant loss function for training to handle the problems of depth scale and offset between different data sets; the affine invariant loss function is as shown in formula (10):
[0191]
[0192] where H and W are the height and width of the airborne monocular image respectively, d i are the true and predicted crop canopy depths of the i-th pixel respectively, are the loss functions between the predicted crop canopy depth value and the true crop canopy depth value respectively, are the normalized true crop canopy depth value and predicted crop canopy depth value respectively, t(d) is the median of the crop canopy depth map d, and s(d) is the average of the absolute deviations of the crop canopy depth map d;
[0193] S33. Use the trained teacher model T′ to predict the crop canopy depth of the unlabeled data set D u to generate a pseudo-labeled data set, denoted as T(ui ) represents the predicted crop canopy depth map of the teacher model T′ for the unlabeled airborne monocular image u i ;
[0194] S34. Combine the labeled dataset D l and the pseudo-labeled dataset to form a comprehensive training set. Train the student model S on the comprehensive training set. During the training process, inject strong perturbations into the unlabeled airborne monocular images, including color distortion and spatial distortion, such as CutMix, to make the student model S actively seek additional visual knowledge and obtain invariant representations from the unlabeled airborne monocular images; First, randomly insert a pair of unlabeled airborne monocular images u a and u b , specifically represented by formula (11):
[0195] u ab = u a ⊙M + u b ⊙(1 - M) (11)
[0196] where u a , u b is a pair of randomly unlabeled airborne monocular images, and M is a binary mask with the rectangular region set to 1:
[0197] Unlabeled loss is obtained by calculating the affine invariant loss in the effective regions defined by M and 1 - M respectively, specifically represented by formula (12):
[0198]
[0199] The two losses are added together through weighted average to obtain the unlabeled loss specifically represented by formula (13)
[0200]
[0201] where S is the student model, ρ is the loss function, ∑M and ∑(1 - M) are the sums of the masks;
[0202] S35. Use the semantic segmentation model DINOv2 as auxiliary supervision in the depth estimation task, adopt the feature alignment loss to ensure that the student model S retains more semantic information in the pre-trained model; Set a tolerance margin to allow flexibility in feature alignment, focusing on the most relevant semantic features; The final loss function is the affine invariant loss function Unlabeled loss Feature alignment loss
[0203] The feature alignment loss is formula (14):
[0204]
[0205] where f i and f' i are the features of the student model and the DINOv2 model;
[0206] The final loss function is represented by Equation (15):
[0207]
[0208] S36. Input the feature map generated by the DINOv2 encoder into the DPT decoder to generate an airborne monocular crop canopy depth map; the DPT decoder first reorganizes the feature map into a multi-scale image representation, and then gradually fuses these representations to generate a dense depth prediction; the feature reorganization operation includes three steps: Read operation, Concatenate operation, and Resample operation;
[0209] The Read operation processes the feature map containing the readout token, converting it into a form suitable for subsequent processing, including ignoring the readout token, directly discarding the readout token, and only retaining the tokens;
[0210]
[0211] Add the readout token, adding the value of the readout token to each image token;
[0212]
[0213] Project the readout token: Concatenate the readout token with each image token, and then project it through a multi-layer perceptron (MLP);
[0214]
[0215] where t0 is the readout token, t1 to are the other tokens, MLP represents the multi-layer perceptron, and cat represents the concatenation operation;
[0216] Reorganize the tokens into a two-dimensional feature map according to their positions in the original image, t reshaped represents the feature map arranged by spatial position; use the Resample operation to perform spatial resampling on the reorganized feature map, and use convolution operations to adjust the resolution of the feature map to match the resolution required for subsequent processing;
[0217] Resample s (t) = Conv3x3(Conv1x1(t))
[0218] Perform a convolution operation on the feature map using a residual convolution unit and maintain the resolution of the feature map; gradually increase the resolution of the feature map through upsampling operations so that the finally generated feature map has a resolution gradually approaching that of the input image;
[0219] The output head receives the high-resolution feature map from the feature fusion module, further increases the resolution to the same as that of the input image through multiple convolutional layers and upsampling operations; finally, generates the final absolute depth map D of the crop canopy based on airborne monocular vision through a 1x1 convolutional layer p , the depth map D p Each value d(t) in it represents the distance from the agricultural UAV to the crop canopy at time t;
[0220] x conv = ConvLayer(x input )
[0221] x output = Upsample(x conv )
[0222] D p = Conv1x1(x output )
[0223] Further explanation, the specific step S4 is as follows:
[0224] Take the crop canopy depth map D obtained based on airborne monocular vision p as a constraint of the objective function, and comprehensively consider other beam conditions during the flight of the agricultural UAV to construct an objective function to optimize the trajectory planning to cope with multiple constraints;
[0225] S41. Comprehensively consider other constraint conditions such as the maneuverability of the UAV, the growth of crops, and the terrain undulation
[0226] The maneuverability constraint considers the speed constraint and acceleration constraint of the UAV. The acceleration constraint and speed constraint of the UAV are represented by formulas (16) and (17) respectively:
[0227] ||a(t)|| ≤ a max (16)
[0228] ||v(t)|| ≤ v max (17)
[0229] That is, at any time t, the magnitude of the acceleration and speed of the UAV cannot exceed the predetermined maximum acceleration a max and the predetermined maximum speed v max; where ||a(t)|| represents the magnitude of the acceleration of the UAV at time t, i.e., the norm of the acceleration vector a(t); similarly for ||v(t)||.
[0230] Crop canopy depth map D based on on-board monocular vision p Each value d(t) in it represents the distance from the agricultural UAV to the crop canopy at time t; the flight altitude of the UAV is z u (t), then the crop height is represented by formula (18):
[0231] z crop (t) = z u (t) - d(t) (18)
[0232] It is necessary to ensure that the flight altitude z u (t) is higher than the crop height z crop (t) plus the minimum safety distance h min , so:
[0233] z u (t) ≥ z crop (t) + h min
[0234] z u (t) ≥ z u (t) - d(t) + h min
[0235] d(t) ≥ h min
[0236] where h min is the minimum safety distance between the UAV and the crop;
[0237] In a mountainous environment, the terrain undulates greatly, so the flight trajectory of the UAV also needs to consider the constraints brought by the terrain undulation, which is specifically represented by formula (19):
[0238]
[0239] Among them, is the magnitude of the gradient vector, representing the size of the terrain slope; θ max is a pre-set maximum slope threshold, representing the maximum terrain slope at which the UAV can fly safely;
[0240] S42. Analyze the coupling characteristics of multiple constraint conditions, construct a penalty function, and use it to handle relatively complex and non-linear constraint conditions such as terrain undulation and crop canopy height:
[0241] The speed constraint C1(r(t)) is constraint condition ①. When the speed exceeds the maximum allowable speed v max , a penalty is generated:
[0242] C1(r(t)) = max(0, ||v(t)|| - v max ) ①
[0243] If the speed ||v(t)|| exceeds v max , the value of the penalty function is ||v(t)|| - v max , otherwise it is 0; the same applies to other constraints;
[0244] The acceleration constraint C2(r(t)) is the constraint condition ②. When the acceleration exceeds the maximum allowable acceleration a max , a penalty is generated:
[0245] C2(r(t)) = max(0, ||a(t)|| - a max ) ②
[0246] The crop canopy height constraint C3(r(t)) is the constraint condition ③. When the UAV height is lower than the crop canopy height plus the minimum safety distance z crop (t) + h min , a penalty is generated:
[0247] C3(r(t)) = max(0, z crop + h min - z u (t)) ③
[0248] The terrain undulation constraint C4(r(t)) is the constraint condition ④. When the magnitude of the terrain slope exceeds the preset maximum slope threshold θ max , a penalty is generated:
[0249]
[0250] Combining the above penalty terms together, the comprehensive penalty function Φ(r(t)) is represented by formula (20):
[0251] Φ(r(T)) = λ1C1(r(t)) + λ2C2(r(t)) + λ3C3(r(t)) + λ4C4(r(t))
[0252] That is:
[0253] C i (r(t)) represents the penalty function of the i-th constraint condition, and λ i is the penalty coefficient of the i-th constraint condition;
[0254] S43. By introducing cubic Bezier curves, establish the flight trajectory of the agricultural UAV, fully consider the flight requirements of the UAV and environmental constraints, and construct an objective function to optimize the trajectory planning;
[0255] Based on the mission requirements of agricultural drones, four control points of the cubic Bézier curve are set, namely P0, P1, P2, and P3; P0 is the take-off location of the drone, P1 is selected as a point slightly deviating from the crop row, which can be above the crop canopy to ensure that the drone has sufficient height during operation, P2 can be set between the crop rows to adjust the curvature and make the flight path of the drone smoother, and P3 is selected as the target location of the mission to ensure that the drone operates in this area. For the trajectory planning of the drone, the cubic Bézier curve can represent the position information of the drone at different time points;
[0256] The equation of the cubic Bézier curve is represented by formula (21):
[0257] B(t ′ ) = (1 - t ′ ) 3 P0 + 3t ′ (1 - t ′ ) 2 P1 + 3t ′2 (1 - t ′ )P2 + t ′3 P3 (21)
[0258] where P0 and P3 are the starting position and the final target position of the drone respectively; P1 and P2 are the initially selected intermediate control points; the parameter t ′ is used to represent the position of points on the curve and is a parameter from 0 to 1;
[0259] By discretizing t ′ , a series of trajectory points B(t ′ ) can be generated; the trajectory points B(t ′ ) are used as the flight trajectory r(t) of the drone; considering the smoothness, energy consumption, and various constraint conditions of the drone flight trajectory, a multi-constraint trajectory objective optimization function is constructed, which is specifically represented by formula (22):
[0260] J = ∫(w1||r″(t)|| 2 + w2||r′(t)|| 2 + w3Φ(r(t)))dt (22)
[0261] That is:
[0262] J = ∫0 1 (w1||B″(t′)|| 2 + w2||B′(t′)|| 2 + w3Φ(B(t′)))dt′
[0263] Among them, t′ is a parameter ranging from 0 to 1; w1, w2, and w3 are weight coefficients;
[0264] The constraint condition is formula (23):
[0265]
[0266]
[0267] Further explanation, the specific content of step S5 is as follows:
[0268] S5 selects the genetic algorithm in the optimization algorithm to optimize the objective function J, and takes the optimized curve as the flight path of the UAV to achieve terrain-following flight; the specific steps are as follows:
[0269] Coding: Encode the positions of the control points P1, P2, and the end point P3 into chromosomes;
[0270] Randomly generate the initial population: Randomly generate a group of initial populations, and each individual represents a different combination of control points and end points;
[0271] Fitness function: Calculate the fitness value of each chromosome, that is, calculate the value of the objective function J. The lower the fitness value, the better the path of the individual (i.e., this group of control points and end points);
[0272] Perform the selection operation: Use the roulette wheel selection method to select the individuals with better performance for mating according to the fitness value;
[0273] Perform the crossover operation: Perform the crossover operation on the selected individuals to generate new individuals, that is, new combinations of control points and end points;
[0274] Perform the mutation operation; Apply the mutation operation to the newly generated individuals to introduce new individuals and prevent the algorithm from falling into local optimum;
[0275] Replacement: Add the newly generated individuals to the population and replace the individuals with poorer fitness to keep the population size;
[0276] Termination condition: Stop the iteration when the maximum number of iterations of the algorithm runs or the value of the fitness function converges to a certain threshold;
[0277] Optimal solution: Generate the final curve according to the optimized control points and end points, and this curve is used as the flight path for the UAV to achieve terrain following.
[0278] Based on the above method, an embodiment of the present invention further provides an agricultural drone device. The drone device is equipped with a vision sensor, which is used to obtain infrared images and visible light images. The drone device includes the trained GAN network model in the above S1. The drone device can use this GAN network model to implement the content of step S2. The drone device realizes terrain-following flight according to the content of steps S3, S4, and S5.
[0279] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent manners or changes that do not deviate from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A vision-based terrain following method for agricultural drones, characterized in that, It includes the following: S1. Use infrared and visible light images to perform offline adversarial training on the GAN network; S2. Use the trained GAN network generator to fuse the infrared and visible light images obtained in real time by the agricultural drone to enhance terrain information; S3. Use the depth estimation model to perform monocular vision crop canopy distance estimation; estimate the distance between the drone and the crop canopy through the images with enhanced terrain information taken by a single camera and fused by the pre-trained GAN generator; S4. Use the crop canopy depth map D obtained based on airborne monocular vision p as one of the constraints of the objective function, and construct the objective function J to optimize the trajectory planning to cope with multiple constraints; S5. Use the genetic algorithm to optimize the objective function J, and use the optimized curve as the flight path of the drone to achieve terrain following flight.
2. The vision-based agricultural drone terrain following method according to claim 1, characterized in that The specific steps of step S1 are as follows: S11. Use the visual sensors carried by agricultural drones to obtain visible light images and infrared images of hilly and mountainous areas, conduct multiple flights under different lighting conditions and time periods to ensure the diversity and comprehensiveness of the data, and define the infrared image I ir and the visible light image I vis ; S12. Normalize the obtained infrared image and visible light image, specifically as formula (1) I is the original image, I norm is the normalized image, I min 、I max are the minimum and maximum pixel values of the image, respectively; S13. Input the processed infrared image and visible light image into the generator G of the GAN network, use the convolutional neural network (CNN) architecture, and extract features layer by layer; S14. Insert a channel attention (SE) module after each convolutional layer to enhance the attention to important features; S15. Perform the Squeeze operation in the SE module, perform global average pooling on the features output by each convolutional layer to generate a global feature vector Z, specifically as formula (2): Among them, Z C is the global feature of channel C. Z represents the global feature vector obtained through the global average pooling operation, with a dimension of C, where C represents the number of channels of the feature map; x i,j,c is the value of the input feature map at position (i, j) and channel C. x represents the input feature map, with a dimension of H×W×C, where i and j represent the spatial positions of the feature map, and H and W respectively represent the height and width of the feature map; Perform the Excitation operation in the SE module. Through two fully connected layers and activation functions (ReLU and Sigmoid), generate a feature weight S. First, pass through the first fully connected layer and the Relu activation function to generate the output vector e of the first fully connected layer, as formula (3): e = ζ(W1·Z) (3) Then pass through the second fully connected layer and the Sigmoid activation function to generate the output vector S of the second fully connected layer, specifically represented by formula (4): S = σ(W2·e) (4) Where W1 and W2 represent the weight matrices of the first and second fully connected layers respectively, ζ is the Relu activation function, and σ is the Sigmoid activation function; Finally, reapply the weight S of the Excitation operation to the original feature map to perform recalibration of the feature map, specifically represented by formula (5): is the feature map after reweighting, represents the feature map after recalibration, S c is the weight of channel c; S16. After the SE module operations inserted in each convolutional layer, generate a fused image with enhanced crop canopy information; input the fused image and the real image into the discriminator D, use multiple convolutional layers to extract features layer by layer, perform binary classification through the fully connected layer, and determine whether it is real or generated, and reflect the authenticity of the image according to the real score (between 0 and 1) of the output image; S17. Use the processed infrared image and visible light image to make a data set, use this data set to train the generative adversarial network (GAN), and use the cross-entropy loss function to measure the performance of the generator and the discriminator during the training process. The loss function of the generator is defined as formula (6): In order to enhance the crop canopy information in the generated image, a feature enhancement loss can be additionally added to the generator loss function. Assume a feature extraction network F, and define its feature loss intensity as formula (7): where x ∼ p data (x) indicates that the sample x is drawn from the true data distribution p data (x); z ∼ p z (z) indicates that the random noise vector is drawn from the noise distribution p z (z); Combining the cross-entropy loss and the feature enhancement loss, the total loss function of the generator is formula (8): The cross-entropy loss function is used to measure the discriminator D's ability to distinguish between real images and generated images. The loss function of the discriminator can be defined as formula (9): Among them, G is the generator that generates images based on the input noise z; D is the discriminator that discriminates whether the input image is real or generated; E represents the expectation; x is the real image; G(z) is the generated image, F is the feature extraction network that outputs the feature representation of the image; represents the L2 norm; λ is the balance weight parameter; p z (z) represents the noise distribution; p data (x) represents the real data distribution; S18. During the training process, fix the parameters of the discriminator to train the generator, and fix the parameters of the generator to train the discriminator, and train alternately until the generated images are both real and contain enhanced terrain information; S19. After the training is completed, save the weights of the generator for use during online processing.
3. The vision-based agricultural drone terrain following method according to claim 1, wherein The specific steps of S2 are as follows: S21. Use the vision sensor carried by the drone to obtain infrared images and visible light images in real time, and define the infrared image as I ir , and the visible light image as I vis ; S22. Use the pre-trained GAN network generator to fuse the real-time acquired infrared image and visible light image to generate an image I containing enhanced crop canopy information F .
4. The vision-based agricultural drone terrain following method according to claim 1, wherein The specific steps of S3 are as follows: S31. Use an agricultural drone equipped with a vision sensor to collect infrared images and visible light images in different scenarios such as farmland, hills, and plains. At the same time, use the depth sensor LiDAR to synchronously collect depth information in different scenarios to form an annotated data set, denoted as where x i is the i-th image in the annotated data set, d i is the corresponding depth annotation, and M is the total number of samples in the annotated data set; at the same time, collect a large amount of unannotated agricultural drone image data, denoted as S32. Use the labeled dataset D l Train the initial monocular depth estimation model, i.e., the teacher model T′, using an affine invariant loss function for training to handle the problems of depth scale and offset between different data; the affine invariant loss function is shown in formula (10): where H and W are the height and width of the airborne monocular image, respectively, d i are the true and predicted crop canopy depths of the i-th pixel, respectively, are the loss functions between the predicted crop canopy depth value and the true crop canopy depth value, respectively, are the normalized true crop canopy depth value and predicted crop canopy depth value, respectively, t(d) is the median of the crop canopy depth map d, and s(d) is the average of the absolute deviations of the crop canopy depth map d; S33. Use the trained teacher model T' to predict the crop canopy depth for the unlabeled dataset D u to generate a pseudo-labeled dataset, denoted as T(u i ) represents the crop canopy depth prediction map of the teacher model T' for the unlabeled airborne monocular image u i ; S34. Combine the labeled dataset D l and the pseudo-labeled dataset to form a comprehensive training set. Train the student model S on the comprehensive training set. During the training process, inject strong perturbations, including color distortion and spatial distortion, into the unlabeled airborne monocular images, so that the student model S actively seeks additional visual knowledge and obtains invariant representations from the unlabeled airborne monocular images. First, insert a pair of random unlabeled airborne monocular images u a and u b spatially, which is specifically represented by Equation (11): u ab = u a ⊙M + u b ⊙(1 - M)(11) where u a and u b are a pair of randomly unlabeled airborne monocular images, M is a binary mask with the rectangular region set to 1, and u ab represents the fused image. Unlabeled loss Obtained by calculating the affine invariant loss in the effective regions defined by M and 1-M respectively, specifically represented by formula (12): The unlabeled loss is obtained by adding two losses through weighted averaging Specifically, it is represented by formula (13) Where S is the student model, ρ is the loss function, and ∑M and ∑(1 - M) are the sums of the masks; S35. In the depth estimation task, the semantic segmentation model DINOv2 is used as auxiliary supervision, and the feature alignment loss is adopted to ensure that the student model S retains more semantic information; a tolerance margin is set to allow flexibility in feature alignment, and the most relevant semantic features are focused on; the final loss function includes an affine invariant loss function Unlabeled loss and feature alignment loss The feature alignment loss is formula (14): where f i and f i ′ are the features of the student model and the DINOv2 model; Finally The loss function is represented by formula (15): S36. Input the feature map generated by the DINOv2 encoder into the DPT decoder to generate an airborne monocular crop canopy depth map; the DPT decoder first reorganizes the feature map into a multi-scale image representation, and then gradually fuses these representations to generate a dense depth prediction; The feature reorganization operation includes three steps: Read operation, Concatenate operation, and Resample operation; The Read operation processes the feature map containing the readout token and converts it into a form suitable for subsequent processing, including ignoring the readout token, directly discarding the readout token, and only retaining the tokens; a token refers to a small block extracted from an image, and each block is converted into a feature vector. Add the readout token, and add the value of the readout token to each image token; Project the readout token: Concatenate the readout token with each image token, and then project it through a multi-layer perceptron (MLP); Among them, t0 is the read-out token, and t1 to are other tokens. MLP represents a multi-layer perceptron, and cat represents a concatenation operation; Recombine the tokens into a two-dimensional feature map according to their positions in the original image, t reshaped denotes the feature map arranged by spatial positions; perform spatial resampling on the recombined feature map using the Resample operation, and adjust the resolution of the feature map using the convolution operation to match the resolution required for subsequent processing; Resample s (t) = Conv3x3(Conv1x1(t)) Perform a convolution operation on the feature map using a residual convolution unit and maintain the resolution of the feature map; gradually increase the resolution of the feature map through an upsampling operation, so that the finally generated feature map gradually approaches the resolution of the input image; The output head receives the high-resolution feature map from the DPT decoder feature fusion module, further enhances the resolution to the same as that of the input image through multiple convolutional layers and upsampling operations; finally, generates the final crop canopy depth map D based on airborne monocular vision through a 1x1 convolutional layer p , the depth map D p Each value d(t) in represents the distance from the agricultural drone to the crop canopy at time t; x conv = ConvLayer(x input ) x output = Upsample(x conv ) D p = Conv1x1(x output )。 5. The vision-based terrain following method for an agricultural drone according to claim 1, characterized in that, The specific steps of S4 are as follows: S41. Considering the mobility of the UAV, the growth of crops, and the terrain undulation, etc., establish constraint conditions, including: Mobility constraints: including the acceleration constraint and speed constraint of the UAV. The acceleration constraint and speed constraint of the UAV are represented by formulas (16) and (17) respectively: ∥a(t)∥≤a max (16) ∥v(t)∥≤v max (17) That is, at any time t, the acceleration and speed magnitude of the drone cannot exceed the predetermined maximum acceleration a max and the predetermined maximum speed v max ; where ∥a(t)∥ represents the magnitude of the acceleration of the drone at time t, that is, the norm of the acceleration vector a(t); the same applies to ∥v(t)∥; Safety distance constraint: Crop canopy depth map D p Each value d(t) in it represents the distance from the agricultural drone to the crop canopy at time t; the flight altitude of the drone is z u (t), then the crop height is represented by Equation (18): z crop z(t) = z u z(t) - d(t) (18) It is necessary to ensure that the flight altitude z of the drone u (t) is higher than the crop height z crop (t) plus the minimum safety distance h min , so: z u (t) ≥ z crop (t) + h min z u (t) ≥ z u (t) - d(t) + h min d(t) ≥ h min where h min is the minimum safe distance between the UAV and the crop; Terrain constraints: In a mountainous environment, the terrain undulation is relatively large, so the flight trajectory of the UAV also needs to consider the constraints brought by the terrain undulation, which is specifically represented by formula (19): Among them, is the magnitude of the gradient vector, representing the magnitude of the terrain slope; θ max is a preset maximum slope threshold, representing the maximum terrain slope at which the drone can fly safely; S42. Analyze the coupling characteristics of multiple constraint conditions and construct a penalty function to handle more complex and non-linear constraint conditions, specifically as follows: The speed constraint is Constraint ①. When the speed exceeds the maximum allowable speed v max a speed penalty C1(r(t)) is generated: C1(r(t)) = max(0, ∥v(t)∥ - v max ) ① If the speed ∥v(t)∥ exceeds v max , the value of the penalty function is ∥v(t)∥ - v max , otherwise it is 0; The acceleration constraint is Constraint ②. When the acceleration exceeds the maximum allowable acceleration a max , an acceleration penalty C2(r(t)) is generated: C2(r(t)) = max(0, ∥a(t)∥ - a max ) ② The crop canopy height constraint is Constraint ③. When the UAV height is lower than the crop canopy height plus the minimum safety distance z crop (t)+h min , a crop canopy height penalty C3(r(t)) is generated: C3(r(t)) = max(0, z crop + h min - z u (t)) ③ The terrain undulation constraint is constraint ④, and the magnitude of the terrain slope exceeds the pre-set maximum slope threshold θ max , resulting in a terrain undulation penalty C4(r(t)): Combine the above penalty terms to form a comprehensive penalty function Φ(r(t)), which is represented by formula (20): Φ(r(t)) = λ1C1(r(t)) + λ2C2(r(t)) + λ3C3(r(t)) + λ4C4(r(t)) That is: C i (r(t)) represents the penalty function of the i-th constraint condition, and λ i is the penalty coefficient of the i-th constraint condition; S43. Establish the flight trajectory of the agricultural UAV by introducing a cubic Bézier curve, fully considering the flight requirements and constraints of the UAV, and construct an objective function to optimize the trajectory planning; specifically as follows: Based on the mission requirements of the agricultural drone, four control points of the cubic Bézier curve are set, namely P0, P1, P2, and P3; P0 is the take-off location of the drone, P1 is selected as a point slightly deviating from the crop row, which can be above the crop canopy to ensure that the drone has sufficient height during operation, P2 can be set between the crop rows to adjust the curvature and make the flight path of the drone smoother, and P3 is selected as the target location of the mission to ensure that the drone operates in this area; for the trajectory planning of the drone, the cubic Bézier curve can represent the position information of the drone at different time points; The equation of the cubic Bézier curve is represented by formula (21): B(t′) = (1 - t′) 3 P0 + 3t′(1 - t′) 2 P1 + 3t′ 2 (1 - t′)P2 + t′ 3 P3 (21) where P0 and P3 are the starting position and the final target position of the drone respectively; P1 and P2 are the initially selected intermediate control points; the parameter t′ is used to represent the position of the point on the curve and is a parameter from 0 to 1; By discretizing t′, a series of trajectory points B(t′) can be generated; the trajectory points B(t′) are used as the flight trajectory r(t) of the drone; considering the smoothness, energy consumption and various constraint conditions of the drone flight trajectory, a multi-constraint trajectory objective optimization function is constructed, which is specifically represented by formula (22): J = ∫(w1∥r″(t)∥ 2 + w2∥r′(t)∥ 2 + w3Φ(r(t)))dt (22) That is: where t′ is a parameter from 0 to 1; w1, w2, and w3 are weight coefficients; The constraint condition is formula (23):
6. The vision-based agricultural drone terrain following method according to claim 1, wherein The specific step S5 is as follows: Encoding: Encode the positions of the control points P1, P2, and the end point P3 into chromosomes; Randomly generate the initial population: Randomly generate a group of initial populations, and each individual represents a different combination of control points and end points; Fitness function: Calculate the fitness value of each chromosome, that is, calculate the value of the objective function J. The lower the fitness value, the better the path of the individual; Perform the selection operation: Use the roulette wheel selection method to select better-performing individuals for mating according to the fitness value; Perform the crossover operation: Perform the crossover operation on the selected individuals to generate new individuals, that is, new combinations of control points and end points; Perform the mutation operation; Apply the mutation operation to the newly generated individuals to introduce new individuals and prevent the algorithm from falling into a local optimum; Replacement: Add the newly generated individuals to the population and replace the individuals with poorer fitness to maintain the population size; Termination condition: Stop the iteration when the maximum number of iterations of the algorithm runs or the value of the fitness function converges to a certain threshold; Optimal solution: Generate the final curve according to the optimized control points and end points, and this curve is used as the flight path for the drone to achieve terrain following.
7. An agricultural drone device, characterized in that, The drone device is based on a vision sensor for acquiring infrared images and visible light images; the GAN network model trained in S1 of claim 1 is set in the drone device; the drone device can use this GAN network model to implement the content of S2 in claim 1; the drone device realizes terrain following flight according to the content of S3, S4, and S5 in claim 1.
Citation Information
Cited By
Efficient leaf area estimation method
CN121010779A
A high-efficiency leaf area estimation method
CN121010779B
Low-altitude pollination control method based on closed-loop multi-objective enhanced optimization algorithm
CN121411265A