A real-time positioning method for agricultural drones in row sowing based on visual tracking

By constructing a multi-frame constraint function for visual tracking and fusing RTK observation data, abnormal feature points are eliminated, solving the problem of low positioning accuracy of drones in farmland scenes and achieving high-precision and stable real-time positioning.

CN116958835BActive Publication Date: 2025-09-16GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310590861.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-09-16
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

Existing UAV positioning methods have low positioning accuracy in farmland scenes, especially in scenes with repeated or weak textures, and cannot achieve accurate real-time positioning, resulting in positioning failure or excessive errors.

Method used

A real-time positioning method for agricultural drones in row planting based on visual tracking is adopted. By collecting binocular images of farmland scenes, IMU data and RTK observation data, a multi-frame visual constraint function is constructed. The RTK observation data is integrated for joint optimization, abnormal feature points are eliminated, and positioning accuracy and stability are improved.

Benefits of technology

It achieves higher positioning accuracy and stability in farmland scenarios, has strong adaptability, and can achieve real-time and accurate positioning on drones with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958835B_ABST
    Figure CN116958835B_ABST
Patent Text Reader

Abstract

The present invention proposes a real-time positioning method for agricultural drones in row sowing based on visual tracking, which relates to the technical field of drone positioning. The method collects binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of corresponding positions of the agricultural drones. According to the initial posture information between the binocular image frames of the IMU data, a multi-frame visual constraint function is constructed with the current frame in a sliding window composed of multiple frames to estimate the current real-time posture of the agricultural drone, thereby achieving more stable inter-frame motion estimation. In order to improve scene adaptability, the current frame corresponding to the feature point that meets the stable tracking conditions is inserted into the sliding window as a new key frame, and is jointly optimized with the RTK observation data to improve positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) positioning, and more particularly to a real-time positioning method for a row-sowing agricultural UAV based on visual tracking. Background Art

[0002] Agricultural drones use ground-based remote control or navigation flight control to spray pesticides, seeds, powders, and more. Drones used for row seeding of rice require extremely high-precision positioning, with an error of no more than 2 cm. Otherwise, farmers would need to replant rice after sowing, which defeats the original purpose of agricultural drone row seeding, which is to improve agricultural production efficiency.

[0003] The IMU (Inertial Measurement Unit) is a component of a drone's flight control system. A drone's flight control system consists of a main control MCU and an inertial measurement module (IMU). The IMU provides raw sensor data on the drone's spatial attitude, typically using a gyroscope, accelerometer, or electronic compass to provide 9DOF data. RTK is used to continuously observe the drone during flight. By importing RTK data into differential post-processing software, the 3D geographic information captured by the drone can be obtained.

[0004] At present, the positioning of drones is mainly achieved based on the visual odometry method. One of them is an indirect positioning algorithm based on feature points, such as the ORB-SLAM algorithm. This algorithm obtains image data of farmland scenes in real time, extracts ORB feature points for each image, and performs motion estimation based on the inter-frame matching relationship and the constraint of minimizing the reprojection error equation. Then, the posture of the drone is optimized through the local sub-map method to achieve real-time positioning of agricultural drones. This method requires that the farmland scene facing the drone has good texture to ensure accurate and rich data management relationship between frames, so as to achieve accurate real-time positioning. However, in farmland operation scenes, unsown fields and rice fields are usually repetitive textures and weak texture extreme scenes. Faced with such farm scenes, the ORB-SLAM algorithm cannot work well. In actual operation, positioning mutations and positioning failures will occur, and accurate real-time positioning of agricultural drones cannot be achieved.

[0005] Another type of direct positioning algorithm is based on minimizing photometric errors, such as the DSO algorithm (Direct Sparse Odometry). This algorithm acquires image data of farmland scenes in real time and extracts points with large gradient values ​​evenly distributed across a grid for each image. It then constructs a motion model between frames, uses photometric error minimization for motion estimation, and optimizes the drone's posture using multi-frame constraints within a sliding window, thereby achieving real-time positioning of agricultural drones. Compared to the ORB-SLAM algorithm, the DSO algorithm does not require data association and has low requirements for image texture. Based on fixed parameters, it can achieve highly accurate positioning in farmland scenes. However, in the DSO algorithm, to conserve computational resources, keyframe-based optimization is employed. Multiple residual parameters are introduced to determine whether the current frame is a keyframe. However, these multiple parameters can introduce significant scene variations, leading to inconsistencies and failures in motion tracking, thus compromising positioning accuracy and precision. Summary of the Invention

[0006] In order to solve the problem of low positioning accuracy of traditional methods for real-time positioning of drones, the present invention provides a real-time positioning method for agricultural drones in row sowing based on visual tracking, which has high adaptability to agricultural scenes and more accurate positioning.

[0007] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:

[0008] A real-time positioning method for a row-sowing agricultural drone based on visual tracking includes the following steps:

[0009] S1. Collect binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of the corresponding positions of the agricultural drones;

[0010] S2. Determine the key frame and sliding window of the binocular image based on the initial posture information between binocular image frames provided by the farmland scene binocular image and IMU data;

[0011] S3. Project the feature points corresponding to all key frames in the sliding window to the current frame and construct a multi-frame visual constraint function to estimate the current real-time posture of the agricultural UAV;

[0012] S4. Insert the current frame corresponding to the feature point that meets the stable tracking conditions into the sliding window as a new key frame;

[0013] S5. Initialize the content of the new key frame and remove abnormal feature points in the key frame;

[0014] S6. Integrate RTK observation data to construct and solve the visual positioning-RTK joint optimization objective function to achieve real-time positioning of agricultural UAVs.

[0015] According to the above technical means, binocular images of farmland scenes taken by agricultural UAVs, IMU data synchronized with the agricultural UAVs, and RTK observation data of the corresponding positions of the agricultural UAVs are collected. Based on the initial posture information between the IMU data binocular image frames, a multi-frame visual constraint function is constructed with the current frame in a sliding window composed of multiple frames to estimate the current real-time posture of the agricultural UAV, thereby achieving more stable inter-frame motion estimation. In order to improve scene adaptability, the current frame corresponding to the feature point that meets the stable tracking conditions is inserted into the sliding window as a new key frame, and jointly optimized with the RTK observation data to improve positioning accuracy.

[0016] Preferably, according to the difference between the frame images of the binocular image of the farmland scene, the key frame is determined, and the first n key frames are intercepted to form a sliding window. The sliding window includes n key frames K and m feature points L that can characterize the characteristics of the binocular image corresponding to the key frames, satisfying: K = {K1, K2, ..., K n}、L={l1,l2,…,l m}.

[0017] Preferably, based on the data of the IMU synchronized with the agricultural drone, an integration operation is performed to obtain the current frame I of the binocular image of the farmland scene. c The initial relative motion between the key frame K of the sliding window is denoted as T = {T c1 ,T c2 ,…,T cn}, T ci Indicates the time from the i-th key frame to the current frame I c The initial relative motion between, i = 1, 2, ..., n.

[0018] Preferably, in step S3, the process of projecting the feature points corresponding to all key frames in the sliding window to the current frame is:

[0019] Assume that the three-dimensional coordinates of any feature point among all the key frames are expressed as (X, Y, Z), and the intrinsic parameter matrix of the UAV's binocular camera is: Among them, (f x ,f y ) is the focal length, (c x ,c y ) is the main point of the binocular camera; let the feature point pixel u k The reprojected coordinates in the current frame are u ′ k =(u,v), then:

[0020]

[0021] The constructed multi-frame visual constraint function is:

[0022]

[0023] in, ρ k Represents the feature point pixel u k The corresponding inverse depth value, represents the projection function; (a t ,b t ) represents the photometric affine parameter at time t, e (·) Represents the exponential map, I t represents the reference frame, ω k Represents the residual term; N u Represents the set of feature points corresponding to all key frames.

[0024] According to the above technical means, the feature points corresponding to all key frames K of the sliding window are projected into the current frame, which can achieve more stable inter-frame motion estimation and prevent the current frame I from c There is less overlapping information with the most recent key frame K1, resulting in tracking loss.

[0025] Preferably, the weighted method based on the fusion of Gaussian distribution and t-distribution is used to calculate the residual term ω k Make improvements to meet the following requirements:

[0026]

[0027] in, Indicates that the Gaussian distribution is in the residual term ω k The weight on Represents the Gaussian distribution in the residual term ω k The weight on Represents the feature point pixel u k The corresponding gradient value, c represents a constant, v, σ k Represents the variance and degree of freedom of t-distribution, r k Represents the residual.

[0028] According to the above technical means, the weighted method of fusing Gaussian distribution and t-distribution avoids the influence of abnormal feature points on the optimization target when using the DSO algorithm for positioning.

[0029] Preferably, the stable tracking condition is:

[0030]

[0031] Among them, m represents the total number of feature points, num represents the number of feature points that can be stably tracked, and Q min It represents the minimum tracking ratio threshold, which improves the scene adaptability of the positioning method and further improves the positioning accuracy.

[0032] Preferably, the contents of initializing the content of the new key frame include: selection of gradient points, extraction of image blocks and depth initialization of feature points, wherein the selection of gradient points and extraction of image blocks adopt the DSO algorithm. When initializing the depth of the feature points, the fixed baseline of the binocular camera is used to provide an initial depth range for the depth value of the feature points, which can ensure that the depth value of the feature points converges quickly, reduce the occupation of computing resources, and improve positioning accuracy.

[0033] Preferably, the process of removing abnormal feature points in the key frame is:

[0034] Based on the residual r k Gradient value corresponding to the feature point The constraint relationship between them is used to determine whether an abnormal feature point is an abnormal feature point. The constraint relationship expression is:

[0035]

[0036] If the constraint relationship is satisfied, the feature point is an abnormal feature point; otherwise, the feature point is a normal feature point.

[0037] According to the above technical means, considering that abnormal feature points may destroy the positioning estimation results, the abnormal points are proposed to ensure the robust and accurate execution of the positioning algorithm.

[0038] Preferably, the visual positioning-RTK joint optimization objective function is:

[0039]

[0040] in, Indicates that except for key frame K i All other keyframes except Indicates key frame I j Effective projection to keyframe I i The feature points in E p is the visual constraint term, E g are RTK constraints, which satisfy:

[0041]

[0042] in, represents the photometric error produced by the left camera, C i Indicates key frame I i The corresponding RTK position, Indicates key frame I i Corresponding agricultural drone posture.

[0043] Based on the above technical means, visual constraints are integrated with RTK constraints to obtain positioning in the absolute geographic coordinate system, thereby ensuring more accurate positioning accuracy.

[0044] Preferably, considering the binocular constraint, all feature points are projected into the right eye image, and the same residual function is constructed as for the left eye:

[0045]

[0046] Then, the DSO algorithm is used to solve the visual positioning-RTK joint optimization objective function to achieve real-time positioning of agricultural UAVs.

[0047] Based on the above technical means, through static binocular residuals, each feature point contains only one binocular constraint value, and a binocular weighted hyperparameter is introduced to project all feature points into the right eye image. Accurate scale estimation can be achieved without introducing additional hyperparameters, which has better adaptability to scenes and better stability, thereby achieving higher UAV positioning accuracy.

[0048] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0049] The present invention proposes a real-time positioning method for agricultural drones in row sowing based on visual tracking. The method collects binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of the corresponding positions of the agricultural drones. According to the initial posture information between the IMU data binocular image frames, a multi-frame visual constraint function is constructed with the current frame in a sliding window composed of multiple frames to estimate the current real-time posture of the agricultural drone, thereby achieving more stable inter-frame motion estimation. In order to improve scene adaptability, the current frame corresponding to the feature point that meets the stable tracking conditions is inserted into the sliding window as a new key frame, and jointly optimized with the RTK observation data to improve positioning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic diagram showing a flow chart of a real-time positioning method for a row-sowing agricultural drone based on visual tracking proposed in Example 1 of the present invention;

[0051] Figure 2 A schematic diagram showing the flight route of the agricultural UAV seeding operation proposed in Example 1 of the present invention;

[0052] Figure 3 A schematic diagram showing the overlap between multiple frames proposed in embodiment 2 of the present invention. DETAILED DESCRIPTION

[0053] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0054] In order to better illustrate this embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the actual size;

[0055] It is understandable to those skilled in the art that descriptions of certain well-known contents may be omitted in the drawings.

[0056] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0057] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent;

[0058] Example 1

[0059] This embodiment proposes a real-time positioning method for agricultural drones in row planting based on visual tracking. The implementation flow chart of this method is as follows: Figure 1 , see Figure 1 , the method comprises the following steps:

[0060] S1. Collect binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of the corresponding positions of the agricultural drones;

[0061] S2. Determine the key frame and sliding window of the binocular image based on the initial posture information between binocular image frames provided by the binocular image and IMU data of the farmland scene. When implemented based on visual tracking, an agricultural drone will produce 30 frames of binocular images of the farmland scene in one second. Since the computing resources on the drone are relatively limited, it is impossible to process each frame of the image in detail. Therefore, it is necessary to determine the key frame of the binocular image, such as Figure 2 As shown in the figure, during the actual operation of agricultural drones, each route corresponding to each sowing strip has multiple key frames, that is, multiple positioning points, and there will also be multiple routes. The conventional calculation method is to combine all the information collected before and calculate together to locate the current agricultural drone. However, this approach is bound to fail when the computing resources of agricultural drones are limited. Therefore, it is necessary to intercept key frames to form a sliding window to reduce the amount of calculation and realize real-time calculation on agricultural drones.

[0062] Based on the differences between the frames of the binocular image of the farmland scene, it is determined which frames require coarse processing and which frames require fine processing. The frames requiring fine processing are determined as key frames, and the others are non-key frames. In this embodiment, if the scene displayed by the current frame image is different from the scene displayed by the previous frame image, the current frame image needs to be finely processed for the sake of visual continuity. If the scene displayed by the current frame image is very similar to the previous one, it means that the drone has not moved much, and no additional computing resources are needed to calculate some motion information. The current frame only needs coarse processing.

[0063] S3. Project the feature points corresponding to all key frames in the sliding window to the current frame and construct a multi-frame visual constraint function to estimate the current real-time posture of the agricultural UAV;

[0064] When processing key frames, not all pixels in the key frames will be processed. Instead, more representative points in the key frames will be extracted to reduce the amount of calculation. These representative points can well represent the entire image. These points are the feature points corresponding to the key frames, and feature points can be extracted through feature extraction.

[0065] S4. Insert the current frame corresponding to the feature point that meets the stable tracking conditions into the sliding window as a new key frame;

[0066] S5. Initialize the content of the new key frame and remove abnormal feature points in the key frame;

[0067] S6. Integrate RTK observation data to construct and solve the visual positioning-RTK joint optimization objective function to achieve real-time positioning of agricultural UAVs.

[0068] In general, this embodiment collects binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of the corresponding positions of the agricultural drones. Based on the initial posture information between the IMU data binocular image frames, a multi-frame visual constraint function is constructed with the current frame in a sliding window composed of multiple frames to estimate the current real-time posture of the agricultural drone, thereby achieving more stable inter-frame motion estimation. In order to improve scene adaptability, the current frame corresponding to the feature point that meets the stable tracking conditions is inserted into the sliding window as a new key frame, and jointly optimized with the RTK observation data to improve positioning accuracy.

[0069] Example 2

[0070] In this embodiment, the first n key frames are intercepted to form a sliding window, which includes n key frames K and m feature points L that can characterize the characteristics of the binocular image corresponding to the key frames, satisfying: K = {K1, K2, ..., K n}、L={l1,l2,…,l m}.

[0071] Based on the data of the IMU synchronized with the agricultural UAV, an integration operation is performed to obtain the current frame I of the binocular image of the farmland scene c The initial relative motion between the key frame K of the sliding window is denoted as T = {T c1 ,T c2 ,…,T cn}, T ci Indicates the time from the i-th key frame to the current frame I cThe initial relative motion between, i = 1, 2, ..., n.

[0072] In step S3, the process of projecting the feature points corresponding to all key frames in the sliding window to the current frame is as follows:

[0073] Assume that the three-dimensional coordinates of any feature point among all the key frames are expressed as (X, Y, Z), and the intrinsic parameter matrix of the UAV's binocular camera is: Among them, (f x ,f y ) is the focal length, (c x ,c y ) is the main point of the binocular camera; let the feature point pixel u k The reprojected coordinates in the current frame are u ′ k =(u,v), then:

[0074]

[0075] At this time, you need to know the depth information corresponding to the feature point pixels to obtain the complete three-dimensional space coordinates Otherwise, only the 3D space coordinates of the normalized plane (-,-,1) can be obtained, that is, the 3D space coordinates at a distance of 1m from the binocular camera.

[0076] The constructed multi-frame visual constraint function is:

[0077]

[0078] in, ρ k Represents the feature point pixel u k The corresponding inverse depth value, represents the projection function; (a t ,b t ) represents the photometric affine parameter at time t, e (·) Represents the exponential map, I t represents the reference frame, ω k Represents the residual term; N u Represents the set of feature points corresponding to all key frames.

[0079] To prevent the current frame I v There is less overlapping information with the most recent key frame K1, which leads to tracking loss. Projecting the feature points corresponding to all key frames K in the sliding window into the current frame can achieve more stable inter-frame motion estimation.

[0080] like Figure 3As shown, the only overlapping areas between the current frame and keyframe 1 are areas 2 and 4. Therefore, when optimizing the visual constraint function, the corresponding constraint relationships are only provided by these two parts. After introducing keyframe 2, the overlapping areas between the current frame and keyframes 1 and 2 are divided into four parts, areas 1-2-3-4. That is, by establishing constraint relationships between all keyframes in the sliding window and the current frame, more observation information can be obtained. Compared to the case where only a single keyframe is constrained, the feature points corresponding to all keyframes are projected onto the current frame, achieving overlap between the current frame and multiple keyframes. This results in more constraint relationships and makes it easier to achieve stable and accurate positioning estimation. Because a single keyframe may have no constraint relationship with the current frame, resulting in positioning failure, introducing multiple keyframe constraints can better compensate for this shortcoming.

[0081] Example 3

[0082] This embodiment considers the traditional DSO algorithm for positioning, which only uses the weights related to the gradient values ​​of the feature points to adjust the confidence of each point. This implicit assumption is that the distribution of each point conforms to the Gaussian distribution. However, when the observation point is an inlier, the Gaussian distribution is satisfied, and the weight reflected in the residual term is recorded as Where c is a constant. Represents the feature point pixel u k The corresponding gradient value; when the feature point is an outlier, its probability distribution is more consistent with the t-distribution, and the weight reflected in the residual term is recorded as where v,σ k represents the variance and number of degrees of freedom of the t-distribution, r k Therefore, in this embodiment, the residual term ω is weighted based on the fusion of Gaussian distribution and t-distribution. k Improvements are made to avoid the impact of abnormal feature points on the optimization target when using the DSO algorithm for positioning. After the improvements, the following conditions are met:

[0083]

[0084] in, Represents the Gaussian distribution in the residual term ω k The weight on Represents the Gaussian distribution in the residual term ω k The weight on Represents the feature point pixel u k The corresponding gradient value, c represents a constant, v, σ k Represents the variance and degree of freedom of t-distribution, r k Represents the residual.

[0085] In addition, the traditional DSO algorithm uses a keyframe-based optimization method to optimize the visual constraint function in order to save computing resources. In order to determine whether the current frame is a keyframe, the following constraint relationship needs to be determined:

[0086]

[0087] Among them, f,f t , a represent the mean square optical flow error, the mean square optical flow error of pure translation and the change of photometric affine parameters, and the parameter w f , w a Represents the weight corresponding to each error, T kf Indicates the threshold for key frame determination.

[0088] The above three residuals correspond to the inter-frame motion transformation evaluation situations respectively:

[0089] (1) The translation and rotation between frames are more intense, which will introduce large scene changes;

[0090] (2) The rotational motion between frames is relatively intense, which introduces a large pure translation error and a large field of view change;

[0091] (3) The illumination between frames changes significantly, resulting in a significant change in the photometric affine parameters between frames, which will cause inconsistency between the scenes;

[0092] The above three situations are likely to cause the visual mark of agricultural drone motion tracking, thereby affecting the success rate of drone positioning. The key frame judgment of the DSO algorithm needs to consider the above four hyperparameters. The adjustment is highly dependent on experience and farmland scenes, and has low flexibility.

[0093] In order to improve the adaptability of farmland scenes, a key frame judgment method that only considers one hyperparameter is proposed. By projecting all features of the sliding window to the current frame, the number of feature points that can be stably tracked is counted. The stable tracking condition for statistical judgment is:

[0094]

[0095] Among them, m represents the total number of feature points, num represents the number of feature points that can be stably tracked, and Q min Represents the minimum tracking ratio threshold. If the above conditions are met, the current frame will be inserted into the sliding window as the latest keyframe for joint optimization to further improve the positioning accuracy.

[0096] After the new key frame is determined, the content of the new key frame is initialized. The initialization content includes: gradient point selection, image block extraction, and depth initialization of feature points. The gradient point selection and image block extraction use the DSO algorithm, Gaussian filtering the input image, and calculating the Gaussian pyramid to obtain images at multiple different scales. Specifically, the steps for selecting gradient points for an image are as follows:

[0097] (1) Calculate pixel gradient values ​​(d x ,d y ), and the sum of squared gradients

[0098] (2) Divide each image into multiple blocks of 32*32 size, and calculate the gradient histogram of each block based on the gradient square sum d, and determine the median of the gradient square sum d in each block. Perform mean filtering on all medians to obtain the final screening threshold th;

[0099] (3) Divide the screening window into four sizes, and the corresponding threshold constraints are as follows:

[0100] In a 1*1 window, the threshold is set to 2*th;

[0101] In the 4*4 window, the threshold is set to th;

[0102] In the 8*8 window, the threshold is set to 0.75*th;

[0103] In a 16*16 window, the threshold is set to 0.75*0.75*th;

[0104] Pixels are extracted from small to large. If no gradient points that meet the constraints can be extracted in a 1*1 window, search in a 4*4 window, and so on.

[0105] After filtering out the gradient points, in order to enhance the resolution of the points, it is achieved by introducing the neighborhood pixels of their fixed patterns. This operation is called image block extraction.

[0106] When initializing the depth of feature points, the fixed baseline of the binocular camera provides an initial depth range that is relatively close to the true value for the depth value of the feature points, which can ensure that the depth value of the feature points converges quickly, reduce the occupation of computing resources, and improve positioning accuracy.

[0107] In order to ensure that positioning can be performed robustly and accurately, the strategy of eliminating abnormal feature points is particularly important. Abnormal feature points will destroy the consistency of the entire system and cause the entire estimation result to develop in the wrong direction.

[0108] In this embodiment, based on the residual r k Gradient value corresponding to the feature point The constraint relationship between them is used to determine whether an abnormal feature point is an abnormal feature point. The constraint relationship expression is:

[0109]

[0110] If the constraint relationship is satisfied, the feature point is an abnormal feature point; otherwise, the feature point is a normal feature point.

[0111] Finally, in order to obtain more accurate positioning accuracy, the visual constraints are fused with the RTK constraints, that is, the RTK observation data is fused to obtain positioning in the absolute geographic coordinate system. The visual positioning-RTK joint optimization objective function is constructed as:

[0112]

[0113] in, Indicates that except for key frame K i All other keyframes except Indicates key frame I j Effective projection to keyframe I i The feature points in E p is the visual constraint term, E g are RTK constraints, which satisfy:

[0114]

[0115] in, represents the photometric error produced by the left camera, C i Indicates key frame I i The corresponding RTK position, Indicates key frame I i Corresponding agricultural drone posture.

[0116] In this embodiment, considering the binocular constraint, all feature points are projected into the right eye image. As with the left eye, an equivalent residual function is constructed:

[0117]

[0118] Then, the DSO algorithm is used to solve the visual positioning-RTK joint optimization objective function to achieve real-time positioning of agricultural UAVs.

[0119] Through static binocular residuals, each feature point contains only one binocular constraint value. A binocular weighted hyperparameter is introduced to project all feature points into the right eye image. Accurate scale estimation can be achieved without introducing additional hyperparameters, which has better adaptability to the scene and achieves higher UAV positioning accuracy.

[0120] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A real-time positioning method for agricultural drones in row sowing based on visual tracking, characterized in that: The following steps are involved: S1: Collect binocular images of farmland scenes taken by agricultural drones, IMU data synchronized with the agricultural drones, and RTK observation data of the corresponding positions of the agricultural drones; S2: Determine the key frame and sliding window of the binocular image based on the initial posture information between binocular image frames provided by the farmland scene binocular image and IMU data; S3: Project the feature points corresponding to all key frames in the sliding window to the current frame, and construct a multi-frame visual constraint function to estimate the current real-time posture of the agricultural UAV; In step S3, the process of projecting the feature points corresponding to all key frames in the sliding window to the current frame is as follows: Assume that the three-dimensional coordinates of any feature point among all the key frames are expressed as (X, Y, Z), and the intrinsic parameter matrix of the UAV's binocular camera is: Among them, (f x ,f y ) is the focal length, (c x ,c y ) is the main point of the binocular camera; let the feature point pixel u k The reprojected coordinates in the current frame are u ′ k =(u,v), then: The constructed multi-frame visual constraint function is: in, ρ k Represents the feature point pixel u k The corresponding inverse depth value, represents the projection function; (a t ,b t ) represents the photometric affine parameter at time t, e (·) Represents the exponential map, I t represents the reference frame, ω k Represents the residual term; N u Represents the set of feature points corresponding to all key frames; S4: insert the current frame corresponding to the feature point that meets the stable tracking conditions into the sliding window as a new key frame; S5: Initialize the content of the new key frame and remove abnormal feature points in the key frame; S6: Integrate RTK observation data, construct and solve the visual positioning-RTK joint optimization objective function, and realize real-time positioning of agricultural UAVs.

2. The real-time positioning method for agricultural drones based on visual tracking according to claim 1 is characterized in that: According to the difference between the frame images of the binocular image of the farmland scene, the key frame is determined, and the first n key frames are intercepted to form a sliding window. The sliding window includes n key frames K and m feature points L that can characterize the characteristics of the binocular image corresponding to the key frames, satisfying: K = {K1, K2, ..., K n }、L={l1,l2,…,l m }.

3. The real-time positioning method for agricultural drones based on visual tracking according to claim 2 is characterized in that: Based on the data of the IMU synchronized with the agricultural UAV, an integration operation is performed to obtain the current frame I of the binocular image of the farmland scene c The initial relative motion between the key frame K of the sliding window is denoted as T = {T c1 ,T c2 ,…,T cn }, T ci Indicates the time from the i-th key frame to the current frame I c The initial relative motion between, i = 1, 2, ..., n.

4. The real-time positioning method for agricultural drones based on visual tracking according to claim 3 is characterized in that: The weighted method based on the fusion of Gaussian distribution and t-distribution is used to calculate the residual term ω. k Make improvements to meet the following requirements: in, Represents the Gaussian distribution in the residual term ω k The weight on Represents the Gaussian distribution in the residual term ω k The weight on Represents the feature point pixel u k The corresponding gradient value, c represents a constant, v, σ k Represents the variance and degree of freedom of t-distribution, r k Represents the residual.

5. The real-time positioning method for agricultural drones based on visual tracking according to claim 4 is characterized in that: The stable tracking conditions are: Among them, m represents the total number of feature points, num represents the number of feature points that can be stably tracked, and Q min Indicates the minimum tracking ratio threshold.

6. The real-time positioning method for agricultural drones based on visual tracking according to claim 5 is characterized in that: Initialization of the new keyframe content includes: selection of gradient points, extraction of image blocks, and depth initialization of feature points. The selection of gradient points and extraction of image blocks use the DSO algorithm. When initializing the depth of feature points, the fixed baseline of the binocular camera is used to provide an initial depth range for the depth value of the feature points.

7. The real-time positioning method for agricultural drones based on visual tracking according to claim 5 is characterized in that: The process of removing abnormal feature points in key frames is as follows: Based on the residual r k Gradient value corresponding to the feature point The constraint relationship between them is used to determine whether an abnormal feature point is an abnormal feature point. The constraint relationship expression is: If the constraint relationship is satisfied, the feature point is an abnormal feature point; otherwise, the feature point is a normal feature point.

8. The real-time positioning method for agricultural drones based on visual tracking according to claim 5 is characterized in that: The visual positioning-RTK joint optimization objective function is: in, Indicates that except for key frame K i All other keyframes except Indicates key frame I j Effective projection to keyframe I i The feature points in E p is the visual constraint term, E g are RTK constraints, which satisfy: in, represents the photometric error produced by the left camera, C i Indicates key frame I i The corresponding RTK position, Indicates key frame I i Corresponding agricultural drone posture.

9. The real-time positioning method for agricultural drones based on visual tracking according to claim 8, characterized in that: Considering the binocular constraint, all feature points are projected into the right eye image. As with the left eye, an equivalent residual function is constructed: Then, the DSO algorithm is used to solve the visual positioning-RTK joint optimization objective function to achieve real-time positioning of agricultural UAVs.

Citation Information

Patent Citations

  • Mapping method and system based on GPS, IMU and binocular vision

    CN109991636A

  • Quad-rotor unmanned aerial vehicle visual target tracking method based on binocular camera

    CN110222581A