Unmanned ground vehicle three-dimensional navigation method based on low-altitude remote sensing and deep learning
By combining low-altitude remote sensing and deep learning technology in agricultural vehicle navigation, the three-dimensional navigation path of the orchard is generated, which solves the problems of low navigation accuracy and safety hazards in the existing technology, and realizes high-precision and safe three-dimensional navigation path planning.
Patent Information
- Application Number
- CN202510133347.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
AI Technical Summary
In autonomous navigation of agricultural vehicles, especially in large-scale and complex orchard terrain, the prior art is difficult to generate reliable and safe global three-dimensional paths, resulting in low navigation accuracy and safety risks.
The three-dimensional navigation method of unmanned ground vehicles based on low-altitude remote sensing and deep learning is adopted. The low-altitude remote sensing data of orchard plots is collected through drones, converted into digital surface model and digital orthograph image, and the 3D orchard scene image is reconstructed. The multi-scale sliding window and improved LS-YOLO algorithm are used to detect the canopy target, integrate the detection anchor frame, extract navigation points, and generate the vehicle's three-dimensional navigation path planning results.
It realizes high-precision three-dimensional navigation path planning in complex orchard terrain, improves navigation accuracy and security, and solves the safety hazards of traditional two-dimensional path extraction methods in uneven terrain.
Smart Images

Figure CN119958569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional navigation path planning, and in particular to a three-dimensional navigation method for an unmanned ground vehicle based on low-altitude remote sensing and deep learning. Background Art
[0002] Global food security is currently affected by many factors, but the current reliance on human labor and the combination of human labor and machinery is not enough to meet the urgent needs of agricultural food production. Autonomous navigation technology is the basis for agricultural robots to operate autonomously in the field. Automated navigation vehicles can improve the efficiency and accuracy of agricultural activities such as crop growth monitoring, pesticide spraying, weeding and fertilization. The latest technology can guide robots to safely cross farmland without damaging crops.
[0003] With the development of global satellite positioning systems and the continuous improvement of positioning accuracy, satellite systems including global navigation satellite systems (GNSS), GPS and BeiDou have been widely integrated into agricultural autonomous driving navigation operations. Although the global navigation satellite system (GNSS) itself can provide meter-level accuracy, this is not enough for many professional applications. By combining GNSS with real-time kinematic (RTK) technology, centimeter-level accuracy can be achieved to meet the needs of most agricultural tasks. The navigation system is an important component of agricultural robots and directly affects the operation efficiency. Automatic navigation of agricultural machinery based on the global satellite positioning system is a key aspect of the development of smart agriculture.
[0004] Autonomous navigation technology plays a vital role in various agricultural operations and covers a variety of methods. These methods include satellite navigation, SLAM-based lidar navigation, multi-sensor fusion navigation, and artificial intelligence-based visual navigation. However, there are multiple limitations in these autonomous driving navigation methods. Satellite navigation using real-time kinematic differential (RTK) sensors can provide high-precision positioning for agricultural machinery. However, in open areas, signal reflections can cause positioning errors, ignore crop information, and cannot avoid damage to crops during operation. This method is difficult to meet the needs of high-precision agricultural operations and is susceptible to signal interference. Although SLAM lidar navigation is effective, it is expensive and requires frequent updates and map generation. For large-scale scenarios, it also places high demands on on-board equipment. Multi-sensor fusion navigation usually integrates global navigation satellite systems (GNSS), inertial navigation systems (INS), inertial measurement units (IMUs), laser scanners, and cameras to achieve more reliable navigation.
[0005] In recent years, deep learning has become a disruptive technology that has revolutionized multiple fields including machine vision, natural language processing, and autonomous systems. Object detection is a key branch of machine vision that involves identifying and locating target objects in images or video frames. This technology has a wide range of applications. Currently, object detection technology is mainly divided into two methods: single-stage detection and two-stage detection. Single-stage detection methods, such as SSD, YOLO, and RetinaNet, can directly predict object locations without a separate candidate region generation stage. In contrast, two-stage detection methods such as Fast R-CNN, Faster R-CNN, and Cascade R-CNN generate region proposals before object detection. In general, single-stage methods are faster than two-stage methods, although the latter are superior in accuracy. The latest YOLO model achieves fast detection speed while maintaining high accuracy. Due to its strong overall performance, YOLO is widely used in various tasks in different fields.
[0006] Machine vision has been widely used in agricultural robot navigation due to its irreplaceable visual information and low hardware cost. In agricultural machinery management operations, visual navigation plays a vital role in improving operation efficiency, reducing field crop damage, and promoting unmanned agricultural operations. Xu et al. proposed a new vision-based autonomous navigation system that uses low-cost stereo cameras and inertial measurement units (IMUs) to achieve fully autonomous navigation of tractors. In addition, drones have the advantages of high spatial resolution, low application cost, strong timeliness, and reusability. When it is combined with computer vision technology, machine vision can extract accurate automatic navigation route information from centimeter-level aerial images. This method combines the position and orientation system (POS) with a visual sensor to simultaneously collect data containing crop information and spatial position details.
[0007] UAV-based crop row detection technology is widely used in field management. By utilizing relative positioning, crop row detection methods can extract guidance information from field roads and provide field navigation data for autonomous robots. Many researchers have proposed row detection navigation methods for various crops and terrains. In terms of object detection, many methods have been used to extract navigation lines, such as using enhanced YOLOv5 to identify and extract feature points representing pineapple rows, combining UAV survey technology with YOLOv5s object detection algorithm, and proposing a navigation method for paddy field management based on an improved CS-YOLOv5 model for rice seedling detection to extract seedling strip lines, and an enhanced YOLOv5 model for real-time rice seedling recognition to extract seedling navigation lines. To overcome the limitations of traditional machine vision technology in crop row recognition, many studies have explored advanced segmentation algorithms to identify crop rows and navigation paths. For example, a Transformer-based semantic segmentation model, a vision-based autonomous navigation system that utilizes stereo cameras and inertial measurement units (IMUs), an enhanced U-Net fused with ResNet-50 and attention mechanism, a method that combines vegetation index with ridge segmentation, an improved multi-scale efficient residual factorization convolutional neural network (MS-ERFNet) model, a multi-perturbation semi-supervised learning model, an optimized UNet, and a neural network integrated with pixel scanning for navigation line extraction. However, previous drone-based crop row extraction methods mainly rely on 2D plane extraction. In complex terrains such as mountains, hills, or orchards with uneven terrain, 2D extraction of navigation paths may pose safety risks to agricultural vehicle operations.
[0008] During the autonomous cruising of agricultural machinery, vehicles often operate in uneven and complex terrain, especially in hilly areas. Traditional autonomous driving trajectory prediction methods are mainly designed for urban roads. However, applying 2D path extraction methods to agricultural environments may lead to unsafe autonomous driving navigation of vehicles during operation. Recent studies have attempted to address this problem by developing deep learning (DL)-based models that use bird's-eye view (BEV) techniques to infer 3D lane information from 2D planar data. 3D-LaneNet is a new neural network that is able to predict the 3D layout of lanes directly from a single image of a road scene. It represents the first attempt to use on-board sensors for lane detection without assuming constant lane width or relying on pre-drawn environment maps. The BEV-LaneDet method uses techniques such as keypoint representation and spatial transformation pyramids to achieve 3D lane detection from monocular images through a virtual camera. Another important aspect of 3D autonomous driving navigation is to predict and obtain navigation routes directly in 3D space. LATR is a novel end-to-end 3D lane detection model specifically designed to detect lanes directly from monocular images. In addition, there are other different methods to extract 3D routes, such as generating globally consistent 3D maps and extracting lane lines from high-definition aerial images. However, when these methods are applied to agricultural scenarios, their complexity and expressiveness are limited. In addition, they rely on on-board local visual sensors, which limits prediction accuracy and perception capabilities. By leveraging the advantages and perspectives of drones in building 3D maps, plan and elevation maps of orchards can be generated, thereby enabling 3D path planning. Compared with 2D path detection, 3D path detection has two key advantages. First, it provides more accurate path location and direction information, thereby improving the navigation accuracy of agricultural vehicles. Second, 3D path detection provides additional data on road geometry, such as curvature and slope, which can be used to predict vehicle behavior and enhance safety, especially when road conditions change rapidly.
[0009] Despite significant progress in autonomous driving navigation technology, agricultural vehicle navigation, especially in large-scale and complex orchard terrains, still faces several unresolved challenges. These challenges include the following aspects:
[0010] (1) Large-scale planting makes the traditional ground-based manual waypoint navigation method insufficient, while the drone-based path extraction method is more suitable for large-scale applications.
[0011] (2) The equipment performance is limited and cannot process high-resolution, large-scale aerial images from large planting areas, which hinders the extraction of navigation routes.
[0012] (3) The background of aerial images of orchards is complex and contains a lot of redundant information, which is not conducive to the detection of detection algorithms.
[0013] (4) In large-scale aerial images, the targets are often small, and the images need to be compressed into a smaller size and input into the detection model, which will result in a smaller receptive field and lower detection accuracy.
[0014] (5) Traditional large image slice-assisted reasoning methods may lead to loss of spatial continuity when processing large images, as well as false detection of edge detection objects when slices are stitched together.
[0015] (6) Compared with the wide field of view provided by drones, ground robots have limited perception capabilities.
[0016] (7) In hilly terrain, two-dimensional path extraction methods pose safety risks during agricultural vehicle operations.
[0017] In summary, in the field of autonomous navigation of agricultural vehicles, there is a lack of methods that can generate reliable and safe global three-dimensional paths. Summary of the invention
[0018] In view of the above-mentioned deficiencies in the prior art, the present invention provides a three-dimensional navigation method for unmanned ground vehicles based on low-altitude remote sensing and deep learning.
[0019] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0020] A three-dimensional navigation method for an unmanned ground vehicle based on low-altitude remote sensing and deep learning comprises the following steps:
[0021] Use drones to collect low-altitude remote sensing data of orchard plots;
[0022] Convert low-altitude remote sensing data of orchard plots into digital surface models and digital orthophotos, and reconstruct 3D orchard scene images;
[0023] The 3D orchard scene image is segmented using a multi-scale sliding window to obtain a multi-scale segmented image;
[0024] Perform tree crown target detection on multi-scale segmented images respectively to obtain canopy detection anchor frames of segmented images at different scales;
[0025] The canopy detection anchor frames of segmented images at different scales are fused and mapped back to the 3D orchard scene image to obtain the final canopy detection anchor frame.
[0026] Navigation points are extracted based on the final canopy detection anchor frame to generate the vehicle's three-dimensional navigation path planning results.
[0027] Furthermore, the 3D orchard scene image is segmented using a multi-scale sliding window to obtain a multi-scale segmented image, including:
[0028] Using the 3D orchard scene image reconstructed from the image collected by the drone as the slice base, setting the first sliding window size to correspond to the size of the 3D orchard scene image, and obtaining a first-scale segmented image;
[0029] Using the 3D orchard scene image reconstructed from the drone image as the baseline, the size of the second sliding window is set to correspond to the size of the drone image to obtain a second-scale segmented image;
[0030] Using the 3D orchard scene image reconstructed from the image collected by the UAV as a baseline, setting the size of the third sliding window to a first ratio of the size of the image collected by the UAV, and obtaining a third-scale segmented image;
[0031] The 3D orchard scene image reconstructed by the image collected by the UAV is used as the baseline, and the size of the fourth sliding window is set to the second ratio of the size of the image collected by the UAV to obtain a fourth-scale segmented image.
[0032] Furthermore, when a multi-scale sliding window is used for segmentation, a set overlap ratio is maintained between consecutive segmented images.
[0033] Furthermore, tree crown target detection is performed on the multi-scale segmented images respectively to obtain the canopy detection anchor frames of the segmented images at different scales, including:
[0034] Constructing a tree crown target detection model; the tree crown target detection model backbone network, neck network and prediction network;
[0035] The backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps;
[0036] The neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map;
[0037] The prediction network is used to generate canopy detection anchor boxes of the corresponding scale segmented image based on the enhanced feature map.
[0038] Furthermore, the backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps, including:
[0039] A first deep convolution module, a second deep convolution module, a first C2F module, a first feature enhancement module, a second C2F module, a third deep convolution module, a second feature enhancement module, a fourth deep convolution module, a third feature enhancement module and an SPPF module are used in sequence to perform feature extraction on segmented images of different scales, and a first-scale feature map is output through the second C2F module, a second-scale feature map is output through the second feature enhancement module, and a third-scale feature map is output through the SPPF module.
[0040] Furthermore, the neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map, including:
[0041] Performing an upsampling operation on the third scale feature map using the first upsampling module;
[0042] Using the first feature enhancement and splicing module, the output feature map of the first upsampling module and the second scale feature map are subjected to feature enhancement and splicing operations;
[0043] Using the first lightweight feature extraction module to extract features from the output feature map of the first feature enhancement splicing module;
[0044] Using the second upsampling module to perform an upsampling operation on the output feature map of the first lightweight feature extraction module;
[0045] Using a second feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second upsampling module and the first scale feature map;
[0046] Using the second lightweight feature extraction module to extract features from the output feature map of the second feature enhancement splicing module, the first enhanced feature map is obtained and output to the prediction network;
[0047] Performing a convolution operation on the first enhanced feature map using a first 2D convolution module;
[0048] Using a third feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the first 2D convolution module and the output feature map of the first lightweight feature extraction module;
[0049] Using the third lightweight feature extraction module to extract features from the output feature map of the third feature enhancement splicing module, the obtained second enhanced feature map is output to the prediction network;
[0050] Performing a convolution operation on the second enhanced feature map using a second 2D convolution module;
[0051] Using the fourth feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second 2D convolution module and the third scale feature map;
[0052] The fourth lightweight feature extraction module is used to extract features from the output feature map of the fourth feature enhancement splicing module to obtain a third enhanced feature map, which is output to the prediction network.
[0053] Furthermore, the canopy detection anchor frames of segmented images of different scales are fused to obtain the final canopy detection anchor frame, including:
[0054] The canopy detection anchor frames of segmented images of different scales are mapped to the 3D orchard scene image for multi-scale fusion, and the non-maximum suppression method is used to screen the fused anchor frames to obtain the final canopy detection anchor frames.
[0055] Furthermore, the navigation points are extracted according to the final canopy detection anchor frame, including:
[0056] According to the final canopy detection anchor frame, the center point is calculated, a center point is selected as the origin, a circular area is established with the set threshold as the radius, the center point of the anchor frame in the circular area is selected as the associated point, and the midpoint between the origin and the associated point is calculated as the navigation point.
[0057] Furthermore, a vehicle three-dimensional navigation path planning result is generated, including:
[0058] The navigation points are connected in sequence to generate a navigation route, the navigation route is calculated and converted into real-world geographic coordinates, and the elevation channel value of a single pixel is extracted from the digital surface model, and the elevation channel value of a single pixel is linked to the navigation point to generate a vehicle three-dimensional navigation path planning result.
[0059] The present invention has the following beneficial effects:
[0060] The present invention adopts a multi-scale fusion sliding window method to process large-scale images and combines it with an enhanced LS-YOLO algorithm for orchard canopy detection, using data acquired by drones to accurately predict canopy center points, navigation points, and comprehensive 3D paths. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 A flowchart of a three-dimensional navigation method for unmanned ground vehicles based on low-altitude remote sensing and deep learning;
[0062] Figure 2 This is a schematic diagram of the workflow for orthophoto modeling and reconstruction based on low-altitude UAV remote sensing;
[0063] Figure 3 Schematic diagram of large-scale fusion sliding window for tree crown detection method;
[0064] Figure 4 Schematic diagram of sliding windows with different overlap ratios;
[0065] Figure 5 This is a schematic diagram of the data set processing flow;
[0066] Figure 6 This is a schematic diagram of the LS-YOLO detection model structure;
[0067] Figure 7 Extract process diagrams for navigation points;
[0068] Figure 8 Schematic diagram of post-processing of canopy midpoints and orchard roadpoints extraction. DETAILED DESCRIPTION
[0069] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0070] For fruit tree navigation, the generated path should avoid the main body of the fruit tree and pass between the trees. By combining low-altitude UAV remote sensing technology, it is expected to achieve global three-dimensional path planning for agricultural vehicles. The present invention proposes a three-dimensional navigation method for unmanned ground vehicles based on low-altitude remote sensing and deep learning. The method first uses drones to collect low-altitude remote sensing data of orchard plots; then converts the collected data into DSM and TDOM formats; then constructs an algorithm training data set, combines the improved LS-YOLO algorithm with the sliding window method, and predicts large-scale orchard crown information; then, uses self-developed software to extract three-dimensional navigation lines; finally, the navigation data is transmitted to the autonomous driving vehicle in KML format for navigation.
[0071] like Figure 1 As shown, an unmanned ground vehicle three-dimensional navigation method based on low-altitude remote sensing and deep learning provided by an embodiment of the present invention includes the following steps S1 to S6:
[0072] S1. Collect low-altitude remote sensing data of orchard plots using drones;
[0073] In an optional embodiment of the present invention, step S1 is for citrus trees. From the perspective of the drone, the crown of citrus trees is oval and is usually planted on a large scale in plains or rugged hilly areas. The experimental field is located in Ji'an City, Jiangxi Province, China. Aerial images were collected from an orchard with a total area of 46,783 square meters under clear and cloudy conditions. Figure 2As shown in the figure, (A) is a satellite image of a plot in Ji'an City, Jiangxi Province, (B) is an aerial image of an orchard taken by a drone, (C) is the planning information of drone aerial photography operations, (D) is the posture of the drone's onboard camera, (E) is the type and configuration of the drone, (F) is the 3D orchard scene reconstructed by the drone, (G) is the TDOM map, and (H) is the DSM map. A drone equipped with a 20-megapixel visible light camera flew over the orchard according to a preprogrammed flight path. After manually determining the coordinates of the boundary of the survey area, the flight trajectory was generated using DJI flight planning software. The drone then took images regularly at each specified point along this path. Image acquisition was completed in a single flight, while recording position and orientation (POS) data. To ensure the best ground resolution, the drone flew at an altitude of 25 meters with an 80% overlap of the side and front. After data collection, images with exposure problems, focus errors, or motion blur were manually discarded.
[0074] S2, converting the low-altitude remote sensing data of the orchard into digital surface models and digital orthophotos, and reconstructing 3D orchard scene images;
[0075] In an optional embodiment of the present invention, step S2 uses photogrammetry technology to generate a digital surface model (DSM) and a true digital orthophoto map (TDOM) of the orchard. The DSM provides digital elevation data, including ground elevation and tree height, while the TDOM provides detailed information on the appearance of the terrain and tree crowns. Finally, a 3D orchard scene with latitude, longitude and altitude is generated.
[0076] S3, segmenting the 3D orchard scene image using a multi-scale sliding window to obtain a multi-scale segmented image;
[0077] In an optional embodiment of the present invention, step S3 uses a multi-scale sliding window to segment the 3D orchard scene image to obtain a multi-scale segmented image, including:
[0078] Using the 3D orchard scene image reconstructed from the image collected by the drone as the slice base, setting the first sliding window size to correspond to the size of the 3D orchard scene image, and obtaining a first-scale segmented image;
[0079] Using the 3D orchard scene image reconstructed from the drone image as the baseline, the size of the second sliding window is set to correspond to the size of the drone image to obtain a second-scale segmented image;
[0080] Using the 3D orchard scene image reconstructed from the image collected by the UAV as a baseline, setting the size of the third sliding window to a first ratio of the size of the image collected by the UAV, and obtaining a third-scale segmented image;
[0081] The 3D orchard scene image reconstructed by the image collected by the UAV is used as the baseline, and the size of the fourth sliding window is set to the second ratio of the size of the image collected by the UAV to obtain a fourth-scale segmented image.
[0082] In particular, when this patent uses a multi-scale sliding window for segmentation, a set ratio of overlap is maintained between consecutive segmented images.
[0083] In this example, recent advances in deep learning have enabled many tasks to be performed efficiently. However, automated navigation of large orchards requires accurate identification of tree crowns in large maps. The YOLO object detection algorithm requires the input image to be resized to 640×640 pixels for efficient processing. Unfortunately, small object crowns in large format images are often poorly predicted. Therefore, detecting small tree crowns from large format images from the perspective of a drone is a major challenge.
[0084] In order to solve the problem of detecting small tree crowns in large-format images, this embodiment proposes a multi-scale fusion sliding window method. The method mainly consists of two parts: sliding window segmentation and multi-scale fusion. The method divides the large-format image into smaller images as the input of the LS-YOLO algorithm, and outputs the canopy detection results through post-processing. The LS-YOLO algorithm can detect tree crowns more effectively by inputting small fragments; in these smaller fragments, the canopy of small objects occupies a larger pixel area. The large-scale fusion sliding window of the tree crown detection method is as follows: Figure 3 As shown in the figure, (A) is the tree crown detection process of large-scale sliding window fusion, (B) is the multi-scale fusion sliding window method, (C) is the LS-YOLO algorithm reasoning, and (D) is the post-processing operations such as reasoning and NMS processing.
[0085] The 46,783 square meter orchard large-scale map stitched by drone aerial photography was used as the slicing base, and sliding windows of different sizes were set on the 40099×37272 pixel large-scale map for tree crown detection. The LS-YOLO algorithm was trained on a 5472x 3648 pixel dataset captured by the drone, which served as a baseline for setting the sliding window size. The window size was set to correspond to the large-format image pixels, as well as 90%, 100%, and 110% of the drone image size. The method uses a sliding window method to systematically move the window from left to right and from top to bottom to extract fixed-size fragments.
[0086] In order to extract the route more accurately, setting an appropriate slice window overlap ratio can improve the performance of canopy detection. Setting an unreasonable window overlap ratio will result in predicting multiple anchor box targets for a single fruit tree canopy, which is not conducive to route extraction. Sliding windows with different overlap ratios, such as Figure 4As shown. To solve the problem of inaccurate canopy detection due to gaps between windows, this embodiment maintains a 20% overlap between consecutive segments. Afterwards, the LS-YOLO algorithm compresses the slice image to 640*640 and predicts the canopy detection anchor box for each segment.
[0087] S4, performing tree crown target detection on the multi-scale segmented images respectively, and obtaining the canopy detection anchor frames of the segmented images at different scales;
[0088] In an optional embodiment of the present invention, step S4 performs tree crown target detection on the multi-scale segmented images respectively to obtain canopy detection anchor frames of the segmented images at different scales, including:
[0089] Constructing a tree crown target detection model; the tree crown target detection model backbone network, neck network and prediction network;
[0090] The backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps;
[0091] The neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map;
[0092] The prediction network is used to generate canopy detection anchor boxes of the corresponding scale segmented image based on the enhanced feature map.
[0093] Among them, the backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps, including:
[0094] A first deep convolution module, a second deep convolution module, a first C2F module, a first feature enhancement module, a second C2F module, a third deep convolution module, a second feature enhancement module, a fourth deep convolution module, a third feature enhancement module and an SPPF module are used in sequence to perform feature extraction on segmented images of different scales, and a first-scale feature map is output through the second C2F module, a second-scale feature map is output through the second feature enhancement module, and a third-scale feature map is output through the SPPF module.
[0095] Among them, the neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map, including:
[0096] Performing an upsampling operation on the third scale feature map using the first upsampling module;
[0097] Using the first feature enhancement and splicing module, the output feature map of the first upsampling module and the second scale feature map are subjected to feature enhancement and splicing operations;
[0098] Using the first lightweight feature extraction module to extract features from the output feature map of the first feature enhancement splicing module;
[0099] Using the second upsampling module to perform an upsampling operation on the output feature map of the first lightweight feature extraction module;
[0100] Using a second feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second upsampling module and the first scale feature map;
[0101] Using the second lightweight feature extraction module to extract features from the output feature map of the second feature enhancement splicing module, the first enhanced feature map is obtained and output to the prediction network;
[0102] Performing a convolution operation on the first enhanced feature map using a first 2D convolution module;
[0103] Using a third feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the first 2D convolution module and the output feature map of the first lightweight feature extraction module;
[0104] Using the third lightweight feature extraction module to extract features from the output feature map of the third feature enhancement splicing module, the obtained second enhanced feature map is output to the prediction network;
[0105] Performing a convolution operation on the second enhanced feature map using a second 2D convolution module;
[0106] Using the fourth feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second 2D convolution module and the third scale feature map;
[0107] The fourth lightweight feature extraction module is used to extract features from the output feature map of the fourth feature enhancement splicing module to obtain a third enhanced feature map, which is output to the prediction network.
[0108] This embodiment uses a deep learning-based target detection method to extract navigation paths by identifying tree crowns. The process starts with data collection, using drones to capture aerial images of citrus tree crowns in an orchard. The canopy data is then manually filtered, and the tree crowns are annotated with bounding boxes using the Labelme tool. The dataset is randomly divided into training, validation, and test sets in a ratio of 8:1:1. Prior to training, the data is preprocessed using various techniques to improve data quality and increase sample size. A dataset consisting of 1,446 images taken by a drone is generated, containing 58,762 tree crown target objects. The dataset processing flow is as follows: Figure 5 shown.
[0109] The extraction of autonomous navigation lines in this embodiment is closely related to the tree planting pattern in the orchard, because the path must be planned between the trees without causing damage. The key step of this process is the detection of tree crowns based on deep learning and drones. Many studies have applied the YOLO series of algorithms to various tasks because of their high performance and fast detection speed.
[0110] The YOLO framework consists of several parts, including the input layer, backbone, neck, and prediction head. Although the YOLO series shows satisfactory performance on general datasets, its application in agricultural environments and single-target detection tasks still faces challenges such as low detection accuracy and redundant model parameters.
[0111] Based on the powerful YOLO series framework, this embodiment uses YOLOv8 as the basic model and enhances it to accurately detect tree crowns. This embodiment proposes a tree crown detection LS-YOLO model based on the perspective of a drone, such as Figure 6 As shown. Given the low proportion of small targets in the drone acquisition data set, this embodiment integrates the FEM module to improve the detection of small targets from the perspective of the drone. In addition, this embodiment optimizes the DWConv module in the backbone network and the neck network, and introduces a new module fastnext. The integration of fastnext and DWConv modules into the backbone and neck significantly reduces model redundancy. In addition, this embodiment introduces the FFM_Concat2 module, which is a custom neural network module for feature fusion. It accepts two input feature maps, then connects the weighted feature maps along the specified dimension, and outputs fused features, which improves the information fusion ability of the neck network, maximizes the integration of features from different sources, and improves the overall performance of the model. Finally, this embodiment uses the CIoU loss function to fit the true value as the loss function, which improves the convergence speed and accuracy of the model. These enhancements improve the detection capability of low-altitude remote sensing for small targets while maintaining a lightweight structure.
[0112] The deep learning-based target detection method relies on the backbone to extract high-dimensional features. In order to reduce redundant calculations and improve the detection capability of small targets, this embodiment introduces DWConv and FEM modules in both the backbone network and the neck network. In the backbone, this embodiment alternately applies DWConv, FEM, and C2F modules to the canopy dataset to better match the aspect ratio of the canopy anchor frame and facilitate feature learning. The introduction of the DWConv module simplifies the backbone network and reduces redundant structures, while the FEM module enhances the extraction and learning of small target features, such as the crown of a tree from the perspective of a drone. The lightweight FEM module captures richer local contextual information through a multi-branch structure that combines standard convolution and dilated convolution. To further optimize feature learning, the DWConv module is repeated once in the backbone and the FEM module is repeated once in a 1-6-3 pattern. In the neck network, the FFM_Concat2 module is repeated once and the FasterNeXt module is repeated 3 times. In order to make full use of the features of all channels, 1×1conv is integrated into the FasterNet block within the residual structure to form a feature pyramid, so that the improved FasterNeXt maintains lightweight features and powerful feature extraction capabilities. These enhancements improve the ability of the LS-YOLO model to learn tree crown features, enabling it to effectively handle the complexity of the orchard environment while remaining lightweight.
[0113] In the YOLO series of models, the anchor box is usually determined by clustering on the COCO dataset, and the size of the anchor box affects the convergence speed and prediction accuracy of the model. This embodiment adopts an anchor-free mechanism and abandons the traditional predefined anchor box, thereby greatly simplifying the anchor-based detection process. The anchor-free method bypasses the preset anchor points and directly predicts the position and size of objects on the feature map. The network only predicts the center, width, and height of the object instead of adjusting the bounding box based on the anchor point. This anchor box simplification not only simplifies the structure of the canopy detection model, but also improves training efficiency and enhances adaptability.
[0114] The loss function quantifies the difference between the model's prediction and the manually determined value. In the tree crown detection task, the loss function is used to estimate the deviation between the predicted bounding box and the ground truth annotation. Two loss calculation methods are used to evaluate the LS-YOLO loss: localization loss and classification loss. The localization loss usually uses the sum of squared errors to evaluate the position deviation of the bounding box, while the classification loss measures the difference between the predicted object category and the actual category. The calculation method is:
[0115]
[0116] Where W coord is the positioning loss weight, F represents the feature map, and B represents the number of bounding boxes at each predicted position. and correspond to the center of the actual bounding box, and the coordinates and Corresponding to the predicted bounding box respectively. Indicator function I ij Indicates whether the i position and j bounding box contain an object. In addition, C represents the total number of classes, p i (c) is the model’s predicted probability that location i and bounding box j belong to class C, and is the probability of the actual class. Finally, the two positioning losses and classification losses are added together to get the total loss.
[0117] The improved detection network was trained on a desktop computer running Ubuntu 23.04 equipped with an Intel i9-13900K CPU, 64gb RAM, and an NVIDIA 4090 GPU with 24gb memory. The experimental code uses CUDA 12.1, cuDNN 9.4.0, Python 3.10, PyTorch 2.1.0, and VSCode. The input size of the network is set to 640×640 and the batch size is 16. The Adam optimizer with stochastic gradient descent is used, with an initial learning rate of 0.01, a weight decay of 0.0005, and a momentum of 0.937.
[0118] The performance of the improved LS-YOLO model was evaluated using precision, recall, and mean average precision (mAP) as key indicators for tree crown detection. Precision, recall, and mAP were calculated according to the following formulas. Wherein, true positive (TP) represents the correctly detected canopy, false positive (FP) represents the incorrectly detected canopy, and false negative (FN) represents the undetected canopy. Mean average precision (AP) is the area under the precision-recall curve and is used to evaluate the performance of LS-YOLO. In addition, mAP represents the average value of AP in all detection categories. Since only one category (tree crown) was detected in this study, N = 1. In addition, Param and weight size were selected as lightweight evaluation indicators for the LS-YOLO model. Parameter refers to the number of parameters in the network, and weight is the size of the algorithm weight. Wherein, O represents a constant order, K represents the size of the convolution kernel, C is the number of channels, M represents the size of the input image, and i is the number of iterations. The main purpose of this study is to verify the performance of the improved LS-YOLO model in tree crown detection.
[0119] Precision=TP / (TP+FP)
[0120] Recall=TP / (TP+FN)
[0121]
[0122] S5, fusing the canopy detection anchor frames of the segmented images at different scales, mapping them back to the 3D orchard scene image, and obtaining the final canopy detection anchor frame;
[0123] In an optional embodiment of the present invention, step S5 fuses the canopy detection anchor frames of the segmented images of different scales to obtain the final canopy detection anchor frame, including:
[0124] The canopy detection anchor frames of segmented images of different scales are mapped to the 3D orchard scene image for multi-scale fusion, and the non-maximum suppression method is used to screen the fused anchor frames to obtain the final canopy detection anchor frames.
[0125] After the LS-YOLO algorithm predicts sliding windows of different sizes, this embodiment integrates anchor frames of different sizes and maps them back to a larger orchard image for multi-scale fusion. By fusing sliding windows of different sizes, this method can fuse information across multiple scales, enhance the representation of canopy features, and improve robustness. This method improves the positioning accuracy of fruit tree canopy detection and can better identify canopies of different sizes.
[0126] When the algorithm predicts and fuses images of different sizes, a single crown in a large-scale image may generate multiple anchor frames. In order to solve the problem of anchor frame redundancy, this embodiment adopts an anchor frame-based method to merge multiple anchor frames into a large-scale image. Non-maximum suppression (NMS) is used to remove redundant frames, and finally the most accurate anchor frame prediction for each crown is achieved.
[0127] S6. Extract navigation points based on the final canopy detection anchor frame and generate a three-dimensional navigation path planning result for the vehicle.
[0128] In an optional embodiment of the present invention, step S6 extracts navigation points according to the final canopy detection anchor frame, including:
[0129] According to the final canopy detection anchor frame, the center point is calculated, a center point is selected as the origin, a circular area is established with the set threshold as the radius, the center point of the anchor frame in the circular area is selected as the associated point, and the midpoint between the origin and the associated point is calculated as the navigation point.
[0130] Step S6 generates the vehicle three-dimensional navigation path planning result, including:
[0131] The navigation points are connected in sequence to generate a navigation route, the navigation route is converted into real-world geographic coordinates, and the elevation channel value of a single pixel is extracted from the digital surface model, and the elevation channel value of a single pixel is linked to the navigation point to generate a vehicle three-dimensional navigation path planning result.
[0132] In order to accurately identify the center point between trees in the navigation route, navigation point prediction is performed after obtaining the canopy box. The centroid is calculated based on the center of the canopy, and LS-YOLO uses the size of the predicted bounding box to determine the midpoint. Given that there are multiple trees in a large-scale image, the associated points within the defined range are calculated by traversing the tree.
[0133] First, a tree is selected as the center point and its midpoint is used as the starting point. A predefined threshold is used as the radius to establish a circular area. All tree crowns within this circular range are regarded as associated points, and the midpoint between the associated point and the center point is calculated to derive the navigation point. By applying this method and selecting a suitable threshold, this embodiment can traverse each tree in the large-scale image and obtain associated points as candidate navigation points. The navigation point extraction process is as follows: Figure 7 As shown, (A) is the extraction of the center point through the prediction box, (B) is the screening of candidate crown-related points through threshold selection, and (C) is the calculation of the route point.
[0134] Then, combined with the LS-YOLO algorithm, the sliding window slicing method is used to generate the navigation point map of the large-scale image. Figure 8 As shown in the figure, (A) is the large-scale image extraction result, (B) is the enlarged large-scale image extraction result, (C) is the selection of automatic navigation waypoints, (D) is the automatic navigation route, and (E) is a three-dimensional map. The red dot represents the center of the tree crown, and the blue dot represents the predicted centroid between trees as the navigation point. Post-processing and path planning are performed based on the waypoint map to plan the navigation path of the autonomous vehicle.
[0135] In order to further improve the navigation accuracy of agricultural vehicles on uneven terrain and ensure the safety of driving behavior, autonomous navigation path planning is performed based on navigation points. First, a continuous sequence of navigation points is drawn, and then these navigation points are automatically connected to generate a navigation route. Finally, the generated 2D path is converted into real-world geographic coordinates and combined with the corresponding elevation data. The planned 2D path points are aligned with the digital surface model (DSM) coordinates based on latitude and longitude values. The elevation channel values of individual pixels are then extracted from the DSM and linked to these 2D path points. Finally, the generated 3D map information can be executed by the generated automatic guided vehicle. The resulting 3D map.
[0136] The main advantages of the present invention are as follows:
[0137] (1) Utilizing the high maneuverability, flexibility, and convenience of low-altitude remote sensing UAVs, the UAVs and ground robots are coordinated to solve the problem of limited sensing capabilities of ground robots in large-scale orchard navigation.
[0138] (2) Utilizing the high maneuverability, flexibility, and convenience of low-altitude remote sensing UAVs, coordinate UAVs and ground robots to solve the problem of limited sensing capabilities of ground robots in large-scale orchard navigation.
[0139] (3) A new multifunctional agricultural robot automatic navigation workflow is proposed. The workflow integrates canopy target detection, multi-scale sliding window and path planning software to generate a three-dimensional navigation path with a latitude and longitude error of only 9.89 cm.
[0140] (4) A low-altitude remote sensing dataset for small target tree crowns in orchards is introduced. A multi-scale fusion method is proposed, which improves the accuracy by 2.52% compared with the single-scale fusion method, effectively solving the challenges associated with large images, small targets, and tree segmentation.
[0141] (5) The LS-YOLO detection algorithm was enhanced using lightweight modules (FEM, DWConv, and FasterNext). This method improved the precision (P) from 83.10% to 84.90%, an improvement of 1.8%. In addition, the mAP50, mAP75, and mAP50-95 values increased from 83.10% to 84.00%, 62.12% to 63.93%, and 57.00% to 58.30%, respectively, an improvement of 0.9%, 1.81%, and 1.3%. This method enhanced the algorithm's ability to predict orchard canopies while adapting to different devices.
[0142] (6) The multi-scale sliding window slicing technique solves the problem of large-area images being difficult to process, namely, images with pixels that are too large to be processed directly, the problem of small targets in tree crowns, and the problem of segmenting individual tree bodies.
[0143] (7) A 3D path extraction software based on RTK vehicle was developed, which can extract latitude, longitude, and elevation data in different orchard environments.
[0144] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0145] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0147] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0148] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A three-dimensional navigation method for unmanned ground vehicles based on low-altitude remote sensing and deep learning, characterized in that: The following steps are involved: Use drones to collect low-altitude remote sensing data of orchard plots; Convert low-altitude remote sensing data of orchard plots into digital surface models and digital orthophotos, and reconstruct 3D orchard scene images; The 3D orchard scene image is segmented using a multi-scale sliding window to obtain a multi-scale segmented image; Perform tree crown target detection on multi-scale segmented images respectively to obtain canopy detection anchor frames of segmented images at different scales; The canopy detection anchor frames of segmented images at different scales are fused and mapped back to the 3D orchard scene image to obtain the final canopy detection anchor frame. Navigation points are extracted based on the final canopy detection anchor frame to generate the vehicle's three-dimensional navigation path planning results.
2. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1, characterized in that: The 3D orchard scene image is segmented using a multi-scale sliding window to obtain a multi-scale segmented image, including: Using the 3D orchard scene image reconstructed from the image collected by the drone as the slice base, setting the first sliding window size to correspond to the size of the 3D orchard scene image, and obtaining a first-scale segmented image; Using the 3D orchard scene image reconstructed from the drone image as the baseline, the size of the second sliding window is set to correspond to the size of the drone image to obtain a second-scale segmented image; Using the 3D orchard scene image reconstructed from the image collected by the UAV as a baseline, setting the size of the third sliding window to a first ratio of the size of the image collected by the UAV, and obtaining a third-scale segmented image; The 3D orchard scene image reconstructed by the image collected by the UAV is used as the baseline, and the size of the fourth sliding window is set to the second ratio of the size of the image collected by the UAV to obtain a fourth-scale segmented image.
3. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1 or 2, characterized in that: When using a multi-scale sliding window for segmentation, a set overlap ratio is maintained between consecutive segmented images.
4. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1, characterized in that: The tree crown target detection is performed on the multi-scale segmented images respectively, and the canopy detection anchor frames of the segmented images at different scales are obtained, including: Constructing a tree crown target detection model; the tree crown target detection model backbone network, neck network and prediction network; The backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps; The neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map; The prediction network is used to generate canopy detection anchor boxes of the corresponding scale segmented image based on the enhanced feature map.
5. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 4 is characterized in that: The backbone network is used to extract features from segmented images of different scales to obtain multi-scale feature maps, including: A first deep convolution module, a second deep convolution module, a first C2F module, a first feature enhancement module, a second C2F module, a third deep convolution module, a second feature enhancement module, a fourth deep convolution module, a third feature enhancement module and an SPPF module are used in sequence to perform feature extraction on segmented images of different scales, and a first-scale feature map is output through the second C2F module, a second-scale feature map is output through the second feature enhancement module, and a third-scale feature map is output through the SPPF module.
6. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 5, characterized in that: The neck network is used to enhance the multi-scale feature map to obtain an enhanced feature map, including: Performing an upsampling operation on the third scale feature map using the first upsampling module; Using the first feature enhancement and splicing module, the output feature map of the first upsampling module and the second scale feature map are subjected to feature enhancement and splicing operations; Using the first lightweight feature extraction module to extract features from the output feature map of the first feature enhancement splicing module; Using the second upsampling module to perform an upsampling operation on the output feature map of the first lightweight feature extraction module; Using a second feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second upsampling module and the first scale feature map; Using the second lightweight feature extraction module to extract features from the output feature map of the second feature enhancement splicing module, the first enhanced feature map is obtained and output to the prediction network; Performing a convolution operation on the first enhanced feature map using a first 2D convolution module; Using a third feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the first 2D convolution module and the output feature map of the first lightweight feature extraction module; Using the third lightweight feature extraction module to extract features from the output feature map of the third feature enhancement splicing module, the obtained second enhanced feature map is output to the prediction network; Performing a convolution operation on the second enhanced feature map using a second 2D convolution module; Using the fourth feature enhancement and splicing module to perform feature enhancement and splicing operations on the output feature map of the second 2D convolution module and the third scale feature map; The fourth lightweight feature extraction module is used to extract features from the output feature map of the fourth feature enhancement splicing module to obtain a third enhanced feature map, which is output to the prediction network.
7. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1, characterized in that: The canopy detection anchor frames of segmented images at different scales are fused to obtain the final canopy detection anchor frame, including: The canopy detection anchor frames of segmented images of different scales are mapped to the 3D orchard scene image for multi-scale fusion, and the non-maximum suppression method is used to screen the fused anchor frames to obtain the final canopy detection anchor frames.
8. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1, characterized in that: Extract navigation points based on the final canopy detection anchor frame, including: According to the final canopy detection anchor frame, the center point is calculated, a center point is selected as the origin, a circular area is established with the set threshold as the radius, the center point of the anchor frame in the circular area is selected as the associated point, and the midpoint between the origin and the associated point is calculated as the navigation point.
9. The method for three-dimensional navigation of an unmanned ground vehicle based on low-altitude remote sensing and deep learning according to claim 1, characterized in that: Generate vehicle 3D navigation path planning results, including: The navigation points are connected in sequence to generate a navigation route, the navigation route is calculated and converted into real-world geographic coordinates, and the elevation channel value of a single pixel is extracted from the digital surface model, and the elevation channel value of a single pixel is linked to the navigation point to generate a vehicle three-dimensional navigation path planning result.
Citation Information
Cited By
Orchard operation robot visual navigation image splicing method
CN120672564A
An image stitching method for visual navigation of an orchard working robot
CN120672564B