Rapid spruce height measurement method based on improved YOLOv11 model and unmanned aerial vehicle image
By improving the method of combining the YOLOv11 model with UAV imagery, the problems of high error, poor real-time performance, and high cost in spruce height measurement were solved, achieving high-precision, low-cost, and real-time spruce height measurement to meet the needs of forestry surveys.
Patent Information
- Application Number
- CN202511100998.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for measuring spruce height suffer from high errors, poor real-time performance, and high costs, making it difficult to meet the real-time and accuracy requirements of forestry surveys.
An improved YOLOv11 model was adopted, combined with UAV imagery, and the model's feature extraction capability for small targets was enhanced through camera distortion correction and global attention mechanism (GAM). A UAV remote sensing imagery dataset was constructed, and a target detection model was deployed on the Flask framework to achieve automated measurement of spruce tree height.
It significantly improves the accuracy and efficiency of spruce height measurement, reduces measurement costs, and meets the real-time requirements of forestry surveys.
Smart Images

Figure CN120913110A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical fields of unmanned aerial vehicle remote sensing, computer vision and forestry measurement, and relates to but is not limited to a spruce height rapid measurement method based on an improved YOLOv11 model and unmanned aerial vehicle images. BACKGROUND
[0002] Spruce is an important ecological and economic tree species in high-altitude areas of China. As an important distribution area of the Tianshan Mountain spruce, the forest area in the southern mountain area of Urumqi accounts for 57% of the Urumqi forest land, and is a unique ecological barrier in the arid region of northwest China. As a evergreen tree, the spruce can grow up to 20-30 meters high, and the highest can reach 60-70 meters, with significant water conservation, carbon sequestration and oxygen release, and biodiversity protection functions. According to the research in 2025, the biomass density of spruce forest in this area is 91.21 t / hm 2 , and the total carbon storage is 4.1 x 10 4 t, which is a key object of forest carbon sink assessment in Xinjiang. The height measurement is of great significance to forest resource management, carbon sink assessment and ecological protection.
[0003] In the prior art, spruce height measurement methods include manual measurement, laser ranging height measurement, satellite remote sensing measurement, and laser radar measurement. Manual measurement requires climbing or using a height measuring rod for individual operation, which is extremely inefficient in complex terrain and has an error rate as high as 18%; laser ranging height measurement has a large deviation in finding the single crown top and tree roots in complex terrain and dark light environment, and is also affected by personnel, with an error of about 5% from the true tree height; satellite remote sensing is limited by spatial resolution and cloud cover, making it difficult to accurately capture the single crown top; ground-based laser radar and unmanned aerial vehicle laser radar can generate millimeter-level precision three-dimensional point clouds, but the device cost is more than 200,000 yuan, and data processing requires professional software support, and the measurement of tree height requires a large amount of point cloud processing process, which cannot meet the real-time demand of forestry census. You Only Look Once (YOLO) v11, as a newer iteration model in the YOLO series, through the deep fusion of Cross-stage Partial Connection with Kernel size 2 (C3K2) modules, Spatial Pyramid Feature Fusion (SPFF) modules, and Convolutional Pyramid Spatial Attention (C2PSA) mechanisms, achieves more accurate local feature focusing in complex scenes (such as occlusion or dense targets), and to address the sample imbalance problem, the model integrates the focal loss function and dynamic class weight adjustment strategy, which improves the small sample class recall rate to 92%, but due to the 2-meter standard rod in the larger resolution unmanned aerial vehicle image, the pixel is small, which belongs to small target detection, resulting in poor detection accuracy of the YOLOv11 model for the standard rod.
[0004] Therefore, there is an urgent need for a more accurate spruce height measurement method to solve the problems of high error, poor real-time performance, and high cost in existing measurement methods, and to greatly improve the real-time performance and accuracy of spruce height measurement. SUMMARY
[0005] The embodiments of the present application provide a spruce height rapid measurement method based on an improved YOLOv11 model and unmanned aerial vehicle images.
[0006] The technical solution of the embodiments of the present application is as follows: In a first aspect, the embodiments of the present application provide a method for quickly measuring the height of spruce based on an improved YOLOv11 model and unmanned aerial vehicle images, which comprises: collecting unmanned aerial vehicle remote sensing images of spruce and standard rods in a sample area, correcting camera distortion of the unmanned aerial vehicle remote sensing images of the spruce and the standard rods, constructing an unmanned aerial vehicle remote sensing image dataset, and dividing the unmanned aerial vehicle remote sensing image dataset into an unmanned aerial vehicle remote sensing image training set, an unmanned aerial vehicle remote sensing image test set, and an unmanned aerial vehicle remote sensing image validation set; embedding a global attention mechanism (GAM) in a YOLOv11 model to construct an improved YOLOv11 model, training the unmanned aerial vehicle remote sensing image dataset through the improved YOLOv11 model to obtain a target detection model; deploying a front end based on an application programming interface (API) of a developer platform on an unmanned aerial vehicle remote controller, taking complete images of spruce and standard rods through the unmanned aerial vehicle, and controlling the unmanned aerial vehicle remote controller to transmit the complete images of the spruce and the standard rods to a back end database; receiving the complete images of the spruce and the standard rods by the back end database and correcting camera distortion, calling an API of the target detection model deployed on a Flask framework to detect the complete images of the spruce and the standard rods, obtaining the height of the spruce, and transmitting the height of the spruce back to a front end page and recording.
[0007] The technical scheme provided in the application collects unmanned aerial vehicle remote sensing images of spruces and standard rods in a sample area, corrects camera distortion of the unmanned aerial vehicle remote sensing images of the spruces and the standard rods to eliminate system errors and avoid proportion distortion of tree height calculation caused by distortion, guarantees image quality, constructs an unmanned aerial vehicle remote sensing image dataset, divides the unmanned aerial vehicle remote sensing image dataset into an unmanned aerial vehicle remote sensing image training set, an unmanned aerial vehicle remote sensing image test set and an unmanned aerial vehicle remote sensing image verification set to optimize model generalization ability and prevent overfitting, embeds GAM in a YOLOv11 model to construct an improved YOLOv11 model to strengthen feature extraction ability of the model on small targets, solve the problem of missed detection of the YOLOv11 model due to small pixel proportion of the standard rod, improve small target detection precision and reduce the probability of false detection in a complex background, train the unmanned aerial vehicle remote sensing image dataset through the improved YOLOv11 model, quickly converge to an optimal model through hyperparameter optimization to save parameter tuning time cost, optimize training efficiency, and obtain a target detection model, deploy a front end based on a developer platform API on an unmanned aerial vehicle remote controller, do not need additional equipment to reduce hardware cost, capture complete images of spruces and standard rods through the unmanned aerial vehicle, transmit the complete images of the spruces and the standard rods to a back end database through the unmanned aerial vehicle remote controller, can select manual or automatic upload modes to adapt to different scene requirements, improve operation flexibility, realize operation closed loop and guarantee data real-time performance, the back end database receives the complete images of the spruces and the standard rods and corrects camera distortion, calls an API of the target detection model deployed on a Flask framework to detect the complete images of the spruces and the standard rods, obtains spruce height, transmits the spruce height back to a front end page and records, meets traceability of detection results, finally realizes full-process automation of distortion correction-model detection-height calculation, greatly improves detection efficiency, and the spruce height is calculated based on a pixel proportion method combined with a camera pinhole imaging principle, ensures accurate physical scale conversion and greatly improves calculation precision of the spruce height. The technical scheme provided in the application finally realizes improvement of measurement precision and efficiency and reduction of measurement cost of spruce height and meets real-time performance requirements of forestry census.
[0008] Optionally, the unmanned aerial vehicle remote sensing images of spruces and standard rods in the sample area are collected by: making the unmanned aerial vehicle face the spruces and ascending to a position at the middle left or right of the spruces, capturing complete unmanned aerial vehicle remote sensing images of the spruces and the standard rods when a camera optical axis of the unmanned aerial vehicle is horizontal, and the unmanned aerial vehicle remote sensing images of the spruces and the standard rods contain multi-angle spruce and standard rod images, wherein the standard rod is vertically attached to a trunk of the spruce, and when the standard rod is blocked by leaves of the spruce, the standard rod is placed on the left or right side of the spruce and is at a right angle with a flight direction of the unmanned aerial vehicle.
[0009] Optionally, the camera distortion correction of the spruce and standard pole unmanned aerial vehicle remote sensing image comprises: reading the spruce and standard pole unmanned aerial vehicle remote sensing image and obtaining image original resolution parameters, locating a drone-dji:DewarpData field in an Extensible Metadata Platform (XMP) metadata segment of the spruce and standard pole unmanned aerial vehicle remote sensing image, and converting a string parameter in the drone-dji:DewarpData field into a floating-point number value; constructing a camera intrinsic parameter matrix and a distortion coefficient array required by an Open Source Computer Vision Library (OpenCV) based on the converted floating-point number value; and performing repositioning calculation on image pixels according to a bilinear interpolation method, eliminating lens distortion, and obtaining a non-distortion image.
[0010] Optionally, the constructing of the unmanned aerial vehicle remote sensing image dataset and the dividing of the unmanned aerial vehicle remote sensing image dataset into the unmanned aerial vehicle remote sensing image training set, the unmanned aerial vehicle remote sensing image test set and the unmanned aerial vehicle remote sensing image verification set comprises: annotating spruce outlines and standard poles in the unmanned aerial vehicle remote sensing image through Labeling for Efficient Machine Learning (Labelme) software to construct the unmanned aerial vehicle remote sensing image dataset; and dividing the unmanned aerial vehicle remote sensing image dataset into the unmanned aerial vehicle remote sensing image training set, the unmanned aerial vehicle remote sensing image test set and the unmanned aerial vehicle remote sensing image verification set according to a 7:2:1 ratio.
[0011] Optionally, the YOLOv11 model includes a backbone feature extraction network (Feature Extraction Backbone Network, Backbone) and a feature fusion neck network (Feature Fusion Neck Network, Neck), the GAM is embedded in the YOLOv11 model to construct an improved YOLOv11 model, and the improved YOLOv11 model is used to train the unmanned aerial vehicle remote sensing image dataset to obtain a target detection model, comprising: embedding the GAM at the end of the Backbone of the YOLOv11 model, the GAM being part of the feature extraction process; configuring the GAM to complete channel attention weighting and spatial attention focusing; capturing the context information of large-scale targets in the spruce and standard pole unmanned aerial vehicle remote sensing image through the GAM; establishing a progressive feature learning path from the low-level convolutional layer of the Backbone to the GAM, so that the Backbone gradually learns the image features of the spruce and standard pole unmanned aerial vehicle remote sensing image from low-level texture to high-level semantics; outputting the attention-enhanced feature image to the Neck to construct the improved YOLOv11 model; for the improved YOLOv11 model, setting the learning rate parameter to 0.01, 0.001 and 0.0001 three momentum gradients, setting the batch size parameter to 8, 16 and 32 three gradients, setting the model training rounds to 300, performing training for each group of hyperparameter combinations, comparing the detection accuracy indicators of all hyperparameter combinations on the unmanned aerial vehicle remote sensing image validation set, and selecting the hyperparameter combination with the optimal detection accuracy indicator for model training to obtain the target detection model.
[0012] Optionally, the backend database receives the spruce and standard pole complete image and performs camera distortion correction, calls the API of the target detection model deployed on the Flask framework to detect the spruce and standard pole complete image, obtains the spruce height, and returns the spruce height to the front-end page and records, comprising: the backend database receives the spruce and standard pole complete image, extracts the drone-dji:DewarpData field from the XMP metadata segment of the spruce and standard pole complete image for metadata analysis, generates a correction matrix, applies a bilinear interpolation method for pixel relocation calculation to the spruce and standard pole complete image, completes camera distortion correction, and obtains a non-distorted complete image; deploying the improved YOLOv11 model on the Flask framework, making an API for the target detection model through the Flask framework, inputting the non-distorted complete image into the target detection model through the API, detecting the non-distorted complete image, obtaining the spruce height and returning it to the front-end page and recording.
[0013] Optionally, the detecting the distortionless complete image to obtain the spruce height comprises: detecting the distortionless complete image to obtain a standard pole detection frame pixel height and a spruce detection frame pixel height; and calculating the spruce height according to a camera imaging principle and a pixel proportion method, wherein a calculation formula of the spruce height is represented by the following formula: ; In the formula, represents the spruce height; represents the spruce detection frame pixel height; represents the standard pole detection frame pixel height; represents a standard pole actual height.
[0014] In a second aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the steps of the above method for quickly measuring the spruce height based on the improved YOLOv11 model and the unmanned aerial vehicle image when executing the program.
[0015] In a third aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above method for quickly measuring the spruce height based on the improved YOLOv11 model and the unmanned aerial vehicle image.
[0016] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects: The application provides a spruce height rapid measurement method based on an improved YOLOv11 model and a UAV image. UAV remote sensing images of spruces and standard poles in a sample area are collected, and camera distortion correction is performed on the UAV remote sensing images of the spruces and the standard poles to eliminate system errors and avoid proportion distortion in tree height calculation caused by distortion, thereby ensuring image quality. A UAV remote sensing image dataset is constructed, and the UAV remote sensing image dataset is divided into a UAV remote sensing image training set, a UAV remote sensing image test set and a UAV remote sensing image verification set to optimize model generalization ability and prevent overfitting. GAM is embedded in the YOLOv11 model to construct an improved YOLOv11 model, thereby strengthening the feature extraction ability of the model for small targets, solving the problem of missed detection of the YOLOv11 model due to a small proportion of standard pole pixels, improving small target detection accuracy, and reducing the probability of false detection in a complex background. The improved YOLOv11 model is used to train the UAV remote sensing image dataset, and the optimal model is quickly converged through hyperparameter optimization, thereby saving parameter tuning time and cost, optimizing training efficiency, and obtaining a target detection model. The front end based on the developer platform API is deployed on a UAV remote controller, and no additional equipment is needed, thereby reducing hardware cost. The UAV captures complete images of spruces and standard poles, the UAV remote controller transmits the complete images of the spruces and the standard poles to a back end database, manual or automatic upload modes can be selected to adapt to different scene requirements, improve operation flexibility, realize operation closed loop, and ensure data real-time performance. The back end database receives the complete images of the spruces and the standard poles and performs camera distortion correction, calls the API of the target detection model deployed on the Flask framework to detect the complete images of the spruces and the standard poles, obtains the height of the spruces, and transmits the height of the spruces back to the front end page and records it, so that the detection result can be traced back. Finally, distortion correction-model detection-height calculation automation is realized, and the detection efficiency is greatly improved. Moreover, the height of the spruces is calculated based on the pixel proportion method combined with the camera pinhole imaging principle, thereby ensuring the accuracy of physical scale conversion and greatly improving the calculation accuracy of the height of the spruces. The technical scheme provided in the application finally realizes the improvement of the measurement accuracy, the measurement efficiency and the reduction of the measurement cost of the height of the spruces, and meets the real-time requirement of forestry census. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor. Figure 1 A flowchart of a spruce height rapid measurement method based on an improved YOLOv11 model and a UAV image provided in the embodiments of the application; Figure 2 A schematic diagram of spruce and standard pole images collected at different camera tilt angles provided for an embodiment of the present application; Figure 3 A schematic diagram of spruce and standard pole contours labeled by Labelme software provided for an embodiment of the present application; Figure 4 A schematic diagram of a login web page provided for an embodiment of the present application; Figure 5 A schematic diagram of a model detection diagram provided for an embodiment of the present application; Figure 6 A schematic diagram of a camera pinhole imaging model provided for an embodiment of the present application; Figure 7 A schematic diagram of a UAV device management page provided for an embodiment of the present application; Figure 8 A schematic diagram of a detection result history record page provided for an embodiment of the present application; Figure 9 A schematic diagram of a detection result display page provided for an embodiment of the present application; Figure 10 A schematic diagram of a hardware entity of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in a clear and complete manner with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. The following embodiments are used to describe the present application but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0019] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0020] It should be noted that the terms “first\second\third” involved in the embodiments of the present application are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, “first\second\third” can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0021] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field of the embodiments of the present application. It should also be understood that terms such as those defined in a general dictionary have meanings consistent with those in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as such herein.
[0022] The embodiments of the present application are further described below with reference to the accompanying drawings.
[0023] In view of the problems in the field of unmanned aerial vehicle remote sensing, computer vision and forestry measurement technology for measuring the height of spruce, the embodiments of the present application provide a method for quickly measuring the height of spruce based on an improved YOLOv11 model and unmanned aerial vehicle images.
[0024] The technical solutions of the present application are described below. First, the method embodiments of the present application are described.
[0025] Please refer to Figure 1 , which shows a flowchart of the method for quickly measuring the height of spruce based on an improved YOLOv11 model and unmanned aerial vehicle images provided by the embodiments of the present application, as shown in Figure 1 , the method comprises at least the following steps S110-S140.
[0026] Step S110: Collecting unmanned aerial vehicle remote sensing images of spruce and standard poles in a sample area, correcting camera distortion of the unmanned aerial vehicle remote sensing images of spruce and standard poles, constructing an unmanned aerial vehicle remote sensing image dataset, and dividing the unmanned aerial vehicle remote sensing image dataset into an unmanned aerial vehicle remote sensing image training set, an unmanned aerial vehicle remote sensing image test set and an unmanned aerial vehicle remote sensing image verification set.
[0027] In the embodiments of the present application, first, unmanned aerial vehicle remote sensing images of spruce and standard rods in the sample area are collected. Specifically, the sample area is selected to be close to the edge of the spruce forest, for example, a 30m x 30m area in a certain region of the Nanshan Shagou is selected as the sample area of the spruce forest, and the camera carried by the unmanned aerial vehicle is used to take pictures in the sample area of the spruce forest. The unmanned aerial vehicle is directed at the spruce and rises to the middle or left and right positions of the spruce, and the complete spruce and standard rod unmanned aerial vehicle remote sensing images are taken when the optical axis of the unmanned aerial vehicle camera is horizontal (i.e., the camera pitch angle is 0°) to obtain high-resolution remote sensing images. In an optional embodiment, -10°, -5°, 0°, 5° and 10° are used as the pitch angle of the unmanned aerial vehicle camera to take pictures of the spruce. For different angles of the spruce, a set of pictures with the above pitch angles are taken, and finally 5 sets of 25 images are taken for one spruce to determine the influence of the camera pitch angle on the measurement of the height of the spruce. The final spruce and standard rod unmanned aerial vehicle remote sensing images contain spruce and standard rod images at multiple angles. For example, please refer to Figure 2 which shows a schematic diagram of spruce and standard rod images collected by the camera at different pitch angles according to the embodiments of the present application. The diagram shows spruce and standard rod images collected at pitch angles of -10°, 0° and 10°. It should be particularly noted that during the shooting process, the standard rod is placed vertically and attached to the spruce trunk to ensure that the standard rod is completely displayed in the unmanned aerial vehicle remote sensing image. When it is observed that the standard rod is blocked by the spruce leaves in the unmanned aerial vehicle remote sensing image, the standard rod is placed on the left or right side of the spruce at a right angle to the flight direction of the unmanned aerial vehicle. The length of the standard rod can be 2 meters. The technical solutions provided by the embodiments of the present application do not limit the type of unmanned aerial vehicle and the sensor carried by the unmanned aerial vehicle. Optionally, DJI Spark visible light version, DJI Spark multispectral version, DJI Mavic 3 or DJI Matrice 4 are used to collect spruce and standard rod unmanned aerial vehicle remote sensing images in the sample area.
[0028] In the embodiment of the present application, camera distortion correction is performed on the spruce and standard pole unmanned aerial vehicle remote sensing images. Specifically, due to the radial distortion, attitude distortion and barrel distortion of the images captured by the unmanned aerial vehicle, according to the camera imaging principle, if the original spruce and standard pole unmanned aerial vehicle remote sensing images are directly input into the model, it is easy to cause large error in the measured spruce height, therefore, it is necessary to perform camera distortion correction on the original collected spruce and standard pole unmanned aerial vehicle remote sensing images, eliminate radial distortion, barrel distortion and attitude distortion, thereby greatly reducing the problem that the ratio of the actual height of the standard pole to the pixel height of the standard pole detection frame and the ratio of the spruce height to the pixel height of the spruce detection frame are greatly different due to camera distortion. The specific camera distortion correction process is as follows: first, read the original spruce and standard pole unmanned aerial vehicle remote sensing images and obtain the image original resolution parameters, locate the drone-dji:DewarpData field in the XMP metadata section of the original spruce and standard pole unmanned aerial vehicle remote sensing images, and convert the string parameters in the drone-dji:DewarpData field into floating point values, then, based on the converted floating point values, construct the camera intrinsic parameter matrix and distortion coefficient array required by OpenCV, finally, according to the bilinear interpolation method, calculate the pixel repositioning of the image pixels to eliminate lens barrel distortion and attitude distortion, and finally obtain the non-distortion image.
[0029] In the embodiment of the present application, an unmanned aerial vehicle remote sensing image dataset is constructed, and the unmanned aerial vehicle remote sensing image dataset is divided into an unmanned aerial vehicle remote sensing image training set, an unmanned aerial vehicle remote sensing image test set and an unmanned aerial vehicle remote sensing image verification set. Specifically, the spruce outline and the standard pole in the non-distortion unmanned aerial vehicle remote sensing image are labeled by using the Labelme software to construct the unmanned aerial vehicle remote sensing image dataset. For example, please refer to Figure 3 which shows a schematic diagram of the spruce and standard pole outlines labeled by the Labelme software provided in the embodiment of the present application, the unmanned aerial vehicle remote sensing image dataset is divided into the unmanned aerial vehicle remote sensing image training set, the unmanned aerial vehicle remote sensing image test set and the unmanned aerial vehicle remote sensing image verification set according to the ratio of 7:2:1.
[0030] Step S120, embedding a global attention mechanism in the YOLOv11 model to construct an improved YOLOv11 model, training the unmanned aerial vehicle remote sensing image dataset by using the improved YOLOv11 model to obtain a target detection model.
[0031] In the embodiments of the present application, for the unmanned aerial vehicle remote sensing image, an improved model based on the YOLOv11 model is used to complete target detection. YOLO is a one-stage target detection algorithm, that is, only one scan is needed to identify the category and boundary box of the object in the image. YOLOv11 is the latest YOLO series target detection algorithm released by the Ultralytics company, which is used to complete image classification, object detection and instance segmentation and other tasks. The YOLOv11 model includes Backbone and Neck. Optionally, the YOLOv11 model architecture includes YOLOv11n and YOLOv11s, etc., and appropriate model size can be selected for use according to actual project requirements and resource limitations. GAM is composed of two attention mechanisms, namely channel attention mechanism and spatial attention mechanism. The channel attention mechanism is responsible for allocating resources between each convolutional channel, and the spatial attention mechanism is responsible for transforming the spatial information in the original image to another space and retaining key information, so that the neural network pays more attention to the area that plays a decisive role in image classification.
[0032] In the embodiments of the present application, GAM is embedded in the YOLOv11 model to construct an improved YOLOv11 model, so that the model can more accurately identify the contour of the standard pole and improve the identification performance of small targets. The improved YOLOv11 model is trained on the unmanned aerial vehicle remote sensing image dataset to obtain a target detection model. Specifically, first, a GAM is embedded at the end of the Backbone of the YOLOv11 model, and the GAM is part of the feature extraction process. Second, the GAM is configured to complete channel attention weighting and spatial attention focusing. The GAM captures the context information of large-scale targets in spruce and standard pole unmanned aerial vehicle remote sensing images. An incremental feature learning path is established from the low-level convolutional layer of the Backbone to the GAM. The Backbone gradually learns the image features of the spruce and standard pole unmanned aerial vehicle remote sensing images from low-level textures to high-level semantics. Finally, the feature image with enhanced attention is output to the Neck to construct an improved YOLOv11 model, thereby improving the target detection accuracy in complex scenes. The technical scheme provided in the embodiments of the present application embeds the GAM at the end of the Backbone, and forms a collaborative feature enhancement process with the C3k2 module. In view of the characteristics of the large-scale target scattered distribution and complex background interference in the unmanned aerial vehicle remote sensing image, the GAM is placed at the end of the Backbone to preferentially capture the high-level semantic features after cross-layer fusion. The three-dimensional (3D) arrangement mechanism of the GAM retains the channel-space-scale multi-dimensional information, strengthens the context correlation, and solves the small target missing detection and background false detection problems caused by the single feature level in the traditional attention mechanism in the unmanned aerial vehicle scene. In terms of incremental multi-granularity feature learning, the GAM is embedded at the end of the Backbone to meet the incremental learning needs of the unmanned aerial vehicle data from low-level texture details to high-level semantic features. The channel-space dual attention series structure of the GAM can avoid the loss of spatial details caused by the traditional maximum pooling.
[0033] In the embodiments of the present application, for the improved YOLOv11 model, the learning rate parameter is set to three momentum gradients of 0.01, 0.001 and 0.0001, the batch size parameter is set to three gradients of 8, 16 and 32, and the model training rounds are set to 300. Each set of hyperparameter combinations is trained, the detection accuracy indicators of all hyperparameter combinations on the unmanned aerial vehicle remote sensing image validation set are compared, the hyperparameter combination with the optimal detection accuracy indicator is selected for model training, the last saved model weight file is selected as the final training result, and a target detection model is obtained. The trained target detection model is verified, and the verification result is that the error between the tree height measured by the target detection model and the actual tree height is within 3%, which meets the national standard that the error is less than or equal to 3% when the tree height is less than 10 meters. In particular, when the tree height is greater than or equal to 10 meters, the national standard is that the error is less than or equal to 5%.
[0034] Step S130, deploy the front end based on the developer platform application programming interface on the unmanned aerial vehicle remote controller, take a complete image of the spruce and the standard pole by the unmanned aerial vehicle, and control the unmanned aerial vehicle remote controller to transmit the complete image of the spruce and the standard pole to the back end database.
[0035] In the embodiment of the present application, the front end based on the developer platform API is deployed on the unmanned aerial vehicle remote controller. The developer platform API can be DJI Developer Platform Application Programming Interface (DJI-API). The DJI-API mainly adopts the Message Queuing Telemetry Transport (MQTT) protocol, the Hypertext Transfer Protocol Secure (HTTPS) protocol and the WebSocket protocol, abstracts the aircraft capability into a thing model of an Internet of Things device, and develops a business based on the thing model. In DJI Pilot 2, a customized login web page is developed on the unmanned aerial vehicle remote controller through an embedded webview engine, which is used for logging in to the back end. For example, please refer to Figure 4 which shows a schematic diagram of a login web page provided by the embodiment of the present application. The image taken by the unmanned aerial vehicle can be uploaded to the back end database manually or automatically in advance. When the manual upload is set, after the complete image of the spruce and the standard pole is taken by the unmanned aerial vehicle, the unmanned aerial vehicle remote controller is controlled to manually select the complete image of the spruce and the standard pole to be measured, and the image is transmitted to the back end database by clicking the upload cloud button in the image preview interface.
[0036] Step S140, the back end database receives the complete image of the spruce and the standard pole and performs camera distortion correction, calls the application programming interface of the target detection model deployed on the Flask framework to detect the complete image of the spruce and the standard pole, obtains the spruce height, and transmits the spruce height back to the front end page and records it.
[0037] In the embodiment of the present application, the backend database receives and stores the complete image of the spruce and the standard pole, the backend adopts the springboot technology framework of Java, the database adopts My Structured Query Language (MySQL), the storage scheme adopts Minimal Input / Output (MinIO), and when the backend receives the complete image of the spruce and the standard pole and the image processing request sent by the front end, the backend is responsible for recording and completing the detection, calling the API of Flask, and processing the camera distortion through python. Specifically, first, the drone-dji:DewarpData field in the XMP metadata segment of the complete image of the spruce and the standard pole is extracted for metadata analysis to generate a correction matrix, then the bilinear interpolation method is applied to the complete image of the spruce and the standard pole for pixel relocation calculation to complete the camera distortion correction, and finally the complete image without distortion is obtained.
[0038] Further, the API of the target detection model deployed on the Flask framework is called to detect the complete image without distortion to obtain the height of the spruce, and the height of the spruce is returned to the front-end page and recorded. Specifically, the improved YOLOv11 model is deployed on the Flask framework, the API is made for the target detection model through the Flask framework, the complete image without distortion is input into the target detection model through the API, and the complete image without distortion is detected. For example, please refer to Figure 5 which shows a model detection diagram provided by an embodiment of the present application. After detecting the spruce and the standard pole, the pixel height of the standard pole detection frame and the pixel height of the spruce detection frame are obtained, and then the height of the spruce is calculated according to the camera imaging principle and the pixel proportion method. For example, please refer to Figure 6 which shows a schematic diagram of a camera pinhole imaging model provided by an embodiment of the present application. Based on the pinhole imaging model, after the light passes through the lens or the pinhole, an inverted and reduced real image is formed on the imaging plane. This process follows the straight line propagation law of light. The object distance (the distance from the object to the optical center) and the image distance (the distance from the image to the optical center) establish a proportion through the similar triangle relationship, and the proportion formula is represented by the following formula (1): Formula (1); In the formula, represents the object distance; represents the focal length; and represent the object size and the image size respectively. Similarly, the calculation formula of the height of the spruce is represented by the following formula (2): Formula (2); In the formula, represents the height of the spruce; Indicates the pixel height of the spruce detection box; Indicates the pixel height of the standard pole detection frame; This indicates the actual height of the standard pole, which is used to obtain the final spruce height. The test results are then sent back to the front-end page for users to view and are recorded. For an example, please refer to [link / reference needed]. Figure 7 to Figure 9 The illustration shows a schematic diagram of a drone equipment management page, a schematic diagram of a detection result history page, and a schematic diagram of a detection result display page provided in the embodiments of this application.
[0039] In summary, the rapid spruce height measurement method based on an improved YOLOv11 model and UAV imagery provided in this application collects UAV remote sensing images of spruce trees and standard poles within a sample area. Camera distortion correction is applied to these images to eliminate systematic errors, prevent distortion from causing inaccurate tree height calculations, and ensure image quality. A UAV remote sensing image dataset is constructed and divided into a training set, a test set, and a validation set to optimize model generalization and prevent overfitting. A Gaussian Animation Model (GAM) is embedded in the YOLOv11 model to construct an improved YOLOv11 model, enhancing its feature extraction capability for small targets. This addresses the issue of missed detections due to the small pixel proportion of standard poles in the YOLOv11 model, improving small target detection accuracy and reducing false detection probability in complex backgrounds. The improved YOLOv11 model is used to train the UAV remote sensing image dataset, and hyperparameter optimization quickly converges to the optimal model, saving on parameter tuning costs. This approach optimizes training efficiency and reduces time costs to obtain a target detection model. The front-end, based on the developer platform API, is deployed on a drone remote controller, eliminating the need for additional equipment and thus reducing hardware costs. The drone captures complete images of a spruce tree and a standard pole, which are then transmitted to a back-end database via the drone remote controller. Manual or automatic upload modes are available to adapt to different scenarios, enhancing operational flexibility, achieving a closed-loop operation, and ensuring real-time data transmission. The back-end database receives the complete images of the spruce tree and standard pole, performs camera distortion correction, and calls the API of the target detection model deployed on the Flask framework to detect the images, obtaining the spruce tree height. This height is then transmitted back to the front-end page and recorded, ensuring traceability of detection results. Ultimately, the entire process of distortion correction, model detection, and height calculation is automated, significantly improving detection efficiency. Furthermore, the spruce tree height is calculated based on pixel ratios combined with the camera's pinhole imaging principle, ensuring accurate physical scale conversion and greatly improving the accuracy of the spruce tree height calculation. The technical solution provided in this application ultimately improves the measurement accuracy and efficiency of spruce height and reduces measurement costs, meeting the real-time requirements of forestry surveys.
[0040] It should be noted that, in the embodiments of this application, if the above-mentioned method for rapid measurement of spruce height based on the improved YOLOv11 model and UAV imagery is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0041] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by a processor, this computer program implements the steps in any of the above-described methods for rapid measurement of spruce height based on an improved YOLOv11 model and UAV imagery. Correspondingly, embodiments of this application also provide a computer program product that, when executed by a processor of an electronic device, implements the steps in any of the above-described methods for rapid measurement of spruce height based on an improved YOLOv11 model and UAV imagery.
[0042] Based on the same technical concept, this application provides an electronic device for implementing the rapid spruce height measurement method based on an improved YOLOv11 model and UAV imagery described in the above method embodiments. Figure 10 This is a hardware entity diagram of an electronic device provided in an embodiment of this application, such as... Figure 10 As shown, the electronic device 1000 includes a memory 1010 and a processor 1020. The memory 1010 stores a computer program that can run on the processor 1020. When the processor 1020 executes the program, it implements the steps in any of the embodiments of this application of the rapid measurement method for spruce height based on an improved YOLOv11 model and UAV imagery.
[0043] The memory 1010 is configured to store instructions and applications executable by the processor 1020, and can also cache data to be processed or already processed by the processor 1020 and various modules in the electronic device (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).
[0044] The processor 1020 implements the steps of the improved YOLOv11 model and unmanned aerial vehicle image-based spruce height rapid measurement method of any one of the above when executing a program. The processor 1020 generally controls the overall operation of the electronic device 1000.
[0045] The processor described above can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device that implements the functions of the above processor can also be other, and the embodiments of the present application are not limited specifically.
[0046] The computer storage medium / memory described above can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM) memory, etc. It can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0047] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0048] It should be understood that every feature, structure, or characteristic described in relation to one embodiment is applicable to at least one other embodiment, in combination or in isolation, unless specifically noted otherwise. Furthermore, the initial and subsequent appearances of "in one embodiment" are not necessarily referring to the same embodiment nor to one specific embodiment. Moreover, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that the sequences of processes in the various embodiments can be performed in any order without departing from the scope of the application unless otherwise specifically noted. The sequences of processes in the various embodiments are provided for purposes of example and illustration and are not intended to be limiting, unless otherwise specifically noted. There is no intention to be bound by any indicated or implied theory presented in the preceding description(s).
[0049] It should be noted that, as used in this document, the terms "include," "includes," "including," "has," "have," "having," or the like are used inclusively, in a manner that will also cover excludability. As used herein, the term "or" as used herein is generally intended to mean "and / or" unless otherwise indicated. As used herein, the term "about" means approximately or nearly as understood by one of ordinary skill in the art. As used herein, the term "comprising" or "comprise" means "including, but not limited to" as understood by one of ordinary skill in the art. The terms "program" or software" are used herein in a generic sense to refer to any type of computer code (e.g., software or microcode) that can be employed to program a computer or other processor. When implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium, such as one or more non-transitory computer-readable media. Computer-readable media include computer storage media, which are media that computer-readable instructions, data structures, program code, or the like are particularly adapted or configured to be maintained. By way of example, computer storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other hardware storage device that can be used for computer storage. Computer storage media does not include a modulated data signal or carrier wave that is transmitted over a network or other communication medium. The term "modulated data signal" or "carrier wave" means a signal that has one or more of its characteristics changed or set in a manner so as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as wireless networks, cellular telephone connections, RF data links, Bluetooth or other wireless communication links, and the like.
[0050] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The described device embodiments are merely illustrative, and the division into units is merely logical function division, and there can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0051] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place or distributed on a plurality of network units; and some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0052] In addition, each functional unit in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in the form of hardware or hardware plus software functional units.
[0053] Alternatively, the above-mentioned integrated units of the present application, if realized in the form of software function modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to make the equipment test line execute all or part of the method described in the embodiments of the present application. The foregoing storage medium includes: mobile storage devices, ROM, magnetic discs or optical discs and various media that can store program codes.
[0054] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0055] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0056] The above is only an implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for rapid measurement of spruce height based on improved YOLOv11 model and unmanned aerial vehicle image, characterized in that, The method comprises: Collecting unmanned aerial vehicle remote sensing images of spruces and standard poles in a sample area, performing camera distortion correction on the unmanned aerial vehicle remote sensing images of spruces and standard poles, constructing an unmanned aerial vehicle remote sensing image dataset, and dividing the unmanned aerial vehicle remote sensing image dataset into an unmanned aerial vehicle remote sensing image training set, an unmanned aerial vehicle remote sensing image test set and an unmanned aerial vehicle remote sensing image verification set; Embedding a global attention mechanism in a YOLOv11 model to construct an improved YOLOv11 model, training the unmanned aerial vehicle remote sensing image dataset through the improved YOLOv11 model to obtain a target detection model; Deploying a front end based on an application programming interface of a developer platform on an unmanned aerial vehicle remote controller, taking complete images of spruces and standard poles through the unmanned aerial vehicle, and transmitting the complete images of spruces and standard poles to a back end database through the unmanned aerial vehicle remote controller; The back end database receives the complete images of spruces and standard poles and performs camera distortion correction, calls an application programming interface of the target detection model deployed on a Flask framework to detect the complete images of spruces and standard poles, obtains spruce heights, and transmits the spruce heights back to a front end page and records them.
2. The method of claim 1, wherein, The collecting of the unmanned aerial vehicle remote sensing images of spruces and standard poles in the sample area comprises: The unmanned aerial vehicle is directly opposite the spruce and rises to the left and right positions of the middle part of the spruce, and the complete unmanned aerial vehicle remote sensing images of the spruce and the standard pole are taken when the optical axis of the unmanned aerial vehicle camera is horizontal, and the unmanned aerial vehicle remote sensing images of the spruce and the standard pole contain multi-angle spruce and standard pole images, wherein the standard pole is vertically attached to the trunk of the spruce, and when the standard pole is blocked by the leaves of the spruce, the standard pole is placed on the left or right side of the spruce, and is perpendicular to the flight direction of the unmanned aerial vehicle.
3. The method of claim 1, wherein, The camera distortion correction on the unmanned aerial vehicle remote sensing images of spruces and standard poles comprises: Reading the unmanned aerial vehicle remote sensing images of spruces and standard poles and obtaining image original resolution parameters, locating a drone-dji:DewarpData field in an extensible metadata platform metadata segment in the unmanned aerial vehicle remote sensing images of spruces and standard poles, and converting string parameters in the drone-dji:DewarpData field into floating point values; Based on the converted floating point values, a camera intrinsic matrix and a distortion coefficient array required by an open source computer vision library are constructed; According to a bilinear interpolation method, the image pixels are repositioned and calculated to eliminate lens distortion, and a non-distortion image is obtained.
4. The method of claim 1, wherein, The construction of the unmanned aerial vehicle remote sensing image dataset and the division of the unmanned aerial vehicle remote sensing image dataset into the unmanned aerial vehicle remote sensing image training set, the unmanned aerial vehicle remote sensing image test set and the unmanned aerial vehicle remote sensing image verification set comprise: The spruce outlines and standard poles in the unmanned aerial vehicle remote sensing images are labeled through an image labeling tool software to construct the unmanned aerial vehicle remote sensing image dataset; The unmanned aerial vehicle remote sensing image dataset is divided into the unmanned aerial vehicle remote sensing image training set, the unmanned aerial vehicle remote sensing image test set and the unmanned aerial vehicle remote sensing image verification set according to a ratio of 7:2:
1.
5. The method of claim 1, wherein, The YOLOv11 model includes a backbone feature extraction network and a feature fusion neck network, a global attention mechanism is embedded in the YOLOv11 model, an improved YOLOv11 model is constructed, the unmanned aerial vehicle remote sensing image dataset is trained through the improved YOLOv11 model, and a target detection model is obtained, including: A global attention mechanism is embedded at the end of the backbone feature extraction network of the YOLOv11 model, and the global attention mechanism is part of the feature extraction process; The global attention mechanism is configured to complete channel attention weighting and spatial attention focusing; The global attention mechanism captures the context information of large-scale targets in the spruce and standard pole unmanned aerial vehicle remote sensing image; An incremental feature learning path is established from the low-level convolutional layer of the backbone feature extraction network to the global attention mechanism, so that the backbone feature extraction network gradually learns the image features of the spruce and standard pole unmanned aerial vehicle remote sensing image from low-level texture to high-level semantics; The attention-enhanced feature image is output to the feature fusion neck network to construct the improved YOLOv11 model; For the improved YOLOv11 model, the learning rate parameter is set to 0.01, 0.001 and 0.0001 three momentum gradients, the batch size parameter is set to 8, 16 and 32 three gradients, the model training rounds are set to 300, each group of hyperparameter combinations is trained, the detection accuracy indexes of all hyperparameter combinations on the unmanned aerial vehicle remote sensing image validation set are compared, and the hyperparameter combination with the optimal detection accuracy index is selected for model training to obtain the target detection model.
6. The method of claim 1, wherein, The backend database receives the spruce and standard pole complete image and performs camera distortion correction, calls the application programming interface of the target detection model deployed on the Flask framework to detect the spruce and standard pole complete image, obtains the spruce height, and returns the spruce height to the front-end page and records, including: The backend database receives the spruce and standard pole complete image, extracts the drone-dji:DewarpData field from the extensible metadata platform metadata segment of the spruce and standard pole complete image for metadata analysis, generates a correction matrix, applies a bilinear interpolation method to the spruce and standard pole complete image for pixel relocation calculation, completes camera distortion correction, and obtains a distortion-free complete image; The improved YOLOv11 model is deployed on the Flask framework, the Flask framework is used to make an application programming interface for the target detection model, the distortion-free complete image is input into the target detection model through the application programming interface, the distortion-free complete image is detected, the spruce height is obtained, and the spruce height is returned to the front-end page and recorded.
7. The method of claim 6, wherein, The distortion-free complete image is detected to obtain the spruce height, including: The distortion-free complete image is detected to obtain the standard pole detection frame pixel height and the spruce detection frame pixel height; According to the camera imaging principle, the pixel proportion method is used to calculate the spruce height, and the calculation formula of the spruce height is represented by the following formula: ; In the formula, represents the height of the spruce; represents the height of the spruce detection frame in pixels; represents the height of the standard pole detection frame in pixels; represents the actual height of the standard pole.
8. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, The processor, when executing the program, implements the steps in the method of any one of claims 1 to 7.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps in the method of any one of claims 1 to 7.