Unmanned aerial vehicle ortho-video-based dead and dead wood quantity counting and positioning method

By building a synchronous acquisition system for drone video streaming and SRT positioning data, a dual detection framework and target motion state machine model are designed to achieve high-precision positioning and statistics of dead wood, the problems of poor real-time, low positioning accuracy, and insufficient adaptability of dynamic scenes in traditional forestry monitoring methods are solved, and automated forestry pest monitoring technology is provided.

CN120147904APending Publication Date: 2025-06-13FUJIAN NORMAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510212892.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional forestry monitoring methods have problems such as poor real-time, low positioning accuracy, and insufficient adaptability to dynamic scenes. Especially when dealing with high-resolution aerial videos, it is difficult to achieve automated identification, continuous tracking and precise positioning of dead wood.

Method used

By building a synchronous acquisition system for drone video stream and SRT positioning data, a dual detection framework for full-frame tracking and slice detection is designed, a target motion state machine model is established, and a target trajectory smoothing is achieved using velocity prediction and trajectory buffer management, and a pixel-GPS coordinate mapping model is constructed based on perspective projection transformation to achieve high-precision positioning and statistics of dead wood.

Benefits of technology

It realizes high-precision continuous tracking of dead wood in complex forest areas, solves the missed inspection problems caused by target occlusion and scale changes, reduces resource consumption of manual inspections, and provides automated technical assistance for forestry pest monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147904A_ABST
    Figure CN120147904A_ABST
Patent Text Reader

Abstract

The invention provides a dead wood quantity counting and positioning method based on an unmanned aerial vehicle ortho-video. The method comprises the following steps: constructing a data set according to an original video frame of the unmanned aerial vehicle ortho-video and synchronous positioning data; a deep learning model based on a single-stage detection architecture is trained, target detection is performed on frame-by-frame images, basic tracking is realized by detecting an associated multi-target tracker, periodic enhanced detection is performed in combination with a slice detection strategy, a redundant detection frame is eliminated through an overlapping region fusion algorithm, and a high-confidence target detection result is formed; based on a target motion track and space-time continuity analysis, appearance-tracking-loss-recovery period management of a dead wood target is realized; based on the mapping relation between the pixel coordinate system and the geographic coordinate system, the target position is converted into a GPS coordinate in combination with the flight parameters of the unmanned aerial vehicle, and a dead wood statistical result containing the space-time distribution characteristics and geographic positioning information are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision, UAV remote sensing, video target detection and tracking, geospatial positioning technology, and adaptive target recognition in dynamic scenarios, and particularly relates to a method for counting and positioning dead trees based on UAV ortho-video. Background Art

[0002] Traditional forest pest and disease monitoring mainly relies on manual ground inspections or visual interpretation of aerial images, which has problems such as low efficiency, high cost, and poor real-time performance. Existing dead tree detection methods based on deep learning are mostly designed for static remote sensing images and have the following limitations: (1) Insufficient utilization of spatio-temporal correlation features of consecutive frames of video streams, resulting in poor cross-frame target tracking stability; (2) There are dynamic interference factors such as resolution fluctuations, illumination changes, and target occlusions in UAV ortho-videos, and the misdetection rate of existing single-frame detection methods is relatively high; (3) Lack of a real-time fusion mechanism with geospatial information, making it difficult to achieve precise positioning; (4) Insufficient model generalization ability, unable to adapt to vegetation feature changes in different forest areas.

[0003] Current mainstream object detection algorithms (such as the YOLO series) face a video memory bottleneck when processing 4K-level high-resolution videos, and direct downsampling will cause loss of small target features. Although traditional sliding window detection methods can maintain the original resolution, their computational efficiency is low. In terms of target tracking, algorithms based on Kalman filtering have poor adaptability to non-consecutive moving targets, while deep learning trackers are prone to identity drift in long-term tracking. In terms of geolocation, existing methods mostly use post-artificial matching and cannot achieve real-time spatial coordinate mapping of detection results. Summary of the Invention

[0004] Therefore, aiming at the defects and deficiencies of the existing technology, the purpose of the present invention is to propose a method for counting and positioning dead trees based on UAV ortho-video, so as to solve the technical problems such as poor real-time performance, low positioning accuracy, and insufficient adaptability to dynamic scenarios existing in traditional forestry monitoring methods, and realize automatic recognition, continuous tracking, and precise positioning of dead trees in high-resolution aerial videos.

[0005] It constructs a system for synchronously collecting UAV video streams and SRT positioning data; designs a dual-detection framework of full-frame tracking and slice detection, and realizes the fusion of detection results through dynamic IoU matching; establishes a target motion state machine model, and adopts speed prediction and trajectory buffer management to smooth the target trajectory; constructs a pixel-GPS coordinate mapping model based on perspective projection transformation to achieve the geolocation of dead trees; and outputs statistical results with spatio-temporal distribution characteristics. Through a periodic enhancement detection and trajectory continuity maintenance mechanism, the present invention realizes high-precision continuous tracking of dead trees in complex forest area scenes, solves the problem of missed detection caused by target occlusion and scale change in traditional methods, reduces the resource consumption of manual inspections at the same time, and provides automated technical assistance for forest pest monitoring.

[0006] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:

[0007] A method for counting and locating dead trees based on UAV orthophoto video, which constructs a data set according to the original video frames of the UAV orthophoto video and the synchronized positioning data; trains a deep learning model based on a single-stage detection architecture, performs target detection on each frame of image, and realizes basic tracking through a multi-object tracker associated with the detection. Combining with a slice detection strategy for periodic enhancement detection, redundant detection boxes are eliminated through an overlapping area fusion algorithm to form a high-confidence target detection result; based on the analysis of the target motion trajectory and spatio-temporal continuity, the cycle management of the appearance-tracking-loss-recovery of the dead tree target is realized; based on the mapping relationship between the pixel coordinate system and the geographic coordinate system, combined with the UAV flight parameters, the target position is converted into GPS coordinates, and the dead tree statistical results and geolocation information including spatio-temporal distribution characteristics are output.

[0008] Further, the process of constructing the data set specifically includes: by parsing the synchronized positioning data of the UAV, obtaining the timestamp, longitude and latitude coordinates, relative elevation and flight attitude parameters, and establishing a spatio-temporal reference coordinate system; initializing the video stream processing pipeline, constructing a video stream input pipeline through the video capture interface, and setting a circular buffer to store frame images; constructing a geographic coordinate conversion matrix, and based on the UAV lens parameters and the initial longitude and latitude coordinates recorded in the synchronized positioning data, establishing a mapping relationship from the pixel coordinate system to the GPS coordinate system through perspective transformation.

[0009] Further, the training of the deep learning model based on the single-stage detection architecture performs object detection on each frame of the image, and realizes basic tracking through the multi-object tracker associated with the detection. Combining the slice detection strategy for periodic enhanced detection, redundant detection boxes are eliminated through the overlapping region fusion algorithm to form a high-confidence object detection result, which specifically includes the following steps: training the deep learning model based on the single-stage detection architecture, performing object detection on each frame of the image, and realizing basic tracking through the multi-object tracker associated with the detection, and outputting the basic detection box and confidence; triggering frame image slice detection every N frames, dividing the input frame into multiple overlapping sub-regions, performing object detection on each sub-region, and generating a local detection result set; calculating the spatial similarity of adjacent slice detection boxes, constructing an IoU matrix, eliminating redundant detection boxes through non-maximum suppression, and fusing the full-frame detection result and the slice detection result based on the IoU threshold, where the fusion process includes: uniformly converting the detection boxes of the full-frame detection result and the slice detection result into the center point coordinates and width-height format; sorting all detection boxes in descending order of confidence score; adopting the non-maximum suppression strategy, processing each high-score detection box in turn, calculating its intersection-over-union ratio with other detection boxes, and retaining the detection box with a higher confidence and removing the low-score detection box with a high overlap when the intersection-over-union ratio is greater than the preset threshold, and finally obtaining a non-redundant fusion detection result; finally, feeding back the enhanced detection result to update the tracker state.

[0010] Further, based on the analysis of the target motion trajectory and spatio-temporal continuity, the cycle management of the appearance - tracking - loss - recovery of dead wood targets is realized, which specifically includes the following steps: generating a unique identifier for each detected target, and when a new target is detected, establishing a state machine instance containing spatio-temporal feature vectors, where the specific process of the state machine includes: initializing the state dictionary, including feature vectors such as target position confidence, continuous hit counter, lost frame counter, motion vector, and predicted trajectory window; calculating the instantaneous velocity through the displacement difference of the target center point, setting the exponential decay coefficient to update the motion vector, and realizing trajectory smoothing; recording auxiliary information such as the frame number when the target is first detected and the position of the previous frame for state transition judgment; then realizing state transition by comparing the tracking threshold and the recovery threshold. When the target is continuously hit for N frames, the stable tracking state is activated, the motion vector is updated using the exponential decay model, and the sliding window smoothing of the position coordinates is synchronously performed; when the target is lost, the following prediction matching recovery process is executed:

[0011] (1) Predicting the target position based on the historical motion vector: According to the center position P(t) of the target at the t-th frame and the historical motion vector V(t), calculate the predicted position P'(t + 1) = P(t) + α·V(t) at the t + 1-th frame through the momentum decay coefficient α ∈ (0, 1), where V(t) = P(t) - P(t - 1);

[0012] (2) Constructing the prediction box B predIoU matrix with the current frame detection box set {Bdet}, calculate the intersection over union When the maximum IoU value exceeds the set threshold, it is determined that the match is successful;

[0013] (3) When the state is restored, execute: reset the lost counter lost_count to 0, update the target position to the center P of the matched detection box matched , and use the linear interpolation algorithm to compensate for the trajectory points during the lost period. The interpolation formula is where t lost ≤k≤t recover represents the frame number during the lost period;

[0014] (4) When the number of consecutive lost frames satisfies lost_count > N, move the target to the inactive list and retain the feature vectors {position, velocity, confidence, timestamp} of the last N frames for trajectory backtracking.

[0015] Furthermore, the state transition is realized by comparing the tracking threshold and the recovery threshold. When the target is continuously hit for N frames, the stable tracking state is activated. The motion vector is updated using the exponential decay model, and the sliding window smoothing of the position coordinates is performed synchronously. The motion trajectory is optimized specifically through the speed prediction model and the trajectory buffer management. The implementation process includes: establishing a target speed prediction model. When two consecutive frames of the target are detected, calculate the displacement difference of the center points as the instantaneous speed, and update the speed vector using the exponential weighted average method; based on the circular trajectory buffer, maintain the target motion trajectory using the first-in-first-out strategy. When a new trajectory point is added, the earliest historical point is automatically eliminated; perform sliding window trajectory smoothing, and perform exponential decay weighted average based on the trajectory points within the current window; execute speed-constrained trajectory correction. The speed-constrained trajectory correction process includes: obtaining the sequence of trajectory points, calculating the center position difference between adjacent trajectory points, and generating a sequence of speed vectors representing the motion characteristics of the target; using the exponential moving average algorithm to smooth the sequence of speed vectors, where the smoothing coefficient α ranges from 0 to 1 and is used to adjust the weight of the historical speed information; taking the starting point of the trajectory as the reference, successively accumulate the smoothed speed vectors to the position coordinates of the previous trajectory point to generate a new sequence of trajectory points that satisfies the motion continuity constraint; update the original trajectory cache with the newly generated sequence of trajectory points to realize the trajectory optimization based on the physical kinematic model and suppress the trajectory jitter caused by unstable target detection.

[0016] Further, the conversion of the target position into GPS coordinates by combining the mapping relationship between the pixel coordinate system and the geographic coordinate system and the UAV flight parameters specifically includes the following steps: constructing a pixel-geospatial mapping model based on the synchronous positioning data parameters recorded during the UAV flight, and establishing a two-way conversion relationship between the image pixel coordinate system and the geographic coordinate system through perspective projection transformation; calculating the GPS longitude and latitude coordinates through the geospatial interpolation algorithm according to the relative position of the target center point in the video frame image pixels; the geospatial interpolation algorithm uses the bilinear interpolation method.

[0017] Further, the output of the dead tree statistical result and geographic positioning information containing spatio-temporal distribution characteristics specifically includes the following steps: establishing a spatio-temporal coordinate mapping relationship for each detection target by aligning the video frame timestamp with the time of the synchronous positioning data; rendering a floating window containing the target ID and GPS coordinates in real time in the video stream, recording the total number of dead trees, the first appearance time of each target, the GPS longitude and latitude coordinates, and the confidence level, and generating a statistical result.

[0018] And, a system for counting and positioning the number of dead trees based on UAV orthophoto video;

[0019] Including:

[0020] A UAV video stream acquisition module, used to acquire the original video frames and synchronous positioning data of the UAV orthophoto video;

[0021] A dual detection module, which trains a deep learning model based on a single-stage detection architecture to perform target detection on each frame of the image, and realizes basic tracking through a multi-target tracker associated with the detection, combines a slice detection strategy for periodic enhanced detection, and eliminates redundant detection frames through an overlapping area fusion algorithm to form a high-confidence target detection result;

[0022] A target state machine management module, based on the analysis of the target motion trajectory and spatio-temporal continuity, to realize the cycle management of the appearance - tracking - loss - recovery of the dead tree target;

[0023] An output module, based on the mapping relationship between the pixel coordinate system and the geographic coordinate system, combines the UAV flight parameters to convert the target position into GPS coordinates; outputs the dead tree statistical result and geographic positioning information containing spatio-temporal distribution characteristics.

[0024] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of a method for counting and positioning the number of dead trees based on UAV orthophoto video as described above.

[0025] A non-transitory computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of a method for counting and locating dead trees based on ortho-video of an unmanned aerial vehicle as described above.

[0026] Compared with the prior art, the present invention and its preferred solutions construct a collaborative inference mechanism for full-frame detection and slice enhancement, and balance detection accuracy and computational efficiency through a dynamic weight fusion strategy; construct an intelligent state machine model based on motion prediction compensation to achieve cross-frame continuity maintenance of target trajectories and identity consistency maintenance; design an end-to-end real-time mapping pipeline for geographical coordinates, and deeply embed the traditional post-processing positioning process into the main loop of detection and tracking.

[0027] Through ring buffer management, abnormal fusing mechanism and adaptive resource scheduling strategy, it improves the long-term operation stability in complex forest area scenarios to solve the problems of insufficient detection coverage, frequent trajectory breaks and excessive positioning delay in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0029] Figure 1 It is a general flowchart for constructing and implementing the solution of the embodiment of the present invention;

[0030] Figure 2 It is a flowchart for fusing detection results of the embodiment of the present invention;

[0031] Figure 3 It is a state transition diagram of the target state machine of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0033] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0035] Such as Figures 1 - 3As shown in the figure, this embodiment provides a construction and implementation process for a dead tree quantity statistics and positioning solution based on orthophoto video of an unmanned aerial vehicle (UAV), including the following steps;

[0036] Step S1: Obtain the original video frames and synchronous positioning data (SRT) in real time through the UAV video stream acquisition module;

[0037] Step S2: Construct a dual detection system, use a deep learning model based on a single-stage detection architecture (such as YOLO) for full-frame object detection, and implement basic tracking based on the multi-object tracker associated with the detection. Combine the slice detection strategy for periodic enhanced detection, and eliminate redundant detection frames through the overlapping area fusion algorithm to form a high-confidence object detection result;

[0038] Step S3: Establish a target state machine management system, and based on the analysis of the target motion trajectory and spatio-temporal continuity, implement the cycle management of the appearance - tracking - loss - recovery of dead tree targets;

[0039] Step S4: Based on the mapping relationship between the pixel coordinate system and the geographic coordinate system, combine the UAV flight parameters to convert the target position into GPS coordinates;

[0040] Step S5: Output the dead tree statistics results and geographic positioning information including spatio-temporal distribution characteristics.

[0041] As a preferred embodiment, the implementation of step S1 specifically includes the following steps;

[0042] Step A1: Obtain the timestamp, longitude and latitude coordinates, relative elevation, and flight attitude parameters by parsing the SRT data of the UAV, and establish a spatio-temporal reference coordinate system.

[0043] Step A2: Initialize the video stream processing pipeline, construct a video stream input pipeline through the video capture interface of the computer vision library (such as OpenCVVideoCapture), and set a circular buffer to store frame images.

[0044] Step A3: Construct a geographic coordinate transformation matrix, and based on the UAV lens parameters and the initial longitude and latitude coordinates recorded by SRT, establish a mapping relationship from the pixel coordinate system to the GPS coordinate system through perspective transformation.

[0045] As a preferred embodiment, step S2 specifically includes the following steps;

[0046] Step B1: Train a deep learning model based on a single-stage detection architecture, perform object detection on each frame image, and implement basic tracking through the multi-object tracker associated with the detection, and output the basic detection frame and confidence level.

[0047] Step B2: Trigger slice detection every N frames. Preferably, N is 2 frames. Divide the input frame into multiple overlapping sub-regions, perform object detection on each sub-region, and generate a set of local detection results.

[0048] Step B3: Calculate the spatial similarity of adjacent slice detection boxes, construct an IoU matrix, eliminate redundant detection boxes through non-maximum suppression, and fuse the full-frame detection results and slice detection results based on the IoU threshold, and feedback the enhanced detection results to update the tracker state.

[0049] As a preferred embodiment, step S3 specifically includes the following steps;

[0050] Step C1: Generate a unique identifier for each detected object. When a new object is detected, establish a state machine instance containing spatio-temporal feature vectors. The specific process of the state machine includes: initializing a state dictionary containing feature vectors such as target position confidence, consecutive hit counter, lost frame counter, motion vector, and predicted trajectory window; calculating the instantaneous velocity through the displacement difference of the target center point, setting an exponential decay coefficient to update the motion vector to achieve trajectory smoothing; recording auxiliary information such as the frame number when the target is first detected and the position of the previous frame for state transition judgment.

[0051] Step C2: Achieve state transition by comparing the tracking threshold and the recovery threshold. When the target is continuously hit for N frames, activate the stable tracking state, update the motion vector using an exponential decay model, and synchronously perform sliding window smoothing of the position coordinates.

[0052] Step C3: Construct a prediction box generation model based on motion extrapolation. When the target is lost, execute the following prediction matching recovery process:

[0053] (1) Predict the target position based on the historical motion vector: According to the center position P(t) of the target at frame t and the historical motion vector V(t), calculate the predicted position P'(t + 1) = P(t) + α·V(t) at frame t + 1 through the momentum decay coefficient α ∈ (0, 1), where V(t) = P(t) - P(t - 1);

[0054] (2) Construct the IoU matrix of the prediction box B pred with the set of current frame detection boxes {Bdet}, and calculate the intersection over union When the maximum IoU value exceeds the set threshold, it is determined that the match is successful;

[0055] (3) When the state is restored, execute: Reset the lost counter lost_count to 0, update the target position to the center P of the matched detection box matched , and use the linear interpolation algorithm to compensate for the trajectory points during the lost period. The interpolation formula is where t lost ≤ k ≤ trecover Indicates the frame sequence number during loss;

[0056] Loss

[0057] (4) When the number of consecutive lost frames satisfies lost_count > N, move the target to the inactive list and retain the feature vectors {position, velocity, confidence, timestamp} of the last N frames for trajectory backtracking;

[0058] Step C4: Establish a lost frame count fusing mechanism. When more than N consecutive frames are lost, remove the inactive targets, where N ∈ [20, 30], preferably 30 frames. While ensuring the continuity of tracking, it can effectively prevent long-term lost targets from occupying system resources; retain the feature vectors of the last N frames for trajectory backtracking; for the targets with successful recovery, use the linear interpolation algorithm for trajectory compensation.

[0059] As a preferred embodiment, step C2 specifically includes the following steps;

[0060] Step D1: Establish a target speed prediction model. When two consecutive frames of the target are detected, calculate the displacement difference of the center point V(t) = P(t) - P(t - 1) as the instantaneous speed; use the exponential weighted average method to update the speed vector

[0061] V'(t) = α·V(t - 1) + (1 - α)·V(t), where the historical speed weight coefficient α ∈ [0.6, 0.8], preferably α is 0.7, which can achieve a balance between the stability of the historical speed and the responsiveness of the current speed, and can effectively suppress the speed fluctuations caused by detection jitter;

[0062] Step D2: Design a circular trajectory buffer, set the buffer length N ∈ [20, 50] frames, preferably 30 frames, which can achieve a balance between memory occupancy and the time window length of trajectory analysis, facilitating time series analysis; use the first-in-first-out strategy to maintain the target motion trajectory {P(t - N + 1), ···, P(t)}; when the new trajectory point P(t + 1) is added, automatically eliminate the earliest historical point P(t - N + 1);

[0063] Step D3: Implement sliding window trajectory smoothing, perform exponential decay weighted averaging based on the trajectory point sequence {P(t - w + 1), ···, P(t)} within the current window; the formula for the smoothed position is where i ∈ [0, w - 1], the decay factor β ∈ [0.8, 0.95], preferably 0.9, which provides a good smoothing effect while maintaining sensitivity to trajectory changes; the window size w is preferably 5, and it will not introduce too much delay when performing local smoothing;

[0064] Step D4: Perform speed constraint trajectory correction, and perform backward correction on the historical trajectory according to the smoothed speed vector V'(t); the correction formula is P corrected (k) = P'(t) + V'(t)·(k - t), where k ∈ [t - w + 1, t]; in this way, the trajectory mutation caused by detection jitter is eliminated.

[0065] As a preferred embodiment, step S4 specifically includes the following steps;

[0066] Step E1: Construct a pixel-geospatial mapping model based on the SRT parameters recorded during the flight of the drone, and establish a bidirectional conversion relationship between the image pixel coordinate system and the geodetic coordinate system through perspective projection transformation.

[0067] Step E2: According to the relative image pixel position of the target center point in the video frame, calculate the GPS longitude and latitude coordinates through the geospatial interpolation algorithm, and the geospatial interpolation algorithm uses the bilinear interpolation method.

[0068] As a preferred embodiment, step S5 specifically includes the following steps;

[0069] Step F1: Establish a spatio-temporal coordinate mapping relationship for each detected target by aligning the video frame timestamp with the SRT positioning data;

[0070] Step F2: Render a floating window containing the target ID and GPS coordinates in real time in the video stream, record the total number of dead trees, the first appearance time of each target, the GPS longitude and latitude coordinates, and the confidence level, generate statistical results, and store them as structured JSON format data.

[0071] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0072] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0073] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0074] The above shows and describes the basic principles, main features, and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

[0075] This patent is not limited to the above best implementation manner. Anyone can obtain other various forms of a method for counting and positioning the number of dead trees based on orthophoto video of unmanned aerial vehicles under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.

Claims

1. A method for counting and locating dead trees based on drone orthophoto video, characterized in that: A dataset constructed based on the original video frames and synchronous positioning data of the drone's orthophoto video; Train a deep learning model based on a single-stage detection architecture, perform target detection frame by frame, and implement basic tracking through a multi-target tracker that detects associations. Combine the slice detection strategy for periodic enhanced detection, and eliminate redundant detection boxes through an overlapping region fusion algorithm to form a high-confidence target detection result. Based on the target movement trajectory and spatiotemporal continuity analysis, the emergence-tracking-loss-recovery cycle management of dead wood targets can be achieved; Based on the mapping relationship between the pixel coordinate system and the geographic coordinate system, the target position is converted into GPS coordinates in combination with the UAV flight parameters, and the dead wood statistical results and geographic positioning information containing the temporal and spatial distribution characteristics are output.

2. The method for counting and locating dead trees based on drone orthophoto video according to claim 1 is characterized by: The process of constructing the data set specifically includes: parsing the UAV's synchronous positioning data to obtain the timestamp, longitude and latitude coordinates, relative elevation and flight attitude parameters, and establishing a time-space reference coordinate system; initializing the video stream processing pipeline, building a video stream input pipeline through the video capture interface, and setting up a circular buffer to store frame images; constructing a geographic coordinate conversion matrix, based on the UAV lens parameters and the initial longitude and latitude coordinates recorded by the synchronous positioning data, and establishing a mapping relationship from the pixel coordinate system to the GPS coordinate system through perspective transformation.

3. The method for counting and locating dead trees based on drone orthophoto video according to claim 1 is characterized by: The training is based on a deep learning model of a single-stage detection architecture, performing target detection frame by frame, and realizing basic tracking through a multi-target tracker associated with detection, performing periodic enhanced detection in combination with a slice detection strategy, eliminating redundant detection frames through an overlapping region fusion algorithm, and forming a high-confidence target detection result. Specifically, the following steps are included: A deep learning model based on a single-stage detection architecture is trained to detect targets frame by frame, and basic tracking is achieved through a detection-associated multi-target tracker, and basic detection frames and confidences are output; frame image slice detection is triggered at intervals of N frames, the input frame is divided into multiple overlapping sub-regions, target detection is performed on each sub-region, and a local detection result set is generated; the spatial similarity of adjacent slice detection frames is calculated, an IoU matrix is ​​constructed, redundant detection frames are eliminated through non-maximum suppression, and full-frame detection results and slice detection results are fused based on an IoU threshold, wherein the fusion process includes: uniformly converting the detection frames of the full-frame detection results and the slice detection results into center point coordinates and width-height formats; sorting all detection frames in descending order according to confidence scores; adopting a non-maximum suppression strategy, processing each high-scoring detection frame in turn, calculating the intersection-over-union ratio with other detection frames, retaining the detection frame with higher confidence and removing the low-scoring detection frame with high overlap when the intersection-over-union ratio is greater than a preset threshold, and finally obtaining a non-redundant fused detection result; and finally feeding back the enhanced detection results to update the tracker state.

4. The method for counting and locating dead trees based on drone orthophoto video according to claim 1 is characterized in that: The target movement trajectory and spatiotemporal continuity analysis are used to achieve the management of the dead wood target appearance-tracking-loss-recovery cycle. The steps include: A unique identifier is generated for each detected target. When a new target is detected, a state machine instance containing a spatiotemporal feature vector is established. The specific process of establishing the state machine instance includes: initializing a state dictionary containing a target position confidence, a continuous hit counter, a lost frame counter, a motion vector, and a feature vector of a predicted trajectory window; calculating the instantaneous speed through the displacement difference of the target center point, setting an exponential decay coefficient to update the motion vector, and achieving trajectory smoothing; recording auxiliary information including the frame number and the position of the previous frame when the target is first detected, which is used for state transition judgment; and then comparing the tracking threshold with the recovery threshold. To achieve state transfer, when the target hits N frames continuously, the stable tracking state is activated, the motion vector is updated using an exponential decay model, and the sliding window smoothing of the position coordinates is performed synchronously; when the target is lost, the following prediction matching recovery process is executed: (1) Predict the target position based on the historical motion vector: According to the target’s center position P(t) in the tth frame and the historical motion vector V(t), the predicted position of the t+1 frame is calculated by the momentum decay coefficient α∈(0,1): P'(t+1)=P(t)+α·V(t), where V(t)=P(t)-P(t-1); (2) Construct the prediction box B pred Calculate the IoU matrix with the current frame detection box set {Bdet} and the intersection over union ratio When the maximum IoU value exceeds the set threshold, the match is considered successful; (3) When the state is restored, the lost counter lost_count is reset to 0, and the target position is updated to the center of the matching detection box P matched , and a linear interpolation algorithm is used to compensate for the trajectory points during the loss period. The interpolation formula is: where t lost ≤k≤t recover Indicates the frame number during the loss period; (4) When the number of consecutive lost frames satisfies lost_count>N, the target is moved to the inactive list and the feature vectors of the most recent N frames are retained {position, velocity, confidence, timestamp} is used for trajectory backtracking.

5. The method for counting and locating dead trees based on drone orthophoto video according to claim 4 is characterized in that: The state transfer is achieved by comparing the tracking threshold and the recovery threshold. When the target hits N frames continuously, the stable tracking state is activated, the motion vector is updated by using the exponential decay model, and the sliding window smoothing of the position coordinates is performed synchronously. Specifically, the motion trajectory is optimized through the speed prediction model and the trajectory buffer management. The implementation process includes: establishing a target speed prediction model, when two consecutive frames of targets are detected, the center point displacement difference is calculated as the instantaneous speed, and the speed vector is updated by using the exponential weighted average method; based on the annular trajectory buffer, the first-in-first-out strategy is used to maintain the target motion trajectory, and the earliest historical point is automatically eliminated when a new trajectory point is added; Implement sliding window trajectory smoothing, and perform exponential decay weighted averaging based on the trajectory points in the current window; Execute speed-constrained trajectory correction, the speed-constrained trajectory correction process includes: obtaining a trajectory point sequence, calculating the center position difference of adjacent trajectory points, and generating a velocity vector sequence that characterizes the target motion characteristics; using an exponential moving average algorithm to smooth the velocity vector sequence, wherein a smoothing coefficient α ranges from 0 to 1 and is used to adjust the weight of historical velocity information; taking the trajectory starting point as a reference, sequentially accumulating the smoothed velocity vector to the position coordinates of the previous trajectory point to generate a new trajectory point sequence that satisfies the motion continuity constraint; and updating the original trajectory cache with the newly generated trajectory point sequence to achieve trajectory optimization based on a physical kinematic model to suppress trajectory jitter caused by unstable target detection.

6. The method for counting and locating dead trees based on drone orthophoto video according to claim 1 is characterized by: The method of converting the target position into GPS coordinates based on the mapping relationship between the pixel coordinate system and the geographic coordinate system and combining the flight parameters of the drone specifically includes the following steps: constructing a pixel-geographic space mapping model based on the synchronous positioning data parameters recorded during the flight of the drone, and establishing a bidirectional conversion relationship between the image pixel coordinate system and the geographic coordinate system through perspective projection transformation; calculating the GPS latitude and longitude coordinates through a geographic space interpolation algorithm according to the relative position of the image pixels of the target center point in the video frame; the geographic space interpolation algorithm adopts a bilinear interpolation method.

7. The method for counting and locating dead trees based on drone orthophoto video according to claim 1 is characterized by: The output includes the dead wood statistical results and geographic positioning information with temporal and spatial distribution characteristics, and specifically includes the following steps: establishing a temporal and spatial coordinate mapping relationship for each detected target by aligning the video frame timestamp with the time of the synchronous positioning data positioning data; rendering a floating window containing the target ID and GPS coordinates in real time in the video stream, recording the total number of dead trees, the first appearance time of each target, the GPS longitude and latitude coordinates, and the confidence level, and generating statistical results.

8. A dead tree counting and positioning system based on drone orthophoto video, characterized by: include: The UAV video stream acquisition module is used to obtain the original video frames and synchronous positioning data of the UAV orthographic video; The dual detection module performs target detection on a frame-by-frame basis by training a deep learning model based on a single-stage detection architecture, and implements basic tracking through a multi-target tracker that detects associations. It combines the slice detection strategy for periodic enhanced detection, and eliminates redundant detection boxes through an overlapping region fusion algorithm to form a high-confidence target detection result. The target state machine management module is based on the target motion trajectory and spatiotemporal continuity analysis to achieve the appearance-tracking-loss-recovery cycle management of dead wood targets; The output module converts the target position into GPS coordinates based on the mapping relationship between the pixel coordinate system and the geographic coordinate system and combines the flight parameters of the UAV; it outputs the statistical results of dead trees and geographic positioning information containing the temporal and spatial distribution characteristics.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a method for counting and locating the number of dead trees based on unmanned aerial vehicle orthophoto video as described in any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for counting and locating dead trees based on unmanned aerial vehicle orthophoto video as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Overlapped frame optimization processing method based on remote sensing image partitioning

    CN120823521A

  • Automatic target detection and tracking method and system based on deep learning

    CN121305052A

  • Deep learning-based automatic target detection and tracking method and system

    CN121305052B