Image recognition method based on multi-rotor unmanned aerial vehicle
By establishing an inertial reference frame attitude reconstruction model and an image entropy change rate enhancement mechanism on a multi-rotor UAV, and combining it with a deep cross-attention neural network, the problems of recognition delay and accuracy of multi-rotor UAV image recognition system in complex environments are solved, achieving stable and efficient target recognition and tracking.
Patent Information
- Application Number
- CN202511110608.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multi-rotor UAV image recognition systems suffer from delays between image acquisition and recognition during flight. Recognition accuracy is significantly affected by changes in flight attitude and lighting interference, making it difficult to achieve stable recognition in complex scenarios. In particular, targets are prone to being lost or misidentified when they are moving rapidly or in areas with complex terrain.
By acquiring image frame sequences and attitude information during the flight of a multi-rotor UAV, an attitude reconstruction model based on an inertial reference frame is established for dynamic distortion correction. Combined with an adaptive enhancement mechanism for the rate of change of image entropy, a deep cross-attention neural network is used for multi-scale recognition and dynamic tracking, and a local adaptive re-inspection strategy is triggered when the target behavior is unstable.
It significantly improves the stability and accuracy of image recognition, enhances the system's robustness and self-recovery capability in complex environments, and achieves efficient target recognition and tracking, making it suitable for inspection and aerial sensing in complex environments.
Smart Images

Figure CN120976800A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image recognition and autonomous perception technology for unmanned aerial vehicles (UAVs), and specifically to an image recognition method based on a multi-rotor UAV. Background Technology
[0002] With the rapid development of artificial intelligence and drone technology, multi-rotor drones, with their advantages of maneuverability, vertical takeoff and landing, and strong hovering capabilities, have been widely used in various fields such as environmental monitoring, disaster assessment, agricultural inspection, and intelligent security. In practical applications, drones equipped with image acquisition equipment acquire high-altitude images, and through image recognition technology, they achieve automatic identification and analysis of ground targets, becoming an important means to improve operational efficiency and information acquisition accuracy.
[0003] Existing multi-rotor UAV image recognition systems primarily employ static acquisition followed by post-processing. This results in a significant delay between image acquisition and recognition, and recognition accuracy is greatly affected by factors such as changes in flight attitude and lighting interference, making it difficult to achieve stable recognition results in complex scenarios. Furthermore, when the target is moving rapidly or flying in areas with complex terrain, traditional recognition methods are prone to target loss or misidentification, reducing the system's practicality and intelligence level.
[0004] Therefore, there is an urgent need to provide an image recognition method based on multi-rotor UAVs that can acquire image information in real time during flight and improve recognition accuracy by combining attitude compensation, image enhancement and other processing methods, so as to meet the requirements of real-time performance and stability of image recognition in complex environments. This has become a key technical problem that needs to be solved in the field of UAV intelligent perception. Summary of the Invention
[0005] The purpose of this invention is to provide an image recognition method based on a multi-rotor unmanned aerial vehicle (UAV) to address the shortcomings of the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an image recognition method based on a multi-rotor unmanned aerial vehicle, comprising: S100: Acquire the image frame sequence I and its corresponding attitude information P that are acquired in real time by the image acquisition module during the flight of the multi-rotor UAV. The attitude information includes the current three-axis acceleration, three-axis angular velocity and global positioning coordinates of the aircraft. S200. Based on the image frame sequence I and the attitude information P, a dynamic distortion correction is performed on the image frames by establishing an attitude reconstruction model based on an inertial reference frame, and the attitude-stable image frame sequence I' after time synchronization processing is obtained. S300. An adaptive enhancement mechanism based on the image entropy change rate is introduced into the attitude-stabilized image frame sequence I' to dynamically adjust the image brightness, contrast and edge sharpness, and generate a multi-resolution image pyramid sequence I''. S400 employs a deep cross-attention neural network that combines an image pyramid sequence I'' with a trajectory prediction model to perform multi-scale recognition and dynamic tracking of targets in images, and outputs the target recognition result O and its temporal position distribution vector V. S500: Input the recognition result O and the flight trajectory vector V into the behavior evaluation module, and calculate the target behavior stability factor R based on the target continuity, confidence distribution and speed consistency. S600 When the target behavior stability factor R is lower than the set threshold, the UAV local adaptive re-inspection strategy is activated. The inspection path is reconstructed based on the current image change trend, and the aircraft is controlled to perform directional return flight and re-acquire images of the target area.
[0007] Preferably, S100 includes: S101. Obtain the image frame sequence I and the attitude information P; S102. Combining the attitude estimation matrix with the global positioning coordinates, the image frame I is labeled with location encoding and the trajectory sequence is sorted to construct an image attitude temporal structure G = {I, P, L, t} that perfectly matches the flight state, where L is the geographic label and t is the acquisition timestamp.
[0008] Preferably, S200 includes: S201. Based on the image attitude temporal structure G = {I, P, L, t}, an inertial reference frame is introduced as a unified reference coordinate system. The three-axis acceleration and angular velocity data in the attitude information P are transformed to form the attitude input set in the inertial coordinate domain. S202. Establish an attitude reconstruction model to estimate the heading angle, pitch angle and roll angle corresponding to the image frame in the time domain. The model is based on the improved quaternion filter and the image temporal continuity to jointly solve the image attitude state variables. S203. The reconstructed attitude state variables are used as transformation parameters to input the nonlinear geometric distortion correction function. Dynamic geometric correction is performed frame by frame on the image frame sequence I. At the same time, the image sequence is synchronized with high precision by combining the image acquisition timestamp, and the attitude-stable image frame sequence I′ is output.
[0009] Preferably, S300 includes: S301. Divide the attitude-stabilized image frame sequence I′ into local regions, calculate the image entropy change rate of each region in the image frame using a sliding window method, and construct the entropy gradient distribution matrix. S302. Based on the entropy gradient distribution matrix and the image enhancement mechanism based on regional priority control, the brightness compensation coefficient, local contrast stretching parameter and edge sharpening kernel size of the corresponding region are dynamically adjusted to achieve targeted image enhancement processing. S303. After enhancement processing, a multi-resolution image pyramid sequence I″ is constructed using a multi-scale decomposition algorithm. This pyramid sequence generates image copies of different scales centered on high-entropy blocks. S304. Feed the changes in local parameters and entropy weights during the image enhancement and pyramid construction process back to the enhancement mechanism to form a closed-loop adjustment structure based on the dynamic response of image entropy.
[0010] Preferably, S400 includes: S401. Based on the generated multi-resolution image pyramid sequence I″, construct a multi-scale feature map set F; S402. Based on the historical flight path data of UAVs, a trajectory prediction model T is constructed. The position change trend is extracted using a long short-term memory network, and the predicted trajectory vector sequence T′ is output, which is used as the time-series dynamic prior input. S403. Construct a deep cross-attention neural network M, feed F and T′ into the spatial branch and the temporal branch respectively, and use a multi-layer cross-attention mechanism to realize the fusion reasoning of image spatial features and trajectory dynamic priors to obtain the target position saliency response map. S404. Extract the target recognition result O based on the saliency response map, and generate its temporal position distribution vector V in the image sequence. At the same time, attach its fusion confidence score to each recognition result.
[0011] Preferably, S500 includes: S501. Associate the target category, image coordinates, and confidence information in the recognition result O with the corresponding temporal position information in the trajectory vector V to generate a fusion structure alignment set. , where t is the collection timestamp sequence; S502. Based on the structure alignment set D, construct a behavior evaluation module that includes a target continuity analysis submodule, a confidence fluctuation discrimination submodule, and a speed consistency evaluation submodule. S503, the target continuity analysis submodule calculates the continuity score Rc based on the target's spatial movement distance in the image frame sequence and the inter-frame matching success rate; the confidence fluctuation discrimination submodule calculates the stability score Rs based on the variance of the confidence sequence change; the velocity consistency assessment submodule compares the target velocity vector with the flight velocity vector and outputs the dynamic matching score Rv. S504. A multi-factor fusion strategy is adopted to combine Rc, Rs and Rv in a weighted manner to form the target behavior stability factor R, which is used to measure the overall credibility and dynamic consistency of the identified target.
[0012] Preferably, S600 includes: S601. When the target behavior stability factor R is lower than the set threshold Rth, the re-inspection strategy is started, and the image change analysis module is called to perform time-series analysis on the most recent image frame sequence I″, and extract the image texture change vector ΔG and the edge offset trajectory ΔE. S602. Input ΔG and ΔE into the backtracking path prediction model. This model integrates the image difference heatmap and the current heading angle of the UAV to generate a locally reconstructed inspection path vector P′, and optimizes the turning angle and flight speed between path nodes. S603: Call the flight control interface module to send the path vector P′ to the multi-rotor navigation control unit, so that the aircraft can fly back to the optimal shooting point of the suspected target area and trigger the image acquisition module to reacquire the image frame of the target area; S604. The reacquired image frames are automatically fed back to the image acquisition and attitude synchronization process in step S100 to realize a closed-loop flight control recognition cycle based on dynamic image judgment.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention provides an image recognition method based on a multi-rotor unmanned aerial vehicle (UAV), which fully integrates flight attitude modeling, image enhancement processing, multi-scale recognition networks, and flight control path adaptive mechanisms to establish a closed-loop intelligent recognition system encompassing image acquisition, recognition, behavior evaluation, and re-inspection. This method effectively improves image quality and spatiotemporal consistency by introducing an attitude reconstruction model in an inertial reference frame and an enhancement mechanism based on image entropy change rate, significantly enhancing the stability and recognizability of image input under complex flight conditions.
[0014] 2. This invention combines an image pyramid structure with a trajectory prediction model to construct a deep cross-attention recognition network, enabling the fusion of multi-scale spatial features and temporal prior information for recognition and reasoning. Simultaneously, by constructing a target behavior stabilization factor and triggering an adaptive return-to-flight patrol strategy, the system possesses highly robust recognition capabilities, self-recovery ability, and dynamic interference tolerance. Its overall performance is significantly superior to traditional image recognition paths, making it widely applicable to inspection, monitoring, and aerial sensing scenarios in complex environments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0016] Figure 1 This is a mind map of the method of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1, please refer to Figure 1 As shown in this embodiment, an image recognition method based on a multi-rotor unmanned aerial vehicle (UAV) includes: S100: Acquire the image frame sequence I and its corresponding attitude information P that are acquired in real time by the image acquisition module during the flight of the multi-rotor UAV. The attitude information includes the current three-axis acceleration, three-axis angular velocity and global positioning coordinates of the aircraft. S200. Based on the image frame sequence I and the attitude information P, a dynamic distortion correction is performed on the image frames by establishing an attitude reconstruction model based on an inertial reference frame, and the attitude-stable image frame sequence I' after time synchronization processing is obtained. S300. An adaptive enhancement mechanism based on the image entropy change rate is introduced into the attitude-stabilized image frame sequence I' to dynamically adjust the image brightness, contrast and edge sharpness, and generate a multi-resolution image pyramid sequence I''. S400 employs a deep cross-attention neural network that combines an image pyramid sequence I'' with a trajectory prediction model to perform multi-scale recognition and dynamic tracking of targets in images, and outputs the target recognition result O and its temporal position distribution vector V. S500: Input the recognition result O and the flight trajectory vector V into the behavior evaluation module, and calculate the target behavior stability factor R based on the target continuity, confidence distribution and speed consistency. S600 When the target behavior stability factor R is lower than the set threshold, the UAV local adaptive re-inspection strategy is activated. The inspection path is reconstructed based on the current image change trend, and the aircraft is controlled to perform directional return flight and re-acquire images of the target area.
[0019] To achieve the image recognition method based on a multi-rotor UAV as described in this invention, the following technical solutions are adopted in the image acquisition and attitude information acquisition stages to ensure the accuracy and stability of subsequent image compensation and recognition processing: In step S100, an integrated dual-mode sensor acquisition module installed under the multi-rotor UAV fuselage is used to synchronously acquire image frame sequence I and attitude information P. The image acquisition module continuously acquires image frames at a preset frame rate to form an image stream; the attitude information acquisition module is a high-frequency inertial measurement unit (IMU) used to output triaxial acceleration and triaxial angular velocity data in real time, thereby forming the original attitude vector set.
[0020] To address the impact of inertial disturbances and vibrations during flight, the system introduces a dynamic fusion algorithm based on the coupling of a Kalman filter and a flight state model. This algorithm effectively filters out high-frequency noise and disturbance errors in the acquired attitude information by establishing a state prediction-observation update mechanism. The output of the fusion algorithm is an attitude estimation matrix, which is used for time synchronization and registration of subsequent image frames, ensuring consistency and accurate matching between image frames and attitude data.
[0021] To further enhance the spatial information representation capability of image frames, the attitude estimation matrix is fused with location information provided by the Global Positioning System (GPS). A geographic label L is added to each image frame I, and the image sequence is sorted according to the trajectory order based on the acquisition timestamp t, thereby constructing an image attitude temporal structure G = {I, P, L, t}. This structure fully expresses the mapping relationship between image frames, attitude states, location information, and time information.
[0022] Finally, the aforementioned image attitude temporal structure G is input into the subsequent attitude compensation module, providing fundamental data support for dynamic distortion correction and image enhancement. By introducing this structure, the spatial consistency of image frames under different flight states is enhanced, thereby significantly improving the generation accuracy of the attitude-stabilized image frame sequence I′ and the overall robustness of image processing.
[0023] The above technical approach effectively solves the problems of asynchronous image and attitude data, significant image distortion, and inaccurate registration during the flight of multi-rotor UAVs, providing a stable input foundation for achieving high-precision image recognition in complex scenarios.
[0024] In the image recognition method based on a multi-rotor UAV described in this invention, to improve the accuracy and robustness of image recognition, especially addressing issues such as image distortion and offset caused by attitude disturbances in image frames during flight, the system introduces an attitude compensation process during the image acquisition and preprocessing stages. This process is specifically embodied in step S200, the core of which involves constructing an attitude reconstruction model based on an inertial reference frame and performing dynamic geometric distortion correction on the image frames, ultimately outputting an attitude-stabilized image frame sequence I′ as input for subsequent image enhancement and target recognition models.
[0025] In step S201, the system first parses the image attitude temporal structure G = {I, P, L, t} constructed in the previous stage, where I represents the original image frame sequence acquired by the image acquisition module, P represents the attitude information corresponding to the image frame, including the three-axis acceleration and three-axis angular velocity vectors; L is the global positioning geographic tag corresponding to the time of image acquisition, and t is the acquisition timestamp. To ensure that the attitude compensation processing between each image frame is uniform and effective, the system introduces an inertial reference frame as the global attitude processing coordinate framework. Specifically, the three-axis acceleration (Ax, Ay, Az) and three-axis angular velocity (Gx, Gy, Gz) in the attitude information P are transformed from the local coordinate system of the flight control body to the inertial coordinate system to establish an attitude input set consistent with the real space physical state, thereby enhancing the global consistency and physical interpretability of subsequent attitude estimation.
[0026] In step S202, the system constructs an attitude reconstruction model based on the attitude input set in the inertial coordinate domain, which is used to accurately estimate the spatial attitude at the time of image acquisition in the time domain. This model incorporates the dynamic state change characteristics of the aircraft during continuous frame acquisition and introduces an improved quaternion filter as the core solution unit. Compared with traditional algorithms, the quaternion filter improves the update step size and error weight processing mechanism, better filtering out high-frequency fluctuation interference and improving the stability of attitude estimation. To further improve estimation accuracy, the model also integrates image temporal continuity constraints. That is, when calculating the attitude of a certain image frame, it not only relies on the current acceleration and angular velocity inputs but also refers to the smoothing trend of the attitude state in previous and subsequent frames, thus forming a state prediction-update mechanism within the time window. Through this modeling method, corresponding three-axis attitude angle values can be generated for each image frame: yaw, pitch, and roll, forming a complete sequence of attitude state variables.
[0027] After attitude state reconstruction is completed, in step S203, the system uses the obtained attitude state variables as geometric transformation parameters and inputs them into a preset nonlinear geometric distortion correction function. This function is based on an image projection perturbation model caused by multi-rotor flight attitude changes, and corrects affine distortion and rotational distortion caused by viewpoint deflection in the image frame-by-frame through a nonlinear function mapping relationship. Specifically, using the pixel coordinates of the image frame as a reference, the quaternion-form attitude transformation is mapped to a geometric transformation matrix in image space, thereby correcting key areas such as image edges, corners, and texture structures to ensure that the spatial structure of the image remains consistent under flight perturbations.
[0028] Furthermore, considering the potential time shift issues that may arise from differences in sensor sampling frequencies during flight, this step further incorporates the image acquisition timestamp t to perform high-precision synchronization processing of the image frame sequence in the time domain. Specifically, the system ensures that image frame I and its corresponding attitude state are synchronously registered within microseconds through interpolation reconstruction and time window alignment mechanisms, avoiding the cumulative effect of attitude compensation errors caused by time misalignment.
[0029] After distortion correction and time synchronization are completed, the system finally outputs a sequence of attitude-stabilized image frames, I′. The image frames in this sequence have been compensated in spatial geometry and are highly consistent in temporal structure. They can serve as a high-quality input source for subsequent modules of the image recognition model, significantly improving the overall recognition system's adaptability to dynamic disturbances and the extractability of image features.
[0030] To further improve the detection accuracy and robustness of image recognition models under different environmental conditions, this invention, after completing image frame pose compensation, designs an adaptive image enhancement mechanism driven by image entropy change rate to address issues such as unstable image quality, unclear target details, and uneven distribution of image region information. Furthermore, it constructs a pyramid image structure by combining a multi-scale image representation strategy, significantly improving the hierarchical representation capability of image features. This processing corresponds to step S300 of this invention and specifically includes the following technical implementation path: In step S301, the input pose-stabilized image frame sequence I′ is first divided into local regions. A fixed-size sliding window strategy is adopted to scan the image frames with a set step size, and the entropy change rate of the image is calculated within each sliding window region. This entropy change rate is used to measure the dispersion of pixel gray-level distribution and its information density within the local region. By traversing the entire image frame, an entropy gradient distribution matrix corresponding one-to-one with the image pixel space is finally constructed. This matrix not only characterizes the changes in information density between image regions, but also reflects features such as texture complexity and edge change trends, serving as the priority decision basis for subsequent image enhancement operations.
[0031] In step S302, based on the generated entropy gradient distribution matrix, the system introduces a region-priority controlled image enhancement mechanism. This mechanism performs brightness, contrast, and sharpness enhancement processing on local image regions according to the entropy weights of different regions. Specifically, it includes: For low-entropy regions (regions with sparse information and concentrated grayscale distribution), the system automatically increases the brightness compensation coefficient and appropriately increases the local contrast stretching amplitude to improve visual recognition. For high-entropy regions (regions with dense edges and complex structures), the system introduces an adaptive sharpening kernel, which improves the clarity of image structural details by adjusting the weight parameters of the edge enhancement filter. The enhancement parameters for all regions are dynamically adjusted based on the entropy distribution to avoid local over-enhancement or loss of detail that may be caused by a globally uniform enhancement strategy.
[0032] This region-weighted enhancement intensity allocation mechanism ensures that image enhancement is highly targeted and retains information effectively, greatly improving the visual feature contrast and texture integrity of key target areas in the image.
[0033] Next, in step S303, the system performs a multi-scale decomposition algorithm on the enhanced image frame sequence to generate a multi-resolution image pyramid structure. Unlike traditional image pyramid construction methods, this invention uses high-entropy blocks in the image as the central region for pyramid construction, and generates image copies of different resolutions layer by layer through a non-linear scale sampling strategy, forming a pyramid image sequence I″. This structure preserves the expressive power of high-information-density regions at different scales, while enhancing the model's multi-scale recognition capability during feature extraction, effectively solving the problem of small targets being easily ignored or large targets having blurred boundaries at a single scale.
[0034] In step S304, the system's adaptive adjustment capability is further enhanced. The system feeds back the changes in local processing parameters (such as brightness compensation coefficient, contrast adjustment range, and sharpening intensity) of each region during the entire image enhancement and pyramid construction process, along with their corresponding image entropy values, to the image enhancement mechanism, thus constructing an image entropy-driven closed-loop enhancement control structure. This structure optimizes the next round of enhancement strategy based on the previous round's enhancement output and forms an enhancement self-adjustment path based on image content distribution through dynamic adjustment of entropy weights.
[0035] The entire S300 process forms a four-stage closed loop of image structure optimization, from "image information distribution analysis" → "adaptive enhancement of regional features" → "pyramid-style multi-scale construction" → "enhanced feedback regulation". This not only significantly enhances the recognizability of the target region in the image, but also ensures that the model has consistent processing adaptability under different scene conditions (such as backlight, low light, sparse texture or complex interference).
[0036] In an image recognition method based on a multi-rotor UAV of the present invention, to effectively improve the accuracy, robustness, and temporal tracking capability of image target recognition, especially in application scenarios with multi-scale image input and complex flight trajectory disturbances, the system, after completing image enhancement and pyramid construction, uses a deep cross-attention neural network model that integrates image spatial features and dynamic prior knowledge of the flight path to perform multi-scale target recognition and dynamic tracking processing. This processing flow corresponds to step S400 and includes four sub-steps: image feature construction, trajectory prediction modeling, fusion recognition network construction, and target extraction output, as detailed below: In step S401, the system first extracts image structure information at different resolutions based on the multi-resolution image pyramid sequence I″ constructed in step S300, in order to construct a multi-scale feature map set F. At each image scale, the system uses a lightweight convolutional module to process the image, extracting key visual features such as texture features, edge gradients, corner structures, and color distribution. A scale normalization mechanism is then used to unify the representation dimensions of the feature maps at different resolutions, ultimately constructing a multi-scale image feature map set. , where n is the total number of images. This set can fully express the feature response information of images at different spatial levels, and is an important input foundation for achieving accurate target recognition.
[0037] In step S402, to introduce dynamic prior information about the flight trajectory's impact on target position changes, the system utilizes historical flight path data from a multi-rotor UAV to construct a trajectory prediction model T. This model is based on a Long Short-Term Memory (LSTM) network structure. The input is a sequence of the UAV's flight positions over a past period (timestamped and normalized), and the output is a sequence of predicted trajectory vectors T′ for several future moments. The prediction process captures speed trends, turning patterns, and attitude change characteristics along the flight path through time-series modeling, thereby providing a dynamic prior distribution of the target's possible locations in space. This prior T′ serves as a feature input in the time dimension, guiding the image feature inference process to focus on regions with spatial evolutionary significance.
[0038] In step S403, the system constructs a deep cross-attention neural network model M based on a fusion structure to integrate multi-scale spatial features F of the image with the trajectory prediction vector T′, so as to achieve cross-dimensional feature linkage reasoning and fusion recognition. The network includes two core input branches: one is a spatial branch, which receives the image feature set F and extracts spatial hierarchical semantics through a convolutional residual module; the other is a temporal branch, which receives the trajectory vector T′ and encodes it into a temporal prior feature tensor using a fully connected mapping layer.
[0039] In the network backbone, the system introduces a multi-layer cross-attention mechanism to model the saliency association between spatial feature maps and temporal priors. The attention module assigns dynamic weights from the predicted trajectory prior to each region response in the image feature map, enabling the network to perform spatial attention weighting by combining the historical trajectory trend of "potential target occurrence" when making target region judgments, thereby forming a temporally sensitive saliency response map S.
[0040] The cross-attention mechanism not only effectively enhances the adaptability of image features to trajectory trends, but also propagates target location information across different scales through a multi-scale fusion path, so that small target detection is no longer limited by the single-scale expression limitation, greatly improving the system's sensitivity to complex targets and recognition coverage.
[0041] In step S404, the system extracts the target location from the image based on the saliency response map S and outputs the target recognition result O. This recognition result includes: the target's category label, center coordinates in the image, bounding box, confidence value, and other information. Simultaneously, based on the temporal variation trend of the recognition result throughout the entire image sequence, the system combines the timestamp and trajectory prediction vector to generate a temporal location distribution vector V for the target, describing the dynamic distribution path of the target in consecutive frames.
[0042] In addition, to enhance the reliability of identification, the system adds a fusion confidence score to each identification result in the output stage. This score is calculated by weighting multiple factors such as spatial feature matching degree, track prior consistency and response map significance value, and can be used in subsequent steps to determine behavior stability and trigger re-inspection.
[0043] In the image recognition method based on a multi-rotor UAV described in this invention, to ensure the reliability of the recognition results and the consistency of tracking in a dynamic flight environment, the system further introduces a behavioral stability evaluation mechanism after completing target recognition and preliminary position distribution extraction. This mechanism, based on the target recognition results and flight trajectory information, calculates a comprehensive behavioral stability factor from three dimensions: spatial continuity, confidence stability, and velocity matching degree. This factor is used to measure whether the target possesses the stable characteristics to be identified as a highly reliable target. This processing flow corresponds to step S500 and includes the following four sub-steps: First, the system fuses the recognition result O with the information in the trajectory vector V in a one-to-one correspondence. Specifically, the recognition result O includes the target's category label, image coordinates (such as center point or bounding box coordinates), and recognition confidence value; the trajectory vector V represents the target's temporal position trajectory in the image sequence, output by step S404, and has temporal order. The system uses the collected timestamp sequence t as the main index to pair the recognition result in each frame of the image with the predicted flight path position at the corresponding time, forming a unified structure, denoted as the fusion structure alignment set D = {O, V, t}.
[0044] This structure alignment set not only integrates spatial location information and confidence information, but also introduces a time synchronization dimension, providing basic data support for subsequent stability analysis based on time-series change characteristics.
[0045] Based on the structural alignment set D, the system constructs a behavior evaluation module, which contains three functionally independent sub-modules: Target continuity analysis submodule: responsible for measuring the smoothness and consistency of the spatial location of the identified target in the image frame sequence; Confidence fluctuation discrimination submodule: used to determine the stability of target recognition confidence over time. Speed Consistency Assessment Submodule: Used to compare the consistency between the target's actual moving speed and the aircraft's own motion state.
[0046] This module avoids ambiguous judgments about behavioral stability through multidimensional modeling, ensuring that the evaluation process is fully interpretable and numerically quantifiable.
[0047] Within each submodule, the system calculates three core behavioral scoring indicators: Continuity score Rc: The target continuity analysis submodule calculates the consistency score in the spatial dimension based on the target's spatial displacement distance (such as Euclidean distance difference) in consecutive image frames and the success rate of cross-frame recognition matching (such as IOU overlap or positional correlation coefficient). If the target is always present in multiple consecutive frames, its position changes smoothly, and the matching is accurate, the score Rc is high; otherwise, it is low.
[0048] Stability Score Rs: The confidence fluctuation discrimination submodule uses the variance of the confidence sequence as the core parameter to construct a stability scoring model based on statistical volatility. If the confidence of a target changes little in each frame, the recognition system is considered to have a stable judgment of that target, and the score Rs is high; if the confidence fluctuates drastically or jumps frequently, it is considered an unstable target, and the score is reduced.
[0049] Velocity Matching Score Rv: The velocity consistency assessment submodule compares the target's identified velocity vector with the current flight velocity vector recorded by the aircraft's navigation system to determine whether the target might be a false target caused by the autonomous movement of the flight platform, or whether it matches the actual trajectory. This score is based on a comprehensive judgment of features such as the velocity vector angle, velocity amplitude ratio, and trend similarity, and outputs a dynamic matching score Rv.
[0050] After scoring across the three dimensions, the system enters the fusion judgment stage, employing a multi-factor weighted fusion strategy to combine the three scoring indicators Rc, Rs, and Rv, ultimately outputting the target behavior stability factor R. This fusion strategy allows for setting weighting coefficients based on different application scenarios to meet varying requirements for spatial continuity, confidence stability, and speed matching in different contexts. The formula is expressed as: ; where α, β, and γ are empirical weighting coefficients, satisfying α + β + γ = 1. This factor reflects the overall identifiability, temporal stability, and physical consistency of the target during the current flight phase, and is an important decision variable for determining whether to conduct retrospective data collection, secondary identification, or behavioral intervention.
[0051] In the image recognition method based on a multi-rotor UAV described in this invention, to improve the system's robustness and target re-capture capability in low-confidence scenarios, especially in complex dynamic environments where recognition results suffer from insufficient confidence or discontinuous behavior, the system introduces an image-driven local adaptive re-survey strategy. Furthermore, through intelligent path prediction and closed-loop control mechanisms, dynamic linkage control between image acquisition and the flight control system is achieved. This process, step S600 in the main claim, includes four key sub-steps, as detailed below: When the target behavior stability factor R calculated in step S500 is lower than the set confidence threshold Rth, the system determines that there is a risk of insufficient stability in the current target identification and immediately triggers a re-inspection strategy to enhance the identification reliability.
[0052] The system first invokes the image change analysis module, selects the image frame sequence I″ acquired within a recent period, and performs temporal analysis on it. During this analysis, the system uses the changes in texture features and edge structures between consecutive frames as the basis for calculation, obtaining the following two dynamic image features: Image texture change vector ΔG: By calculating texture feature indicators such as gray-level co-occurrence matrix and histogram of oriented gradient in each image frame, and comparing the change trends between frames, the local change direction and intensity of image texture are extracted. Edge offset trajectory ΔE: Use edge detection operators (such as Canny or Sobel) to extract the image edge structure, analyze the spatial offset path of edge lines in consecutive frames, and establish an edge point trajectory model.
[0053] These two features reflect the subtle dynamic changes that occur in the target region in the image, and are important evidence for identifying target position drift and uncertainty.
[0054] The two features ΔG and ΔE mentioned above are input into a pre-defined backtracking path prediction model. This model employs a multi-mode fusion path modeling structure, combining the following information sources for path reconstruction: Image difference heatmap: A map of significant image changes generated jointly by ΔG and ΔE, used to highlight the areas in an image frame where target state changes are most likely to occur; Current UAV heading angle information: used to correct the path prediction direction and avoid the path deviating excessively from the current flight attitude direction.
[0055] The model generates a reconstructed inspection path vector P′ within a local region by fusing the spatial variation trend of the image with the heading angle prediction control parameters. The path vector contains a set of spatial coordinate points ordered by time and their corresponding optimal heading angle adjustment information.
[0056] In addition, the model performs smooth optimization of the steering angle between each node in the path vector, and combines the characteristics of the aircraft's current speed curve to perform speed compensation at the path execution point, thereby generating a flight path that can be executed in real time under controllable physical conditions, ensuring optimal flight continuity and energy efficiency.
[0057] After the path vector P′ is generated, the system calls the flight control interface module to send the path command to the multi-rotor navigation control unit in the flight control system. The navigation system plans the flight path based on the spatial coordinates of the path points, the desired speed, and the heading angle change parameters, and controls the UAV to fly back to the target area along the predicted path.
[0058] When the aircraft reaches the end of its path or the system determines that it is at the optimal shooting point in the suspected target area, the system automatically triggers the image acquisition module to re-acquire image frames within the current field of view. This process provides the system with updated image data that highly overlaps with the original target identification area.
[0059] The reacquired image frames are immediately sent to the image recognition process and automatically fed back to the image acquisition and attitude synchronization processing module corresponding to step S100. The system treats them as a new input starting point and re-performs the attitude estimation, distortion compensation, enhancement processing, and recognition judgment of the image frames, thereby forming an image-recognition-path control-reacquisition closed-loop control structure triggered by target stability judgment.
[0060] This mechanism enables the image recognition system to adaptively adjust to low-quality targets during flight, avoiding erroneous outputs of unreliable results and reducing the risk of over-reliance on the confidence level of the initial recognition model. Furthermore, this closed-loop mechanism is highly automated, enabling a continuous process of recognition, evaluation, resampling, and re-recognition during flight without human intervention.
[0061] Example 2: To verify the effectiveness of the "image recognition method based on multi-rotor UAV" described in this invention, the following experimental system was constructed and field tests were conducted.
[0062] This embodiment is built on the following hardware and software platform: Drone platform: DJI Matrice 300 RTK multi-rotor drone, equipped with a dual-mode sensor acquisition module (including a 4K RGB camera and a 9-axis IMU inertial navigation unit). Image recognition terminal: NVIDIA Jetson AGX Xavier edge computing platform, running the deep cross-attention neural network described in this invention; Flight control system interface: Open ROS flight control framework, used to receive path vectors and control the patrol; Test scenario: Open farmland area (including shading, changes in terrain features, and light interference); Baseline comparison: Traditional YOLOv5 recognition model + fixed path flight recognition strategy.
[0063] Experimental procedure: Acquisition Phase: The aircraft flies along the set path, acquiring image frames and attitude information in real time during the flight, completing image synchronization and distortion compensation in steps S100–S200; Image Enhancement and Pyramid Construction: A comparison of the results of a group using the traditional CLAHE enhancement method and a group using the S300 adaptive enhancement + entropy-driven pyramid structure of this invention; Target recognition: Target detection was performed using both the traditional YOLOv5 and the deep cross-attention neural network model of this invention; Behavioral stability analysis and re-inspection: When the confidence level is below 0.6, the traditional group has no inspection response, while the group of this invention triggers flyback and reacquires images; Evaluation metrics include detection accuracy (mAP), temporal consistency (IoU sequence standard deviation), flight back response latency, and processing frame rate.
[0064] Experimental results: The results show that the local enhancement strategy based on the image entropy change rate in this invention can significantly improve image clarity compared to traditional unified enhancement methods, especially under conditions of blurred edges and uneven illumination, where the target texture is preserved more completely; the multi-resolution image pyramid structure provides rich context for small target detection, and the cross-attention network significantly improves detection accuracy and temporal coherence; the re-patrol strategy that introduces a behavior stabilization factor as a trigger mechanism effectively compensates for recognition failures caused by attitude disturbances or temporary occlusions, improving the system's target retention capability; although the fusion model slightly reduces the frame rate, the overall system processing is stable and meets the real-time requirements of conventional flight recognition.
[0065] In summary, this embodiment fully verifies the practicality of the present invention in terms of recognition robustness, target retention capability, and intelligent feedback behavior in complex flight environments, demonstrating significant technological advancements and achieving the aforementioned beneficial effects.
[0066] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A multi-copter drone based image recognition method, characterized by: Comprise: S100, acquire the image frame sequence I and its corresponding attitude information P collected by the image acquisition module in the flight process of the multi-rotor unmanned aerial vehicle, the attitude information includes the current three-axis acceleration, three-axis angular velocity and global positioning coordinates of the aircraft; S200, according to the image frame sequence I and the attitude information P, through the establishment of the attitude reconstruction model based on the inertial reference frame, the dynamic distortion correction of the image frame is carried out, and the attitude stable image frame sequence I' after time synchronization processing is obtained; S300, the adaptive enhancement mechanism based on image entropy change rate is introduced to the attitude stable image frame sequence I', the image brightness, contrast and edge definition are dynamically adjusted, and a multi-resolution image pyramid sequence I'' is generated; S400, a deep cross attention neural network combining the image pyramid sequence I'' and the flight path prediction model is used to carry out multi-scale recognition and dynamic tracking of the target in the image, and output the target recognition result O and its time sequence position distribution vector V; S500, the recognition result O and the flight trajectory vector V are input into the behavior evaluation module, and the target behavior stability factor R is calculated according to the target continuity, confidence distribution and speed consistency; S600, when the target behavior stability factor R is lower than the set threshold, the local adaptive re-patrol strategy of the unmanned aerial vehicle is started, the reconstruction patrol path is calculated according to the current image change trend, the aircraft is controlled to fly back and reacquire the image of the target area. 2.The method of claim 1, wherein: The S100 comprises: S101, acquire the image frame sequence I and the attitude information P; S102, combine the attitude estimation matrix and the global positioning coordinates, introduce the position label coding to the image frame I, and carry out the trajectory sequence sorting, construct the image attitude time sequence structure body G = {I, P, L, t} which is completely matched with the flight state, wherein L is the geographical label, and t is the acquisition time stamp. 3.The method of claim 1, wherein: The S200 comprises: S201, based on the image attitude time sequence structure body G = {I, P, L, t}, introduce the inertial reference frame as a unified reference coordinate system, carry out coordinate conversion on the three-axis acceleration and angular velocity data in the attitude information P, so as to form the attitude input set in the inertial coordinate domain; S202, establish an attitude reconstruction model for estimating the heading angle, pitch angle and roll angle corresponding to the image frame in the time domain, wherein the model solves the image attitude state quantity based on the improved quaternion filter and the image time sequence continuity; S203, input the reconstructed attitude state quantity as a transformation parameter into a nonlinear geometric distortion correction function, execute dynamic geometric correction on the image frame sequence I frame by frame, and carry out high-precision synchronization on the image sequence in combination with the image acquisition time stamp, output the attitude stable image frame sequence I'. 4.The method of claim 1, wherein: The S300 comprises: S301, divide the attitude stable image frame sequence I' into local regions, calculate the image entropy change rate of each region in the image frame through the sliding window method, and construct an entropy gradient distribution matrix; S302, according to the entropy gradient distribution matrix, based on the image enhancement mechanism of regional priority control, dynamically adjust the brightness compensation coefficient, local contrast stretching parameter and edge sharpening kernel size of the corresponding region, realize the targeted image enhancement processing; S303, after the enhancement processing, a multi-resolution image pyramid sequence I" is constructed by a multi-scale decomposition algorithm, which generates image copies of different scales centered on the high-entropy block; S304, the image enhancement and the local parameter change in the pyramid construction process are fed back to the enhancement mechanism with the entropy weight, forming a closed-loop adjustment structure based on the dynamic response of image entropy.
5. The method of claim 1, wherein: The S400 includes: S401, based on the generated multi-resolution image pyramid sequence I", a multi-scale feature map set F is constructed; S402, based on the historical data of the flight path of the unmanned aerial vehicle, a track prediction model T is constructed, the long short-term memory network is used to extract the position transformation trend, and a predicted trajectory vector sequence T' is output, which is used as a time sequence dynamic prior input; S403, a deep cross-attention neural network M is built, F and T' are sent into the spatial branch and the time branch respectively, and a multi-layer cross-attention mechanism is used to realize the fusion reasoning of the image spatial features and the track dynamic prior, so as to obtain a target position saliency response map; S404, according to the saliency response map, a target recognition result O is extracted, and a time sequence position distribution vector V of the target recognition result in the image sequence is generated, and a fusion confidence score is added to each recognition result. 6.The method of claim 1, wherein: The S500 includes: S501. Associate the target category, image coordinates, and confidence information in the recognition result O with the corresponding temporal position information in the trajectory vector V to generate a fusion structure alignment set. , where t is the collection timestamp sequence; S502, based on the structure alignment set D, an behavior evaluation module containing a target continuity analysis submodule, a confidence fluctuation discrimination submodule and a speed consistency evaluation submodule is constructed; S503, the target continuity analysis submodule calculates a continuity score Rc according to the spatial moving distance of the target in the image frame sequence and the inter-frame matching success rate; the confidence fluctuation discrimination submodule calculates a stability score Rs based on the variance of the confidence sequence change; the speed consistency evaluation submodule compares the target speed vector and the flight speed vector, and outputs a dynamic matching score Rv; S504, a multi-factor fusion strategy is used to weight and combine Rc, Rs and Rv into a target behavior stability factor R, which is used to measure the overall credibility and dynamic consistency of the identified target.
7. The multi-copter drone based image recognition method of claim 1, wherein: The S600 includes: S601, when the target behavior stability factor R is lower than a set threshold Rth, a re-inspection strategy is started, the image change analysis module is called to perform time sequence analysis on the recent image frame sequence I", and an image texture change vector ΔG and an edge offset trajectory ΔE are extracted; S602, ΔG and ΔE are input into the backtracking path prediction model, which fuses the image difference heat map and the current heading angle of the unmanned aerial vehicle to generate a local reconstructed inspection path vector P', and optimizes the turning angle and flight speed between path nodes; S603, the path vector P' is sent to the multi-rotor navigation control unit through the flight control interface module, the aircraft is directed to fly back to the optimal shooting point of the suspected target area, and the image acquisition module is triggered to reacquire the image frames of the target area; S604, the reacquired image frames are automatically fed back to the image acquisition and attitude synchronization process of step S100, realizing a closed-loop flight control recognition cycle based on dynamic image judgment.
Citation Information
Cited By
Millimeter-level high-precision autonomous landing docking method and system for unmanned aerial vehicle logistics docking
CN122219596A