Unmanned aerial vehicle visual perception processing method, system and device based on motion event optical flow and medium

By employing a UAV visual perception method based on motion event optical flow, and utilizing optical flow confidence assessment and search area adjustment, the problems of target drift and scale inaccuracy in UAV visual tracking are solved, achieving stable target localization in complex scenarios.

CN121937490APending Publication Date: 2026-04-28XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-01-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing UAV visual tracking methods are prone to target drift or loss in high-speed motion and complex backgrounds, and lack effective evaluation and utilization of motion information, resulting in inaccurate updates of the search area's location or scale.

Method used

The UAV visual perception method based on motion event optical flow calculates the average optical flow and divergence of the search area through optical flow credibility assessment and constraints, realizes dynamic adjustment of the search area, and constructs a spatial attention score map in the adjusted area to jointly represent the spatial position of the target.

Benefits of technology

It improves the adaptability of UAV visual perception methods in complex motion scenarios, enhances the stability and accuracy of target localization, and is suitable for tracking tasks of high-speed moving and small-scale targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937490A_ABST
    Figure CN121937490A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle visual perception processing method, system and device based on a motion event optical flow, and a medium. The method comprises the following steps: extracting the optical flow of a motion event target; performing quality evaluation on the optical flow of the motion event target to obtain optical flow credibility; based on the optical flow of the motion event target and the optical flow credibility, calculating the average optical flow and divergence in the previous frame of target frame to obtain an adjusted search area of the current frame; calculating a score chart of the search area based on the adjusted search area of the current frame to obtain a final position of the target; the system, the equipment and the medium are used for implementing the method. The method has the advantages of high dynamic modeling capability, high universality, high estimation precision and the like, the tracking performance of the small-target unmanned aerial vehicle in a complex scene is remarkably improved, and a key technical support is provided for an intelligent visual system of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent visual perception technology, and specifically relates to a method, system, device and medium for processing visual perception of unmanned aerial vehicles based on optical flow of motion events. Background Technology

[0002] In UAV target tracking tasks, accurately acquiring continuous target motion information is crucial for dealing with high-speed maneuvers, small-scale targets, and complex background interference. Motion cues not only supplement appearance features but also provide dynamic priors such as target orientation and scale changes, which are important factors in improving the stable tracking capability of UAVs in aerial scenarios. Current Transformer-based single-target tracking methods, relying on powerful self-attention mechanisms, effectively enhance the feature association ability between the template and the search area, becoming the mainstream paradigm in the current tracking field. However, these methods mainly rely on appearance information for matching, and their modeling of continuous target motion remains relatively weak. Therefore, in scenarios with high-speed movement and scale changes, target drift or loss is prone to occur.

[0003] Patent application CN116309683B discloses a motion-information-assisted visual single-target tracking method. It predicts the target's future position by combining historical motion cues with a convolutional long short-term memory network, thereby improving robustness under occlusion and lighting changes. Patent application CN120259368B discloses a moving target tracking method, device, electronic device, and storage medium. It embeds historical position information into a feature sequence and inputs it along with current visual features into a motion perception fusion module, thereby improving prediction accuracy under complex motion conditions. Patent application CN106296743A discloses an adaptive moving target tracking method and a UAV tracking system. It predicts the target's position in the next frame by using the centroid displacement direction and tracks the target by keeping it at the center of the search window; its essence is also the utilization of historical motion trends.

[0004] While the aforementioned methods have validated the effectiveness of introducing motion priors in improving tracking performance to some extent, most still rely on frame-level positional information between adjacent appearance frames to estimate target motion. This results in a relatively coarse utilization of motion information and a lack of effective assessment of its reliability. In high-speed UAV scenarios, target displacement between adjacent frames is often large and changes rapidly. Motion modeling based solely on frame-level positional changes struggles to reflect the target's true motion trend in a timely manner. Furthermore, background interference, noise, or local occlusion can easily introduce unstable or misleading motion cues. Without constraints on the quality of motion information, inaccurate updates to the search area's position or scale can easily occur, making the target localization process susceptible to interference from sudden changes in direction, velocity variations, and complex background motion. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention aims to provide a UAV visual perception processing method, system, device, and medium based on optical flow of moving events. The present invention uses the optical flow information of moving event targets as its core, introducing optical flow credibility to quantify and constrain the optical flow of moving events. This allows only motion information that meets preset conditions to participate in subsequent search area adjustment and target positioning calculations, thereby solving the problem of low-quality or unstable optical flow indiscriminately participating in decision-making and easily interfering with target positioning. Based on the motion information constrained by optical flow credibility, the present invention calculates the average optical flow and divergence of the search area, jointly estimating the center position of the search area and the target size, realizing dynamic adjustment of the search area's position and size. This solves the problem in existing UAV visual perception methods where the search area relies on a fixed scale or static prediction and is difficult to adapt to changes in the target's motion state. Furthermore, within the adjusted search area, this invention constructs a spatial attention score map based on motion optical flow to jointly represent and score the target's spatial position, thereby solving the problems of the single spatial response form and the difficulty in effectively distinguishing the target from complex background motion in the existing target localization process. Through the above technical solution, this invention achieves the coordinated utilization of motion information, spatial position information, and scale information under a unified processing framework, improving the adaptability of UAV visual perception methods in complex motion scenarios.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for visual perception processing of unmanned aerial vehicles (UAVs) based on optical flow of motion events includes the following steps: Step 1: Extract the optical flow of the moving event target; Step 2: Evaluate the quality of the optical flow of the moving event target extracted in Step 1 to obtain the optical flow confidence level; Step 3: Based on the optical flow of the moving event target obtained in Step 1 and the optical flow confidence obtained in Step 2, calculate the average optical flow and divergence within the target box region of the previous frame to obtain the adjusted search region for the current frame; Step 4: Based on the adjusted search area of ​​the current frame obtained in Step 3, calculate the score map of the search area to obtain the final position of the target.

[0007] The specific methods for step one include: Step 1.1: Input the UAV video sequence into the pulsed retinal network to obtain motion event data at multiple time steps between frames; Step 1.2: Input the motion event data and UAV video sequence obtained in Step 1.1 into the temporal optical flow network that fuses event and image information to obtain the optical flow of the motion event target.

[0008] The specific method for step two includes: Step 2.1: Let the target bounding box of the previous frame be: , Within the target bounding box of the previous frame, convert the optical flow of the moving event target obtained in step 1.2 into angles: in, Represents the pixels within the target bounding box in the previous frame. exist and Velocity in the direction; Step 2.2: Calculate the angle transformed in Step 2.1 using circumferential statistical methods. direction and Average component in direction and ; Where N represents the number of pixels; Step 2.3, based on the angle obtained in step 2.2 direction and Average component in direction and Calculate the directional consistency score to obtain the optical flow confidence level. : .

[0009] The specific methods for step three include: Step 3.1: Calculate the optical flow confidence level obtained in Step 2.3. Set threshold Comparison is performed when the optical flow confidence level is less than a threshold. When the target box center position and bounding box size remain unchanged from the previous frame, proceed to step 4.1; when the optical flow confidence level is greater than or equal to the threshold At that time, using the motion event target optical flow extracted in step one, the average optical flow is calculated within the target bounding box region of the previous frame. The center position of the target in the current frame is predicted. ; in, The x-coordinate of the target center position in the current frame is the predicted x-coordinate. The ordinate of the predicted target center position in the current frame. The target bounding box of the previous frame The size of the region; Calculate the target bounding box of the previous frame Internal optical flow field The divergence; When the optical flow field divergence When F>0, it indicates that the local optical flow is spreading outward, and the target is getting larger or closer to the camera; when the optical flow field... divergence When F < 0, it indicates that the local optical flow converges inward, and the target is shrinking or moving away from the camera; when the optical flow field... divergence When F=0, it indicates that the local motion remains consistent, the target is translated, and there is no scale change; Step 3.2: Calculate the divergence obtained in Step 3.1. F is transformed into a scale factor that can be directly used for target box adjustment through the following linear mapping model. : in, It is the average divergence value within the target bounding box region of the previous frame; This is the scaling adjustment factor; It serves as a scale constraint boundary, used to suppress scale anomalies caused by noise; Step 3.3: Utilize the scaling factor obtained in Step 3.2 The updated target width and height are calculated as follows: Step 3.4: Combine with step 3.1 to obtain the x-coordinate of the target's center position. and ordinate The target width obtained in step 3.3 and height Finally, the adjusted target box position is obtained; .

[0010] Step 3.5, using the target box from Step 3.4 Centered on the target bounding box, the width and height are expanded to obtain the adjusted search area for target localization.

[0011] The specific methods for step four include: Step 4.1: Based on the adjusted search area obtained in Step 3, calculate the optical flow amplitude within the search area. Defined as: in, Indicates the location of the target box optical flow within; Step 4.2: Normalize the optical flow amplitude obtained in Step 4.1 to obtain a dimensionless relative optical flow amplitude. ; Step 4.3: Adjust the dimensionless relative optical flow amplitude obtained in Step 4.2 using an adaptive intensity adjustment method. Adjustments were made to match the optical flow attention with the tracking score map based on the visual image from the original tracking algorithm, resulting in a spatial attention score map. : (p) (γ ), in, γ is the mean of the tracking score map based on the visual image in the original tracking algorithm, and γ is the weight coefficient. Step 4.4: Spatial attention score map obtained in Step 4.3 Tracking score map of the original tracking algorithm based on visual images Determine the final score chart : = + in, The point with the highest score is the final center position of the target. ; Step 4.5: Determine the final center position of the target obtained in Step 4.4. The target width obtained in step 3.3 and height By combining these, we can obtain the final location of the target: .

[0012] The present invention also provides a UAV visual perception processing system based on motion event optical flow, comprising: The motion feature extraction module is used to extract the optical flow of moving event targets. The optical flow confidence calculation module is used to evaluate the quality of the optical flow of the moving event target extracted by the motion feature extraction module and obtain the optical flow confidence. The search region adjustment module is used to calculate the optical flow of the moving event target obtained by the optical flow confidence calculation module and the optical flow confidence calculated by the optical flow confidence calculation module, calculate the average optical flow and divergence within the target box area of ​​the previous frame, and obtain the adjusted search region of the current frame. The target position output module is used to calculate the score map of the search area after adjustment based on the current frame obtained by the search area adjustment module, and obtain the final position of the target.

[0013] The present invention also provides a UAV visual perception processing device based on motion event optical flow, comprising: Memory: A computer program that stores the above-mentioned UAV visual perception processing method based on motion event optical flow, and is a computer-readable device; Processor: Used to implement the UAV visual perception processing method based on motion event optical flow when executing the computer program.

[0014] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the aforementioned UAV visual perception processing method based on motion event optical flow.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a unified UAV visual perception processing method based on optical flow of moving events. This method uses the optical flow information of moving event targets as its core, evaluates and constrains motion information by calculating optical flow confidence, and then calculates the average optical flow and divergence within the target bounding box of the previous frame based on motion information that meets preset conditions to determine the center of the search region and the target size in the current frame. Furthermore, a spatial attention score map is calculated within the adjusted search region to jointly represent and score the target's spatial position, thereby achieving stable perception and localization of UAV targets. The overall advantages are as follows: (1) This invention constructs a unified UAV visual perception processing framework based on the optical flow of moving events, organically combining the optical flow of moving event targets, optical flow quality assessment, and their utilization in search area adjustment and target localization, realizing the coordinated processing of motion information, spatial location information, and scale information. This framework avoids the fragmented use of motion cues between different processing stages, enabling the visual perception process to perform overall modeling around the motion state of the target, and is suitable for target perception tasks in UAV scenarios.

[0016] (2) This invention quantifies and evaluates the optical flow quality of moving event targets by calculating optical flow confidence, thereby effectively constraining motion information, i.e., the optical flow of moving event targets, before it participates in subsequent search area adjustment and target localization calculations, thus reducing the impact of low-quality or unstable optical flow on the visual perception process. This method avoids the indiscriminate use of all motion information (optical flow of moving event targets), improving the rationality and stability of motion information participating in the decision-making process.

[0017] (3) Based on the motion information (optical flow of the moving event target) after optical flow credibility assessment, the present invention calculates the average optical flow and divergence of the search area, realizes the joint estimation of the center position of the search area and the target size, overcomes the problem that the search area depends on fixed scale or static prediction in the existing UAV visual perception method, and enables the search area to change dynamically with the target motion state, thereby improving the adaptability to complex motion scenes.

[0018] (4) The method proposed in this invention has good versatility and can be used as an independent motion prior module in various existing target tracking algorithms without retraining the existing tracking model. Under the premise of keeping the original model structure and parameters unchanged, this invention can introduce motion prior based on motion event optical flow into the position and scale adjustment process of the target prediction box, thus effectively supplementing the existing UAV target tracking methods. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the UAV visual perception processing method based on motion event optical flow according to the present invention.

[0020] Figure 2 This is a schematic diagram of the pulsed retinal network structure of the present invention.

[0021] Figure 3 This is a schematic diagram of the temporal optical flow estimation network structure that integrates event and image information according to the present invention.

[0022] Figure 4(a) is a pair of infrared UAV images against the background of centralized airspace in the ANTI-UAV2021 dataset used in the embodiment of the present invention.

[0023] Figure 4(b) shows the motion event optical flow results (left) and color wheel diagram (right) of infrared UAV images under the centralized airspace background of the ANTI-UAV2021 dataset used in the embodiment of the present invention. Different colors in the color wheel diagram represent different directions and speeds.

[0024] Figure 4(c) shows a pair of UAV infrared images against a complex forest background in the ANTI-UAV2021 dataset used in this embodiment of the invention.

[0025] Figure 4(d) shows the motion event optical flow results (left) and color wheel diagram (right) of the UAV infrared image under a complex forest background in the ANTI-UAV2021 dataset used in the embodiment of the present invention. Different colors in the color wheel diagram represent different directions and speeds.

[0026] Figure 5 These are image pairs in the ANTI-UAV2021 dataset where the target size changes, as used in this embodiment of the invention. Detailed Implementation

[0027] The technical solution adopted by the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0028] like Figure 1 As shown, a UAV visual perception processing method based on motion event optical flow includes the following steps: Step 1: Extract the optical flow of the moving event target; The specific method of step one includes: Step 1.1: Input the UAV video sequence into the pulsed retinal network to obtain motion event data at multiple time steps between frames; The pulsed retinal network comprises a frequency coding module, a photoreceptor layer, a bipolar cell layer, and a ganglion cell layer. The following section will combine these components... Figure 2 The network structure and its corresponding workflow in step 1.1 are described in detail.

[0029] The frequency encoding module is used to process the input image data of the previous frame. Input image data of the next frame Frequency coding is performed using T analog time steps, where To obtain the frequency-encoded data of the previous frame With the next frame of data ,in ; The photoreceptor layer includes a convolutional layer (Conv), a regularization layer (BN), and a LIF neuron model. The photoreceptor layer first encodes the frequency-encoded data from the previous frame. With the next frame of data The input is fed into a convolutional layer (Conv) for spatiotemporal feature extraction, followed by batch normalization through a regularization layer (BN), and finally, the spatiotemporal dynamic information of the impulse response at the previous time step is generated by a LIF neuron model. Ph t-1 and the spatiotemporal dynamic information of the impulse response at the next moment Ph t : Ph t-1 = LIF(BN(Conv(It-1 ))), Ph t-1 Ph t = LIF(BN(Conv(I t ))), Ph t The bipolar cell layer comprises two types of cells: ON-type and OFF-type. Both types of cells are constructed using a LIF neuron model. The ON-type cells are used to process the spatiotemporal dynamic information of the impulse response generated by the photoreceptor layer at the next moment. Ph t Spatiotemporal dynamic information of the impulse response at the previous moment Ph t-1 After subtraction, input into LIF In the neuron model, regions with increased illumination are filtered out; OFF-type cells are used to process the spatiotemporal dynamics of the impulse response generated by the photoreceptor layer in the previous moment. Ph t-1 Spatiotemporal dynamic information of the impulse response at the next moment Ph t After subtraction, input into LIF In the neuron model, regions with reduced illumination are filtered out; ON =LIF( Ph t+1 Ph t ) OFF = LIF( Ph t Ph t+1 ) The ganglion cell layer includes a linear layer ( Linear ), regularized layer (BN) and LIF neuron model; among which, linear layer ( Linear Linear transformations were performed on the output data of ON-type cells and OFF-type cells, respectively. The linearly transformed data were then input into a regularization layer (BN) for batch normalization. Finally, the normalized data was input into the LIF neuron model to obtain the final output data. ON out and OFF out motion event data; ON out =LIF(BN(Linear(ON))) OFF out =LIF(BN(Linear(OFF))) .

[0030] Step 1.2: Input the motion event data and UAV video sequence obtained in Step 1.1 into the temporal optical flow network that fuses event and image information to obtain the optical flow of the moving event target; like Figure 3 As shown, the temporal optical flow network that integrates event and image information includes a temporal feature extraction module and a dual-branch collaborative iterative optical flow module for events and images. The network structure and its corresponding workflow in step 1.2 are described in detail.

[0031] The temporal feature extraction module includes an event branch and an image branch. The event branch includes an event data processing module, a first convolutional feature extraction module, and a first feature encoding module; the image branch includes a second convolutional feature extraction module and a second feature encoding module.

[0032] The event data processing module accumulates the positive and negative polarity event data corresponding to each image frame, obtaining multi-timestep dual-channel event data of size (T×2×H×W), where T represents the number of time steps, 2 represents the ON / OFF polarity channels, and H×W is the image resolution. Then, the first convolutional feature extraction module extracts preliminary features from the multi-timestep dual-channel event data output by the event data processing module through its convolutional layer, obtaining multi-timestep dual-channel event data features. Next, the Patch Embedding module divides the multi-timestep dual-channel event data features into multiple fixed-size non-overlapping event feature blocks. These blocks are then input into the first feature encoding module to obtain the event context features between T-1 time steps. Event space characteristics Quantities related to event motion ,in ; The image branch includes a second convolutional feature extraction module and a second feature encoding module. The convolutional layers of the second convolutional feature extraction module perform preliminary feature extraction on the input image data to obtain image data features. Subsequently, the image data features are divided into multiple fixed-size non-overlapping image feature blocks by the Patch Embedding module. These blocks are then input into the second feature encoding module to obtain the image motion correlation between T-1 time steps. Image spatial features Image context features ,in .

[0033] The dual-branch collaborative iterative optical flow module for events and images includes: an event feature extraction network (ME) and an image feature extraction network (MF), a first spatiotemporal attention module, a second spatiotemporal attention module, a first motion update module, and a second motion update module. In the zeroth iteration, the initial optical flow for both the event and image branches is defined as a zero vector. =0, =0; During the first iteration update, the optical flow is initialized with events. The event motion correlation quantity obtained from the first feature encoding module in step one The input is fed into the event feature extraction network (ME) to obtain the event motion features at k time steps updated in the first iteration. ; Initialize optical flow for the image Image motion correlation coefficient obtained by the second feature encoding module The input is fed into the image feature extraction network MF to obtain the image motion features at k time steps updated in the first iteration. ; The first spatiotemporal attention module includes a lightweight first-time learning layer. The first spatial cross-attention module A1 and the first connection operation C1; the second spatiotemporal attention module includes a lightweight second temporal learning layer. The second spatial cross attention module A2 and the second connection operation C2; The event motion features at the k time steps updated in the first iteration are obtained from the event feature extraction network (ME). Input into the first time-of-flight learning layer of the first spatiotemporal attention module In the first iteration, temporal information is extracted from the motion features to obtain the temporally embedded event features extracted from the event motion features of all time steps. ; The image motion features obtained from the first iteration of the image feature extraction network (MF) are updated at k time steps. Input into the second temporal learning layer of the second spatiotemporal attention module In the first iteration, temporal information is extracted from the motion features to obtain the temporally embedded image features extracted from the image motion features at all time steps. ; The event space features obtained in the first feature encoding module With the characteristics of event motion The input is fed into the first spatial cross-attention module A1 to achieve cross-domain spatial alignment and complementary perception, resulting in the first... In the next iteration update Spatial cross-attention event characteristics at each time step ; The image spatial features obtained in the second feature encoding module With image motion features The input is fed into the second spatial cross-attention module A2 to achieve cross-domain spatial alignment and complementary perception, resulting in the first... In the next iteration update Spatial cross-attention image features at each time step ; Embedding time into event features Features of Spatial Intersection Attention Events Perform the first connection operation C1 to enhance the model's expressive power and contextual understanding, and obtain the spatiotemporal features of events in the first iteration. ; Embedding time into image features Spatial Cross-Attention Image Features A second connection operation C2 is performed to enhance the model's expressive power and contextual understanding, yielding the spatiotemporal features of the image in the first iteration. ; The event context features obtained by the first feature extraction module Event motion features obtained by the Event Feature Extraction Network (ME) And the event spatiotemporal features obtained by the first spatiotemporal attention module The data is further input into the first motion update module to obtain the optical flow residual values ​​between multiple time steps of the events during the first iteration. ; Image context features obtained by the second feature extraction module Image motion features obtained by the image feature extraction network MF And the image spatiotemporal features obtained by the second spatiotemporal attention module The data is further input into the second motion update module to obtain the optical flow residual values ​​of the image across multiple time steps during the first iteration. ; The optical flow residual values ​​between multiple time steps of the event during the first iteration. With initial optical flow Add them together to get the first one. Event Optical Flow after the next iteration update The optical flow residual values ​​of the image at multiple time steps during the first iteration are obtained. With initial optical flow Add them together to get the first one. Image optical flow updated in the next iteration The above process is iterated and updated continuously to obtain event optical flow and image optical flow at multiple time steps. Then, the summation and averaging operations are performed to obtain the final optical flow of the moving event target.

[0034] Step 2: Evaluate the quality of the optical flow of the moving event target extracted in Step 1 to obtain the optical flow confidence level; The specific method for step two includes: Step 2.1: Let the target bounding box of the previous frame be: , In the target bounding box of the previous frame, the optical flow of the moving event target obtained in step 1.2 is converted into angle; in, Represents the pixels within the target bounding box in the previous frame. exist and Velocity in the direction; Step 2.2: Calculate the angle transformed in Step 2.1 using circumferential statistical methods. direction and Average component in direction and ; Where N represents the number of pixels; Step 2.3, based on the angle obtained in step 2.2 direction and Average component in direction and Calculate the directional consistency score to obtain the optical flow confidence level. : .

[0035] Step 3: Based on the optical flow of the moving event target obtained in Step 1 and the optical flow confidence obtained in Step 2, calculate the average optical flow and divergence within the target box region of the previous frame to obtain the adjusted search region for the current frame; The specific methods for step three include: Step 3.1: Calculate the optical flow confidence level obtained in Step 2.3. Set threshold Comparison is performed when the optical flow confidence level is less than a threshold. When the target box center position and bounding box size remain unchanged from the previous frame, proceed to step 4.1; when the optical flow confidence level is greater than or equal to the threshold At that time, using the motion event target optical flow extracted in step one, the average optical flow is calculated within the target bounding box region of the previous frame. The center position of the target in the current frame is predicted. ; in, The x-coordinate of the target center position in the current frame is the predicted x-coordinate. The ordinate of the predicted target center position in the current frame. The target bounding box of the previous frame The size of the region; Calculate the target bounding box of the previous frame Internal optical flow field divergence: When the optical flow field divergence When F>0, it indicates that the local optical flow is spreading outward, and the target is getting larger or closer to the camera; when the optical flow field... divergence When F < 0, it indicates that the local optical flow converges inward, and the target is shrinking or moving away from the camera; when the optical flow field... divergence When F=0, it indicates that the local motion remains consistent, the target is translated, and there is no scale change; Step 3.2: Calculate the divergence obtained in Step 3.1. F is transformed into a scale factor that can be directly used for target box adjustment through the following linear mapping model. : in, It is the average divergence value within the target bounding box region of the previous frame; This is the scaling adjustment factor; It serves as a scale constraint boundary, used to suppress scale anomalies caused by noise; Step 3.3: Utilize the scaling factor obtained in Step 3.2 The updated target width and height are calculated as follows: Step 3.4: Combine with step 3.1 to obtain the x-coordinate of the target's center position. and ordinate The target width obtained in step 3.3 and height Finally, the adjusted target box position is obtained; .

[0036] Step 3.5, using the target box from Step 3.4 Centered on the target bounding box, the width and height are expanded by a factor of 1.5 to obtain the adjusted search area for subsequent target localization.

[0037] Step 4: Based on the adjusted search area of ​​the current frame obtained in Step 3, calculate the score map of the search area to obtain the final position of the target.

[0038] The specific methods for step four include: Step 4.1: Based on the adjusted search area obtained in Step 3, calculate the optical flow amplitude within the search area. Defined as: in, Indicates the location of the target box optical flow within; Step 4.2: Normalize the optical flow amplitude obtained in Step 4.1 to obtain a dimensionless relative optical flow amplitude. ; Step 4.3: Adjust the dimensionless relative optical flow amplitude obtained in Step 4.2 using an adaptive intensity adjustment method. Adjustments were made to match the optical flow attention with the tracking score map based on the visual image from the original tracking algorithm, resulting in a spatial attention score map. : (p) (γ ), in, γ is the mean of the tracking score map based on the visual image in the original tracking algorithm, and γ is the weight coefficient. Step 4.4: Spatial attention score map obtained in Step 4.3 Tracking score map of the original tracking algorithm based on visual images Determine the final score chart : = + in, The point with the highest score is the final center position of the target. ; Step 4.5: Determine the final center position of the target obtained in Step 4.4. The target width obtained in step 3.3 and height By combining these, we can obtain the final location of the target: .

[0039] In UAV target tracking missions, accurately acquiring continuous target motion information is crucial for dealing with high-speed maneuvers, small-scale targets, and complex background interference. Motion cues not only supplement appearance features but also provide dynamic priors such as changes in target orientation and scale, which are important factors in improving the stable tracking capability of UAVs in aerial scenarios. To address this, this invention proposes a UAV visual perception processing method based on motion event optical flow, the overall process of which is as follows: Figure 1 As shown, this invention first inputs a UAV video sequence into a pulsed retinal network to obtain motion event data at multiple time steps between frames. Since this type of data only characterizes whether the target is moving, it is binary data (pixels with motion have a value of 1, and pixels without motion have a value of 0), lacking texture information and unable to reflect the target's direction and speed. Therefore, this motion event data and the UAV video sequence are further input into a temporal optical flow network that fuses event and image information to obtain the optical flow of the moving event target. In this optical flow method, the image branch can enhance motion perception capabilities with the help of dynamic information provided by the event branch, while the event branch can obtain static structural information from the image branch to compensate for insufficient perception of events in sparse regions.

[0040] Optical flow estimation quality is affected by factors such as occlusion and illumination variations. Directly using unfiltered optical flow can lead to tracking offset and degrade tracker performance. In the target motion field, the direction of motion should be consistent. If the direction distribution diverges, it indicates optical flow failure or background interference. Therefore, this invention determines the reliability of optical flow by calculating the direction consistency score. When the optical flow reliability is... Exceeding the threshold In this case, optical flow can be used to predict the target location.

[0041] When the optical flow confidence level is greater than or equal to the threshold At that time, based on the extracted motion event target optical flow, the average optical flow is calculated within the target bounding box of the previous frame. The center position of the target in the current frame is predicted. However, simply adjusting the center position is insufficient to handle scale changes common to targets in real-world scenarios, such as targets moving closer to / away from the camera or changes in viewpoint due to accelerated motion. Therefore, it is necessary to further extract structural information reflecting scale changes from the optical flow field to achieve adaptive adjustment of the target search area. In terms of motion geometry, divergence describes the difference between "outflow" and "inflow" within a region and is an indicator of whether the local motion field exhibits a spreading or contracting trend. To determine target size changes, the divergence of the optical flow motion field within the target bounding box in the previous frame is calculated.

[0042] The intensity of motion varies across different regions within the target. Regions of dominant motion are often highly correlated with the main target, while regions of weak motion may originate from the background or edges. Constructing optical flow amplitude into a spatial attention map can enhance the network's ability to respond to regions of significant motion.

[0043] This invention proposes a UAV visual perception processing method based on motion event optical flow, which effectively fuses optical flow and visual information at different locations. Specifically, it ensures that only reliable optical flow is used by calculating confidence; scale adaptation is achieved through divergence calculation based on physical constraints; and motion amplitude attention enhances focus on the motion region at the feature level. During experiments, results can be directly optimized based on existing tracking models and weights without additional training, making it a plug-and-play optical flow guidance module.

[0044] Experimental Analysis 1. Experimental conditions The experiments were conducted in a Linux-based computing environment. The algorithm was implemented in Python and run on a computing platform equipped with a graphics processing unit. The publicly available UAV target dataset ANTI-UAV2021 was used to verify the applicability of the method in UAV scenarios. ANTI-UAV2021 is a long-range infrared UAV dataset. Its test set contains 140 video clips, including complex backgrounds such as clouds, buildings, and trees, and various types and sizes of UAVs. The dataset covers multi-scale target variations and motion blur conditions.

[0045] To further verify the algorithm's general performance, this invention tested the algorithm on the FE108 dataset. The FE108 dataset consists of 108 video sequences captured by a DAVIS346 event camera. The test set contains 32 video sequences. This dataset contains 21 different types of objects, which can be divided into three categories: animals, vehicles, and everyday objects (such as bottles and boxes). FE108 includes four challenging scenes: low light, high dynamic range, and fast motion with and without motion blur.

[0046] 2. Experiment Content (1) Verification of the optical flow calculation effect of motion events The experiment first verifies the motion event-based optical flow calculation method proposed in this invention. By processing UAV infrared images under different scenarios in the ANTI-UAV2021 dataset, the corresponding motion event optical flow results are obtained. The changes in the motion direction and velocity of the UAV in consecutive frames are compared, and the response distribution of the moving target in the optical flow field is observed to verify whether the calculated optical flow can reflect the motion direction and magnitude characteristics of the UAV target.

[0047] (2) Verification of the effect of dynamic adjustment of search area Based on the optical flow calculation for motion events, the dynamic adjustment effect of the search region based on motion information in this invention is further verified. By introducing motion information evaluated with optical flow reliability, the average optical flow and divergence of the search region are calculated, and the center position and scale of the search region are updated. During the experiment, the changes in the position and size of the search region of the target in consecutive frames are observed to verify the ability of this invention to adaptively adjust the search region under conditions of high-speed target motion and scale changes.

[0048] (3) Target localization performance verification Furthermore, the target localization process of the proposed method is validated in scenarios containing complex backgrounds and background motion interference. Within the adjusted search area, a spatial attention score map is constructed based on the optical flow of motion events to jointly represent and score the spatial location of the target. The stability of the proposed method in complex scenes is verified by testing on the ANTI-UAV2021 dataset and a subset of small target video frames after upsampling.

[0049] 3. Experimental Results Figure 4(a) shows a pair of infrared UAV images against a clear background in the ANTI-UAV2021 dataset, as used in an embodiment of the present invention. The left image of Figure 4(b) is the optical flow result of motion events obtained based on the infrared UAV image pair against a clear background in the ANTI-UAV2021 dataset shown in Figure 4(a). Different colors in the figure represent different directions and speeds, and the direction and speed information of different colors is shown in the color wheel diagram on the right side of Figure 4(b). Figure 4(c) shows a pair of UAV infrared images against a complex forest background in the ANTI-UAV2021 dataset. The left image of Figure 4(d) is the optical flow result of motion events obtained based on the infrared UAV image pair against a complex forest background in Figure 4(c). Different colors in the figure represent different directions and speeds, and the direction and speed information of different colors is shown in the color wheel diagram on the right side of Figure 4(d). As can be seen from Figures 4(b) and 4(d), under both clear background and complex background conditions, the present invention can extract optical flow information corresponding to the changes in the direction and speed of the UAV movement from infrared images, so that the moving UAV target forms a relatively concentrated response area in the optical flow field, indicating that the optical flow calculation method based on motion events can effectively characterize the motion characteristics of the moving UAV target.

[0050] Figure 5 This illustration shows image pairs in the ANTI-UAV2021 dataset where the target size has changed, selected according to an embodiment of the present invention. The red boxes represent the search regions. Figure 5 It can be observed that the size of the search area can be adjusted accordingly as the size of the target changes.

[0051] The UAV visual perception processing method based on motion event optical flow proposed in this invention was evaluated on the ANTI-UAV2021 dataset, with the OSTrack method selected as the base tracking method. During testing, the tracking performance of the entire dataset and small-scale targets (Pixel ≤ 20×20) was tested and compared with several mainstream tracking algorithms. The table lists detailed results for three metrics: Area Under the Curve (AUC), Regularized Precision (P-Norm), and Precision (P), used to comprehensively evaluate the effectiveness of the optical flow guidance strategy on targets at various scales. Specific results are shown in Table 1.

[0052] As shown in Table 1, on the overall test set, the proposed UAV visual perception processing method (referred to as FG-OSTrack in this invention) outperforms the basic method OSTrack in all metrics. This method achieves stable and consistent performance improvement through optical flow guidance without changing the backbone structure, demonstrating that the hierarchical fusion strategy proposed in this invention can effectively enhance the model's motion perception capability. The proposed method shows significant performance improvement on small targets (Pixel ≤ 20×20). Compared to the basic OSTrack method, FG-OSTrack improves AUC, P-Norm, and P by 1.2%, 1.5%, and 1.8%, respectively, in small target scenarios. This indicates that the local motion prior provided by optical flow can effectively compensate for the feature weakening problem caused by low resolution and blurred edges of small targets in the visual branch, enabling the tracking algorithm to obtain a more prominent and stable response, thereby improving the tracking robustness under small target conditions.

[0053] Overall, the FG-OSTrack invention achieves stable improvements across all metrics and demonstrates significant advantages in the most challenging small-target scenarios, validating the effectiveness and universality of the proposed UAV visual perception processing method.

[0054] Table 1 shows the comparison results of the proposed UAV visual perception processing method based on motion event optical flow with other advanced methods on the ANTI-UAV2021 dataset.

[0055] Table 2 compares the performance of the proposed motion event optical flow-based UAV visual perception processing method FG-OSTrack with the baseline method OSTrack on the FE108 dataset. It can be observed that when simultaneously adjusting the search region based on optical flow confidence and calculating the spatial attention score map based on the adjusted search region, the model achieves the highest improvement in AUC and P, indicating that the collaborative work of the proposed method helps enhance the accuracy of target localization. FG-OSTrack improves the AUC and P metrics by 3.4% and 6.2% respectively compared to the OSTrack baseline, demonstrating the positive effect of direction and velocity information on target tracking performance. The experimental results further verify the effectiveness and robustness of the proposed motion event optical flow-based UAV visual perception processing method FG-OSTrack in improving tracking performance in complex motion scenarios.

[0056] Table 2 compares the proposed UAV visual perception processing method based on motion event optical flow with the baseline method OSTrack on the FE108 dataset.

[0057] The search area is adjusted based on optical flow reliability. Based on adjusting the search region, a spatial attention score map is calculated. AUC P OSTrack 0.423 0.637 FG-OSTrack(Ours) √ 0.455 0.693 FG-OSTrack(Ours) √ √ 0.457 0.699 The present invention also provides a UAV visual perception processing system based on motion event optical flow, comprising: The motion feature extraction module is used to extract the optical flow of moving event targets in step one. The optical flow confidence calculation module is used to perform quality assessment on the optical flow of the moving event target extracted in step one in step two, and obtain the optical flow confidence. The search region adjustment module is used to implement the optical flow of the moving event target obtained in step one and the optical flow confidence obtained in step two in step three, calculate the average optical flow and divergence within the target box area of ​​the previous frame, and obtain the adjusted search region of the current frame. The target position output module is used to implement the search area adjusted based on the current frame obtained in step three in step four, calculate the score map of the search area, and obtain the final position of the target.

[0058] The present invention also provides a visual perception processing device for unmanned aerial vehicles, comprising: Memory: A computer program that stores the above-mentioned UAV visual perception processing method based on motion event optical flow, and is a computer-readable device; Processor: Used to implement the UAV visual perception processing method based on motion event optical flow when executing the computer program.

[0059] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the aforementioned UAV visual perception processing method based on motion event optical flow.

Claims

1. A method for visual perception processing of unmanned aerial vehicles based on optical flow of motion events, characterized in that, Includes the following steps: Step 1: Extract the optical flow of the moving event target; Step 2: Evaluate the quality of the optical flow of the moving event target extracted in Step 1 to obtain the optical flow confidence level; Step 3: Based on the optical flow of the moving event target obtained in Step 1 and the optical flow confidence obtained in Step 2, calculate the average optical flow and divergence within the target box region of the previous frame to obtain the adjusted search region for the current frame; Step 4: Based on the adjusted search area of ​​the current frame obtained in Step 3, calculate the score map of the search area to obtain the final position of the target.

2. The UAV visual perception processing method based on motion event optical flow according to claim 1, characterized in that, The specific methods for step one include: Step 1.1: Input the UAV video sequence into the pulsed retinal network to obtain motion event data at multiple time steps between frames; Step 1.2: Input the motion event data and UAV video sequence obtained in Step 1.1 into the temporal optical flow network that fuses event and image information to obtain the optical flow of the motion event target.

3. The UAV visual perception processing method based on motion event optical flow according to claim 1, characterized in that, The specific method for step two includes: Step 2.1: Let the target bounding box of the previous frame be: , Within the target bounding box of the previous frame, convert the optical flow of the moving event target obtained in step 1.2 into angles: in, Represents the pixels within the target bounding box in the previous frame. exist and Velocity in the direction; Step 2.2: Calculate the angle transformed in Step 2.1 using circumferential statistical methods. direction and Average component in direction and ; Where N represents the number of pixels; Step 2.3, based on the angle obtained in step 2.2 direction and Average component in direction and Calculate the directional consistency score to obtain the optical flow confidence level. : 。 4. The UAV visual perception processing method based on motion event optical flow according to claim 1, characterized in that, The specific methods for step three include: Step 3.1: Calculate the optical flow confidence level obtained in Step 2.

3. Set threshold Comparison is performed when the optical flow confidence level is less than a threshold. When the target box center position and bounding box size remain unchanged from the previous frame, proceed to step 4.1; when the optical flow confidence level is greater than or equal to the threshold At that time, using the motion event target optical flow extracted in step one, the average optical flow is calculated within the target bounding box region of the previous frame. The center position of the target in the current frame is predicted. ; in, The x-coordinate of the target center position in the current frame is the predicted x-coordinate. The ordinate of the predicted target center position in the current frame. The target bounding box of the previous frame The size of the region; Calculate the target bounding box of the previous frame Internal optical flow field The divergence; When the optical flow field divergence When F>0, it indicates that the local optical flow is spreading outward, and the target is getting larger or closer to the camera; when the optical flow field... divergence When F < 0, it indicates that the local optical flow converges inward, and the target is shrinking or moving away from the camera; when the optical flow field... divergence When F=0, it indicates that the local motion remains consistent, the target is translated, and there is no scale change; Step 3.2: Calculate the divergence obtained in Step 3.

1. F is transformed into a scale factor that can be directly used for target box adjustment through the following linear mapping model. : in, It is the average divergence value within the target bounding box region of the previous frame; This is the scaling adjustment factor; It serves as a scale constraint boundary, used to suppress scale anomalies caused by noise; Step 3.3: Utilize the scaling factor obtained in Step 3.2 The updated target width and height are calculated as follows: Step 3.4: Combine with step 3.1 to obtain the x-coordinate of the target's center position. and ordinate The target width obtained in step 3.3 and height Finally, the adjusted target box position is obtained; 。 Step 3.5, using the target box from Step 3.4 Centered on the target bounding box, the width and height are expanded to obtain the adjusted search area for target localization.

5. The UAV visual perception processing method based on motion event optical flow according to claim 1, characterized in that, The specific methods for step four include: Step 4.1: Based on the adjusted search area obtained in Step 3, calculate the optical flow amplitude within the search area. Defined as: in, Indicates the location of the target box optical flow within; Step 4.2: Normalize the optical flow amplitude obtained in Step 4.1 to obtain a dimensionless relative optical flow amplitude. ; Step 4.3: Adjust the dimensionless relative optical flow amplitude obtained in Step 4.2 using an adaptive intensity adjustment method. Adjustments were made to match the optical flow attention with the tracking score map based on the visual image from the original tracking algorithm, resulting in a spatial attention score map. : (p) (c ), in, γ is the mean of the tracking score map based on the visual image in the original tracking algorithm, and γ is the weight coefficient. Step 4.4: Spatial attention score map obtained in Step 4.3 Tracking score map of the original tracking algorithm based on visual images Determine the final score chart : = + in, The point with the highest score is the final center position of the target. ; Step 4.5: Determine the final center position of the target obtained in Step 4.

4. The target width obtained in step 3.3 and height By combining these, we can obtain the final location of the target: 。 6. A UAV visual perception processing system based on motion event optical flow according to the method of claim 1, characterized in that, include: The motion feature extraction module is used to extract the optical flow of moving event targets. The optical flow confidence calculation module is used to evaluate the quality of the optical flow of the moving event target extracted by the motion feature extraction module and obtain the optical flow confidence. The search region adjustment module is used to calculate the optical flow of the moving event target obtained by the optical flow confidence calculation module and the optical flow confidence calculated by the optical flow confidence calculation module, calculate the average optical flow and divergence within the target box area of ​​the previous frame, and obtain the adjusted search region of the current frame. The target position output module is used to calculate the score map of the search area after adjustment based on the current frame obtained by the search area adjustment module, and obtain the final position of the target.

7. A UAV visual perception processing device based on motion event optical flow, characterized in that, include: Memory: A computer program for a UAV visual perception processing method based on motion event optical flow as described in any one of claims 1-5, which is a computer-readable device; Processor: Used to implement the UAV visual perception processing method based on motion event optical flow as described in any one of claims 1-5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the UAV visual perception processing method based on motion event optical flow as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Adaptive motion object tracking method and unmanned aerial vehicle tracking system

    CN106296743A

  • A Visual Single Object Tracking Method Assisted by Motion Information

    CN116309683B

  • Moving target tracking method, device, electronic device and storage medium

    CN120259368B