Delay compensation control method and system for unmanned aerial vehicle dynamic target interaction

CN122592977APending Publication Date: 2026-08-18ZHEJIANG HYDROGEN SOURCE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610624320.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种面向无人机动态目标交互的延迟补偿控制方法和控制系统,用以解决视觉语言动作模型输出动作块序列时,由于感知、模型推理、通信传输和飞控执行之间存在全链路时延,导致动作点对应状态与实际执行状态不匹配,以及低频动作块难以直接适配高频飞控的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122592977A_ABST
    Figure CN122592977A_ABST
Patent Text Reader

Abstract

The application discloses a delay compensation control method and system for unmanned aerial vehicle dynamic target interaction. The method acquires multi-source observation data with acquisition time stamps and task instructions, receives a full-link delay formed by execution, and completes timing alignment; the observation data is input into a lightweight visual language action model, and an action block sequence carrying execution time stamps is output; for the expected execution time stamps of each action point in the block, the states of the unmanned aerial vehicle and the target at the corresponding execution time are respectively predicted; based on the predicted states, coordinate system transformation and delay error compensation are carried out to obtain a compensated action sequence; according to the frequency of flight control, a multi-multiple interpolation is carried out to generate a high-frequency continuous set value, and the set value is output to the flight control or a switching safety control signal after safety verification of flight constraints. The application can reduce the state deviation caused by the inconsistency between the perception time and the execution time, and improve the matching of the control instruction and the actual execution state in the dynamic target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) flight control technology, specifically to a delay compensation control method and system for dynamic target interaction of UAVs. Background Technology

[0002] In scenarios such as target tracking, dynamic alignment, inspection interaction, capture and delivery, and collaborative operations, UAVs typically need to simultaneously process visual environmental information, task commands, and their own status information, and convert the upper-level decision results into control setpoints that the flight control mechanism can execute. With the development of multimodal models and visual-language-action models, the upper-level control module can directly generate action sequences based on visual observation and linguistic task intent. However, such models are still limited by computing power, bandwidth, and flight control cycle when running on the airborne end.

[0003] Current UAV control processes mostly employ a sequential execution approach, involving perception, inference, trajectory planning, and flight control. There is a significant time delay between visual data acquisition, state estimation, model inference, action signal transmission, and flight control reception and execution. When the upper-level model outputs action commands, these commands typically correspond to the UAV and target states at the observation or inference time. However, by the time the commands are actually transmitted to the flight control unit and executed, the positions, attitudes, and velocities of the UAV and target have already changed, resulting in inconsistencies between the observation and execution coordinates.

[0004] For traditional trajectory planning or fixed controllers, some compensation can be achieved through conventional filtering, amplitude limiting, or simple prediction. However, when the visual language action model outputs a sequence of action blocks, the model output is usually a low-frequency discrete action point, and there are temporal and task semantic relationships between the action points. If only simple delay compensation is applied to the final control quantity, it is difficult to ensure that each action point matches the UAV state, target state, and coordinate system state at the actual execution time.

[0005] Furthermore, flight control mechanisms typically operate at frequencies ranging from 100Hz to 400Hz or even higher, while visual language action models are limited by inference overhead, resulting in action block output frequencies that are usually lower than the flight control execution frequency. Directly inputting low-frequency discrete action points into the flight control system can easily lead to abrupt changes in setpoints, attitude fluctuations, control overshoot, and increased target tracking errors. These problems can further impact flight safety, especially in environments with high-speed moving targets, strong winds, or dense obstacles.

[0006] Therefore, a UAV control method is needed that can address the characteristics of action block output from visual language action models, while simultaneously handling end-to-end latency, execution moment state prediction, coordinate system compensation, and high-frequency flight control adaptation, in order to improve the matching between action commands and actual execution states. Summary of the Invention

[0007] The purpose of this invention is to provide a delay compensation control method and control system for dynamic target interaction of UAVs, in order to solve the problems that when visual language action models output action block sequences, due to the end-to-end time delay between perception, model inference, communication transmission and flight control execution, the corresponding state of the action point does not match the actual execution state, and low-frequency action blocks are difficult to directly adapt to high-frequency flight control.

[0008] To achieve the above objectives, this invention provides a delay compensation control method for dynamic target interaction of unmanned aerial vehicles (UAVs). The method includes: acquiring raw observation state data with acquisition timestamps from multiple UAV sensors and task commands; estimating the end-to-end delay duration based on the raw observation state data, model input / output timestamps, action signal transmission / reception timestamps, and flight control execution timestamps, and performing time-series alignment of the raw observation state data based on a unified reference time to obtain observation data; inputting the observation data and task commands into a lightweight visual language action model, and outputting an action block sequence carrying execution timestamps; generating UAV predicted state data and target predicted state data corresponding to the expected execution timestamps of each action point in the action block sequence; performing coordinate system transformation and delay error compensation on the action points based on the predicted state data, and outputting a compensated action sequence signal; performing multi-rate interpolation on the compensated action sequence signal based on the ratio of the flight control operating frequency to the action point output frequency to generate a high-frequency continuous setpoint signal matching the flight control frequency; performing flight constraint safety verification on the high-frequency continuous setpoint signal, and outputting a normal control signal or a safety control signal based on the verification result.

[0009] The present invention also provides a control system for an unmanned aerial vehicle (UAV), comprising an acquisition module, a delay and alignment module, a model module, a deduction module, a compensation module, an interpolation module, and a safety output module. Each module is used to implement the corresponding steps in the above-described control method.

[0010] The present invention also provides a computer-readable storage medium and a computer program product, wherein the computer program stored therein, when executed by a processor, is capable of implementing the above-described delay compensation control method for dynamic target interaction of unmanned aerial vehicles.

[0011] Compared with existing technologies, the present invention has at least the following beneficial effects: First, by recording and estimating the timestamps of data acquisition, preprocessing, model inference, communication transmission, and flight control execution, a more accurate end-to-end latency duration can be obtained, thereby reducing the insufficient compensation caused by simply considering model inference latency or communication latency; Second, by binding the expected execution timestamp to each action point within the action block and predicting the states of the UAV and target at the corresponding execution time for each action point, a one-to-one correspondence is formed between the action point, the predicted state, and the execution time, reducing the spatiotemporal misalignment between the perception time and the execution time; Third, by performing coordinate system transformation and delay error compensation based on the predicted state, the model output in the observation coordinate system is converted into a control setpoint that is executable by the flight control and matches the state at the execution time; Fourth, by converting low-frequency action blocks into high-frequency continuous setpoints that match the flight control frequency through multi-rate interpolation, it is beneficial to reduce setpoint mutations, attitude fluctuations, and control overshoot; Fifth, by performing flight constraint safety verification and safety control signal switching, flight safety in complex dynamic scenarios is improved. Attached Figure Description

[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0013] Figure 1 This is a schematic diagram of the process steps of the delay compensation control method for dynamic target interaction of UAVs according to the present invention;

[0014] Figure 2 This is a functional structure diagram of the control system of the UAV of the present invention;

[0015] Figure 3 This is a schematic diagram illustrating the relationship between the end-to-end latency and the action point execution timestamp of this invention;

[0016] Figure 4 This is a schematic diagram illustrating the relationship between coordinate system transformation, delay error compensation, and multi-rate interpolation in this invention. Detailed Implementation

[0017] The embodiments of the present invention will now be described with reference to the accompanying drawings. It should be understood that the following embodiments are used to illustrate the technical solutions of the present invention, and not to limit the scope of protection of the present invention. Those skilled in the art can combine, substitute, or make equivalent modifications to the relevant features without departing from the concept of the present invention.

[0018] In this application, "target" refers to an object that the UAV needs to align with, follow, identify, operate, or avoid during target tracking, dynamic alignment, grasping, delivery, inspection interaction, or other operational tasks. The target can be a moving vehicle, a mobile platform, personnel-carried equipment, an item to be grasped, a delivery location, an obstacle, or other object with a real-time changing state.

[0019] In this application, an "action block sequence" refers to a set of actions output by a lightweight visual language action model in one or multiple consecutive inferences. Each action block includes multiple action points; each action point includes at least a control setting value and a desired execution timestamp, and may also include execution priority, load status, action duration, and validity identifier.

[0020] I. Control Method Implementation Examples

[0021] See Figure 1 This embodiment provides a delay compensation control method for dynamic target interaction of unmanned aerial vehicles (UAVs). This method can run on the UAV's onboard computing unit, or it can be executed collaboratively by the onboard computing unit and a ground computing unit. The onboard computing unit and the flight control mechanism can be connected via serial port, CAN, Ethernet, MAVLink link, or other communication links.

[0022] Step S10: Obtain multi-source observation data with timestamps and task instructions.

[0023] The acquisition module collects multi-source sensor data from the UAV. These multi-source sensors may include one or more of the following: RGB camera, RGB-D camera, event camera, IMU, GNSS module, RTK positioning module, visual odometry, LiDAR, ultrasonic ranging sensor, barometer, motor status acquisition unit, and battery management unit.

[0024] Each piece of sensor data is bound to an acquisition timestamp, sensor identifier, and data validity identifier. For example, visual data may include image frame number, image acquisition timestamp, camera intrinsic and extrinsic parameters, and target detection results; inertial measurement data may include angular velocity, linear acceleration, and attitude angle; positioning data may include position, velocity, and heading; and ranging and altitude data may include altitude relative to the ground, distance relative to obstacles, or distance relative to a target.

[0025] Task instructions can be natural language instructions, structured task commands, or a combination of both. For example, task instructions could be "follow the vehicle in front and maintain a distance of five meters," "aim at the target platform and then perform the deployment," or "go around the obstacle and return to the flight path." Task instructions are used to constrain the sequence of action blocks generated by the lightweight visual language action model.

[0026] Step S20: Estimate end-to-end latency and complete timing alignment

[0027] The delay and alignment module first records the timestamps of each link in the entire chain. Specifically, it records the sensor acquisition timestamp t_cap, the preprocessing input timestamp t_pre_in, the model input timestamp t_model_in, the model output timestamp t_model_out, the action signal transmission timestamp t_send, the flight controller reception timestamp t_recv, and the flight controller forming executable control variables timestamp t_fc.

[0028] In one implementation, the data preprocessing latency is the difference between t_model_in and t_pre_in, the model inference latency is the difference between t_model_out and t_model_in, the communication transmission latency is the difference between t_recv and t_send, and the flight control receiver execution latency is the difference between t_fc and t_recv. The total link latency τ can be expressed as the sum of the latency of each of the above stages and the fixed sensor acquisition latency. To improve robustness, a moving average, exponential smoothing, or outlier removal can be applied to τ obtained from N consecutive measurements.

[0029] After delay estimation is completed, the delay and alignment module performs time-series alignment of multi-source data based on a unified reference time. The unified reference time can be selected from the latest valid visual frame timestamp, the current flight control timestamp, or the execution reference time corrected by the end-link delay. For inertial, positioning, or ranging data earlier than the unified reference time, alignment can be performed using linear interpolation, attitude spherical interpolation, or extrapolation based on the motion model; for data later than the unified reference time, buffering or discarding can be performed.

[0030] For example, if the visual frame acquisition time is 30ms, the IMU data acquisition time is 10ms and 20ms, and the ranging data acquisition time is 25ms, and the unified reference time is 30ms, then the inertial state at 30ms can be obtained by extrapolating or interpolating the IMU data, and the distance state at 30ms can be obtained by interpolating the ranging data, so that the observation data of the input model remains consistent on the time axis.

[0031] Step S30: The lightweight visual language action model outputs a sequence of action blocks with execution timestamps.

[0032] The model module inputs time-aligned observation data and task commands into the lightweight visual-language-action model. The lightweight visual-language-action model can include a visual encoder, a speech encoder, a UAV state encoder, a multimodal fusion layer, and an action decoder. The visual encoder extracts environmental, target, and obstacle features; the speech encoder extracts task intent and constraints; the UAV state encoder extracts position, velocity, attitude, battery level, and payload status; the multimodal fusion layer maps these features to the same feature space; and the action decoder outputs a sequence of action blocks.

[0033] Lightweighting methods can include using lightweight backbone networks like MobileNet, lightweight Transformer architectures, network pruning, weight quantization, knowledge distillation, low-rank adaptation, or a combination of these methods. Lightweighting reduces the number of model parameters and inference computations, making the model more suitable for operation in UAV onboard computing units.

[0034] The action block sequence comprises multiple action blocks, and each action block includes multiple action points. The control settings for each action point can include position increment, velocity setpoint, acceleration setpoint, heading angle, attitude angle, rate of climb, throttle opening, or load action command. The model module binds a desired execution timestamp to each action point. The desired execution timestamp of the i-th action point can be expressed as T_i = T_trigger + τ + i × Δt_a, where T_trigger is the model inference trigger time, τ is the end-to-end latency, Δt_a is the action point control period, and i is the action point number.

[0035] Combination Figure 3 The horizontal axis represents time, showing the complete timing flow from left to right: data acquisition, model inference, communication transmission, and flight control execution. The figure shows that the end-to-end latency τ consists of model inference latency τm, communication latency τc, flight control execution / queuing latency τf, and other fixed processing latency, i.e., τ = τm + τc + τf + other fixed processing latency. The expected execution timestamps for action point i, action point i+1, and action point i+2 are calculated using the formula T_i = T_trigger + τ + i × Δt_a.

[0036] In this way, the model output no longer represents only "the action that should be performed at the current moment," but rather "the action points that should be performed at each expected execution moment in the future." This structure enables subsequent state prediction, coordinate transformation, and interpolation adaptation to be performed around the actual execution time of each action point.

[0037] Step S40: Predict the execution time state for each action point.

[0038] The inference module generates UAV predicted state data and target predicted state data for the expected execution timestamp of each action point in the action block sequence. The UAV state vector may include position, velocity, acceleration, attitude angle, angular velocity, and heading angle; the target state vector may include target position, velocity, acceleration, direction of motion, and recognition confidence.

[0039] When the target's motion is stable, a uniform motion model can be used for forward extrapolation; when the target accelerates or decelerates, a uniformly accelerated motion model can be used; when there is noise, obstruction, or abrupt changes in target observation, an extended Kalman filter model, an unscented Kalman filter model, or a particle filter model can be used. For the UAV's own state, extrapolation can be performed by combining IMU integrals, positioning data, and flight control state estimation results.

[0040] For example, if the target speed is 2 m / s and the end-to-end latency is 200 ms, without execution time prediction, the target position alone may experience a time misalignment of approximately 0.4 m. This embodiment incorporates this time difference into the state deduction of each action point, ensuring that action point compensation is based on the predicted state at the execution time, rather than the old state at the observation time.

[0041] Step S50: Coordinate system transformation and delay error compensation

[0042] The compensation module performs coordinate system transformation and delay error compensation on the action block sequence based on the UAV's predicted state data and the target's predicted state data. The coordinate system can include the camera coordinate system C, the body coordinate system B, the inertial coordinate system I, or the world coordinate system W. The action points output by the visual language action model can be located in the camera coordinate system, the target relative coordinate system, or the body coordinate system at the time of observation; flight control execution usually requires setpoints in the body coordinate system or the inertial coordinate system.

[0043] In one implementation, the compensation module converts the action point in the observation coordinate system into the control setpoint in the body coordinate system or inertial coordinate system at the execution time based on the UAV's predicted position, predicted attitude, and camera extrinsic parameters at the execution time. Subsequently, the position, velocity, heading angle, or attitude angle are corrected based on the change in UAV's predicted pose, the target's predicted displacement, and the relative state change between the observation time and the execution time.

[0044] Through this step, the spatial meaning of the action point is transformed from the relative state at the observation time to the relative state at the execution time. For example, if the target moves forward during the time delay, the compensation module will correct the UAV's expected tracking position or heading based on the target's predicted displacement; if the UAV itself undergoes an attitude change during the time delay, the compensation module will recalculate the velocity or attitude settings in the body coordinate system based on the predicted attitude.

[0045] Step S60: Action Buffer Scheduling and Smoothing

[0046] The compensated action points are written to the action cache queue. The action cache queue is sorted according to the expected execution timestamp and is used to connect action blocks output from different inference cycles. If the expected execution timestamp of an action point is earlier than the current flight control time and exceeds a preset lag threshold, the action point is identified as an expired action point and discarded.

[0047] When action points output by adjacent inference cycles overlap in timestamps, the cache queue can perform overwrite replacement, weighted fusion, or smoothing and limiting. For high-priority actions, such as obstacle avoidance, emergency stop, grab closure, or drop actions, a higher execution priority can be set in the cache queue; for regular tracking actions, smoothing fusion can be performed to reduce action abrupt changes.

[0048] Step S70: Generate high-frequency continuous setpoints using multi-rate interpolation

[0049] The interpolation module reads the compensation action sequence signal and determines the flight control operating frequency f_fc and the action point output frequency f_a. The interpolation factor R can be expressed as R = f_fc / f_a. When the flight control operating frequency is 200Hz and the action point output frequency is 20Hz, the interpolation factor R is 10, the flight control period is 5ms, and the action point period is 50ms.

[0050] When the interpolation factor is an integer, a fixed number of interpolation points can be generated between adjacent compensation action points; when the interpolation factor is a non-integer, time-normalized interpolation can be performed according to each flight control execution timestamp. Interpolation algorithms can employ linear interpolation, spline interpolation, fifth-order polynomial interpolation, minimum jerk interpolation, or model prediction interpolation. For position and velocity setpoints, fifth-order polynomial interpolation or minimum jerk interpolation can be preferentially used to reduce abrupt changes in acceleration and jerk; for heading angles, angle continuity processing can be used before interpolation to avoid jumps across ±π.

[0051] Through multi-rate interpolation, low-frequency discrete action points are converted into high-frequency continuous setpoints that the flight controller can read periodically. These high-frequency continuous setpoints can include position setpoints, velocity setpoints, attitude setpoints, heading setpoints, rate of climb setpoints, and load action setpoints.

[0052] Steps S80 to S90: Flight constraint safety verification and output control

[0053] The safety output module performs flight constraint safety verification on high-frequency continuous setpoint signals. Verification items may include geofence boundary verification, obstacle safety distance verification, flight altitude verification, speed verification, attitude angle limit verification, throttle status compliance verification, climb rate verification, and remaining battery power support verification.

[0054] When all verification items meet the preset safety threshold requirements, the S90B safety output module outputs a high-frequency continuous setpoint signal to the flight control system. When any verification item fails to meet the preset safety threshold requirements, the S90A safety output module outputs a safety control signal. Safety control signals may include hover control signals, automatic return-to-home control signals, safe landing control signals, or amplitude limiting deceleration control signals.

[0055] In one implementation, if the safe distance to an obstacle is insufficient, a hovering or obstacle avoidance control signal is prioritized; if the remaining battery power is lower than a preset return-to-home battery power threshold, an automatic return-to-home or safe landing control signal is output; if the attitude angle, rate of climb, or throttle opening exceeds the safe range, the high-frequency continuous setpoint is limited or switched to a stable hovering mode. Through these methods, the flight risks caused by abnormal control commands can be reduced while preserving normal operational capabilities.

[0056] In one embodiment of this application, combined with Figure 4 The process of delay error compensation and multi-rate interpolation control for UAVs is explained. Figure 4 This diagram illustrates the UAV delay error compensation and multi-rate interpolation control process, showing the complete processing flow from initial observation to high-frequency flight control command generation. First, at the observation time, the camera acquires motion point data based on the camera / relative coordinate system, which is input to the UAV / target state prediction module at the execution time. Based on the observation data, the module predicts the UAV and target states at the actual execution time. Then, the coordinate system transformation module completes the conversion between the camera coordinate system, the body coordinate system, and the inertial coordinate system. Finally, the module enters the delay error compensation module, which corrects the position, velocity, and heading parameters of the predicted state based on the end-to-end latency. The system first obtains the sequence of compensated action points and sorts them by execution timestamp. Then, the multi-rate interpolation module uses the ratio of the flight control frequency to the output frequency of the action points as the interpolation ratio to perform interpolation calculations on the low-frequency compensated action points. Multiple control points are added between adjacent action points, and finally, a high-frequency continuous setpoint is generated and output to the flight control or safety control system. The timing diagram below intuitively shows the distribution relationship between the low-frequency compensated action points and the interpolated high-frequency setpoints of the flight control. The larger black dots are the low-frequency compensated action points, and the dense small black dots are the interpolated high-frequency setpoints to match the high-frequency control cycle of the flight control and ensure the continuity and smoothness of the control commands.

[0057] II. Control System Implementation Examples

[0058] See Figure 2 This embodiment also provides a control system for an unmanned aerial vehicle (UAV). The control system includes an acquisition module 10, a delay and alignment module 20, a model module 30, a deduction module 40, a compensation module 50, a cache scheduling module 55, an interpolation calculation module 60, and a safety output module 70.

[0059] The acquisition module 10 is used to collect raw observation status data and task instructions with acquisition timestamps. The acquisition module 10 can be connected to a camera, IMU, GNSS / RTK module, ranging and altitude sensor, motor status acquisition unit, and battery management unit, and writes the collected data into the observation buffer according to the timestamp.

[0060] The delay and alignment module 20 is used to estimate the end-to-end delay duration and perform time-series alignment of the raw observation data based on a unified reference time. The delay and alignment module 20 can also perform moving average, exponential smoothing, and outlier removal on the delay measurements.

[0061] Model module 30 is used to input observation data and task instructions into a lightweight visual language action model and output a sequence of action blocks carrying execution timestamps. Model module 30 can be deployed on the UAV's onboard computing unit or distributed among the onboard computing unit and edge computing devices.

[0062] The inference module 40 is used to generate UAV predicted state data and target predicted state data for the expected execution timestamp of each action point in the action block sequence. The inference module 40 can select a uniform motion model, a uniformly accelerated motion model, an extended Kalman filter model, or an unscented Kalman filter model according to the target motion state and observation stability.

[0063] The compensation module 50 is used to perform coordinate system transformation and delay error compensation on the action points based on the predicted state data. The cache scheduling module 55 is used to sort the action points according to their expected execution timestamps, remove expired action points, process overlapping action points, and output a continuous and stable compensated action sequence signal.

[0064] The interpolation module 60 performs multi-rate interpolation based on the ratio of the flight control operating frequency to the action point output frequency to generate a high-frequency continuous setpoint signal that matches the flight control frequency. The safety output module 70 performs flight constraint safety verification on the high-frequency continuous setpoint signal and outputs a high-frequency continuous setpoint signal or a safety control signal based on the verification result.

[0065] III. Storage Media and Program Product Examples

[0066] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it can implement the aforementioned delay compensation control method for dynamic target interaction of unmanned aerial vehicles. The computer-readable storage medium may include a read-only memory, a random access memory, flash memory, a solid-state drive, a memory card, or other media capable of storing program code.

[0067] This embodiment also provides a computer program product, including a computer program. When executed by a processor, the computer program can implement the aforementioned delay compensation control method for dynamic target interaction with UAVs. The processor can be a CPU, GPU, NPU, FPGA, DSP, or other processor capable of executing computational instructions in an onboard computing unit.

[0068] IV. Supplementary Examples

[0069] In one example, the UAV performs a dynamic target following task. The flight control operates at a frequency of 200Hz, the visual language action model outputs an action point every 50ms, the model inference latency is 120ms, the communication transmission latency is 25ms, and the total fixed latency for preprocessing and flight control reception and execution is 35ms, resulting in a total end-to-end latency τ of 180ms. The model module binds the expected execution timestamp to each action point in the action block; the inference module predicts the UAV and target states based on each expected execution timestamp; the compensation module corrects the action points based on the predicted states; the interpolation module generates high-frequency continuous setpoints at 5ms intervals using a 10x interpolation; and the safety output module verifies and outputs the qualified setpoints to the flight control.

[0070] In another example, if the target's speed is high and the observation noise is significant, the inference module can select an extended Kalman filter model. If the action blocks output by adjacent inference cycles overlap, the cache scheduling module can retain high-confidence action points from the newer action block and perform weighted fusion on the overlapping intervals. If the safety output module detects that the obstacle's safe distance is below a threshold, it stops the normal setpoint output and switches to hovering or avoidance control signals.

[0071] Furthermore, it should be noted that the calculation of end-to-end latency includes statistical processing of latency sample values. When performing a moving average on N consecutive latency sample values, the window size can be set to 5–20; when performing exponential smoothing, the smoothing coefficient is set to 0.1–0.9; when removing outliers, latency sample values ​​exceeding the mean ± 3 times the standard deviation are identified as outliers and removed, thereby obtaining a stable and reliable end-to-end latency estimation result.

[0072] In non-integer multiplier interpolation, when the interpolation multiplier is non-integer, the actual timestamp of each flight control cycle is used as the reference, the time interval between adjacent action points is normalized to the [0,1] interval, and then linear interpolation or fifth-order polynomial interpolation is used to calculate the control setpoint of the corresponding interpolation point to match the actual control cycle of the flight control.

[0073] The safety threshold values ​​are preset based on the drone model and mission scenario, including: geofence boundary threshold of ≥1m from the boundary; obstacle safety distance of ≥0.5m; flight altitude of ≤120m; attitude angle roll / pitch of ≤30°; and initiating the return-to-home process when the remaining battery power is ≥15%.

[0074] Supplementary Explanation of Fixed Delay in Sensor Data Acquisition

[0075] In this embodiment, the end-to-end latency includes the fixed latency of sensor acquisition, which is the inherent delay of sensor hardware exposure and data output, including visual frame exposure latency, IMU output latency, and GNSS positioning latency. The value is either the fixed latency given in the corresponding sensor manual or the measured average value obtained from multiple samplings.

[0076] In the lightweight model structure and data flow, the data flow of each module of the lightweight model is as follows: the visual encoder outputs visual features, the language encoder outputs task features, and the UAV state encoder outputs state features. The three are input to the multimodal fusion layer for feature splicing and mapping. The fused features are input to the action decoder, and the output is a sequence of action blocks containing control setpoints and expected execution timestamps.

[0077] In the state vector and forward timing extrapolation, the forward timing extrapolation extrapolates the current state of the UAV and the target using a uniform speed or uniform acceleration model based on the time difference ΔT between the observation reference timestamp and the expected execution timestamp. When using a Kalman-type model, the observed state is used as the observation quantity and the motion equation is used as the state equation. The optimal predicted state at time ΔT is output, forming the state vector at the execution time.

[0078] In coordinate system transformation, the transformation is achieved based on camera extrinsic parameters and UAV predicted pose (position, attitude angle): first, the camera coordinate system motion point is transformed to the body coordinate system through camera extrinsic parameters, and then the UAV position and heading are transformed to the inertial coordinate system; or the UAV predicted pose is directly transformed to the body coordinate system at the time of execution, providing a unified coordinate reference for subsequent delay error compensation.

[0079] In the action cache scheduling rules, the processing rules of the action cache scheduling module include: when action point timestamps overlap, high-priority action points cover low-priority action points; regular action points adopt a weighted fusion method, with the weight being positively correlated with the proximity of the timestamp; at the same time, smoothing and limiting processing is adopted, and the position and speed change rate are limited to a preset safety range through first-order low-pass filtering.

[0080] Safety control signal switching rules: when different verification items are not met, the corresponding safety signal is matched: when the distance beyond the geofence or obstacle is insufficient, the hovering control signal is output; when the battery is low, the automatic return signal is output; when the attitude, speed or throttle exceeds the limit, the amplitude-limiting deceleration signal is output; and in extreme abnormal conditions, the safe landing signal is output.

[0081] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. Those skilled in the art can adjust the parameters, models, and thresholds in each step according to the type of UAV, flight control frequency, sensor configuration, mission scenario, and model structure; such adjustments should all fall within the protection scope of the present invention without departing from the essential concept of the present invention.

Claims

1. A delay compensation control method for dynamic target interaction of unmanned aerial vehicles (UAVs), characterized in that, The control method includes: Acquire raw observation data with acquisition timestamps from the UAV's multi-source sensors, as well as mission instructions corresponding to the current flight mission; Based on the original observation data, model input / output timestamps, action signal transmission / reception timestamps, and flight control execution timestamps, the end-to-end latency consisting of data acquisition, data preprocessing, model inference, communication transmission, and flight control reception / execution is estimated. The original observation data is then time-aligned based on a unified reference time to obtain the observation data. The observation data and the task instructions are input into the lightweight visual language action model. Based on the lightweight visual language action model, an action block sequence carrying an execution timestamp is output. The action block sequence includes multiple action points, and each action point includes a control setting value and a corresponding expected execution timestamp. For the expected execution timestamp of each action point in the action block sequence, based on the original observation state data and the full-link delay duration, and combined with the preset motion model, the UAV predicted state data and target predicted state data for the corresponding execution time are generated respectively. Based on the UAV predicted state data and the target predicted state data, the action points in the action block sequence represented by the observation time coordinate system are transformed to the execution time coordinate system, and delay error compensation is performed on the transformed action points to output the compensated action sequence signal. Based on the ratio of the flight control operating frequency to the action point output frequency, a multi-rate interpolation operation is performed on the compensation action sequence signal to generate a high-frequency continuous setpoint signal that matches the flight control frequency. The high-frequency continuous setpoint signal is subjected to flight constraint safety verification. If the verification result meets the preset safety threshold requirements, the high-frequency continuous setpoint signal is output to the flight control system for execution. If the verification result does not meet the preset safety threshold requirements, a safety control signal is output.

2. The control method according to claim 1, characterized in that, The steps for acquiring raw observation data with acquisition timestamps from multiple sources of sensors on a UAV include: Collect at least three types of data from the following sources: visual data, inertial measurement data, positioning data, ranging and altitude data, motor status data, and battery status data. Each type of data is bound to a collection timestamp, sensor identifier, and data validity identifier to form the raw data of the observation state. The visual data includes RGB images, RGB-D images, or event camera data, and the positioning data includes GNSS data, RTK data, or visual odometry data.

3. The control method according to claim 1, characterized in that, The steps for estimating the end-to-end latency include: Record the timestamps of the raw observation data entering the preprocessing module and the timestamps of entering the lightweight visual language action model to determine the data preprocessing delay; Record the timestamps of the lightweight visual language action model receiving observation data and the timestamps of the output action block sequence to determine the model inference latency; Record the timestamps of the action signals sent by the airborne computing unit and the timestamps of the action signals received by the flight control unit to determine the communication transmission delay; Record the time interval between the flight controller receiving the action signal and forming an executable control quantity to determine the flight controller's receive and execute delay; The data preprocessing delay, model inference delay, communication transmission delay, flight control receiving and execution delay, and sensor acquisition fixed delay are summed, and the sums are then subjected to moving average, exponential smoothing, or outlier removal to obtain the end-to-end delay duration.

4. The control method according to claim 1, characterized in that, The steps for time-series alignment of the raw observation data based on a unified reference time include: The latest valid visual frame timestamp, the current flight control timestamp, or the execution reference time corrected by the end-link delay duration shall be used as the unified reference time. Inertial measurement data, positioning data, and ranging and altitude data earlier than the unified reference time are interpolated or extrapolated, while data later than the unified reference time are cached or discarded, so that multi-source data correspond to the same time axis position.

5. The control method according to claim 1, characterized in that, The lightweight visual language action model includes a visual encoder, a language encoder, a UAV state encoder, a multimodal fusion layer, and an action decoder. The visual encoder is used to extract visual features of the environment, targets, and obstacles; the language encoder is used to extract task intent and constraint features; the UAV status encoder is used to extract position, speed, attitude, battery level, and payload status features; the multimodal fusion layer is used to fuse visual features, task intent features, and UAV status features; and the motion decoder is used to generate the motion block sequence based on the fused features. The lightweight visual language action model reduces the number of model parameters and inference computation through at least one of the following methods: lightweight convolutional backbone network, lightweight Transformer, network pruning, weight quantization, knowledge distillation, or low-rank adaptation.

6. The control method according to claim 1, characterized in that, The control setpoint of the action point includes at least one of the following: position increment, speed setpoint, acceleration setpoint, heading angle, attitude angle, climb rate, throttle opening, or load action command; The expected execution timestamp of the i-th action point is determined based on the model inference trigger time, the full-link delay duration, the action point sequence number, and the action point control cycle. The expected execution timestamp of the i-th action point is: T_i = T_trigger + τ + i × Δt_a, where T_i represents the expected execution timestamp of the i-th action point, T_trigger represents the model inference trigger time, τ represents the end-to-end latency duration, Δt_a represents the action point control cycle, and i represents the action point sequence number.

7. The control method according to claim 1, characterized in that, The steps for generating UAV predicted state data and target predicted state data corresponding to the execution time include: Establish a UAV state vector and a target state vector. The UAV state vector includes the UAV's position, velocity, acceleration, attitude angle, and heading angle. The target state vector includes the target's position, velocity, acceleration, and direction of motion. Based on the time difference between the expected execution timestamp and the observation reference timestamp for each action point, a uniform motion model, a uniformly accelerated motion model, an extended Kalman filter model, or an unscented Kalman filter model is selected to perform forward time-series deduction on the UAV state vector and the target state vector. Output the predicted state data of the UAV and the predicted state data of the target, which correspond one-to-one with each action point.

8. The control method according to claim 1, characterized in that, The steps of converting the action points in the action block sequence, represented in the observation time coordinate system, to the execution time coordinate system, and performing delay error compensation on the converted action points, include: The action points, represented by the camera coordinate system, target relative coordinate system, or body coordinate system at the observation time, are converted into control setpoints in the inertial coordinate system or body coordinate system at the execution time, based on the UAV's predicted pose and the target's predicted pose at the execution time. Based on the predicted UAV pose change, the predicted target displacement, and the relative state change between the observation time and the execution time, the position, velocity, heading angle, or attitude angle in the control setpoint are corrected to obtain the compensation action point. Before outputting the compensation action sequence signal, an action buffer scheduling step is also included: Write the compensation action points into the action cache queue and sort them according to the expected execution timestamp; Discard expired action points whose expected execution timestamp is earlier than the current flight control time and exceeds the preset lag threshold; Perform overlay replacement, weighted fusion, or smoothing and limiting processing on action points with overlapping or adjacent timestamps, and output the compensated action sequence signal.

9. The control method according to claim 1, characterized in that, The step of performing multi-rate interpolation on the compensated action sequence signal includes: Determine the flight control operating frequency f_fc and the action point output frequency f_a of the compensation action sequence signal, and determine the ratio of f_fc to f_a as the interpolation factor; When the interpolation factor is an integer, a corresponding number of interpolation points are generated between adjacent compensation action points according to the flight control cycle; when the interpolation factor is not an integer, time normalization interpolation is performed according to the actual flight control timestamp of each interpolation point. The high-frequency continuous setpoint signal is generated using at least one of the following methods: linear interpolation, spline interpolation, fifth-order polynomial interpolation, minimum jerk interpolation, or model prediction interpolation. The flight constraint safety verification includes at least three of the following: geofence boundary verification, obstacle safety distance verification, flight altitude verification, speed verification, attitude angle limit verification, throttle status compliance verification, climb rate verification, and remaining battery power support verification. When any verification item fails to meet the preset safety threshold requirement, a corresponding safety control signal is output according to the type of verification item that fails to meet the threshold. The safety control signal includes a hovering control signal, an automatic return-to-home control signal, a safe landing control signal, or a limited deceleration control signal.

10. A delay compensation control system for dynamic target interaction of unmanned aerial vehicles (UAVs), characterized in that, The control system includes: The acquisition module is used to acquire raw observation data with acquisition timestamps collected by the UAV's multi-source sensors, as well as mission instructions corresponding to the current flight mission. The delay and alignment module is used to estimate the end-to-end delay duration consisting of data acquisition, data preprocessing, model inference, communication transmission, and flight control execution based on the original observation data, model input and output timestamps, action signal transmission and reception timestamps, and flight control execution timestamps, and to perform time-series alignment of the original observation data based on a unified reference time to obtain the observation data. The model module is used to input the observation data and the task instructions into the lightweight visual language action model and output an action block sequence carrying execution timestamps; The deduction module is used to generate UAV predicted state data and target predicted state data at the corresponding execution time for the expected execution timestamp of each action point in the action block sequence. The compensation module is used to convert the action points to the execution time coordinate system and perform delay error compensation based on the UAV predicted state data and the target predicted state data, and output the compensated action sequence signal. The interpolation module is used to perform multi-rate interpolation on the compensation action sequence signal based on the ratio of the flight control operating frequency to the action point output frequency, and generate a high-frequency continuous setpoint signal that matches the flight control frequency. The safety output module is used to perform flight constraint safety verification on the high-frequency continuous setpoint signal and output the high-frequency continuous setpoint signal or safety control signal according to the verification result.