Unmanned aerial vehicle trajectory prediction method, device, equipment, storage medium and program product

CN122841902APending Publication Date: 2026-09-29PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611044602.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本申请提供了一种无人机轨迹预测方法、装置、设备、存储介质及程序产品,以解决无人机轨迹预测效果不佳的问题

Benefits of technology

[0016]本申请实施例提供的无人机轨迹预测方法,通过获取历史轨迹序列与红外图像序列,并利用多模态行为编码网络对轨迹特征与红外视觉特征进行融合编码,得到能同时表征运动状态和飞行意图的行为表征向量,从而比单纯依赖轨迹坐标更精准地捕捉无人机的瞬时机动意图;在此基础上,从预构建的行为模式记忆库中检索出与行为表征向量相匹配的多个行为模式向量,再以各行为模式向量作为生成条件,由条件轨迹解码网络分别进行轨迹解码,最终输出多条轨迹预测结果,能够在缺乏外部强结构约束的条件下,为轨迹生成提供显式的典型飞行模式先验,使预测轨迹既可覆盖多种飞行意图、有效应对高机动无人机的行为不确定性,又因行为模式约束而保持良好的物理合理性,显著提升了预测的多样性与实用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841902A_ABST
    Figure CN122841902A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and program product for predicting UAV trajectory, relating to the field of trajectory prediction technology. The method includes: acquiring historical trajectory sequences and infrared image sequences of the UAV within an observation time window; using a multimodal behavior coding network to fuse and encode the trajectory features of the historical trajectory sequences and the infrared visual features of the infrared image sequences to obtain behavior representation vectors describing the UAV's motion state and flight intention; retrieving multiple behavior pattern vectors matching the behavior representation vectors from a behavior pattern memory; and using a conditional trajectory decoding network to decode the trajectory using each behavior pattern vector as a generation condition, generating multiple trajectory prediction results. By implementing the technical solution of this application, multiple highly accurate, diverse, and physically reasonable trajectories can be generated based on multimodal fusion of behavior representations and behavior pattern priors, effectively covering the flight uncertainties of UAVs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of trajectory prediction technology, specifically to UAV trajectory prediction methods, devices, equipment, storage media, and program products. Background Technology

[0002] In applications such as low-altitude airspace management and drone logistics, accurate prediction of drone trajectories is a key technology for ensuring safety. Related trajectory prediction methods are mainly divided into two categories: deterministic prediction and uncertain prediction. Deterministic prediction can only output a single trajectory and cannot cover the diverse behaviors of highly maneuverable drones in three-dimensional airspace. Uncertain prediction methods are mostly derived from two-dimensional ground scenarios and heavily rely on prior external structures such as road topology and interaction rules to constrain the generation space. In the unconstrained three-dimensional airspace of low altitude, lacking such structured priors, direct transfer can easily lead to multiple generated trajectories that are highly similar or physically unreasonable. Summary of the Invention

[0003] This application provides a method, apparatus, device, storage medium, and program product for predicting drone trajectories, in order to solve the problem of poor performance in drone trajectory prediction.

[0004] In a first aspect, this application provides a method for predicting the trajectory of a UAV, comprising: acquiring a historical trajectory sequence and an infrared image sequence of a target UAV within an observation time window; using a multimodal behavior coding network to fuse and encode the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence to obtain a behavior representation vector of the target UAV, the behavior representation vector being used to represent the motion state and flight intention of the target UAV; retrieving multiple behavior pattern vectors that match the behavior representation vectors from a pre-built behavior pattern memory; and using a conditional trajectory decoding network to perform trajectory decoding with each behavior pattern vector as a generation condition to generate multiple trajectory prediction results for the target UAV.

[0005] In some optional implementations, a multimodal behavior coding network is used to fuse and encode the trajectory features of historical trajectory sequences and the infrared visual features of infrared image sequences to obtain the behavior representation vector of the target UAV. This includes: using a multimodal behavior coding network to perform gated fusion of trajectory features and infrared visual features to obtain a fused feature sequence; and performing temporal coding on the fused feature sequence to obtain the behavior representation vector.

[0006] In some optional implementations, a multimodal behavior coding network is used to perform gated fusion of trajectory features and infrared visual features to obtain a fused feature sequence. This includes: encoding the trajectory features using a trajectory encoder to obtain a trajectory feature vector; encoding the infrared visual features using an infrared visual encoder to obtain an infrared visual feature vector; concatenating the trajectory feature vector and the infrared visual feature vector to obtain a concatenated vector; mapping the concatenated vector to fusion weights using an adaptive gating network, and performing element-wise weighted summation of the trajectory feature vector and the infrared visual feature vector according to the fusion weights to obtain the fused feature sequence; wherein the multimodal behavior coding network includes a trajectory encoder, an infrared visual encoder, and an adaptive gating network.

[0007] In some optional implementations, temporal encoding of the fused feature sequence is performed to obtain a behavior representation vector, including: encoding the fused feature sequence using a recurrent temporal encoder to extract the hidden state of the last time step within the observation time window; performing dimensional transformation on the hidden state to obtain the behavior representation vector; wherein the multimodal behavior encoding network includes a recurrent temporal encoder.

[0008] In some optional implementations, the process of determining the trajectory features of the historical trajectory sequence includes: taking the absolute coordinates of the last frame of the historical trajectory within the observation time window as the observation endpoint, subtracting the absolute coordinates of the observation endpoint from the absolute coordinates of each time step in the historical trajectory sequence to obtain a relative trajectory sequence; for each pair of adjacent time steps in the relative trajectory sequence, subtracting the relative coordinates of the previous time step from the relative coordinates of the subsequent time step to obtain the velocity between adjacent time steps; taking the modulus of the velocity to obtain the velocity intensity; and combining the relative coordinates, velocity, and velocity intensity of the subsequent time step in adjacent time steps to obtain the trajectory features.

[0009] In some optional implementations, the process of determining the infrared visual features of an infrared image sequence includes: for any frame of an infrared image in the infrared image sequence, obtaining the projection coordinates of the target UAV in the infrared image; cropping out a square region centered on the projection coordinates and with a side length of a preset pixel size from the infrared image; and combining the square regions corresponding to each frame of the infrared image in chronological order to obtain the infrared visual features.

[0010] In some optional implementations, multiple behavioral pattern vectors that match the behavioral representation vectors are retrieved from a pre-built behavioral pattern memory, including: embedding the behavioral representation vectors to obtain a pattern query vector; determining the similarity between the pattern query vector and each pattern vector in the behavioral pattern memory, and selecting a preset number of target pattern vectors with the highest similarity; and projecting the target pattern vectors to obtain a preset number of behavioral pattern vectors.

[0011] In some optional implementations, the following steps are taken: First, the historical trajectory sequence, infrared image sequence, and true future trajectory of the sample UAV within the sample observation time window are acquired. Then, a multimodal behavior coding network is used to fuse and encode the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence to obtain the sample behavior representation vector of the UAV. Next, multiple sample behavior pattern vectors matching the sample behavior representation vector are retrieved from the behavior pattern memory. Finally, the posterior distribution of the true future trajectory is obtained through posterior encoding, and for each sample behavior pattern vector, corresponding latent variables are sampled from the posterior distribution. Then, each sample behavior pattern vector, the sample behavior representation vector, and the latent variables corresponding to the sample behavior pattern vector are fused to obtain the fusion condition vector corresponding to each sample behavior pattern vector. Finally, trajectory decoding is performed on each fusion condition vector to obtain the training trajectory prediction result. Based on the difference between the training trajectory prediction result and the true future trajectory, the parameters of the multimodal behavior coding network, the behavior pattern memory, and the conditional trajectory decoding network are adjusted.

[0012] Secondly, this application provides a drone trajectory prediction device, comprising: a first acquisition module for acquiring historical trajectory sequences and infrared image sequences of a target drone within an observation time window; a first encoding module for fusing and encoding the trajectory features of the historical trajectory sequences and the infrared visual features of the infrared image sequences using a multimodal behavior encoding network to obtain a behavior representation vector of the target drone, the behavior representation vector being used to represent the motion state and flight intention of the target drone; a first retrieval module for retrieving multiple behavior pattern vectors matching the behavior representation vectors from a pre-built behavior pattern memory; and a prediction module for using a conditional trajectory decoding network to perform trajectory decoding with each behavior pattern vector as a generation condition, generating multiple trajectory prediction results for the target drone.

[0013] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the UAV trajectory prediction method of the first aspect or any corresponding embodiment described above.

[0014] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the UAV trajectory prediction method of the first aspect or any corresponding embodiment described above.

[0015] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the UAV trajectory prediction method described in the first aspect or any corresponding embodiment.

[0016] The UAV trajectory prediction method provided in this application acquires historical trajectory sequences and infrared image sequences, and uses a multimodal behavior coding network to fuse and encode trajectory features and infrared visual features to obtain behavior representation vectors that can simultaneously represent motion state and flight intention. This allows for more accurate capture of the UAV's instantaneous maneuvering intention than simply relying on trajectory coordinates. Based on this, multiple behavior pattern vectors matching the behavior representation vectors are retrieved from a pre-built behavior pattern memory. Then, each behavior pattern vector is used as a generation condition, and a conditional trajectory decoding network performs trajectory decoding respectively, ultimately outputting multiple trajectory prediction results. This method can provide explicit typical flight pattern priors for trajectory generation even in the absence of external strong structural constraints. This allows the predicted trajectory to cover multiple flight intentions, effectively cope with the behavioral uncertainties of highly maneuverable UAVs, and maintain good physical rationality due to behavior pattern constraints, significantly improving the diversity and practicality of prediction. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the first type of drone trajectory prediction method according to an embodiment of this application; Figure 2 This is a schematic diagram of a second flowchart of the drone trajectory prediction method according to an embodiment of this application; Figure 3 This is a schematic diagram of the third process of the drone trajectory prediction method according to the embodiments of this application; Figure 4 This is a structural block diagram of a drone trajectory prediction device according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0021] Drone trajectory prediction is a key technology in applications such as low-altitude airspace management, drone logistics, and security monitoring. Its main goal is to predict the movement trajectory of drones over a period of time based on historical observation data, providing decision support for tasks such as situational awareness, conflict early warning, and flight path planning.

[0022] Currently, drone trajectory prediction technology still faces several prominent challenges.

[0023] First, the relevant technologies struggle to balance the diversity and reasonableness of prediction results. Some methods can only output a single deterministic trajectory, failing to cover the various flight intentions that highly maneuverable UAVs may undertake in the three-dimensional airspace (such as rapid turns, hovering, and variable-speed flight), which can easily lead to missed or misjudgments in scenarios such as airspace conflict warnings. Other methods attempt to generate multiple trajectory predictions to address the aforementioned uncertainties, but these methods often rely on prior external structures (such as road topology and interaction rules) to constrain the trajectory generation space. Since such fixed constraints do not exist in the low-altitude three-dimensional airspace, UAVs have extremely high degrees of freedom of motion. Directly applying these methods can easily result in multiple generated trajectories converging to the same pattern (mode collapse) or violating kinematic laws (such as trajectory jitter and sudden motion changes), failing to meet the diverse and reasonable prediction needs in practical applications.

[0024] Secondly, the relevant technologies have shortcomings in utilizing multimodal perception information, making it difficult to accurately perceive the instantaneous motion state and flight intentions of UAVs. Most methods rely solely on the coordinate sequences of historical trajectories for modeling, ignoring the instantaneous motion cues such as attitude adjustments and changes in aircraft orientation provided by infrared images. Even those few methods that attempt to incorporate infrared image information often fail to effectively fuse and encode trajectory features with infrared visual features, resulting in the underutilization of infrared modal discrimination capabilities and difficulty in accurately capturing the UAV's real-time maneuvering intentions, thus affecting prediction accuracy.

[0025] In summary, how to effectively integrate multimodal information such as historical trajectories and infrared images in low-altitude 3D scenes lacking prior external structures, accurately represent the motion state and flight intention of UAVs, and generate accurate, diverse, and physically reasonable multiple trajectory prediction results based on this is a technical problem that urgently needs to be solved in the field of UAV trajectory prediction.

[0026] The UAV trajectory prediction method provided in this application simultaneously acquires historical trajectory sequences and infrared image sequences, and uses a multimodal behavior coding network to deeply fuse and encode trajectory features and infrared visual features to obtain behavior representation vectors that accurately represent the UAV's motion state and flight intention. This overcomes the intention perception bias caused by related technologies neglecting instantaneous infrared motion cues and insufficient multimodal information fusion. Furthermore, this application no longer relies on external road topology or interaction rules as structural priors. Instead, it retrieves multiple behavior pattern vectors matching the current behavior representation vector from a pre-built behavior pattern memory. These typical flight patterns learned from the data are used as intrinsic generation priors. A conditional trajectory decoding network then decodes the trajectory using each behavior pattern vector as a condition, generating multiple trajectory prediction results that conform to different flight intentions and are kinematically reasonable. This solves the problems of deterministic prediction failing to cover behavioral diversity and uncertain prediction easily leading to pattern collapse and physically unreasonable generated trajectories in unconstrained low-altitude scenarios, achieving accurate and diversified prediction of highly maneuverable UAV trajectories.

[0027] According to an embodiment of this application, an embodiment of a UAV trajectory prediction method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0028] This embodiment provides a method for predicting the trajectory of a drone, which can be used in electronic devices such as servers and edge computing devices. Figure 1 This is a flowchart of a drone trajectory prediction method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the historical trajectory sequence and infrared image sequence of the target UAV within the observation time window.

[0029] The target drone refers to the specific drone that is the object of prediction in the current prediction task.

[0030] The observation time window refers to a fixed duration used to collect historical flight data of the target UAV, such as 1 second (corresponding to 5 time steps, sampling rate 5Hz).

[0031] Historical trajectory sequence refers to the sequence of three-dimensional coordinate positions recorded in chronological order by the target UAV within the observation time window.

[0032] An infrared image sequence refers to a continuous sequence of image frames captured by an infrared camera, synchronized with the historical trajectory sequence in time.

[0033] Specifically, motion data and visual images of the target UAV are simultaneously acquired through sensor systems deployed within the monitoring area. The trajectory sequence, sourced from positioning devices such as radar, electro-optical tracking systems, or global navigation satellite systems, records the three-dimensional spatial coordinates of the target UAV at various points in time within a set observation window. The infrared image sequence, synchronously sourced from infrared thermal imaging equipment, captures the thermal radiation visual information of the target UAV at corresponding time points. Both are aligned temporally, forming a complete observation sample that is then input into the subsequent prediction process.

[0034] Step S102: Using a multimodal behavior coding network, the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence are fused and encoded to obtain the behavior representation vector of the target UAV. The behavior representation vector is used to represent the motion state and flight intention of the target UAV.

[0035] A multimodal behavior encoding network (MMEC) is a neural network module that receives trajectory features extracted from historical trajectory sequences and infrared visual features extracted from infrared image sequences as input. It fuses and encodes these two different modalities, ultimately outputting a behavior representation vector that characterizes the motion state and flight intention of the target UAV. For example, it could be a motion semantic gating fusion network (MGFNet).

[0036] Trajectory features refer to high-dimensional mathematical vectors extracted from historical trajectory sequences by an encoder, which characterize the kinematic properties of the target UAV.

[0037] Infrared visual features refer to high-dimensional mathematical vectors extracted from infrared image sequences of a target UAV using a visual encoder, representing its appearance and attitude information. These features are projected into the same motion semantic space as the trajectory features.

[0038] The behavior representation vector refers to the final output vector of the multimodal behavior coding network, which integrates complete trajectory motion information and infrared visual semantic information within the observation time window. It is a global encoding of the target UAV's current comprehensive motion state and flight intention.

[0039] Specifically, trajectory features and infrared visual features are extracted from historical trajectory sequences and infrared image sequences, respectively. Trajectory features reflect the kinematic changes of the target UAV, while infrared visual features reflect its morphology and attitude information in the infrared images. These features, belonging to two different sensing modalities, are then input into a multimodal behavior encoding network. This network performs cross-modal semantic fusion, allowing the trajectory features and infrared visual features to interact and complement each other at the semantic level during the encoding process, ultimately integrating and outputting a compact behavior representation vector. This vector, fusing trajectory motion cues and infrared visual cues, can characterize the comprehensive motion state and flight intention of the target UAV within the observation time window.

[0040] Step S103: Retrieve multiple behavior pattern vectors that match the behavior representation vectors from the pre-built behavior pattern memory.

[0041] A behavior pattern memory refers to a pre-built storage module that stores a large number of vector representations of typical flight behavior patterns.

[0042] Behavioral pattern vectors are vectors retrieved from the behavioral pattern memory and mapped through a projection layer. Each vector represents a possible behavioral hypothesis, such as hovering, straight-line cruising, or rapid maneuvering.

[0043] Specifically, a pre-built behavior pattern memory stores a large number of encoded representations of typical flight behavior patterns. Each behavior pattern vector corresponds to a common UAV movement, such as continuous cruise, hovering, and rapid turning maneuvers. Once the behavior representation vector of the target UAV is obtained, this vector is used as a query condition. All stored behavior pattern vectors in the behavior pattern memory are traversed according to a similarity metric, calculating the degree of matching between the target's current behavior state and each typical behavior pattern. Finally, a predetermined number of behavior pattern vectors with the highest matching degree are selected as the search results. These vectors represent several possible behavioral hypotheses most closely related to the target UAV's current state.

[0044] Step S104: Using the conditional trajectory decoding network, trajectory decoding is performed with each behavior pattern vector as the generation condition to generate multiple trajectory prediction results for the target UAV.

[0045] Conditional trajectory decoding networks refer to a class of trajectory generation networks that use behavioral pattern vectors as explicit generation conditions. Their core function is to receive behavioral pattern vectors output from a behavioral pattern memory, and use these vectors as conditions to drive the decoding process, generating future trajectory sequences that conform to the corresponding behavioral assumptions. For example, a conditional variational autoencoder can be used.

[0046] The trajectory prediction result refers to the final output set of future trajectories covering multiple possible behavioral intentions of the target UAV. Each trajectory corresponds to the decoding result under a specific behavioral pattern condition and consists of three-dimensional coordinate points at multiple time steps within the prediction time domain (i.e., the time range for predicting the future, such as the next 5 seconds). Multiple trajectories together characterize the uncertainty and diversity of the target UAV's future movement.

[0047] Specifically, the multiple behavior pattern vectors obtained in the previous step are input one by one into the conditional trajectory decoding network as explicit behavioral constraints for trajectory generation. For each behavior pattern vector, the conditional trajectory decoding network decodes it under the guidance of the behavior assumption, and, combined with the position and state of the target UAV at the observation endpoint, gradually generates a future trajectory that matches the corresponding behavior pattern. Since different behavior pattern vectors represent different flight mode assumptions, the network will output different future motion paths under different conditions. By summing up the decoding results corresponding to all behavior pattern vectors, multiple trajectory prediction results for the target UAV are obtained. These trajectories cover a variety of possible future motion directions for the target UAV, thus effectively characterizing the uncertainty of highly maneuverable UAV behavior.

[0048] The UAV trajectory prediction method provided in this application acquires historical trajectory sequences and infrared image sequences, and uses a multimodal behavior coding network to fuse and encode trajectory features and infrared visual features to obtain behavior representation vectors that can simultaneously represent motion state and flight intention. This allows for more accurate capture of the UAV's instantaneous maneuvering intention than simply relying on trajectory coordinates. Based on this, multiple behavior pattern vectors matching the behavior representation vectors are retrieved from a pre-built behavior pattern memory. Then, each behavior pattern vector is used as a generation condition, and a conditional trajectory decoding network performs trajectory decoding respectively, ultimately outputting multiple trajectory prediction results. This method can provide explicit typical flight pattern priors for trajectory generation even in the absence of external strong structural constraints. This allows the predicted trajectory to cover multiple flight intentions, effectively cope with the behavioral uncertainties of highly maneuverable UAVs, and maintain good physical rationality due to behavior pattern constraints, significantly improving the diversity and practicality of prediction.

[0049] This embodiment provides a method for predicting the trajectory of a drone, which can be used in electronic devices such as servers and edge computing devices. Figure 2 This is a flowchart of a drone trajectory prediction method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the historical trajectory sequence and infrared image sequence of the target UAV within the observation time window. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0050] For example, in a specific implementation, a single input contains N historical trajectory sequences X of drones, and the historical trajectory sequence X can be represented as... Its length is Each time step contains three-dimensional spatial coordinates; simultaneously, the corresponding infrared image sequence is acquired. Infrared image sequence It can be represented as The infrared image sequence contains Each frame contains N infrared images, and the model outputs K predicted future trajectories Y for each drone. The predicted trajectory Y can be represented as... The length of each predicted trajectory is Each time step.

[0051] Step S202: Using a multimodal behavior coding network, the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence are fused and encoded to obtain the behavior representation vector of the target UAV. The behavior representation vector is used to represent the motion state and flight intention of the target UAV.

[0052] Specifically, step S202 includes: Step S2021: Using a multimodal behavior coding network, the trajectory features and infrared visual features are gated and fused to obtain a fused feature sequence.

[0053] The fused feature sequence refers to the temporal sequence of comprehensive features generated at each time step within the observation time window, which simultaneously contains kinematic and visual semantic information, after adaptively weighting and fusing trajectory feature vectors and infrared visual feature vectors within the multimodal behavior coding network. Specifically, after acquiring the trajectory features and infrared visual features at each time step within the observation time window, the multimodal behavior coding network does not fuse them using simple addition or direct concatenation. Instead, it introduces a gating mechanism: the network simultaneously receives feature information from both modalities, dynamically assesses the importance of each of the trajectory motion cues and infrared visual cues in determining the UAV's motion state at the current time step, and assigns different contribution ratios to the two modalities accordingly. Based on this ratio, the trajectory features and infrared visual features are weighted and integrated along the feature dimension, allowing the two types of information to selectively complement each other rather than being indiscriminately superimposed. After performing this operation at each time step within the observation time window, a set of fused features arranged chronologically is obtained, which is the fused feature sequence.

[0054] In some optional implementations, step S2021 above includes: Step a1: Use a trajectory encoder to encode the trajectory features to obtain a trajectory feature vector.

[0055] A trajectory encoder is a subnetwork within a multimodal behavior coding network, used to extract high-dimensional trajectory feature vectors from trajectory features. Its function is to map the raw kinematic parameters to a semantic feature space suitable for subsequent fusion. The trajectory feature vector is a high-dimensional numerical vector output by the trajectory encoder after encoding the trajectory features at a single time step; it is an abstract feature representation of the trajectory modality at that time step. Specifically, trajectory features contain kinematic parameter information extracted from the original coordinate sequence. As a subnetwork in the multimodal behavior coding network, the trajectory encoder's role is to map these kinematic parameters to a high-dimensional feature space. During the encoding process, the trajectory encoder abstracts and compresses the input kinematic parameters layer by layer through internal nonlinear transformations, ultimately outputting a fixed-dimensional vector representation, the trajectory feature vector. This vector retains the key information in the kinematic parameters while transforming them into a semantic format suitable for interaction with visual features.

[0056] For example, in a specific implementation, the trajectory encoder is constructed using a two-layer fully connected network, and its encoding process can be represented as: for the enhanced kinematic features of the i-th UAV at time step t... via trajectory encoder Obtain the trajectory feature vector The dimension of the trajectory feature vector is set to... , The value can be 512.

[0057] Step a2: Encode the infrared visual features using an infrared visual encoder to obtain the infrared visual feature vector.

[0058] An infrared visual encoder is a subnetwork within a multimodal behavior coding network, used to extract high-dimensional infrared visual feature vectors from infrared visual features. Its function is to encode visual information such as appearance and pose from infrared images into feature representations aligned with the motion semantic space.

[0059] Infrared visual feature vectors refer to high-dimensional numerical vectors output by an infrared visual encoder after encoding infrared image data at a single time step in an infrared image sequence. They are abstract feature representations of infrared modes at that time step.

[0060] Specifically, infrared visual features include visual information related to the target UAV extracted from infrared images. The infrared visual encoder, as a sub-network in the multimodal behavior coding network, performs deep feature extraction on this visual information. During the encoding process, the infrared visual encoder gradually extracts abstract semantic features related to the UAV's attitude, orientation, and other maneuvering states from the visual input through multiple layers of convolution or transform operations, ultimately outputting a fixed-dimensional vector representation, namely the infrared visual feature vector. The feature space of this vector maintains semantic consistency with the trajectory feature vector, providing a foundation for subsequent two-modal fusion.

[0061] For example, in a specific implementation, the infrared visual encoder uses a ResNet-50 backbone network truncated to Layer 1 to initially extract infrared visual image features, then uses a spatial pooling network AdaptiveMaxPool(1×1) to compress the dimensions, and finally uses a single-layer fully connected network to project onto the motion semantic space. For the i-th UAV at time step t, the cropped infrared image patch... Infrared visual feature vectors are obtained through an infrared visual encoder. Its dimensions are also set to .

[0062] Step a3: Concatenate the trajectory feature vector and the infrared visual feature vector to obtain the concatenated vector.

[0063] A concatenated vector is a long vector formed by concatenating the trajectory feature vector and the infrared visual feature vector at the same time step, end-to-end. Specifically, the trajectory feature vector output by the trajectory encoder and the infrared visual feature vector output by the infrared visual encoder at the same time step are concatenated end-to-end along the vector dimension. For example, if the dimension of the trajectory feature vector is m and the dimension of the infrared visual feature vector is n, the concatenated vector will have a longer dimension of m+n. This concatenated vector retains complete feature information from both modalities.

[0064] Step a4: The splicing vector is mapped to fusion weights using an adaptive gating network, and the trajectory feature vector and infrared visual feature vector are summed element-wise according to the fusion weights to obtain the fusion feature sequence.

[0065] The multimodal behavior coding network includes a trajectory encoder, an infrared vision encoder, and an adaptive gating network.

[0066] An adaptive gating network is a subnetwork in a multimodal behavior coding network used to dynamically calculate the fusion weights of the infrared and trajectory modes based on the concatenated vectors. Its key characteristic is that the weights are not fixed but adaptively adjusted according to the content of the input data.

[0067] The fusion weights are a set of values ​​output by the adaptive gating network, used to control the proportion of information retained by the trajectory feature vector and the infrared vision feature vector during element-wise weighted summation. The values ​​of the weights reflect the relative importance of the two modes in determining the motion state at the current time step.

[0068] Specifically, the adaptive gating network receives the spliced ​​vector as input, performs internal fully connected or transform operations, and outputs a set of weight coefficients corresponding to the dimensions of the trajectory feature vector and the infrared visual feature vector. These weight coefficients reflect the network's adaptive judgment of the contribution of the trajectory mode and the infrared mode at the current time step. Subsequently, each element of the trajectory feature vector is multiplied by its corresponding weight coefficient, and each element of the infrared visual feature vector is also multiplied by its corresponding weight coefficient. The products of the two sets are then added at their corresponding element positions to obtain the fused feature for that time step. This operation is performed sequentially for all time steps within the observation time window, and the resulting sequence constitutes the fused feature sequence.

[0069] For example, in a practical implementation, an adaptive gating network can employ a two-layer fully connected network. The trajectory feature vector... With infrared visual feature vectors After being spliced, the data is input into a gating network and passed through a multilayer perceptron. The gate control coefficients are mapped by the sigmoid activation function, i.e. ,in This represents the sigmoid function. Then, a gated fusion mechanism is used to calculate the fused features: , where ⊙ denotes element-wise multiplication. is the gating coefficient, used to adaptively adjust the contribution ratio of infrared features and trajectory features.

[0070] In the above implementation, by explicitly dividing the multimodal behavior encoding network into three collaborative components—a trajectory encoder, an infrared visual encoder, and an adaptive gating network—the fusion process acquires end-to-end learning capabilities. The trajectory encoder and infrared visual encoder each focus on extracting highly semantic trajectory feature vectors and infrared visual feature vectors from the original input, ensuring the discriminative power of the single-modal representation. Based on this, by concatenating the two vectors and introducing an adaptive gating network to map them as fusion weights, the model can dynamically adjust the contribution ratio of trajectory information and infrared visual information in each feature dimension according to the internal relationship of the dual-modal features. Then, element-wise weighted summation is performed according to these weights, thereby achieving accurate selection of key motion-related cues and effective suppression of redundant information. The resulting fused feature sequence more accurately reflects the instantaneous motion semantic relationships of the target UAV.

[0071] Step S2022: Perform temporal encoding on the fused feature sequence to obtain the behavior representation vector.

[0072] The fused feature sequence records the fused feature representations of each sampling point within the observation time window in chronological order. To capture the evolution of the target UAV's motion state throughout the entire observation period, temporal modeling of this sequence is necessary. The fused feature sequence is fed into a network module with sequence processing capabilities. This module sequentially reads the fused features at each time step and gradually accumulates and updates its understanding of the motion state internally. After processing the fused features of the last time step within the observation time window, the network's internal state at that moment is used as a summary encoding of the entire observation sequence. A dimensionality transformation operation is then performed to adjust it to the set feature dimensions. The final output is the behavior representation vector, used to represent the target UAV's comprehensive motion state and flight intention within the observation time window.

[0073] The UAV trajectory prediction method provided in this application introduces a gating fusion mechanism into a multimodal behavior coding network, enabling trajectory features and infrared visual features to be adaptively integrated into a fused feature sequence. This dynamically adjusts the contributions of the two types of features at different spatiotemporal points, effectively suppressing redundant information and accurately extracting discriminative representations highly correlated with the motion state. Based on this, temporal encoding of the fused feature sequence can further capture the evolution trend of the motion state over time. The resulting behavior representation vector concisely and orderly embodies the comprehensive motion characteristics and maneuvering intentions of the target UAV within the observation window, providing a high-quality conditional basis for subsequent behavior pattern matching.

[0074] In some optional implementations, step S2022 above includes: Step b1: Encode the fused feature sequence using a cyclic temporal encoder to extract the hidden state of the last time step within the observation time window.

[0075] A recurrent temporal encoder is a subnetwork in a multimodal behavior coding network. It adopts a recurrent neural network structure and is used to perform sequence modeling of fused feature sequences along the time dimension to capture the temporal evolution of the target UAV's motion state within the observation time window.

[0076] The hidden state refers to the internal state vector output by the cyclic timing encoder at each time step when processing sequence data. This vector encodes the cumulative temporal information from the beginning of the sequence to the current time step. Typically, the hidden state of the last time step within the observation window is taken as a summary representation of the entire observation sequence.

[0077] Specifically, a recurrent temporal encoder is a network structure that recursively processes sequential data step by step. Features from each time step in the fused feature sequence are sequentially input into the recurrent temporal encoder. At each time step, the encoder receives the current fused feature and the historical state passed down from the previous step, updates and outputs a new current hidden state. This process proceeds progressively along the time direction, with the state at the start of the observation time window being passed down sequentially to the end time. When the fused features for the last time step are input, the final hidden state output by the encoder encapsulates the temporal information of the entire observation sequence and is extracted as the basis for subsequent processing.

[0078] For example, in a specific implementation, this cyclic timing encoder is a recurrent neural network based on gated recurrent units (GRUs), employing... Layered GRU model ( ), The value can be 3. This will fuse the feature sequences of the entire observation sequence. Input to the GRU timing encoder, process the total... By fusing features at each time step, the temporal coding output of the entire sequence is obtained. Then extract the hidden state of the last time step of the observation time window, and denote it as... .

[0079] Step b2: Perform dimensional transformation on the hidden state to obtain the behavior representation vector.

[0080] Among them, the multimodal behavior coding network includes a cyclic timing encoder.

[0081] The hidden state of the last time step extracted from the recurrent temporal encoder has a dimension determined by the encoder's internal structure. To adapt it to the feature space required by downstream tasks, this hidden state is input into a dimension transformation module, which adjusts the dimension of the hidden state to a specified size through linear mapping or nonlinear transformation. The transformed vector is the behavior representation vector, which preserves the complete temporal motion semantics contained in the hidden state and has a dimensional format that matches the pattern vectors in the behavior pattern memory.

[0082] For example, in a concrete implementation, the dimension transformation module, i.e., the output projection layer, is implemented using a two-layer fully connected network. The extracted hidden states... Input and output projection layers yield global behavior representation vectors. This vector encodes the i-th drone from time step 1 to... The complete motion history, including high-level behavioral semantics such as speed changes and turning trends, provides a query benchmark for subsequent behavioral pattern matching.

[0083] In the above implementation, by using a cyclic temporal encoder to encode the fused feature sequence, the temporal evolution of the UAV's motion state within the observation time window can be effectively captured. Extracting only the hidden state at the last time step allows the resulting features to encapsulate the comprehensive behavioral information at the end of the entire observation process, thus reflecting the target UAV's current motion intention in real time. Furthermore, dimensional transformation of the hidden state compresses redundant intermediate temporal representations and maps the features into standardized behavioral representation vectors suitable for subsequent behavioral pattern retrieval, enhancing the compactness and discriminativeness of the representation.

[0084] In some optional implementations, the process of determining the trajectory features of the historical trajectory sequence includes: Step c1: Using the absolute coordinates of the last frame of the historical trajectory within the observation time window as the observation endpoint, subtract the absolute coordinates of the observation endpoint from the absolute coordinates of each time step in the historical trajectory sequence to obtain the relative trajectory sequence.

[0085] The last frame of the historical trajectory refers to the three-dimensional absolute coordinate position of the target UAV at the last sampling time point within the observation time window, which is also the starting point for prediction. The observation endpoint refers to the spatial position of the target UAV at the end of the observation time window, serving as the reference origin for converting absolute coordinates to relative coordinates. The relative trajectory sequence refers to the coordinate sequence obtained by subtracting the absolute coordinates of the observation endpoint from the absolute coordinates of each time step in the historical trajectory sequence, with the observation endpoint as the zero point. Specifically, the historical trajectory of the last frame (i.e., the last sampling time point) within the observation time window records the three-dimensional absolute coordinates of the target UAV at the end of the observation, which is used as the observation endpoint. For the trajectory coordinates of all sampling time points within the observation time window, including the first frame, all intermediate frames, and the last frame, a translation operation is uniformly performed: the absolute coordinate value of each time step is subtracted from the absolute coordinate value of the observation endpoint. After this operation, the coordinates of the observation endpoint become zero, and the coordinates of other time steps become offsets relative to the endpoint. The entire trajectory sequence is transformed into a relative trajectory sequence with the observation endpoint as the origin.

[0086] Taking the i-th drone as an example, its historical trajectory The absolute coordinates of the last frame within the observation time window Using this origin as a reference point, the absolute coordinates of all other time steps are uniformly subtracted from the coordinates of this origin point, thus transforming the entire trajectory into a relative trajectory sequence. This operation not only eliminates positional discrepancies between different drones, but also allows the model to focus on motion patterns rather than absolute positions.

[0087] Step c2: For each pair of adjacent time steps in the relative trajectory sequence, subtract the relative coordinates of the previous time step from the relative coordinates of the later time step to obtain the velocity between adjacent time steps.

[0088] Each pair of adjacent time steps in the relative trajectory sequence refers to a group consisting of two sampling points that are temporally adjacent in the relative trajectory sequence, such as step t and step t+1.

[0089] The next time step refers to the later sampling point in a pair of adjacent time steps, usually denoted as step t+1. The relative coordinates of the next time step refer to the three-dimensional spatial coordinates of the next time step relative to the observation endpoint.

[0090] The previous time step refers to the earlier sampling point in a pair of adjacent time steps, usually denoted as step t. The relative coordinates of the previous time step refer to the three-dimensional spatial coordinates of the previous time step relative to the observation endpoint.

[0091] The velocity between adjacent time steps refers to the displacement vector obtained by subtracting the relative coordinates of the previous time step from the relative coordinates of the subsequent time step, representing the direction and amount of motion of the target UAV within that unit time interval.

[0092] Specifically, the time steps in the relative trajectory sequence are arranged chronologically, and two temporally adjacent sampling points are considered as a pair. For any pair of adjacent time steps, the relative coordinate value of the later time step is taken, and the relative coordinate value of the earlier time step is subtracted to obtain the displacement change within that time interval. This displacement change is the velocity vector between the two time steps, containing both the direction and amplitude of motion information.

[0093] Step c3: Take the velocity modulo to obtain the velocity intensity, and combine the relative coordinates, velocity, and velocity intensity of the next time step in adjacent time steps to obtain the trajectory features.

[0094] Velocity intensity refers to the scalar value obtained by taking the modulus of the velocity vector. It reflects the speed of the target UAV's movement within that time interval, but does not include directional information. Specifically, the velocity vector between adjacent time steps is moduloed, i.e., its Euclidean length is calculated, resulting in a scalar value. This value characterizes the target UAV's movement rate within that time interval, i.e., the velocity intensity. This velocity intensity is then numerically concatenated with the relative coordinates (three-dimensional values) and velocity vector (three-dimensional values) of the next time step to obtain a comprehensive feature representation that includes coordinate information, movement direction information, and movement rate information; this is the trajectory feature of that time step.

[0095] For example, the specific calculation process can be as follows: for each pair of adjacent time steps in the relative trajectory sequence, calculate the velocity between the adjacent time steps. , where t = 1, 2, …, -1. Then, take the modulus of this velocity vector to obtain the velocity intensity. The final constructed enhanced kinematic features... This step transforms the original absolute trajectory sequence into a relative trajectory sequence and increases the feature dimension from 3D to 7D, explicitly modeling the drone's velocity and velocity intensity at each time step.

[0096] In the above implementation, by converting absolute trajectory coordinates into a relative trajectory sequence relative to the observation endpoint, the distribution bias introduced by the difference in absolute position between different trajectory samples is eliminated, allowing the model to focus on modeling the motion pattern itself. On this basis, the velocity and its corresponding velocity intensity between adjacent time steps are explicitly calculated from the relative trajectory sequence, and the relative coordinates, velocity and velocity intensity of the next time step are combined into trajectory features. This process not only preserves the spatial change information of displacement, but also directly embeds the instantaneous kinematic attributes that reflect the speed and maneuvering amplitude of the UAV, thereby providing a more discriminative underlying motion description for subsequent behavior representation.

[0097] In some optional implementations, the process of determining the infrared visual features of an infrared image sequence includes: Step d1: For any frame of infrared image in the infrared image sequence, obtain the projection coordinates of the target UAV in the infrared image.

[0098] Any frame of infrared image refers to a specific frame in an infrared image sequence that is temporally aligned with a certain time step in the historical trajectory sequence. Projected coordinates refer to the pixel positions on the two-dimensional plane of the infrared image, mapped from the three-dimensional spatial position of the target UAV. These are typically the row and column coordinates of the image. Specifically, each frame in the infrared image sequence is temporally aligned with a time step in the historical trajectory sequence. For any frame of the infrared image, based on the target UAV's three-dimensional spatial position at that moment and the imaging parameters of the infrared imaging device, the corresponding position of the target UAV's three-dimensional coordinates on the two-dimensional plane of the infrared image is calculated through spatial projection transformation. This position is represented by the row index and column index in the image, which are the projected coordinates.

[0099] Step d2: Crop a square region from the infrared image, centered at the projection coordinates and with a side length equal to the preset pixel size.

[0100] The preset pixel size refers to the number of pixels on the side of a square that is pre-set when cropping local blocks from an infrared image. It is used to unify the size of all infrared image blocks to adapt to the input requirements of subsequent coding networks.

[0101] A square region refers to a local image patch cropped from an infrared image, centered on the projected coordinates and with a preset pixel size as its side length. This region contains the infrared radiation information of the target drone and its nearby background. Specifically, after obtaining the projected coordinates of the target drone in the current frame of the infrared image, a square boundary with a preset pixel size is defined by extending the coordinates as the geometric center in each of the four directions (up, down, left, and right). The image area covered by this square boundary is cropped from the entire frame of the infrared image, resulting in a fixed-size local image patch. This image patch, centered on the target drone, contains visual information about the drone and its neighboring areas.

[0102] For example, in a specific implementation, the preset pixel size P can be 64, that is, from the original infrared image A 64×64 pixel local block was obtained by cropping from the middle and used as an infrared visual feature. .

[0103] Step d3: Combine the square regions corresponding to each frame of infrared images in chronological order to obtain infrared visual features.

[0104] The temporal order refers to the sequential arrangement of the infrared images according to their acquisition time, ensuring that the trajectory sequence and the infrared image sequence are aligned one-to-one in the temporal dimension. Specifically, the above-mentioned cropping operation is performed on each frame in the infrared image sequence to obtain square local image regions corresponding to each time step. These local image regions are then arranged sequentially according to the temporal order of the frames in the original infrared image sequence, forming an image sequence with a temporal dimension, which is the infrared visual feature, aligned one-to-one with the historical trajectory sequence in the temporal dimension.

[0105] In the above implementation, by obtaining the projection coordinates of the target UAV in each frame of infrared image and cropping a square area of ​​fixed pixel size based on these coordinates, background, clutter and other interfering targets unrelated to the target can be accurately removed from the image, making the extracted visual information highly focused on the target UAV body and its adjacent motion space. Then, the square areas corresponding to each frame are combined in chronological order to form infrared visual features, which not only completely preserves the continuous temporal dynamics of the target UAV's appearance changing with its attitude, but also provides standardized input with spatial alignment and consistent scale in the time dimension for subsequent feature extraction, significantly improving the purity and information density of infrared visual features.

[0106] Step S203: Retrieve multiple behavior pattern vectors that match the behavior representation vectors from the pre-built behavior pattern memory.

[0107] It should be noted that the behavior pattern memory and the subsequent conditional trajectory decoding network used for trajectory decoding together constitute the Behavior Pattern Conditional Generation Network (BPCNet). The Behavior Pattern Conditional Generation Network comprises two sub-modules: the behavior pattern memory and the conditional trajectory decoding network. The behavior pattern memory stores and retrieves typical flight behavior pattern vectors, providing behavioral prior constraints for trajectory generation. The conditional trajectory decoding network uses these behavior pattern vectors as generation conditions, combining them with behavioral representation vectors and latent variables to decode and generate multiple future trajectories that conform to different behavioral assumptions. Working together, the two enable the model to generate physically plausible predicted trajectories that cover diverse flight intentions without relying on external structured priors.

[0108] Specifically, the behavioral pattern memory bank is a learnable memory module. ,in This represents the number of pattern slots in the memory, for example, 100. Each pattern vector represents a typical UAV flight behavior pattern, such as hovering, straight-line cruising, or rapid maneuvering. The behavior pattern memory is registered as a non-parametric buffer, and the initial pattern vectors follow a standard Gaussian distribution. During the inference phase, the behavioral pattern memory directly loads the content of the pre-trained fixed memory and is no longer updated.

[0109] Specifically, step S203 includes: Step S2031: Embed the behavior representation vector to obtain the pattern query vector.

[0110] A pattern query vector is a vector obtained by embedding a behavior representation vector. Its dimension and semantic space are identical to the pattern vectors in the behavior pattern memory, and it is used for similarity retrieval within the behavior pattern memory. Specifically, the behavior representation vector output by the multimodal behavior encoding network is input into an embedding mapping module. This module uses linear or nonlinear transformations to map the behavior representation vector from the current feature space to the same semantic space as the pattern vectors in the behavior pattern memory. The mapped vector is the pattern query vector, whose dimension and semantic characteristics are consistent with the pattern vectors stored in the memory, providing a unified metric basis for subsequent similarity comparisons.

[0111] For example, in a specific implementation, the pattern embedding network employs a two-layer fully connected network. The global behavioral representation obtained in step S202 is then used... Input Pattern Embedded Network This yields the pattern query vector: This vector is used to retrieve behavioral patterns in the behavioral pattern memory that are similar to the current motion state.

[0112] Step S2032: Determine the similarity between the pattern query vector and each pattern vector in the behavior pattern memory, and select the target pattern vector with the highest similarity (preset number).

[0113] The preset quantity refers to the number of most similar pattern vectors retrieved from the behavior pattern memory, which is the final number of candidate behavior patterns used for trajectory generation. The target pattern vectors refer to the preset number of pattern vectors retrieved from the behavior pattern memory that have the highest similarity to the pattern query vector, representing several typical flight behavior patterns that best match the current state of the target UAV. Specifically, each pattern vector stored in the behavior pattern memory is iterated through one by one, and its similarity metric with the pattern query vector is calculated, typically using cosine similarity or Euclidean distance. After all calculations are completed, all pattern vectors in the behavior pattern memory are sorted from highest to lowest similarity to the query vector, and the top-ranked pattern vectors, equal to the preset quantity, are selected as the target pattern vectors. These represent several typical flight behavior types that are closest to the current motion state of the target UAV.

[0114] For example, in a specific implementation, the pattern query vector is first calculated. Cosine similarity with all pattern vectors in the behavior pattern memory: Then, based on the cosine similarity, the K pattern vectors with the highest similarity are selected to form a candidate set. Where K is the preset number of candidate behavior patterns, which can be 10, i.e., K = 10.

[0115] Step S2033: Project the target pattern vector to obtain a preset number of behavior pattern vectors.

[0116] The retrieved target pattern vectors are input one by one into a projection mapping module. This module remaps each target pattern vector in the feature space, transforming it from the storage space of the behavior pattern memory to an expression space that adapts to the input requirements of the subsequent conditional trajectory decoding network. The vectors after projection mapping are the behavior pattern vectors, with the same number as the target pattern vectors. Each behavior pattern vector will serve as a conditional input for trajectory generation.

[0117] For example, in a specific implementation, the pattern projection layer uses a single-layer fully connected network. The retrieved K memory pattern vectors are then passed through the pattern projection layer. Perform projection mapping to generate K behavioral pattern vectors: ,in These K behavioral pattern vectors will serve as the K conditional inputs to the subsequent conditional variational autoencoder, driving the generation of K future trajectories.

[0118] The UAV trajectory prediction method provided in this application achieves accurate transformation from general behavioral representations to specific behavioral pattern vectors through a multi-stage processing flow. First, the behavioral representation vector is embedded and mapped to obtain a pattern query vector, which is then projected into a retrieval space comparable to the pattern vectors in the memory, creating conditions for efficient matching. Next, by measuring the similarity between the query vector and each pattern vector in the memory, and selecting the most similar target pattern vectors of a preset number, rapid focusing and filtering of the behavioral priors that best fit the current motion state is achieved. Finally, the selected target pattern vectors are projected and mapped, endowing them with the conditional expression capabilities required for subsequent trajectory decoding and generation. This maintains the consistency of behavioral semantics while ensuring the effective differentiation and usability of multiple retrieval results, laying a stable and targeted pattern conditional foundation for the generation of multiple differentiated trajectories.

[0119] Step S204: Using a conditional trajectory decoding network, trajectory decoding is performed with each behavior pattern vector as the generation condition to generate multiple trajectory prediction results for the target UAV. For details, please refer to [link to details]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0120] During the model inference phase, the above steps constitute a complete inference chain. At this point, the pattern vectors in the behavior pattern memory remain fixed and are no longer updated; the latent variables required by the conditional trajectory decoding network to generate multiple trajectories are directly sampled from the standard Gaussian prior distribution, i.e. Subsequently, for each behavior pattern branch k, a behavior pattern condition vector is constructed. Then, the latent variables are fused together through a fusion network. With behavioral pattern condition vector Perform fusion to obtain the fusion condition vector. The trajectory decoder GRU is based on this fusion condition vector, combined with the position coordinates of the observation endpoint. The system gradually decodes and generates the relative displacements at multiple future time steps, then performs coordinate transformation to obtain the final predicted trajectory. By combining different behavioral pattern vectors with different random latent variables, the decoder can output K physically plausible future trajectories that cover a variety of flight intentions. .

[0121] In some alternative implementations, the multiple trajectory results predicted in the above steps can be applied to the low-altitude airspace control system to provide decision support for situational awareness and conflict early warning. Specific applications include: Drone Formation Collision Avoidance: Input the predicted trajectory Y into the airspace control system to detect whether there are overlapping areas between the predicted trajectories of multiple drones. For any two drones i and j, calculate the minimum distance between their respective K predicted trajectories across all time steps: .like If the distance is less than the preset safety threshold (e.g., 5 meters), a collision risk is determined, the system triggers an avoidance warning, and prompts the relevant drone to adjust its flight path.

[0122] Flight Conflict Warning: Determines whether the predicted flight path has entered a pre-defined no-fly zone. For the k-th predicted trajectory of UAV i, at each future time step t, check whether the trajectory point falls within the no-fly zone and define a collision flag: ,like ,in (·) is an indicator function that takes the value true when the condition is met. If the collision flag of any trajectory is true, an early warning will be issued, and the drone will be advised to change course in advance to avoid the no-fly zone.

[0123] Counter-drone interception: In low-altitude defense scenarios, based on the assumption of K trajectories for the same target drone, the optimal interception path or interception point is planned. By analyzing the spatial distribution of multiple potential flight trajectories, the areas that the drone may reach in the future can be estimated, thereby optimizing the deployment of interception resources and improving the interception success rate.

[0124] In some optional implementations, the above-described UAV trajectory prediction method further includes: Step e1: Obtain the sample drone's historical trajectory sequence, sample infrared image sequence, and real future trajectory within the sample observation time window.

[0125] By synchronously collecting motion data and visual images of sample drones through sensor systems deployed in the monitoring area, a preset duration observation window is extracted from the complete flight record. The three-dimensional spatial coordinates of the sample drones within the window are arranged in chronological order to form a historical trajectory sequence of the sample. The continuous image frames collected by the time-synchronized infrared thermal imaging device form a sample infrared image sequence. At the same time, a preset duration sequence of the actual flight coordinates of the sample drones after the end of the observation window is extracted as the real future trajectory. These three parts together form a complete training sample.

[0126] For example, in a practical implementation, training samples are obtained from a paired dataset of drone trajectories and infrared images collected in real or simulated low-altitude scenarios. Each training sample contains: a historical trajectory, the corresponding infrared image sequence, and a real future trajectory. For example, you can set = 5 frames (corresponding to 1 second of history, sampling rate 5Hz) = 25 frames (corresponding to a 5-second prediction time domain), K = 10, that is, the drone's movement within 1 second is observed as the historical trajectory, and 10 trajectories of the drone in the next 5 seconds are predicted.

[0127] Step e2: Using a multimodal behavior coding network, the trajectory features of the sample UAV's historical trajectory sequence and the infrared visual features of the sample infrared image sequence are fused and encoded to obtain the sample UAV's behavior representation vector.

[0128] The historical trajectory sequence and infrared image sequence of the samples are preprocessed to extract trajectory features and infrared visual features, respectively, and then fed into a multimodal behavior coding network. The network performs cross-modal gating fusion of the features from the two modalities, adaptively adjusting the information contribution ratio of the two to generate a fused feature sequence. This fused feature sequence is then temporally encoded to extract a generalized representation of the complete motion history within the observation time window. Finally, a sample behavior representation vector is output, which integrates the kinematic cues and infrared visual cues of the sample UAV, representing its comprehensive motion state and flight intention within the sample observation time window.

[0129] In some alternative implementations, during training, to enable the infrared visual encoder to extract semantic features closely related to the motion state while suppressing static interference such as background thermal radiation in the infrared image, it is also necessary to additionally utilize the infrared visual feature vectors of the samples. Trajectory feature vector Calculate the motion semantic alignment loss. Specifically, first, L2 normalize both sides, and then calculate the pairwise cosine similarity matrix. ,in , This is a learnable temperature parameter; for example, the initial value could be set to 0.07, with a lower limit of 0.01. Each row is considered a... For a class classification problem, where diagonal positions are used as positive samples, the motion semantic alignment loss is defined as: This loss term forces the infrared visual features and trajectory motion features to align with each other in the motion semantic space, improving the quality of modality fusion. It should be noted that this loss term is only calculated during the training phase, and not during the inference phase.

[0130] Step e3: Retrieve multiple sample behavior pattern vectors that match the sample behavior representation vectors from the behavior pattern memory.

[0131] The sample behavior representation vector is used as the query condition. Through embedding mapping, it is transformed into the same semantic space as the pattern vectors in the behavior pattern memory, resulting in the query vector. All behavior pattern vectors stored in the behavior pattern memory are traversed, and the similarity between the query vector and each pattern vector is calculated. After sorting by similarity from high to low, a predetermined number of pattern vectors at the top are selected as the search results. These results are the sample behavior pattern vectors most similar to the current behavior state of the sample drone.

[0132] In some optional implementations, during the training phase, after retrieving the sample behavior pattern vectors, an exponential moving average (EMA) soft-weighted update needs to be performed on the hit memory slots in the behavior pattern memory bank. The specific update process is as follows: First, the update weights are calculated. ,in The retrieval temperature parameter is used to retrieve data from the memory. For example, a value of 0.1 can be set to control the sharpness of the softmax distribution. The smaller the value, the more concentrated the weights are on the memory slots with the highest similarity. Then, this weight is used... This will retrieve all items within this batch that were found to be numbered 1. The projected behavioral pattern vector corresponding to the samples of each memory unit According to the preset EMA update factor Weighted blending into the original memory slot vector: ,in For all items retrieved in this batch The sample set of each memory unit. This update method does not rely on gradient backpropagation, and can continuously optimize the memory bank content during training, improving the utilization efficiency of memory slots. Simultaneously, the behavior pattern alignment loss is also calculated during training. Entropy regularization loss The former is used to jointly optimize the memory bank and the projection layer, ensuring that the memory slots retain effective behavioral features with predictive value, while the latter avoids behavioral pattern collapse by penalizing excessive weight concentration. It should be noted that the EMA update and regularization loss calculation mentioned above are only performed during the training phase; this step is skipped during the inference phase.

[0133] Step e4: Obtain the posterior distribution of the real future trajectory after posterior encoding, and sample the corresponding latent variables from the posterior distribution for each sample behavior pattern vector.

[0134] The true future trajectory is input into a posterior encoding network for encoding. The posterior encoding network extracts the motion features of the true trajectory and outputs its probability distribution parameters in the latent space, typically the mean and variance, thereby determining a posterior probability distribution. For each sample behavior pattern vector, a reparameterization technique is used to randomly sample from this posterior distribution. Each sampling yields a latent variable corresponding to the current behavior pattern vector, which encodes the potential motion features of the future trajectory under the corresponding behavior pattern.

[0135] For example, in a concrete implementation, the posterior encoder A 3-layer GRU (e.g., 3 input dimensions, 512 hidden dimensions, dropout=0.1) is used to analyze the true future trajectory. The original three-dimensional coordinates are encoded to obtain the true future trajectory encoding vector. The encoded vector is then further processed using a posterior multilayer perceptron consisting of two fully connected layers. Subsequently, a mean mapper consisting of a single fully connected layer is used to compute the mean of the posterior distribution. The variance of the posterior distribution is calculated using a variance mapper consisting of a single-layer fully connected network. Therefore, the posterior probability distribution is determined as follows: Finally, for each sample behavior pattern vector The latent variables are obtained by sampling from their posterior distribution using the reparameterization technique: , This step is performed only during the training phase; during the inference phase, the data is obtained directly from the standard Gaussian prior distribution. Sample latent variables.

[0136] Step e5: The sample behavior pattern vector, sample behavior representation vector, and latent variables corresponding to each sample behavior pattern vector are fused to obtain the fusion condition vector corresponding to each sample behavior pattern vector.

[0137] For each sample behavior pattern vector, it is concatenated or projected and fused with the sample behavior representation vector and the latent variables sampled from the posterior distribution for that behavior pattern vector along the vector dimension. Specifically, the sample behavior pattern vector is first combined with the sample behavior representation vector to form a behavior pattern condition vector. Then, the behavior pattern condition vector and the latent variables are integrated through a fusion network to output a fused condition vector that combines the behavior prior, the current motion state, and the latent motion encoding. This process is repeated for each sample behavior pattern vector, resulting in multiple fused condition vectors equal to the number of behavior pattern vectors.

[0138] For example, in a specific implementation, there are K behavior patterns. For the k-th behavior branch, the behavior pattern condition vector is first constructed. Then the latent variables With the condition vector of this behavior pattern The fusion is performed through a fusion network to obtain the fusion condition vector. This fusion network serves as the input layer of the conditional trajectory decoding network, responsible for integrating multi-source conditional information into a unified decoded conditional representation.

[0139] Step e6: Decode the trajectory of each fusion condition vector to obtain the training trajectory prediction result.

[0140] Each fusion condition vector is input into the trajectory decoding network. The decoding network uses the behavioral pattern information and latent motion encoding in the fusion condition vector as generation conditions, and combines them with the position coordinates of the sample UAV at the observation endpoint to progressively decode and generate the predicted coordinate sequence for each future time step. Each fusion condition vector corresponds to a predicted trajectory. By summing up the decoding results of all fusion condition vectors, multiple training trajectory prediction results are obtained, which are the same number as the number of behavioral pattern vectors.

[0141] For example, in a specific implementation, the trajectory decoder uses a 3-layer GRU (e.g., 512 hidden dimensions, dropout=0.1) to generate relative displacement predictions for multiple future time steps, and then combines this with the position coordinates of the observed endpoint. Trajectory decoding is performed using a two-layer fully connected network to obtain the k-th predicted trajectory. Ultimately, for N drones, each drone predicts K trajectories, and the model output is... .

[0142] Step e7: Based on the difference between the training trajectory prediction results and the actual future trajectory, adjust the parameters of the multimodal behavior encoding network, behavior pattern memory, and conditional trajectory decoding network.

[0143] The predicted trajectory results from the training process are compared with the actual future trajectories to calculate the prediction error. Multiple loss terms can be constructed based on this error: for example, divergence loss to constrain the proximity of the latent space distribution to the prior distribution; regression loss calculated only for the prediction closest to the actual trajectory; diversity loss to force different predicted trajectories to maintain a certain distance; and motion semantic alignment loss to align infrared features with trajectory features. These losses are weighted and summed to obtain the total loss. The gradient of the total loss with respect to the learnable parameters in each network module is calculated using the backpropagation algorithm. The parameters of the multimodal behavior encoding network, the relevant memory slots in the behavior pattern memory, and the trajectory decoder and posterior encoder in the conditional trajectory decoding network are updated along the gradient descent direction. Through iterative updates using a large number of training samples, the accuracy of future trajectory predictions is gradually improved.

[0144] For example, to uniformly optimize the above network modules, a joint loss function was constructed during training, which is a weighted sum of multiple loss terms, specifically including the following parts: First, KL divergence loss This is used to constrain the posterior distribution of the latent space to approximate the standard Gaussian prior distribution, ensuring smooth and stable sampling of latent variables. Its calculation formula is: .

[0145] Second, Minimum-over-N regression loss The error is calculated only for the K predicted trajectories that are closest to the actual future trajectory. This encourages the model to have at least one prediction that covers the actual trajectory, thus avoiding model collapse. The calculation formula is as follows: .

[0146] Third, margin diversity loss Force the minimum spacing between different predicted trajectories Meters, to prevent trajectory convergence, the calculation formula is: for each trajectory pair, the distance between them is less than... A penalty is imposed on a portion of the meter, calculated using the following formula: .

[0147] Fourth, loss of potential diversity of intentions This approach strengthens feature differentiation at the latent variable level, forcing the decoder to fully utilize the information from latent variables, avoid ignoring latent variable features, and improve the discriminative power of the latent space representation. The calculation formula is as follows: .

[0148] Fifth, the motion semantic alignment loss calculated in step e2 above. and the behavior pattern alignment loss calculated in step e3. Entropy regularization loss .

[0149] The above loss terms are summed according to preset weight coefficients to obtain the total loss for this training round: The weight parameters are set as follows: .

[0150] Subsequently, the total loss was calculated using the backpropagation algorithm. The gradients of the learnable parameters in the multimodal behavior encoding network and the conditional trajectory decoding network are calculated, and these parameters are updated along the gradient descent direction. For the behavior pattern memory, the EMA method in step e3 is continued to be used for updating. By traversing the entire dataset and repeating multiple training rounds (e.g., 100 rounds, or stopping early when the average displacement error on the validation set shows no improvement for 10 consecutive rounds), the error between the predicted trajectory and the true trajectory is gradually reduced, thereby improving the trajectory prediction accuracy of the model.

[0151] In the above implementation, by introducing the posterior encoding distribution of the real future trajectory as an additional supervision signal during the training phase, the multimodal behavior encoding network, behavior pattern memory, and conditional trajectory decoding network are integrated into a unified end-to-end optimization framework. After obtaining the sample behavior pattern vector, the latent variables corresponding to each behavior pattern vector are sampled based on the posterior distribution of the real future trajectory. Then, the latent variables, sample behavior pattern vectors, and sample behavior representation vectors are fused into a conditional input for trajectory decoding, so that trajectory generation during the training process can be guided by both behavior pattern priors and real future information. Finally, with the difference between the training trajectory prediction result and the real future trajectory as the unified goal, the parameters of the encoding, retrieval, and decoding stages are adjusted simultaneously to promote the collaborative optimization and mutual adaptation of each module, thereby systematically improving the accuracy and generation quality of the entire prediction framework.

[0152] like Figure 3 As shown, the overall architecture of the prediction model adopted in this application comprises two core network modules: a motion semantic gating fusion network and a behavior pattern conditional generation network. Historical trajectory sequences and infrared image sequences are input into the motion semantic gating fusion network to complete multimodal feature gating fusion and temporal encoding, outputting a behavior representation vector representing the UAV's flight state. This vector is then input into the behavior pattern conditional generation network, where it is searched and matched against typical flight patterns using a behavior pattern memory, and then multiple future trajectories are generated by the conditional trajectory decoder. During the training phase, a posterior encoding branch of the real future trajectories is additionally introduced, working in conjunction with the loss function to optimize the parameters of the entire model.

[0153] This embodiment also provides a drone trajectory prediction device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0154] This embodiment provides a drone trajectory prediction device, such as... Figure 4 As shown, it includes: The first acquisition module 401 is used to acquire the historical trajectory sequence and infrared image sequence of the target UAV within the observation time window; The first encoding module 402 is used to fuse and encode the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence using a multimodal behavior encoding network to obtain the behavior representation vector of the target UAV. The behavior representation vector is used to represent the motion state and flight intention of the target UAV. The first retrieval module 403 is used to retrieve multiple behavior pattern vectors that match the behavior representation vectors from a pre-built behavior pattern memory. The prediction module 404 is used to perform trajectory decoding using a conditional trajectory decoding network, with each behavior pattern vector as the generation condition, to generate multiple trajectory prediction results for the target UAV.

[0155] In some alternative implementations, the first encoding module 402 includes: The fusion submodule is used to perform gated fusion of trajectory features and infrared visual features using a multimodal behavior coding network to obtain a fused feature sequence; The encoding submodule is used to perform temporal encoding on the fused feature sequence to obtain the behavior representation vector.

[0156] In some alternative implementations, the fusion submodule includes: The first encoding unit is used to encode the trajectory features using a trajectory encoder to obtain a trajectory feature vector; The second encoding unit is used to encode infrared visual features using an infrared visual encoder to obtain an infrared visual feature vector. The stitching unit is used to stitch together the trajectory feature vector and the infrared visual feature vector to obtain the stitched vector; The mapping unit is used to map the spliced ​​vector into fusion weights using an adaptive gating network, and to perform element-wise weighted summation of the trajectory feature vector and the infrared visual feature vector according to the fusion weights to obtain the fusion feature sequence. The multimodal behavior coding network includes a trajectory encoder, an infrared vision encoder, and an adaptive gating network.

[0157] In some alternative implementations, the encoding submodule includes: The third encoding unit is used to encode the fused feature sequence using a cyclic temporal encoder and extract the hidden state of the last time step within the observation time window; Transformation unit is used to perform dimensionality transformation on the hidden state to obtain the behavior representation vector; Among them, the multimodal behavior coding network includes a cyclic timing encoder.

[0158] In some alternative implementations, the first encoding module 402 further includes: The first processing submodule is used to take the absolute coordinates of the last frame of the historical trajectory within the observation time window as the observation endpoint, and subtract the absolute coordinates of the observation endpoint from the absolute coordinates of each time step in the historical trajectory sequence to obtain the relative trajectory sequence. The second processing submodule is used to subtract the relative coordinates of the previous time step from the relative coordinates of the next time step for each pair of adjacent time steps in the relative trajectory sequence, so as to obtain the velocity between adjacent time steps. The first combination submodule is used to obtain the velocity intensity by taking the modulus of the velocity, and to combine the relative coordinates, velocity and velocity intensity of the next time step in adjacent time steps to obtain the trajectory features.

[0159] In some alternative implementations, the first encoding module 402 further includes: The acquisition submodule is used to obtain the projection coordinates of the target UAV in the infrared image for any frame of the infrared image sequence. The cropping submodule is used to crop out a square region from the infrared image, centered at the projection coordinates and with a side length of a preset pixel size; The second combination submodule is used to combine the square regions corresponding to each frame of infrared images in chronological order to obtain infrared visual features.

[0160] In some alternative implementations, the first retrieval module 403 includes: The first mapping submodule is used to embed and map the behavior representation vector to obtain the pattern query vector; The determination submodule is used to determine the similarity between the pattern query vector and each pattern vector in the behavior pattern memory, and select the target pattern vector with the highest similarity to a preset number of patterns. The second mapping submodule is used to project and map the target pattern vector to obtain a preset number of behavior pattern vectors.

[0161] In some alternative implementations, the drone trajectory prediction device further includes: The second acquisition module is used to acquire the sample UAV's historical trajectory sequence, sample infrared image sequence, and real future trajectory within the sample observation time window; The second encoding module is used to fuse and encode the trajectory features of the sample UAV's historical trajectory sequence and the infrared visual features of the sample infrared image sequence using a multimodal behavior encoding network, so as to obtain the sample UAV's behavior representation vector. The second retrieval module is used to retrieve multiple sample behavior pattern vectors that match the sample behavior representation vectors from the behavior pattern memory. The sampling module is used to obtain the posterior distribution of the real future trajectory after posterior encoding, and to sample the corresponding latent variables from the posterior distribution for each sample behavior pattern vector. The fusion module is used to fuse each sample behavior pattern vector, sample behavior representation vector, and latent variables corresponding to the sample behavior pattern vector to obtain the fusion condition vector corresponding to each sample behavior pattern vector. The decoding module is used to decode the trajectory of each fusion condition vector to obtain the training trajectory prediction result; The adjustment module is used to adjust the parameters of the multimodal behavior encoding network, behavior pattern memory, and conditional trajectory decoding network based on the difference between the training trajectory prediction results and the actual future trajectory.

[0162] The UAV trajectory prediction device provided in this application can execute the UAV trajectory prediction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0163] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0164] The following is a detailed reference. Figure 5 This diagram illustrates a suitable structural schematic for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0165] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0166] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the UAV trajectory prediction method of embodiments of this application.

[0167] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0168] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the UAV trajectory prediction method shown in the above embodiments is implemented.

[0169] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0170] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for predicting the trajectory of an unmanned aerial vehicle (UAV), characterized in that, The method includes: Acquire the historical trajectory sequence and infrared image sequence of the target UAV within the observation time window; Using a multimodal behavior coding network, the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence are fused and encoded to obtain the behavior representation vector of the target UAV. The behavior representation vector is used to represent the motion state and flight intention of the target UAV. In a pre-built behavioral pattern memory, multiple behavioral pattern vectors that match the behavioral representation vector are retrieved. Using a conditional trajectory decoding network, trajectory decoding is performed with each of the aforementioned behavior pattern vectors as generation conditions to generate multiple trajectory prediction results for the target UAV.

2. The method according to claim 1, characterized in that, The method of using a multimodal behavior coding network to fuse and encode the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence to obtain the behavior representation vector of the target UAV includes: The trajectory features and the infrared visual features are gated and fused using the multimodal behavior coding network to obtain a fused feature sequence. The fused feature sequence is temporally encoded to obtain the behavior representation vector.

3. The method according to claim 2, characterized in that, The step of using the multimodal behavior coding network to perform gated fusion of the trajectory features and the infrared visual features to obtain a fused feature sequence includes: The trajectory features are encoded using a trajectory encoder to obtain a trajectory feature vector; The infrared visual features are encoded using an infrared visual encoder to obtain an infrared visual feature vector; The trajectory feature vector and the infrared visual feature vector are concatenated to obtain a concatenated vector; An adaptive gating network is used to map the spliced ​​vector into fusion weights, and the trajectory feature vector and the infrared visual feature vector are summed element-wise according to the fusion weights to obtain the fusion feature sequence. The multimodal behavior coding network includes the trajectory encoder, the infrared visual encoder, and the adaptive gating network.

4. The method according to claim 2, characterized in that, The step of temporally encoding the fused feature sequence to obtain the behavior representation vector includes: The fused feature sequence is encoded using a cyclic temporal encoder to extract the hidden state of the last time step within the observation time window; The hidden state is transformed to obtain the behavior representation vector; The multimodal behavior coding network includes the cyclic timing encoder.

5. The method according to claim 1, characterized in that, The process of determining the trajectory features of the historical trajectory sequence includes: Using the absolute coordinates of the last frame of the historical trajectory within the observation time window as the observation endpoint, the absolute coordinates of each time step in the historical trajectory sequence are subtracted from the absolute coordinates of the observation endpoint to obtain the relative trajectory sequence. For each pair of adjacent time steps in the relative trajectory sequence, the relative coordinates of the previous time step are subtracted from the relative coordinates of the subsequent time step to obtain the velocity between the adjacent time steps. The velocity intensity is obtained by taking the modulus of the velocity. The relative coordinates of the next time step in the adjacent time steps, the velocity, and the velocity intensity are combined to obtain the trajectory features.

6. The method according to claim 1, characterized in that, The process of determining the infrared visual features of the infrared image sequence includes: For any frame of infrared image in the infrared image sequence, obtain the projection coordinates of the target UAV in the infrared image; A square region centered at the projected coordinates and with a side length of a preset pixel size is cropped from the infrared image; The infrared visual features are obtained by combining the square regions corresponding to each frame of the infrared image in chronological order.

7. The method according to claim 1, characterized in that, The step of retrieving multiple behavior pattern vectors that match the behavior representation vector from the pre-built behavior pattern memory includes: The behavior representation vector is embedded and mapped to obtain the pattern query vector; Determine the similarity between the pattern query vector and each pattern vector in the behavior pattern memory, and select the target pattern vector with the highest similarity (preset number). The target pattern vector is projected and mapped to obtain the preset number of behavior pattern vectors.

8. The method according to claim 1, characterized in that, The method further includes: Acquire the historical trajectory sequence, infrared image sequence, and true future trajectory of the sample drone within the sample observation time window; Using the multimodal behavior coding network, the trajectory features of the sample historical trajectory sequence and the infrared visual features of the sample infrared image sequence of the sample UAV are fused and encoded to obtain the sample behavior representation vector of the sample UAV. In the behavior pattern memory, multiple sample behavior pattern vectors that match the sample behavior representation vector are retrieved. Obtain the posterior distribution of the true future trajectory after posterior encoding, and sample the corresponding latent variables from the posterior distribution for each sample behavior pattern vector; Each of the sample behavior pattern vectors, the sample behavior representation vectors, and the latent variables corresponding to the sample behavior pattern vectors are fused to obtain the fusion condition vectors corresponding to each of the sample behavior pattern vectors. Trajectory decoding is performed on each of the fusion condition vectors to obtain the training trajectory prediction results; Based on the difference between the training trajectory prediction results and the actual future trajectory, the parameters of the multimodal behavior encoding network, the behavior pattern memory, and the conditional trajectory decoding network are adjusted.

9. A drone trajectory prediction device, characterized in that, The device includes: The first acquisition module is used to acquire the historical trajectory sequence and infrared image sequence of the target UAV within the observation time window; The first encoding module is used to fuse and encode the trajectory features of the historical trajectory sequence and the infrared visual features of the infrared image sequence using a multimodal behavior encoding network to obtain the behavior representation vector of the target UAV. The behavior representation vector is used to represent the motion state and flight intention of the target UAV. The first retrieval module is used to retrieve multiple behavior pattern vectors that match the behavior representation vector from a pre-built behavior pattern memory. The prediction module is used to perform trajectory decoding using a conditional trajectory decoding network, with each of the aforementioned behavior pattern vectors as the generation conditions, to generate multiple trajectory prediction results for the target UAV.

10. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory storing computer instructions, and the processor executing the computer instructions to perform the UAV trajectory prediction method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the unmanned aerial vehicle trajectory prediction method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, Includes computer instructions for causing a computer to execute the unmanned aerial vehicle trajectory prediction method according to any one of claims 1 to 8.