A target intelligent intention prediction method based on data-knowledge fusion
Patent Information
- Application Number
- CN202611072596.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-10-09
AI Technical Summary
[0004]虽然现有研究在意图推理方面取得了一定进展,但仍存在以下不足:其一,多数方法仍以纯数据驱动为主,侧重从轨迹序列中学习统计模式,而对场景语义、行为规则和趋势知识的利用不足;其二,在轨迹早期观测阶段,由于可用轨迹信息有限,深度模型容易出现判别不稳定、置信度不足的问题;其三,现有融合策略多采用固定权重或简单拼接方式,难以根据不同样本状态和不同观测阶段自适应调整数据证据与知识证据的贡献比例
[0019]本申请通过数据-知识动态融合框架,在完整观测条件下实现高精度的低空飞行目标意图识别,并在早期观测比例极低时,仍能保持较高的判别能力,从而能够有效缓解因轨迹信息不足导致的性能下降问题。此外,本申请还能够通过动态融合模块的自适应调节机制,根据样本状态合理分配数据证据与知识证据的权重,并通过规则知识推理模块增强模型在观测不充分条件下的判别稳定性与可解释性,从而在意图推理的准确性、鲁棒性和可解释性方面获得提升。
Smart Images

Figure CN122885818A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of low-altitude flight target recognition technology, specifically involving a target intelligent intent prediction method based on data-knowledge fusion. Background Technology
[0002] With the increasing prevalence of low-altitude flights, drones and other low-altitude targets are being used more widely in scenarios such as inspection, delivery, security monitoring, and emergency response. Intent reasoning for low-altitude targets is gradually becoming a key issue in intelligent sensing and airspace management. Intent reasoning aims to determine the potential mission objective or behavioral category of a target based on its temporal motion state, spatial position changes, and scene interaction relationships during flight, such as inspection, delivery, loitering, approaching sensitive areas, and returning to base. This task is of great significance for improving low-altitude situational awareness, enabling abnormal behavior early warning, and supporting autonomous decision-making.
[0003] In recent years, scholars both domestically and internationally have conducted extensive research on target intent reasoning. Early methods primarily relied on manually designed features and traditional machine learning models, classifying target behavior through kinematic features such as speed, heading angle, and acceleration. However, these methods had limited ability to characterize complex temporal dependencies and nonlinear behavioral patterns. Subsequently, with the development of deep learning, intent recognition methods based on recurrent neural networks, including Long Short-Term Memory (LSTM) networks and Bidirectional LSM networks, have been widely applied. Furthermore, some studies have attempted to introduce prior knowledge, scene constraints, or rule-based reasoning mechanisms to enhance the semantic expressiveness and interpretability of the models.
[0004] While existing research has made some progress in intent reasoning, it still has the following shortcomings: First, most methods are still purely data-driven, focusing on learning statistical patterns from trajectory sequences, while making insufficient use of scene semantics, behavioral rules, and trend knowledge; Second, in the early observation stage of trajectories, due to the limited available trajectory information, deep models are prone to problems such as unstable discrimination and insufficient confidence; Third, existing fusion strategies mostly adopt fixed weights or simple splicing methods, making it difficult to adaptively adjust the contribution ratio of data evidence and knowledge evidence according to different sample states and different observation stages.
[0005] Therefore, how to improve the robustness and interpretability of the model in early recognition scenarios while ensuring recognition accuracy remains an urgent problem to be solved in current research. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this application provides a target intelligent intent prediction method based on data-knowledge fusion. This application can effectively improve the accuracy, robustness, and interpretability of low-altitude flight target intent prediction, and can maintain stable and reliable prediction performance, especially under conditions of incomplete trajectory observation, early observation stage, and limited training samples.
[0007] To achieve the above objectives, this application provides the following technical solution:
[0008] A target intelligent intent prediction method based on data-knowledge fusion is disclosed. The method includes: acquiring flight trajectory information and scene information of a low-altitude flying target within an observation time window, wherein the flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration; and the scene information includes, for example, the target and its starting point, sensitive areas, delivery points, bases, and inspection channels; preprocessing the flight trajectory information and scene information; constructing a low-altitude flying target intelligent intent prediction model and training the model, wherein the model adaptively adjusts the fusion weights of data-driven information and rule knowledge information based on prediction confidence, prediction entropy, observation completeness, rule coverage, and rule consistency through a dynamic gating fusion mechanism to achieve collaborative decision-making of temporal pattern learning and rule reasoning under different observation conditions; and inputting the preprocessed flight trajectory information and related scene information into the trained low-altitude flying target intelligent intent prediction model to predict the intent category of the low-altitude flying target.
[0009] Optionally, the flight trajectory information and scene information are preprocessed, including: repairing missing or abnormal data in the flight trajectory information and scene information; using uniform length truncation and zero padding to align the repaired flight trajectory information and scene information into a sequence and mark the effective length; and normalizing the continuous features in the repaired, aligned and marked sequence.
[0010] Optionally, the intelligent intent prediction model for low-altitude flying targets includes: a data-driven module, a rule-based knowledge reasoning module, and a dynamic fusion module. The data-driven module is used to extract trajectory temporal features from the flight trajectory information of the flying target and output a data-learned category probability distribution as data evidence. The rule-based knowledge reasoning module is used to perform explicit rule mapping based on behavioral events, scene relationships, and trend rules, and output an interpretable category probability distribution as intent category evidence. The dynamic fusion module is used to adaptively generate fusion coefficients to dynamically balance the contribution ratio of data evidence and intent category evidence.
[0011] Optionally, the data-driven module includes: an input layer, a TCN encoder, an attention mechanism layer, and a classification layer arranged sequentially. The TCN encoder is used to extract local temporal dependencies and multi-scale motion patterns from the flight trajectory information of the flight target input by the input layer. The attention mechanism layer is used to enhance the representation of key moments and key segments in the local temporal dependencies and multi-scale motion patterns, and output a weighted representation vector. The classification layer is used to map the weighted representation vector to the category space and output the data-side category discrimination probability.
[0012] Optionally, the rule knowledge reasoning module includes, in sequence: a rule event extraction submodule, a feature extraction submodule, a rule mapping submodule, and a normalization output submodule.
[0013] Optionally, the dynamic fusion module includes, in sequence: a gated feature extraction submodule, a constrained gated network submodule, and a fusion coefficient generation and output submodule. The gated feature extraction submodule extracts gated feature vectors from the outputs of the data-driven module and the rule-based knowledge reasoning module. The constrained gated network submodule takes the gated feature vectors as input and calculates intermediate variables through a linearly constrained mapping. The fusion coefficient generation and output submodule maps the intermediate variables to a preset interval to generate dynamic fusion coefficients, and performs weighted fusion of the probability distributions of the data-driven module and the rule-based knowledge reasoning module to output the final intent reasoning result.
[0014] This application also provides a target intelligent intent prediction device based on data-knowledge fusion. The device includes: an acquisition module for acquiring flight trajectory information and scene information of a low-altitude flying target within an observation time window, wherein the flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration; and the scene information includes, for example, the target and starting point, sensitive area, delivery point, base, and inspection channel; a preprocessing module for preprocessing the flight trajectory information and scene information; a model construction and training module for constructing a low-altitude flying target intelligent intent prediction model and training the model. This model, through a dynamic gating fusion mechanism, adaptively adjusts the fusion weights of data-driven information and rule knowledge information based on prediction confidence, prediction entropy, observation completeness, rule coverage, and rule consistency to achieve collaborative decision-making of temporal pattern learning and rule reasoning under different observation conditions; and an output module for inputting the preprocessed flight trajectory information and related scene information into the trained low-altitude flying target intelligent intent prediction model to predict the intent category of the low-altitude flying target.
[0015] Optionally, the preprocessing module includes: a repair submodule for repairing missing or abnormal data in flight trajectory information and scene information; a marking submodule for aligning the repaired flight trajectory information and scene information into a sequence and marking the effective length using uniform length truncation and zero padding; and a normalization submodule for normalizing continuous features in the repaired, aligned and marked sequence.
[0016] This application also provides a storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in the preceding claim.
[0017] This application also provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method as described in any of the preceding claims.
[0018] Compared with the prior art, the beneficial effects of this application are as follows:
[0019] This application utilizes a data-knowledge dynamic fusion framework to achieve high-precision low-altitude flight target intent recognition under complete observation conditions. Even with extremely low early observation ratios, it maintains high discrimination capability, effectively mitigating performance degradation caused by insufficient trajectory information. Furthermore, this application employs an adaptive adjustment mechanism in the dynamic fusion module to rationally allocate the weights of data evidence and knowledge evidence based on sample states. A rule-based knowledge reasoning module enhances the model's discrimination stability and interpretability under insufficient observation conditions, thereby improving the accuracy, robustness, and interpretability of intent reasoning. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a target intelligent intent prediction method based on data-knowledge fusion provided in this application;
[0021] Figure 2 This application provides a data-knowledge dynamic fusion intent reasoning framework.
[0022] Figure 3 This is a schematic diagram of the structure of the data-driven module provided in this application;
[0023] Figure 4 This is a schematic diagram of the structure of the rule knowledge reasoning module provided in this application;
[0024] Figure 5 This application provides a schematic diagram of the structure of the dynamic weighted fusion module;
[0025] Figure 6 This is a schematic diagram showing the comparison of model performance under complete observation conditions provided in this application;
[0026] Figure 7 This is a schematic diagram showing the performance comparison of the models under local observation conditions provided in this application;
[0027] Figure 8 This is a diagram illustrating the performance comparison of the models under small sample conditions provided in this application;
[0028] Figure 9 This is a schematic diagram showing the performance comparison of ablation experiments provided in this application. Detailed Implementation
[0029] Specific embodiments of this application will now be described in detail with reference to the accompanying drawings. While specific embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0030] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used to refer to the same component. This specification and claims do not distinguish components based on differences in terminology, but rather on differences in function. The terms "comprising" or "including" used throughout the specification and claims are open-ended and should be interpreted as "comprising but not limited to." The following descriptions in the specification are preferred embodiments for carrying out this application; however, these descriptions are for the purpose of understanding the general principles of the specification and are not intended to limit the scope of this application. The scope of protection of this application shall be determined by the appended claims.
[0031] To facilitate understanding of the embodiments of this application, the following will provide further explanation and description with reference to the accompanying drawings and specific embodiments, and the accompanying drawings do not constitute a limitation on the embodiments of this application.
[0032] Figure 1 This is a flowchart illustrating a target intelligent intent prediction method based on data-knowledge fusion provided in this application, as shown below. Figure 1 As shown, the method includes the following steps:
[0033] S100: Acquire flight trajectory information and scene information of low-altitude flying targets within the observation time window. Flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration. Scene information includes, for example, the target and origin, sensitive areas, delivery points, bases, and inspection channels.
[0034] S200: Preprocess the flight trajectory information and scene information;
[0035] S300: Construct an intelligent intent prediction model for low-altitude flying targets and train the model;
[0036] S400: Input the preprocessed flight trajectory information and related scene information into the trained low-altitude flight target intelligent intent prediction model to predict the intent category of the low-altitude flight target.
[0037] This application acquires flight trajectory and scene information of low-altitude flying targets, preprocesses it to construct and train an intelligent intent reasoning model, and finally predicts the intent category of low-altitude flying targets, forming a complete data-knowledge fusion prediction scheme. This scheme can achieve high-precision intent recognition under complete observation conditions, and can improve the robustness and stability of discrimination through an adaptive fusion mechanism in the early observation stage with limited trajectory information and in small sample scenarios. This effectively alleviates the problem of severely degraded performance of traditional methods for low-altitude flying target intent recognition when data is insufficient.
[0038] In another exemplary embodiment, step S200 involves preprocessing the flight trajectory information and scene information, including the following steps:
[0039] S201: Repair missing or abnormal data in flight trajectory and scene information;
[0040] In this step, when there are missing observation data at individual moments in the flight trajectory or scene information, a nearest-time interpolation strategy is adopted to use valid observation values before and after the missing point for linear interpolation to complete the data. When outliers exceeding the normal kinematic range or physical constraints are detected, a threshold pruning strategy is adopted to limit the outliers to a preset reasonable range. By combining nearest-time interpolation with threshold pruning, the quality problems in the original data are effectively repaired, providing clean and reliable input data for subsequent sequence alignment and normalization.
[0041] S202: The repaired flight trajectory information and scene information are sequence aligned and marked with effective length using uniform length truncation and zero padding;
[0042] In this step, this embodiment employs uniform length truncation and zero-padding to align the repaired flight trajectory and scene information sequences and mark their effective lengths. Since the observed trajectory lengths of different low-altitude flight targets may vary, the sequences are first uniformly truncated according to a preset standard observation length: for sequences exceeding the standard value, the initial continuous segments are retained; for sequences shorter than the standard value, zero-padding is applied at the ends to bring them to the standard length. Simultaneously with the alignment operation, an effective length marker vector is generated to clearly identify the positions of the actual observed data and the padded data in the sequence, ensuring that the padded portion does not participate in subsequent feature calculations, thereby guaranteeing the accuracy of model training and inference.
[0043] S203: Normalize the continuous features in the repaired, aligned and labeled sequence.
[0044] In this step, the kinematic features of the flight trajectory information, such as position, speed, altitude, heading angle, and acceleration, as well as the distance relationships between the target and key elements such as the starting point, sensitive areas, delivery points, bases, and inspection channels in the scene information, are all continuous features with different dimensions. Therefore, to eliminate the influence of different dimensions on model training, this embodiment uses the maximum-minimum normalization or standard deviation normalization method to map the numerical range of all continuous features to a unified interval, so that each feature has a similar scale during model training, thereby improving the model convergence speed and training stability.
[0045] In another exemplary embodiment, in step S200, as Figure 2 As shown, the intelligent intent prediction model for low-altitude flying targets includes a data-driven module, a rule-based knowledge reasoning module, and a dynamic fusion module. The data-driven module extracts trajectory temporal features from the flight trajectory information of the flying target and outputs a data-learned category probability distribution as data evidence. The rule-based knowledge reasoning module performs explicit rule mapping based on behavioral events, scene relationships, and trend rules, and outputs an interpretable category probability distribution as intent category evidence. The dynamic fusion module adaptively generates fusion coefficients to dynamically balance the contribution ratio of data evidence and intent category evidence.
[0046] Below, this application will provide a detailed description of the structure and logical processing of the above three modules.
[0047] like Figure 3As shown, the data-driven module includes, in sequence, an input layer, a motion state differential enhancement layer, a multi-scale extended temporal coding layer, a scene relationship-guided attention layer, an effective length mask layer, and a classification output layer. The system comprises the following layers: an input layer receives the preprocessed flight trajectory sequence and corresponding scene relationship features; a motion state differential enhancement layer constructs velocity changes, heading angle changes, altitude changes, and key area distance changes based on the trajectory states at adjacent time points to characterize the motion change trend of low-altitude flying targets within the observation time window; a multi-scale dilated temporal coding layer extracts short-term maneuver features, medium-term trend features, and long-term behavioral pattern features from the trajectory sequence through causal convolutional branches with different dilation rates, and outputs a multi-scale temporal representation; a scene relationship-guided attention layer combines the multi-scale temporal representation with scene relationship features between the target and sensitive areas, delivery points, bases, and inspection channels to generate attention weights for each time step, thereby strengthening key trajectory segments related to intent discrimination; an effective length mask layer masks zero-padding time steps based on the effective length marker generated during sequence alignment, preventing padding data from participating in attention calculation and classification; and a classification output layer maps the weighted trajectory representation to the intent category space and outputs the category probability distribution of the data-driven branches.
[0048] Furthermore, the specific logic processing flow of the data-driven module with the above structure is as follows:
[0049] First, the preprocessed flight trajectory sequence and scene relationship features are input through the input layer. Let the first... The trajectory state characteristics at each time point are as follows: ,in, This includes the target's position, velocity, heading angle, altitude, acceleration, and spatial relationship characteristics between the target and key scene areas.
[0050] Next, to enhance the model's ability to express the trend of motion changes, the motion state differential enhancement layer constructs differential enhancement features based on the trajectory states at adjacent time points, which are represented as follows:
[0051]
[0052] in, Indicates time Trajectory state characteristics Relative to time Trajectory state characteristics The changes specifically include changes in velocity, heading angle, altitude, and distance from the target to the critical area. Furthermore, the trajectory state characteristics... With differential enhancement features By combining these features, we obtain enhanced trajectory features:
[0053]
[0054] in, Indicates time Enhanced trajectory features, This indicates a feature splicing operation.
[0055] Through the above processing, the data-driven module can not only utilize the target's motion state at a single moment, but also explicitly obtain its motion change trend within the observation time window. This is because the motion state differential enhancement layer constructs differential enhancement features such as velocity change, heading angle change, altitude change, and key area distance change based on the trajectory state at adjacent moments, and integrates these features with the trajectory state features. The data is then spliced and fused to form an enhanced input representation that simultaneously contains instantaneous motion information and temporal evolution information. Trajectory state features. It reflects the static attributes of the target at each sampling moment, such as instantaneous position, velocity, and heading, while the differential enhancement features characterize the magnitude and direction of these attributes changing between adjacent moments. The combination of the two enables the model to perceive the continuous motion process of the target from the past to the present, rather than relying solely on the cross-sectional information of isolated moments. This allows the model to fully grasp the motion change trend of low-altitude flying targets within the entire observation time window.
[0056] Then, the enhanced temporal input features are... The input consists of a multi-scale dilated temporal coding layer. This layer comprises multiple causal convolutional branches with different dilation rates, used to extract enhanced temporal input features from different time scales. The characteristics of short-term maneuvering, medium-term trend, and long-term behavioral patterns in the data.
[0057] In a preferred embodiment, the multi-scale dilated temporal coding layer can be configured, for example, as four parallel causal convolutional branches, with dilation rates of the four branches being respectively... , , and The four expansion rates mentioned above can form multi-scale temporal receptive fields ranging from short to long. Among them, the expansion rate is... The branch is mainly used to capture short-term motion changes between adjacent sampling points, such as sudden velocity changes, small adjustments in heading angle, and other local maneuvering features; the expansion rate is The branch is used to capture continuous motion trends within a shorter time frame; the expansion rate is... The branch is used to extract mid-term flight behavior patterns, such as continuous approach to a key area or loitering within a local area; the expansion rate is... The branch is used to characterize task-level behavior patterns over a longer time period, such as stable flight along the inspection channel, continuous return to base, or continuous approach to the delivery point.
[0058] This embodiment preferably uses four causal convolution branches because four branches can achieve a better balance between receptive field coverage and model complexity. On the one hand, , , , The combination of expansion rates can cover common short-term maneuvers, medium-term trends, and long-term behavioral patterns in low-altitude flight target intent reasoning. On the other hand, compared to setting more branches, four branches can avoid the problems of excessive model parameters, increased feature redundancy, and decreased training stability. Therefore, when the length of the low-altitude flight target trajectory sequence is limited and the sampling frequency is relatively fixed, the four-branch structure can better balance feature representation ability and computational efficiency.
[0059] Furthermore, each of the above causal convolution branches has the same structure, including a causal padding layer, a one-dimensional dilated causal convolution layer, a normalization layer, a nonlinear activation layer, a random deactivation layer, and a residual connection layer. The causal padding layer ensures that the convolutional output at each time step depends only on the trajectory information of the current time step and previous time steps, avoiding the introduction of future time step information into the intention prediction, thus ensuring the temporal rationality of the prediction process. The one-dimensional dilated causal convolutional layer performs convolution operations on the enhanced temporal input features according to the corresponding dilation rate, capturing multi-scale motion features such as short-term maneuvers, medium-term trends, and long-term behavior patterns in parallel with an exponentially expanded receptive field, enabling the model to perceive a wider temporal context without increasing network depth. The normalization layer can effectively accelerate the model's training convergence and alleviate the internal covariate bias problem by stabilizing the feature distribution of each batch of samples. The nonlinear activation layer enhances the model's ability to express nonlinear motion patterns such as sudden velocity changes and sharp heading changes by introducing nonlinear transformations. The random deactivation layer reduces the risk of overfitting by probabilistically discarding neuron activation values, thereby improving the model's generalization robustness under small sample and local observation conditions. The residual connection layer is used to directly fuse the branch input with the convolutional output to alleviate the gradient decay problem in the deep temporal feature extraction process, thus supporting deeper temporal coding structures.
[0060] Let the first The dilation rate of each dilated causal convolution branch is... Its output is represented as:
[0061]
[0062] in, Indicates time Enhanced trajectory features, Indicates the expansion rate The causal convolution branch, Indicates the first Temporal feature representations of the outputs of each causal convolutional branch; , respectively corresponding to expansion rates , , and The four causal convolution branches.
[0063] After feature concatenation or weighted fusion of the outputs of each causal convolution branch, a multi-scale temporal representation is obtained:
[0064]
[0065] in, This represents the short-term maneuver feature representation of the output of a causal convolution branch with an expansion rate of 1; This represents the short-term trend characteristics of the output of a causal convolution branch with an expansion rate of 2. This represents the mid-term behavioral feature representation of the output of a causal convolution branch with an expansion rate of 4. This represents the long-term behavioral pattern feature representation of the output of a causal convolution branch with an expansion rate of 8; This represents the multi-scale temporal representation after fusion.
[0066] Through the multi-scale structure described above, the model can simultaneously characterize the local maneuvering behavior of low-altitude flying targets and mission-level behavioral patterns over a longer time span. This is because the multi-scale expanded temporal coding layer uses convolutional branches with different expansion rates to extract features in parallel: branches with smaller expansion rates have narrower receptive fields and are sensitive to short-term maneuvering behaviors such as sudden speed changes and sharp heading changes; branches with larger expansion rates have wider receptive fields, covering a longer time span, thus extracting mission-level behavioral patterns such as continuously approaching delivery points and stable flight along inspection channels. After the outputs of each branch are fused, both fine-grained motion change information and coarse-grained behavioral evolution trends are preserved, enabling the model to perceive both immediate maneuvering features and understand long-term mission intentions within the same framework.
[0067] After obtaining multi-scale temporal representations, scene relationships guide the attention layer to further calculate the importance of each time step in intent discrimination by combining scene relationship features. Let the 1st time step be... The multi-scale hidden state corresponding to each time step is represented as follows: Scene relationship features are The attention score is then expressed as:
[0068]
[0069] in, Indicates the first Attention score at each time step; Indicates the first Multi-scale temporal hidden state representation corresponding to each time step; Indicates the first Scene relationship features corresponding to each time step; and Respectively represent and and The corresponding weight matrix; Indicates the bias term; This represents the transpose of the attention mapping vector, used to map the nonlinearly transformed feature vector to a scalar attention score; Indicates the time step number.
[0070] Furthermore, to avoid zero-padding time steps from participating in attention calculations, the effective length mask layer modifies the attention score output by the scene relationship-guided attention layer based on the effective length marker as follows:
[0071] set up For the first The effective length marker for the nth time step, when the nth time step When each time step is the actual observed data... When the first When the data is filled with zero at any given time. The corrected attention weights are expressed as follows:
[0072]
[0073] in, Indicates the first The attention weights are corrected for each time step; and They represent the first The time step and the first The effective length marker for each time step; and They represent the first The time step and the first Attention scores at each time step, guided by scene relationships, are output by the attention layer. Indicates the index for summation at time steps; This represents the natural exponential function.
[0074] With the above corrections, the time step corresponding to the data filling will not affect the attention allocation and subsequent classification results.
[0075] Therefore, the weighted representation vector of the trajectory sequence can be obtained as follows:
[0076]
[0077] in, This represents the trajectory representation vector after attention weighting, which can highlight key trajectory segments that contribute significantly to the current intent judgment, such as behavior segments like continuously approaching a sensitive area, entering the neighborhood of a delivery point, flying steadily along an inspection channel, or gradually approaching the base. Indicates the first Multi-scale temporal hidden state representation corresponding to each time step.
[0078] Finally, the weighted representation vector Inputting the classification output layer yields the following class prediction probability distribution for the data-driven module:
[0079]
[0080] in, This represents the category probability distribution output by the data-driven module; This represents the weight matrix of the classification output layer; This represents the bias term of the classification output layer; This represents the normalized exponential function. For the ... Class intent, whose data-side class prediction probability is expressed as ,in, Indicates the intent category index, , Total number of intent categories.
[0081] It should be noted that the data-driven module can learn the non-linear mapping relationship between "trajectory pattern - intent category" from a large number of samples and has a strong pattern extraction capability. However, in the intent reasoning stage under local observation conditions, the prediction may still be unstable due to insufficient trajectory information. Therefore, the rule knowledge reasoning module is needed to compensate for it.
[0082] like Figure 4 As shown, the rule-based knowledge reasoning module includes, in sequence, a sliding window statistics layer, a scene relationship calculation layer, a trend determination layer, a rule event activation layer, a rule mapping layer, a conflict suppression layer, and a normalization output layer. The sliding window statistics layer is used to slice the flight trajectory sequence according to a preset window length and to statistically analyze the window-level motion characteristics of low-altitude flying targets within each sliding window.
[0083] Specifically, let the trajectory sequence of a low-altitude flying target within the observation time window be: ,in, Indicates the first Trajectory state characteristics at each time step This represents the length of the trajectory sequence. Let the preset window length be... The sliding step size is Then the first A sliding window can be represented as:
[0084]
[0085] in, Indicates the first The starting time step index of each sliding window, and ; express The index of the next time step; Indicates the first The end time step index of each sliding window, and , This indicates that the preset window length is reduced by one.
[0086] It should be noted that, The frequency of trajectory sampling can be set according to the needs of the task, for example, 5 time steps, 10 time steps, or 20 time steps; when the trajectory sampling interval is 1 second, This indicates that each sliding window covers approximately 10 seconds of flight footage. In a preferred embodiment, Take 10 time steps, Five time steps are used to reduce redundant computation between adjacent windows while maintaining the expressive power of local behavior.
[0087] For the A sliding window The sliding window statistics layer collects the following window-level motion statistics features:
[0088] First, average speed, which characterizes the overall flight speed of the target within the current window, is calculated as follows:
[0089]
[0090] in, Indicates the first The average speed within a sliding window Indicates the first The speed at each time step Indicates the first The number of real valid trajectory points within each sliding window.
[0091] Second, the velocity change, used to characterize whether the target is accelerating or decelerating within the current window, is calculated as follows:
[0092]
[0093] in, express The corresponding velocity vector; express The corresponding velocity vector; Indicates the first The change in velocity at the beginning and end of each sliding window; when When the speed change exceeds a preset threshold, it indicates that the target has an accelerating trend within that window; when... When the value is less than the negative preset speed change threshold, it indicates that the target has a deceleration trend within the window.
[0094] Third, the average heading change, used to characterize the target's turning strength within the current window, is calculated as follows:
[0095]
[0096] in, Indicates the first Average heading change within a sliding window; Indicates the first The heading angle at each time step; Indicates the first The heading angle at each time step; This represents the normalization function for the heading angle difference, which is used to limit the heading angle difference to a preset angle range to avoid calculation errors caused by angle periodicity. Indicates the first The number of real valid trajectory points within a sliding window.
[0097] Fourth, the altitude change, used to characterize the target's ascent or descent trend within the current window, is calculated as follows:
[0098]
[0099] in, express Corresponding flight altitude; express Corresponding flight altitude; Indicates the first The change in flight altitude for each sliding window.
[0100] Fifth, trajectory spatial dispersion, used to characterize the spatial distribution range of the target within the current window, is calculated as follows:
[0101]
[0102] in, Indicates the first The spatial dispersion of the trajectory within a sliding window; Indicates the first The number of real valid trajectory points within each sliding window; Indicates the first The target location coordinates at each time step; Indicates the first The average of all valid position coordinates within a sliding window. The larger the value, the larger the spatial distribution range of the target within that window; The smaller the value, the more likely the target is to hover, linger at low speed, or circle locally within that window.
[0103] Sixth, the change in the position of the beginning and end of the window, which characterizes the overall displacement of the target within the current window, is calculated as follows:
[0104]
[0105] in, express Corresponding position coordinates; express Corresponding position coordinates; Indicates the first The change in the first and last positions of each sliding window; Let represent the Euclidean norm. When Smaller and trajectory space discreteness When the value is small, it indicates that the target may be in a state of partial lingering or hovering within that window; when... When the value is large and the directional change is small, it indicates that the target may be in a stable navigation state within that window.
[0106] Therefore, the first The window-level motion statistics corresponding to each sliding window can be represented as:
[0107]
[0108] in, Indicates the first The window-level motion statistical feature vectors of each sliding window are calculated. Through the above calculations, the original continuous trajectory can be transformed into a local motion state description oriented towards rule discrimination, providing structured input for the subsequent trend determination layer and rule event activation layer.
[0109] Through this layer, the original continuous trajectory can be transformed into a local motion state description for rule-based discrimination. This is because the sliding window statistical layer divides the long-term continuous trajectory into several time segments according to a preset window length, and calculates statistical features such as average velocity, velocity change, heading change, altitude change, trajectory spatial dispersion, and changes in the beginning and end positions of the window within each window. These window-level statistics can condense the densely sampled points within the window into several dimensions of motion features. This effectively filters out single-point noise and instantaneous fluctuations, and retains motion change information with discriminative significance at the window scale. Thus, the continuous numerical trajectory can be transformed into a structured input representation that is convenient for subsequent rule-based event activation layers to perform threshold comparisons and pattern recognition.
[0110] The scene relationship calculation layer is used to calculate the spatial relationship features between the low-altitude flying target and sensitive areas, delivery points, bases, and inspection channels based on the target's current location and preset scene elements. The specific calculation process is as follows:
[0111] Let the first The target position coordinates at each time step are Sensitive areas are The delivery point is located at The base is located at The center line of the inspection channel is Then the spatial distance between the target and each scene element can be expressed as:
[0112]
[0113]
[0114]
[0115]
[0116] in, Indicates the first The distance from the target to the sensitive area at each time step; Indicates the first The distance from the target to the delivery point at each time step; Indicates the first The distance from the target to the base at each time step; Indicates the first The shortest distance from the target to the centerline of the inspection channel at each time step; A distance function from a point to a region; Indicates the center line of the inspection channel Any sampling point on; Indicates the center line of the inspection channel Take the minimum value from the middle. This represents the Euclidean norm.
[0117] Furthermore, it is determined whether the target has entered the neighborhood of the corresponding critical area based on a preset neighborhood radius. Let the neighborhood radius of the sensitive area be... The radius of the delivery point's neighborhood is The radius of the base's neighborhood is The inspection channel width threshold is The corresponding region entry marker can then be represented as:
[0118]
[0119]
[0120]
[0121]
[0122] in, , , and They represent the first Whether the target at each time step is located within the sensitive area neighborhood, delivery point neighborhood, base neighborhood, and inspection channel range. The above thresholds can be set according to the specific scenario scale, for example, Take 100m, Take 50m, Take 50m, The value is 20m, but it can be adjusted according to the actual airspace monitoring radius and the scope of the mission area.
[0123] For the A sliding window The percentage of trajectory points of the target within the inspection channel range can be expressed as:
[0124]
[0125] in, Indicates the first The percentage of trajectory points of the target within the inspection channel within each sliding window. Indicates the first The number of truly valid trajectory points within each sliding window. Therefore, the number of... The scene relationship features corresponding to each time step can be represented as:
[0126]
[0127] in, Indicates the first Scene relationship feature vectors at each time step.
[0128] Through the above calculations, the scene relationship calculation layer can provide quantifiable spatial semantic basis for intent categories such as delivery, return, approach to sensitive areas, and inspection.
[0129] The trend determination layer is used to determine the continuous approach trend, continuous departure trend, or stable trend of the target relative to the sensitive area, delivery point, and base based on the change in the distance between the target and the key area within the previous and subsequent time windows. It also determines whether the target is in a stable corridor flight state based on the continuous proportion of the target in the inspection channel.
[0130] Specifically, for the first A sliding window The window is divided into a front window and a back window, and the distance from the target to the critical region is calculated separately. The average distance, where, This can represent a sensitive area, distribution point, or base. Average distance in the front half of the window. and the average distance of the second half window They are represented as follows:
[0131]
[0132]
[0133] in, Indicates the first The intermediate time step number of each sliding window This indicates the number of truly valid trajectory points within the first half of the window. This indicates the number of truly valid trajectory points within the second half of the window. Indicates the first Targets at each time step to key areas The distance.
[0134] Furthermore, the change in distance is defined as:
[0135]
[0136] in, Indicates the first Within a sliding window, the target is relative to the key area. The change in distance. When When the value is positive, it indicates that the latter half of the target window is closer to the critical region than the former half. ;when When the value is negative, it indicates that the latter half of the target window is further away from the critical region than the former half. .
[0137] Let the trend determination threshold be... The trend determination rules are as follows:
[0138]
[0139] in, Indicates the first Within a sliding window, the target is relative to the key area. The trend determination results; This indicates a continued approximation of the trend; This indicates a continued departure from the trend; This indicates that the trend remains stable. The setting can be adjusted according to the scene scale or distance variation ratio; for example, it can be 5m, or it can be 5% to 10% of the average distance of the first half of the window. In a preferred embodiment, Pick To ensure the stability of judgment in both close-range and long-range scenarios.
[0140] For a stable flight state within the inspection channel, let the threshold for the proportion of trajectory points within the channel be . The average heading change threshold is If the first The proportion of trajectory points where the target is located within the inspection channel range within each sliding window is not less than And the average heading change is not higher than If the target is within a stable corridor flight state within the sliding window, then the determination rule can be expressed as:
[0141]
[0142] in, Indicates the first Whether the target is in a stable corridor flight state within the sliding window. This indicates the percentage of trajectory points of the target within the inspection channel. Indicates the first Average heading change within a sliding window. A value of 0.8 is acceptable, indicating that at least 80% of the valid trajectory points within the window are located within the inspection channel. A value of 15° is acceptable, indicating a relatively small change in average heading within the window. Furthermore, when continuously... Each sliding window satisfies At that time, it can be determined that the target is in a continuous and stable corridor flight state, among which 2 or 3 are acceptable.
[0143] Through this layer of processing, the rule-based knowledge reasoning module can not only reflect the current location of the target, but also characterize the directionality of the target's behavior evolution over time.
[0144] The rule event activation layer is used to match various rule events in the preset rule base based on the features output by the sliding window statistics layer, scene relationship calculation layer, and trend determination layer, and generate corresponding rule activation intensities. The rule events mainly include three categories: first, motion state events, including hovering, low-speed loitering, sharp turns, deceleration, and altitude changes; second, scene approach events, including approaching sensitive areas, approaching delivery points, approaching the base, and flying along inspection corridors; and third, trend change events, including continuously approaching sensitive areas, continuously approaching delivery points, continuously approaching the base, and flying along stable corridors. Through this layer, continuous trajectory information is converted into a set of rule events with clear semantic meaning.
[0145] Specifically, the rule event activation layer compares window-level motion statistics, scene relationship features, and trend determination results with preset rule conditions; when the rule conditions are met, the corresponding rule event is activated; when the rule conditions are not met, the corresponding rule event is not activated or only generates a low activation intensity.
[0146] set up This is a truncation function used to restrict values to a range. Inside, it is defined as:
[0147]
[0148] Taking the "Approaching Sensitive Area" rule event as an example, this rule event requires two conditions to be met simultaneously: First, the target has a tendency to approach a sensitive area within the current window; second, the target has entered or is approaching the neighborhood of the sensitive area. The activation strength of this rule event can be expressed as:
[0149]
[0150] in, Indicates the first The activation intensity of the "near sensitive area" rule event within a sliding window; Indicates the first Whether the target within each sliding window enters the neighborhood of a sensitive area; Indicates the first The change in distance between the target and the sensitive area within a sliding window; This indicates the threshold for determining a trend towards a sensitive area. When a target gradually approaches a sensitive area and has entered its neighborhood, Approaching 1; when the target is far from the sensitive area or has not entered the vicinity of the sensitive area, Close to 0.
[0151] in, You can access the marker from within the window:
[0152]
[0153] Through the above methods, continuous trajectory information can be converted into a set of rule events with clear semantic meaning. Similarly, the "approaching delivery point" rule event can be determined based on the delivery point's neighborhood markers and the approach trend; the "returning home" rule event can be determined based on the base's neighborhood markers and the approach trend; the "stable corridor flight" rule event can be determined based on the proportion of trajectory points within the inspection channel and the average heading change; and the "low-speed loitering" rule event can be determined based on the average speed, trajectory spatial dispersion, and the changes in the window's beginning and end positions. Therefore, the rule event activation layer can provide interpretable and quantifiable evidence of rule activation for the subsequent rule mapping layer.
[0154] The rule mapping layer maps the rule events and their activation strengths output by the rule event activation layer, combined with the scene relationship features output by the scene relationship calculation layer and the trend judgment results output by the trend judgment layer, to the intent category space based on a predefined knowledge rule base, generating initial support scores for each intent category. The knowledge rule base is divided into five categories: single-event rules, scene rules, trend rules, combination rules, and suppression rules. Single-event rules describe the direct support of local salient behaviors for intent categories; scene rules characterize the category evidence provided by the spatial relationship between the target and key areas; trend rules reflect the directional characteristics of target behavior evolving over time; combination rules strengthen the semantic discrimination ability when multiple pieces of evidence are jointly satisfied; and suppression rules weaken the support strength of mutually exclusive categories to reduce the risk of misjudgment caused by rule conflicts. The specific mapping relationship between rule knowledge and intent categories is shown in Table 1.
[0155] Table 1. Mapping of Rule Knowledge and Intent Categories Single event rule hover Average speed below hovering threshold wander Characterized by localized stop or near-stationary flight Single event rule low_speed_loiter Low speed and small spatial discrete range wander Characterizing small-range, low-speed hovering Single event rule approaching_sensitive_zone The distance to the sensitive area decreases and it enters the near zone. Sensitive areas are close Characterization of active approximation of sensitive regions Single event rule Returning to base The distance to the base has decreased and the area has entered the return flight range. Return The behavior clearly indicates a return journey. Single event rule patrol_corridor_cruise The proportion of trajectory points in the corridor is relatively high. Inspection Characterizes patrol along the inspection channel Scene rules near_Delivery_point Enter the neighborhood of the delivery point Delivery Provide spatial evidence of the mission objective point Scene rules near_base Entering the base's neighborhood Return Provide evidence of the return destination Trend Rules Approaching_delivery_point_trend The second half is closer to the delivery point than the first half. Delivery Depicting the trend of continuously approaching delivery points Trend Rules approaching_base_trend The second half is closer to the base than the first half. Return Depicting the continuous return trend Trend Rules stable_in_corridor The proportion of corridors is high and remains stable. Inspection Depicting stable inspection trajectories Combination rules delivery_combo Approaching the delivery point + slowing down Delivery Jointly depict delivery actions Combination rules strong_return_combo Approaching the base + returning to base event Return Jointly enhance the semantics of return. Combination rules sensitive_probe_combo Approaching a sensitive area + sharp turn Sensitive areas are close Demonstrates tentative approach behavior Suppression rules return_Suppresses_delivery Return to base event + approaching base trend Suppress delivery To avoid misclassifying a return flight as a delivery Suppression rules patrol_Suppresses_sensitive Corridor cruise + stable flight Suppressing the proximity of sensitive areas To avoid misjudging normal inspections as abnormal proximity
[0156] Suppose that the set of rule events that can be extracted for a certain trajectory is Each rule This represents a behavioral event with a clearly defined semantic meaning. For each rule... Its activation strength can be defined. This represents the degree to which the rule is satisfied in the current sample. Further, a mapping relationship is established between rule events and intent categories to obtain the initial knowledge support score for each intent category:
[0157]
[0158] in, Indicates the first Initial knowledge support score for class intent; Indicates the total number of rule events; Indicates the first The activation strength of a rule event; Indicates the first Rule event for the first Support weights for class intents; Indicates the intent category index, , The total number of intent categories.
[0159] To reduce the risk of misjudgment caused by mutually exclusive rules or conflicting evidence, the conflict suppression layer further introduces conflict penalty terms. Let the set of conflict rules or suppression terms be... ,in, Indicates the first The degree of activation of a conflict rule or inhibitory term. Indicates the first The conflict rule or suppression term affects the first The penalty weight for class intent, then the first The conflict penalty term corresponding to the class intent can be represented as:
[0160]
[0161] in, Indicates the first Conflict penalty items for class intentions; Indicates the total number of conflicting rules or suppression terms; Indicates the first The degree of activation of each conflicting rule or inhibitory term; Indicates the first The conflict rule or suppression term affects the first Penalty weights for class intents. The larger the value, the better the conflict rule or suppression term affects the first... The stronger the inhibitory effect of the class intent.
[0162] Therefore, the first The modified knowledge support score for class intent can be expressed as:
[0163]
[0164] Right now:
[0165]
[0166] in, Indicates the first The revised knowledge support score for class intent.
[0167] The above formula shows that the positive support of a rule event for a certain intention category is obtained through the accumulation of the first term, while the inhibitory effect of mutually exclusive rules or conflicting evidence on that intention category is achieved through the subtraction of the second term. Therefore, the so-called "punishment" or "inhibition" is specifically manifested as follows: when a conflicting rule or inhibitory item is activated, the effect is determined according to its degree of activation. and penalty weight The initial knowledge support score for the corresponding intent category The corresponding score is deducted, thereby reducing the probability of that intent category in the subsequent normalized output.
[0168] For example, when a target exhibits both a tendency to approach the base and neighborhood characteristics of the delivery point, if the return-to-base rules are strongly activated, the conflict suppression layer can suppress delivery suppression terms through return-to-base. The knowledge support score corresponding to the delivery intention is deducted; when the target flies stably along the inspection channel, the conflict suppression layer can suppress the approach of the inspection suppression sensitive area through the suppression term. The knowledge support score corresponding to the intent to approach the sensitive area will be deducted.
[0169] The normalized output layer is used to convert the corrected knowledge support scores for each intent category into the category probability distribution of the rule-based knowledge reasoning module. Let the corrected knowledge support score vector be:
[0170]
[0171] The category probability distribution output by the rule-based knowledge reasoning module can then be expressed as:
[0172]
[0173] in, This represents the category probability distribution output by the rule-based knowledge reasoning module; This indicates that the current sample belongs to the first... Rule-based prediction probability of class intent; They represent the first category, the second category, ..., the third category, respectively. The revised knowledge support score for class intent.
[0174] For the The class intent, specifically the rule-based prediction probability, is expressed as follows:
[0175]
[0176] in, Represents the natural exponential function; Indicates the first The knowledge support score after the modification of the class intent; Indicates the summation index of intent categories; Indicates the first The revised knowledge support score for class intent.
[0177] Through the above normalization process, the rule-based knowledge reasoning module can convert the corrected knowledge support scores of different intent categories into probabilistic forms and use them as the knowledge-side input for the subsequent dynamic fusion module.
[0178] It should be noted that, compared with the data-driven module, the rule-based knowledge reasoning module does not rely on complex end-to-end training. Instead, it directly generates interpretable category evidence through explicit rule reasoning. Therefore, when the trajectory is only partially observed and the data-driven module has not yet formed a stable pattern, it can provide more robust auxiliary support. In particular, for intent categories that are closely related to scene semantics, such as delivery, approaching sensitive areas, and returning, the rule-based knowledge reasoning branch can effectively make up for the problem of insufficient utilization of prior knowledge by the pure data-driven model.
[0179] like Figure 5 As shown, the dynamic fusion module includes a gated feature extraction submodule, a constrained gated network submodule, and a fusion coefficient generation and output submodule. The gated feature extraction submodule extracts gated feature vectors from the outputs of the data-driven module and the rule-based knowledge reasoning module to characterize the relative reliability of the two information sources on the current sample. Specifically, the input data for this submodule is the output probability distribution of the data-driven module. The output probabilities of the rule-based knowledge reasoning module and the rule-based knowledge reasoning module are respectively Based on the observation state information of the current sample, the following five gating features are extracted and used to construct the gating input vector. .
[0180] in, This represents the prediction confidence of the data-driven module, which is the maximum class probability in the output probability distribution of the data-driven module. It reflects the degree of certainty with which the data-driven module classifies the current sample. The larger the value, the clearer the data-driven module's judgment of the current sample, and the higher the weight should be assigned in subsequent fusion.
[0181] The information entropy, representing the probability distribution output by the data-driven module, is used to measure prediction uncertainty and is expressed as follows:
[0182]
[0183] in, This represents the category probability distribution output by the data-driven module. This indicates that the data-driven module determines that the current sample belongs to the first... The probability of class intent, and satisfying , ; Total number of intent categories; This represents a very small constant to prevent logarithmic operations from overflowing; for example, it can be taken as... .
[0184] When the category probability distribution output by the data-driven module is relatively concentrated, A smaller value indicates that the model's predictions are relatively certain; when the class probability distribution is more dispersed... A large value indicates high uncertainty in the model's predictions. Therefore, during the dynamic fusion process, The larger the value, the lower the fusion weights of the data-driven module should be. Furthermore, to ensure that the prediction entropy has a similar scale to other gating features, the fusion weights can also be... Divide by Normalization is performed to constrain its value range to within Within the range.
[0185] This represents the ratio of the currently observed trajectory length to the complete trajectory length, used to characterize the sufficiency of the current trajectory information, and is expressed as follows:
[0186]
[0187] in, This indicates the effective length of the observed trajectories in the current sample. This indicates the length of the complete trajectory corresponding to the given trajectory; when the length of the complete trajectory is not directly available... It can be determined by a preset standard observation length. Replacement. Observational completeness Used to characterize the sufficiency of the current trajectory information. The larger the value, the more fully the trajectory pattern unfolds, and the more reliable the judgment result of the data-driven module. Therefore, its fusion weight should be increased. Conversely, it is necessary to rely more on the rule knowledge reasoning module to provide auxiliary judgment.
[0188] This represents the proportion of rules in the current sample whose activation intensity exceeds a preset threshold to the total number of valid rules in the rule base, as shown below:
[0189]
[0190] in, This indicates the total number of valid rules in the preset rule base that participate in the current intent reasoning; Indicates the first The activation strength of each rule in the current sample, and ; This indicates the rule activation threshold, when At that time, it was believed that the first The rule was effectively activated; This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. This indicates the number of rules that are effectively activated in the current sample.
[0191] rule coverage The larger the value, the more effective rules are triggered by the current trajectory sample, and the more sufficient explicit semantic evidence the rule knowledge reasoning module can provide. Therefore, the contribution ratio of the rule knowledge reasoning module should be appropriately increased during the dynamic fusion process.
[0192] The inverse measure of the normalized entropy of the rule knowledge distribution is used to measure the consistency of each activation rule in supporting the direction of the candidate intent. It is represented as follows:
[0193]
[0194] in, For the number of intent categories, The representation rule knowledge reasoning module represents the first... Support probability of class intent To prevent logarithmic overflow, a very small constant is used. This is especially important when multiple activation rules primarily support the same intent category. The distribution is more concentrated, and the normalized entropy is smaller. The distribution of knowledge tends to be more dispersed when different rules support multiple mutually exclusive intentions. A decrease indicates a certain degree of conflict within the rules. Therefore, rule consistency can reflect the stability of knowledge evidence.
[0195] The above five gating features characterize the sample state from three aspects: data reliability (confidence, prediction entropy), observation sufficiency (observation completeness), and rule effectiveness (rule coverage, rule consistency), providing a quantitative basis for the adaptive generation of fusion coefficients.
[0196] The constrained gating network submodule takes the gating feature vector as input and calculates intermediate variables through a linear mapping with signed constraints to ensure that the direction of change of the fusion weights conforms to the task prior.
[0197] The core of this submodule is a constrained linear gating function. Let the intermediate variable output by the gating network be... Then we have:
[0198]
[0199] in, For bias terms, These are learnable parameters. To ensure that each weight is non-negative, and to constrain their direction of action through explicit sign, this application stipulates:
[0200]
[0201] in, These are the original trainable parameters, and The function is defined as:
[0202]
[0203] in, Represent the natural logarithm function; Represents the natural exponential function; This represents the result of a translation of the natural exponential function;
[0204] The above design enables the gate function to satisfy the monotonicity constraints shown in Table 2:
[0205] Table 2 Monotonicity Constraint Description Table
[0206] Unlike unconstrained gating, this submodule not only learns the strength of each factor's influence through training, but also ensures that the parameter learning process always conforms to the preset semantic direction through symbolic constraints. Specifically, it ensures that the weight parameters are non-negative through Softplus mapping, and then combines the sign setting before each feature (such as adding a negative sign before the weight corresponding to rule coverage) to ensure that the features and intermediate variables are aligned. The monotonic relationship between them is mathematically guaranteed, thereby enhancing the interpretability and stability of the fusion process.
[0207] The fusion coefficient generation and output submodule is used to generate intermediate variables Dynamic fusion coefficients are generated by mapping to a preset interval using Sigmoid and linear transformations. The probability distributions of the data-driven module and the rule-based knowledge reasoning module are then weighted and fused to output the final intent reasoning result. This module includes two consecutive transformation steps and a final weighted fusion operation, as detailed below:
[0208] Step 1: For intermediate variables Compress to the (0,1) interval;
[0209] First, regarding intermediate variables Apply the Sigmoid function to compress it to the (0,1) interval. The compression process is as follows:
[0210]
[0211] in, .
[0212] Step 2: Compress the intermediate variables into the (0,1) interval Perform a linear mapping;
[0213] Considering that no module should be completely discarded during the actual fusion process, this embodiment further compresses the intermediate variables to the (0,1) interval. Linear mapping to a preset interval The mapping process is as follows:
[0214]
[0215] in, This design ensures that the final fusion coefficient always meets the requirements. This avoids the occurrence of or The problem of a single module completely dominating decision-making.
[0216] Step 3: For the intermediate variables after linear mapping Weighted fusion output.
[0217] After obtaining the dynamic fusion coefficients related to the samples Finally, the final intention prediction result is expressed as follows:
[0218]
[0219] in, The class probability distribution output by the data-driven branch. The category probability distribution output for the knowledge-driven branch. This is the final fusion result.
[0220] For the first in the batch sample The fusion results for the samples are shown below:
[0221]
[0222] in, By the The gating feature vectors corresponding to each sample are dynamically generated.
[0223] From a mechanistic perspective, when the data-driven module has high confidence, low prediction entropy, and high observation completeness, Often increases, resulting in a larger The model tends to rely more on data-driven branches; when rule coverage is high, rule consistency is strong, or the entropy of the data-driven module is high, Tend to decrease, This also decreases. The data-driven module plays a higher role in the final decision. Therefore, the fusion weights are no longer fixed constants, but dynamic variables determined by the sample state. This sample-by-sample adaptive fusion method allows the model to flexibly adjust the contribution ratio of the two information sources according to different trajectory samples, different observation stages, and different rule states, thereby improving the model's inference stability and generalization ability in complex scenarios.
[0224] In another exemplary embodiment, the intelligent intent prediction model for low-altitude flying targets is trained through the following steps:
[0225] The fused output is used as the final supervision object, and the cross-entropy loss function is used to jointly train the data-driven branch and the dynamic fusion module. For a single sample, the loss function is defined as follows:
[0226]
[0227] in, For real category labels, This is the output probability for fusion.
[0228] During training, the parameters of the data-driven module and the dynamic gating module are jointly updated through backpropagation; the rule-based knowledge reasoning module generates category evidence based on predefined rules and does not participate in gradient updates. To ensure that the gating parameters meet the preset directional constraints, the Softplus function is used to perform non-negative mapping on the original parameters. Model hyperparameters such as learning rate, batch size, number of training epochs, and fusion interval parameters are selected on the validation set.
[0229] The specific training process includes the following steps:
[0230] The first step is to construct a training sample set. Multiple low-altitude flight target trajectory samples are acquired, and each trajectory sample is labeled with an intent category tag. The intent category includes one of the following: inspection, delivery, loitering, approaching a sensitive area, and returning to base. Each trajectory sample includes flight trajectory information, scene information, and a corresponding set of rule events.
[0231] The second step is to partition the dataset. The training sample set is divided into a training set, a validation set, and a test set according to a preset ratio, ensuring that the distribution of samples of each intent category in the training set, validation set, and test set meets a preset equilibrium condition.
[0232] The third step is to preprocess the training samples. Missing and outlier data in each trajectory sample in the training set are repaired, and sequence alignment is performed using a uniform length truncation and zero-padding method to generate effective length markers. Continuous features are then normalized to obtain the model training input.
[0233] The fourth step is to initialize the model parameters. This involves initializing the parameters of the motion state differential enhancement layer, multi-scale extended temporal coding layer, scene relationship guided attention layer, and classification output layer in the data-driven module, as well as the gating parameters in the dynamic fusion module. The rule knowledge reasoning module performs rule matching and category evidence generation based on a preset rule base, and its rule parameters can be preset according to task priors.
[0234] The fifth step is to generate the data-side category probability distribution. The preprocessed trajectory sequence is input into the data-driven module, and the category probability distribution of the data-driven branch is obtained through motion state differential enhancement, multi-scale extended temporal coding, and scene relationship-guided attention calculation.
[0235] The sixth step is to generate the category probability distribution on the knowledge side. The corresponding scene information and rule event set are input into the rule knowledge reasoning module. Through sliding window statistics, scene relationship calculation, trend determination, rule activation, rule-intent mapping, and conflict suppression processing, the category probability distribution output by the rule knowledge reasoning branch is obtained.
[0236] Step 7: Generate dynamic fusion coefficients. A gated feature vector is constructed based on the prediction confidence, prediction entropy, observation completeness, rule coverage, and rule consistency of the data-driven branch. This gated feature vector is then input into a constrained gating network to generate the dynamic fusion coefficients corresponding to the current sample.
[0237] Step 8: Calculate the fusion output. Based on the dynamic fusion coefficient, the category probability distributions output by the data-driven branch and the rule-based knowledge reasoning branch are weighted and fused to obtain the final intent prediction probability distribution.
[0238] Step 9: Calculate the training loss and update the parameters. Calculate the cross-entropy loss based on the final intent prediction probability distribution and the true intent category label, and update the parameters of the data-driven module and the dynamic fusion module through backpropagation. The gating parameters of the dynamic fusion module are constrained by a non-negative mapping function to ensure that the influence direction of each gating feature on the fusion coefficient conforms to the preset semantic constraints.
[0239] Step 10: Validation and Model Selection. After each training epoch, the model's performance is evaluated using the validation set, and the optimal model parameters are selected based on the evaluation metrics on the validation set. Training stops when the validation set performance no longer improves within a preset number of epochs, and the trained intelligent intent prediction model for low-altitude flying targets is output.
[0240] The training process of the above models is shown in Table 3:
[0241] Table 3 Model Training Process
[0242] In this application, the preceding text systematically elaborates on the complete theoretical system and technical implementation of the method described herein, from problem modeling, framework design, module construction to training optimization. This method extracts trajectory temporal features through a data-driven module, generates explicit semantic evidence through a rule-based knowledge reasoning module, and adaptively adjusts the contribution ratio of the two types of information sources using a dynamic fusion module, achieving collaborative decision-making based on data and knowledge. Building upon the completed model construction and training strategy design, to comprehensively evaluate the actual effectiveness of this application in the low-altitude flight target intent reasoning task, this application further designs a systematic comparative experiment to verify and analyze the model performance from four dimensions: complete observation, local observation, few-shot learning, and module contribution.
[0243] I. Comparison Method Settings
[0244] To comprehensively evaluate the effectiveness of the proposed method, this application selects the following four types of methods as comparative baselines:
[0245] (1) Traditional data-driven baseline models: including LSTM and BiLSTM, which only use trajectory sequences for intent reasoning and do not introduce explicit rule knowledge, are used to verify the basic performance of pure time series modeling methods on this task.
[0246] (2) Baseline of the data-driven module (TCN-Attn) in this application: Only the data-driven module consisting of TCN and attention mechanism is retained, and the knowledge-driven module and dynamic fusion module are not introduced. It is used to verify the effect of knowledge enhancement strategy on the overall model performance.
[0247] (3) Fixed Fusion Baseline Model (TCN-Attn+rule+Fixed): It uses both the data-driven module and the rule knowledge reasoning module, but uses fixed weights for result fusion.
[0248] 4) Rule-only inference baseline: Only the rule knowledge inference module is used for intent discrimination, without introducing a data-driven module, to verify the discrimination ability of the rule knowledge itself.
[0249] (5) The model of this application (TCN-Attn+rule+Dynamic): a complete method that simultaneously includes a data-driven module, a rule-based knowledge reasoning module, and a dynamic fusion module.
[0250] II. Experimental Dataset and Condition Settings
[0251] This application uses a "scenario simulation generation + rule constraint annotation" approach to construct an intent reasoning dataset for typical mission scenarios of low-altitude flight targets. The simulation scenarios include key spatial elements such as bases, delivery points, sensitive areas, and inspection corridors. Corresponding trajectory generation constraints are designed according to different mission intents. The dataset contains 2,500 trajectory samples, covering five typical intents: inspection, delivery, loitering, approaching sensitive areas, and returning. There are 500 samples for each type. Each trajectory consists of the target's position, speed, altitude, heading angle, and acceleration within a continuous time window. It also records the spatial relationships between the target and the base, delivery point, sensitive area, and inspection corridor. The experiments were conducted using Python 3.12 and PyTorch 2.3.1. Data was partitioned into training, validation, and test sets in a 6:2:2 ratio, ensuring consistent distribution of samples across different categories across subsets. The data-driven branch used the Adam optimizer for training, with an initial learning rate of 0.001, a batch size of 64, and 100 training epochs. Early stopping was employed to prevent overfitting. Accuracy, precision, recall, and macro-F1 score were used to evaluate model performance. Evaluation metrics: Accuracy, precision, recall, and macro-F1 score were used as the primary evaluation metrics.
[0252] III. Experimental Comparison Procedure
[0253] Comparison Experiment 1: Dynamic Fusion Weight Range Selection
[0254] Before conducting a full comparative experiment, we first determined the reasonable range of values for the dynamic fusion weights, and then compared the Macro-F1 performance of the model under different fusion weight constraint ranges under the conditions of full observation and 20% early observation. The results are shown in Table 3.
[0255] Table 3. Comparison of model performance under different dynamic fusion weight ranges
[0256] As can be seen from Table 3, when the fusion weight is set to At that time, the model achieved the optimal Marco-F1 score under both the complete observation and 20% early observation conditions. Therefore, this application ultimately sets the dynamic fusion weight interval to... This setting was used in all subsequent experiments.
[0257] Comparative Experiment 2: Performance Comparison Analysis of Models Under Complete Observation Conditions
[0258] Under complete observation conditions, the method described in this application was compared with LSTM, BiLSTM, TCN-Attn, Rule-only, and TCN-Attn+rule+Fixed, and evaluated using four metrics: Accuracy, Precision, Recall, and Macro-F1. The evaluation results are shown in Table 4. Figure 6 As shown:
[0259] Table 4. Model inference performance under complete observation conditions LSTM 82.6 80.9 78.8 79.6 BiLSTM 83.9 82.1 81.4 81.2 TCN-Attn 86.3 86.8 84.2 85.2 Rule-only 78.4 77.2 76.9 76.8 TCN-Attn+rule+Fit 87.8 86.5 87.6 86.9 TCN-Attn+rule+Dynamic 89.1 88.9 88.2 88.4
[0260] from Figure 6 As shown in Table 4, the method described in this application achieves optimal performance under complete observation conditions, with Accuracy, Precision, Recall, and Macro-F1 reaching 89.1%, 88.9%, 88.2%, and 88.4%, respectively. Compared to TCN-Attn+rule+Fixed, the method described in this application improves performance by 1.3, 2.4, 0.6, and 1.5 percentage points in the four metrics, respectively, indicating that the dynamic fusion mechanism can further coordinate the collaborative discriminative ability of data-driven features and knowledge.
[0261] Compared to LSTM and BiLSTM, TCN-Attn achieves higher performance, indicating that TCN and attention mechanisms are more suitable for extracting key temporal patterns in the trajectories of low-altitude flight targets. The Macro-F1 score of the rule-only model is 76.8%, lower than that of the data-driven model, suggesting that relying solely on rule knowledge is insufficient to fully characterize complex trajectory patterns; however, it still possesses some discriminative ability, indicating that rule events and scene relationships can provide effective semantic priors. In summary, rule knowledge is more suitable as a supplement to data-driven models, while dynamic fusion can more effectively utilize this complementary relationship.
[0262] Comparative Experiment 3: Performance Comparison Analysis of Models under Local Observation Conditions
[0263] To evaluate the model's intent reasoning ability under partial observation conditions, observation ratios of 20%, 40%, 60%, 80%, and 100% were set, with Macro-F1 as the evaluation metric. For each observation ratio, continuous trajectory segments from the beginning of the complete trajectory were retained to simulate the situation in real online reasoning where only early trajectory information is available. Experimental results are shown in Table 5 and... Figure 7 As shown:
[0264] Table 5. Model inference performance under local observation conditions LSTM 56.8 66.7 74.9 76.8 79.6 BiLSTM 57.7 69.8 78.6 83.4 84.0 TCN-Attn 61.5 74.6 81.8 84.3 85.1 Rule-only 68.5 71.2 73.4 75.1 76.8 TCN-Attn+rule+Fit 66.8 73.8 81.2 84.9 86.0 TCN-Attn+rule+Dynamic 74.2 81.3 85.2 87.1 88.4
[0265] From Table 5 and Figure 7As can be seen, the overall performance of each model increases with the observation ratio from 20% to 100%, indicating that the gradual completion of trajectory information helps improve the intention inference effect. The method described in this application achieves the highest Macro-F1 at all observation ratios, demonstrating its good adaptability at different observation stages.
[0266] Comparative Experiment 4: Model Performance Comparison under Small Sample Conditions
[0267] To verify the model's generalization ability under small sample conditions, four training sample sizes were set: 5-shot, 10-shot, 20-shot, and 50-shot. A small sample training subset was constructed by randomly selecting a small number of labeled samples from the complete training set according to each category. The Macro-F1 performance of TCN-Attn, TCN-Attn+rule+Fixed, and the method described in this application are compared in Table 6. Figure 8 As shown:
[0268] Table 6. Model inference performance under small sample conditions TCN-Attn 61.7 70.8 79.1 83.8 TCN-Attn+rule+Fit 67.9 75.8 82.5 85.9 Ours 72.8 79.9 85.1 87.6
[0269] From Table 6 and Figure 8 As can be seen, the method described in this application achieves the best performance in all few-shot learning scenarios, reaching 72.8% under the 5-shot condition, which is an 11.1 percentage point improvement over TCN-Attn, representing the most significant improvement. This indicates that when training samples are insufficient, the knowledge reasoning module can provide additional structured priors, reducing the model's dependence on large-scale labeled data.
[0270] As the number of samples increases, the performance gap between the models narrows, but the method described in this application still achieves 87.6% under 50-shot conditions, maintaining its leading advantage. This indicates that the proposed method is not only suitable for low-resource scenarios, but also has good scalability and stability.
[0271] Comparative Experiment 5: Ablation Experiment
[0272] To further verify the contribution of each module to the model performance, this application employs a module removal strategy for ablation analysis, removing the dynamic fusion module, knowledge reasoning module, and attention module respectively. The performance degradation of the complete model compared to different ablation models is statistically analyzed across four dimensions: Accuracy, Precision, Recall, and Macro-F1. Experimental results are as follows: Figure 9 As shown. From Figure 9It can be seen that the model degrades most significantly after removing the dynamic fusion module, with Accuracy, Precision, Recall, and Macro-F1 decreasing by 2.1, 2.8, 2.3, and 2.2 percentage points, respectively, indicating that the dynamic fusion mechanism is key to improving model performance. Removing the knowledge reasoning module reduces these four metrics by 1.5, 1.9, 1.7, and 1.5 percentage points, respectively, demonstrating that rule-based knowledge provides effective semantic compensation for intent reasoning. Removing the attention module also leads to a decrease in model performance, indicating that the attention mechanism helps enhance the representation ability of key temporal segments.
[0273] Based on the above five sets of comparative experiments, the following conclusions can be drawn:
[0274] 1) Under complete observation conditions, the method described in this application improves the Macro-F1 score by 1.5 percentage points compared to the fixed fusion model and by 3.2 percentage points compared to the pure data-driven model (TCN-Attn), verifying the effectiveness of knowledge enhancement and dynamic fusion.
[0275] 2) Under local observation conditions, the method described in this application improves the efficiency by 12.7 percentage points compared to TCN-Attn when the observation ratio is 20%, which is the most significant improvement. This indicates that the dynamic fusion mechanism can make better use of knowledge to compensate for insufficient data in the early stage.
[0276] 3) Under small sample conditions, the method described in this application improves by 11.1 percentage points compared with TCN-Attn in 5-shot, indicating that the knowledge reasoning module can provide additional structured priors and reduce the model's dependence on large-scale labeled data.
[0277] 4) Ablation experiments show that the dynamic fusion module is the key factor for performance improvement, and the rule knowledge reasoning module plays an important role in enhancing the intention reasoning ability and model interpretability when observation data is limited.
[0278] This application also provides a target intelligent intent prediction device based on data-knowledge fusion. The device includes: an acquisition module for acquiring flight trajectory information and scene information of a low-altitude flying target within an observation time window, wherein the flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration; and the scene information includes, for example, the target and starting point, sensitive area, delivery point, base, and inspection channel; a preprocessing module for preprocessing the flight trajectory information and scene information; a model construction and training module for constructing a low-altitude flying target intelligent intent prediction model and training the model. This model, through a dynamic gating fusion mechanism, adaptively adjusts the fusion weights of data-driven information and rule knowledge information based on prediction confidence, prediction entropy, observation completeness, rule coverage, and rule consistency to achieve collaborative decision-making of temporal pattern learning and rule reasoning under different observation conditions; and an output module for inputting the preprocessed flight trajectory information and related scene information into the trained low-altitude flying target intelligent intent prediction model to predict the intent category of the low-altitude flying target.
[0279] Optionally, the preprocessing module includes: a repair submodule for repairing missing or abnormal data in flight trajectory information and scene information; a marking submodule for aligning the repaired flight trajectory information and scene information into a sequence and marking the effective length using uniform length truncation and zero padding; and a normalization submodule for normalizing continuous features in the repaired, aligned and marked sequence.
[0280] In another exemplary embodiment, this application also provides a storage medium including instructions that, when executed on a computer, cause the computer to perform the target intelligent intent prediction method based on data-knowledge fusion as described in the preceding embodiments.
[0281] In another exemplary embodiment, this application also provides an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the target intelligent intent prediction method based on data-knowledge fusion as described in the preceding embodiments.
[0282] The above embodiments are only for illustrating the technical concept and features of this application, and are intended to enable those skilled in the art to understand the content of this application and implement it accordingly. They should not be construed as limiting the scope of protection of this application. All equivalent changes or modifications made in accordance with the spirit and essence of this application should be included within the scope of protection of this application.
Claims
1. A target intelligent intent prediction method based on data-knowledge fusion, characterized in that, The method includes: Acquire flight trajectory and scene information of low-altitude flying targets within the observation time window. Flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration; scene information includes, for example, the target and origin, sensitive areas, delivery points, bases, and inspection channels. The flight trajectory information and scene information are preprocessed; A low-altitude flight target intelligent intent prediction model is constructed and trained. The model adaptively adjusts the fusion weight of data-driven information and rule knowledge information based on prediction confidence, prediction entropy, observation completeness, rule coverage and rule consistency through a dynamic gating fusion mechanism, so as to achieve collaborative decision-making of temporal pattern learning and rule reasoning under different observation conditions. The preprocessed flight trajectory information and related scene information are input into the trained intelligent intent prediction model for low-altitude flying targets to predict the intent category of low-altitude flying targets.
2. The method according to claim 1, characterized in that, The flight trajectory information and scene information are preprocessed, including: Repair missing or abnormal data in flight trajectory and scene information; The repaired flight trajectory information and scene information were sequence aligned and marked with effective length by uniform length truncation and zero padding. The continuous features in the repaired, aligned and labeled sequence are normalized.
3. The method according to claim 1, characterized in that, The intelligent intent prediction model for low-altitude flying targets includes: The module consists of a data-driven module, a rule-based knowledge reasoning module, and a dynamic fusion module. in, The data-driven module is used to extract the trajectory temporal features from the flight trajectory information of the flight target and output the category probability distribution based on data learning as data evidence; The rule-based knowledge reasoning module is used to perform explicit rule mapping based on behavioral events, scene relationships, and trend rules, and output interpretable category probability distributions as evidence of intent categories. The dynamic fusion module is used to adaptively generate fusion coefficients to dynamically balance the contribution ratio of data evidence and intent category evidence.
4. The method according to claim 3, characterized in that, The data-driven module includes: The layers are arranged sequentially as follows: input layer, motion state differential enhancement layer, multi-scale extended temporal coding layer, scene relationship guided attention layer, effective length mask layer, and classification output layer. in, The motion state differential enhancement layer is used to construct motion change features based on the trajectory state at adjacent time points; Multi-scale extended temporal coding layers are used to extract short-term maneuver features, medium-term trend features, and long-term behavioral pattern features; The scene relationship-guided attention layer is used to generate attention weights by combining trajectory temporal features and scene relationship features; The effective length mask layer is used to mask zero-padding time steps; The classification output layer is used to output the probability of classifying data.
5. The method according to claim 1, characterized in that, The rule-based knowledge reasoning module includes the following components set sequentially: The system consists of a sliding window statistics layer, a scene relationship calculation layer, a trend determination layer, a rule event activation layer, a rule mapping layer, a conflict suppression layer, and a normalized output layer. in, The sliding window statistical layer is used to slice the flight trajectory sequence into windows and extract window-level motion statistical features; The scene relationship calculation layer is used to calculate the spatial relationship features between the target and sensitive areas, delivery points, bases, and inspection channels; The trend determination layer is used to determine whether a target is approaching, moving away from, or maintaining a stable trend relative to a key area. The rule-based event activation layer is used to generate rule activation strength based on window-level motion statistics, spatial relationship features, and trend determination results. The rule mapping layer is used to generate support scores for each intent category based on the mapping relationship between rule knowledge and intent categories; The conflict suppression layer is used to adjust the support score of the corresponding intent category according to the mutual exclusion rule; The normalized output layer is used to convert the corrected support scores into the category probability distribution of the rule-based knowledge reasoning module.
6. The method according to claim 1, characterized in that, The dynamic fusion module includes the following components configured sequentially: The module includes a gated feature extraction submodule, a constrained gated network submodule, and a fusion coefficient generation and output submodule. in, The gated feature extraction submodule is used to extract gated feature vectors from the outputs of the data-driven module and the rule-based knowledge reasoning module; The constrained gating network submodule is used to compute intermediate variables by taking gating feature vectors as input and using linear mapping with signed constraints. The fusion coefficient generation and output submodule is used to map intermediate variables to a preset interval to generate dynamic fusion coefficients, and to perform weighted fusion of the probability distributions of the data-driven module and the rule knowledge reasoning module to output the final intent reasoning result.
7. A target intelligent intent prediction device based on data-knowledge fusion, characterized in that, The device includes: The acquisition module is used to acquire flight trajectory information and scene information of low-altitude flying targets within the observation time window. The flight trajectory information includes, for example, position, speed, heading angle, altitude, and acceleration; the scene information includes, for example, the target and origin, sensitive areas, delivery points, bases, and inspection channels. The preprocessing module is used to preprocess the flight trajectory information and scene information; The model building and training module is used to build an intelligent intent prediction model for low-altitude flying targets and train the model. The model uses a dynamic gating fusion mechanism to adaptively adjust the fusion weight of data-driven information and rule knowledge information based on prediction confidence, prediction entropy, observation completeness, rule coverage and rule consistency, so as to achieve collaborative decision-making of temporal pattern learning and rule reasoning under different observation conditions. The output module is used to input the preprocessed flight trajectory information and related scene information into the trained low-altitude flight target intelligent intent prediction model in order to predict the intent category of the low-altitude flight target.
8. The apparatus according to claim 7, characterized in that, The preprocessing module includes: The repair submodule is used to repair missing or abnormal data in flight trajectory information and scene information; The marking submodule is used to perform sequence alignment and effective length marking on the repaired flight trajectory information and scene information using uniform length truncation and zero padding. The normalization submodule is used to normalize continuous features in the repaired, aligned, and labeled sequence.
9. A storage medium, characterized in that, It includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 6.
10. An electronic device, characterized in that, The electronic device includes: Memory, processor, and computer programs stored in memory and executable on the processor, wherein, When the processor executes the program, it implements the method as described in any one of claims 1 to 6.