Target classification method based on fusion of motion features and track features
By fusing motion features and track features, and utilizing LSTM networks and adaptive weighted summation algorithms, the problem of identifying low-altitude, small, and slow-moving targets in complex environments was solved, achieving more accurate target classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE 724TH RESEARCH INSTITUTE OF CHINA STATE SHIPBUILDING CORP LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-17
AI Technical Summary
Existing radar technology has difficulty effectively identifying low-altitude, small, and slow-moving targets, especially in complex weather or electromagnetic environments. Traditional classification methods struggle to capture dynamic features and adapt to changes in feature weights, resulting in high false alarm and false false alarm rates.
A method combining motion features and track features is adopted. Features are extracted through a long short-term memory neural network (LSTM) and fused using an adaptive weighted summation algorithm to generate accurate classification results.
It improves the ability to identify low-altitude, small, and slow-moving targets, enhances the robustness of the algorithm, and adapts to target classification in complex environments.
Smart Images

Figure CN121878631A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of radar target processing, specifically relating to a target classification method that fuses motion characteristics and track characteristics. Background Technology
[0002] In modern airport security systems, the detection and classification of low-altitude, small, and slow-moving targets is a pressing technical challenge. These targets typically refer to aircraft such as drones and micro-aircraft that pose a potential threat to airports. Due to their small size, low speed, and weak radar cross-section, they are easily confused with weather clutter or ground interference signals, resulting in high false alarm and false false alarm rates for traditional radar systems. Although existing radar technologies can enhance detection capabilities by improving Doppler resolution or adopting bistatic / multistatic configurations, significant challenges remain in the target identification stage: classification methods based on single features (such as velocity, acceleration, or trajectory shape) struggle to cope with the diversity of features in complex scenarios, while traditional fusion strategies (such as simple weighting or linear combination) often result in the loss of crucial information due to a lack of modeling capabilities for temporal features.
[0003] From a technological evolution perspective, existing classification methods for small, low-speed targets suffer from two prominent shortcomings: First, the contradiction between feature dimensions and dynamism. Traditional methods often rely on motion characteristics of a single time phase (such as instantaneous velocity) or geometric features of the trajectory (such as trajectory curvature). However, small, low-speed targets may exhibit non-steady-state motion patterns (such as sudden hovering or maneuvering) in complex weather or electromagnetic environments, making it difficult to capture their dynamic behavior using static features at a single moment. Second, the static limitations of feature weights. Existing fusion models often employ fixed weight allocations (such as based on expert experience or statistical averaging), but in real-world scenarios, feature importance dynamically changes with the target's state.
[0004] To address the aforementioned issues, existing technologies attempt to improve classification accuracy through multi-feature fusion, but they suffer from the following key bottlenecks: 1) Difficulty in data alignment and feature standardization. The spatiotemporal sampling rate differences between motion features (such as acceleration) and trajectory features (such as 3D coordinate sequences) necessitate data alignment through interpolation or downsampling, a process that can lead to information distortion; 2) Insufficient modeling of temporal dependencies. Traditional classifiers (such as SVM or random forests) struggle to handle long-range dependencies in temporal features. For example, the "spiral convergence" trajectory pattern during the descent phase of a UAV requires capturing the correlation across multiple time steps for effective identification; 3) Poor environmental adaptability. In rainy, foggy, or strong electromagnetic interference scenarios, the amplitude and phase characteristics of radar echoes may fluctuate abnormally. If the classification model lacks an adaptive correction mechanism, its generalization ability will be significantly weakened. Summary of the Invention
[0005] To address the aforementioned problems, the present invention aims to provide a target classification method that fuses motion features and track features. After radar detects low, small, and slow-moving targets in the air and automatically batches them, the target's track information is obtained based on the target batch number. The target is visualized using track information such as timestamp, distance, azimuth, and echo intensity. The input dimensions of the visualized track information and motion features are fixed and then fed into a Long Short-Term Memory (LSTM) neural network to extract features. An adaptive weighted summation algorithm is then used to fuse the data to obtain the final classification result. This method effectively solves the problem of difficulty in distinguishing low, slow, and small targets in airport environments, providing strong support for the engineering application and promotion of low, slow, and small target classification.
[0006] The specific technical solution for achieving the objective of this invention is as follows:
[0007] A target classification method that fuses motion features and track features includes the following steps:
[0008] Step 1: After detecting a low, slow, and small target, the radar automatically begins to acquire the target's trajectory information in batches;
[0009] Step 2: Process motion feature information and visualize track information to obtain motion features and track features;
[0010] Step 3: Convert the motion feature matrix and track feature matrix into fixed sizes;
[0011] Step 4: Feed the fixed-length track features and motion features into the initial classification network built on LSTM to generate the initial classification results;
[0012] Step 5: The target classification results are fused using an adaptive weighted summation strategy to obtain accurate classification results.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0014] The innovation of this invention is reflected in the design of the dual-feature stream spatiotemporal modeling and adaptive weight fusion architecture: by deconstructing the radar raw point data into "motion feature sequence" (including temporal features such as velocity, acceleration, and angular velocity) and "track topology features" (such as trajectory curvature, altitude-range envelope, and other geometric features), respectively input into a bidirectional LSTM network to extract temporal dependencies;
[0015] Compared to existing technologies, this method achieves three breakthroughs: 1) Deep modeling of temporal features: LSTM networks can automatically capture key temporal patterns such as acceleration abrupt changes, improving the ability to identify unconventional maneuvers; 2) Dynamic adaptation of feature importance: In environments with strong interference, the weight of echo intensity features can be reduced, and instead, the robustness of trajectory topology features can be relied upon; 3) Standardization of multi-dimensional data fusion: By using feature space alignment techniques (such as normalization and dimensional expansion), the dimensional differences between motion features and track features are eliminated, ensuring the consistency of the physical meaning of the fused calculations.
[0016] This solution improves the ability to identify moving targets and increases the robustness of the algorithm. It can further support the classification of small, slow-moving targets.
[0017] The present invention will be further described below with reference to specific embodiments. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the target classification method that fuses motion features and track features according to the present invention.
[0019] Figure 2 This is a schematic diagram of the classification network structure of the present invention. Detailed Implementation
[0020] Example
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0023] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of this application. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0024] Combination Figure 1 A target classification method that fuses motion features and track features includes the following steps:
[0025] Step 1: After detecting a low, slow, and small target, the radar automatically begins to acquire the target's trajectory information in batches;
[0026] That is, the radar automatically batch-acquires target trajectory information after detecting a target to determine whether it is a low, slow, and small target;
[0027] Step 2: Process motion feature information and visualize track information to obtain motion features and track features:
[0028] Motion feature processing: Extract 6-dimensional motion parameters, including speed, heading, acceleration, bearing, pitch, and distance, to form an N×6 motion feature matrix, where N represents the number of points in this motion feature;
[0029] Track feature processing: Filtering four-dimensional information—timestamp, distance, bearing, and echo intensity—to construct a two-dimensional track map.
[0030] (1) Each D meter is a pixel. Calculate the pixel in the track map according to the distance and orientation. Construct a complete two-dimensional track map with the first timestamp as the starting point and the last timestamp as the ending point. In this embodiment, D is selected as 20 meters. By setting a certain distance of pixels, the points that the target passes through can be observed intuitively in the two-dimensional data to judge the speed of the target's movement.
[0031] (2) After normalizing the echo intensity to [0,1], map it to the pixel value of each track point and record the variation pattern of the echo intensity in the track map;
[0032] (3) Mark the values 1 to N from the starting point to the ending point in the order of timestamps to record the direction of the target's movement. If the target has not been passed, mark it as 0. If the target is passed repeatedly, only the value of the first pass is retained.
[0033] (4) Form a two-dimensional track feature matrix of size M×M from the two-dimensional track map. The value of each track point in the array represents the track order and echo intensity.
[0034] Step 3: Convert the motion feature matrix and track feature matrix into fixed sizes:
[0035] Step 3-1: Convert the N×6 motion feature matrix into a fixed-size A1×6 input matrix. This means converting the first N-dimensional data into A1-dimensional data, while keeping the second 6-dimensional data unchanged. In this embodiment, A1 is 25.
[0036] When N < A1, the matrix is stretched from N dimensions to A1 dimensions by padding with zeros at both ends.
[0037] When N≥A1, the 1D Spatial Pyramid Network (1D-SPP) is used to stitch spatial features of different scales into a spatial matrix of fixed length A1, including:
[0038] a) Determine the pooling levels of the 1D spatial pyramid network 1D-SPP:
[0039] Suppose that 1D-SPP contains K pooling layers. In this embodiment, K=3. The target output space length of each layer is s1_1,s1_2,…,s1_i,…,s1_k, which must satisfy:
[0040]
[0041] b) Calculate the window size and stride for each pooling layer:
[0042] For the i-th pooling layer, the length s1_i of the output after pooling is controlled by the window and stride:
[0043] Pooling window size: w1_i=ceil(N / s1_i), ceil(a / b) represents the rounding up of a divided by b, ensuring that all inputs are covered;
[0044] Pooling step size: stride1_i=floor(N / s1_i), floor(a / b) represents the floor function of a divided by b, controlling the shift interval;
[0045] c) Perform multi-scale 1D
[0046] For each of the 6 channels of input X1, perform K=3 1D pooling operations respectively:
[0047] The expression for element X1_i in the feature map obtained after the i-th pooling is:
[0048]
[0049] in, `pool` is the pooling function;
[0050] d) Concatenate to obtain output Y1
[0051] Concatenate the spatial dimensions of the k feature maps:
[0052]
[0053] in, For concatenation functions, because Therefore, the Y1 matrix is an A1×6 dimensional eigenma matrix.
[0054] Step 3-2: Convert the M×M track feature matrix into a fixed-size A2×A2 input matrix. In this embodiment, A12 is 25.
[0055] When M < A2, in order to ensure the integrity of the internal structure, the surrounding zero-padding method is used to stretch the M×M dimension to A2×A2 dimension;
[0056] When M≥A2, the Spatial Pyramid Network (SPP) is used to transform spatial features M×M at different scales into a fixed-length A2×A2 spatial matrix, including:
[0057] a) Design SPP pooling layer branches:
[0058] Select three pooling layer branches and determine the output sizes s2_1, s2_2, and s2_3 respectively;
[0059] b) Calculate the stride and window size for each pooling layer.
[0060] For the i-th pooling layer, the length s2_i of the output after pooling needs to be controlled by the window and stride, as shown in the formula:
[0061] Pooling step size: stride2_i=floor(M / s2_i), where i=1,2,3, and floor(a / b) represents the floor function of a divided by b, controlling the shift interval;
[0062] Pooling kernel size: w2_i = M - stride2_i × (s2_i - 1), ensuring coverage of all inputs;
[0063] c) Pooling operations
[0064] Pooling the input X2 yields the minimum-size feature map:
[0065]
[0066] in, Represents the max pooling function;
[0067] d) Fusion output
[0068] Adding the three pooled results together yields the final fused result Y2:
[0069]
[0070] Step 4: Feed the fixed-length track features and motion features into the initial classification network built on LSTM to generate the initial classification results;
[0071] like Figure 2 As shown, the initial classification network includes two independent sub-networks, each of which contains two LSTM layers, one fully connected layer, and one SoftMax layer;
[0072] Two independent sub-networks extract features from different dimensions and generate initial classification results, which are then adaptively weighted and fused to obtain the final recognition result.
[0073] One of the sub-networks takes an A1×6 dimensional motion feature matrix as input, extracts temporal features through two layers of LSTM (128 units per layer), and outputs a 6-dimensional classification result Z1.
[0074] In this branch, by inputting track information in the form of a two-dimensional image, the LSTM network can capture the distribution characteristics and changing trends of the target track in the spatial dimension, complementing the motion feature branch in terms of data dimension and feature extraction angle.
[0075] The other network takes an A2×A2 dimensional track feature matrix as input, extracts spatiotemporal features through two layers of LSTM, and outputs a 6-dimensional classification result Z2.
[0076] During feature extraction, the gating mechanism of the LSTM network can effectively handle long-distance dependencies in the data. The motion feature branch can capture the dynamic changes of the target's motion state over time, while the trajectory feature branch can uncover the spatial evolution of the target's trajectory.
[0077] Step 5: The target classification results are fused using an adaptive weighted summation strategy to obtain accurate classification results.
[0078] Based on the adaptive fusion layer, the weights α and β are optimized through end-to-end learning, and the fusion result is output:
[0079] Z = α*Z1 + β*Z2
[0080] Here, softmax normalization ensures that α+β=1.
[0081] In the design of the adaptive fusion layer, the weights α and β are dynamically adjusted through end-to-end learning, enabling the network to flexibly allocate the contribution of motion features and trajectory features in the final decision based on different scenarios and data characteristics. This adaptive mechanism avoids the limitations of manually setting weights, ensuring that in the complex and ever-changing environment of airports, the model can fully utilize the complementary advantages of the two types of features to achieve more accurate target classification.
[0082] This solution deconstructs raw radar point data into "motion feature sequences" (including temporal features such as velocity, acceleration, and angular velocity) and "track topology features" (such as trajectory curvature and altitude-range envelope geometric features), and inputs them into a bidirectional LSTM network to extract temporal dependencies. This effectively solves the problem of difficulty in distinguishing low, slow, and small targets in airport environments, and provides strong support for the engineering application and promotion of low, slow, and small target classification.
[0083] This solution also provides a target classification system that fuses motion features and track features, including the following modules:
[0084] Data acquisition module: used to automatically acquire the target's trajectory information in batches after detecting a low, slow, and small target;
[0085] Feature information acquisition module: used to process motion feature information and visualize track information to obtain motion features and track features;
[0086] Feature information processing module: used to convert motion feature matrices and track feature matrices into fixed-size matrices;
[0087] The classification module is used to feed fixed-length track features and motion features into an initial classification network based on LSTM to generate initial classification results.
[0088] An adaptive weighted summation strategy is used to fuse the target classification results to obtain an accurate classification result.
[0089] This solution also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0090] Step 1: After detecting a low, slow, and small target, the radar automatically begins to acquire the target's trajectory information in batches;
[0091] Step 2: Process motion feature information and visualize track information to obtain motion features and track features;
[0092] Step 3: Convert the motion feature matrix and track feature matrix into fixed sizes;
[0093] Step 4: Feed the fixed-length track features and motion features into the initial classification network built on LSTM to generate the initial classification results;
[0094] Step 5: The target classification results are fused using an adaptive weighted summation strategy to obtain accurate classification results.
[0095] The embodiments described above are merely one implementation method of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A target classification method that fuses motion features and track features, characterized in that, Includes the following steps: Step 1: After detecting a low, slow, and small target, the radar automatically begins to acquire the target's trajectory information in batches; Step 2: Process motion feature information and visualize track information to obtain motion features and track features; Step 3: Convert the motion feature matrix and track feature matrix into fixed sizes; Step 4: Feed the fixed-length track features and motion features into the initial classification network built on LSTM to generate the initial classification results; Step 5: The target classification results are fused using an adaptive weighted summation strategy to obtain accurate classification results.
2. The target classification method based on the fusion of motion features and track features according to claim 1, characterized in that, The motion characteristics and track characteristics in step 2 are specifically as follows: Motion feature processing: Extract 6-dimensional motion parameters, including speed, heading, acceleration, bearing, pitch, and distance, to form an N×6 motion feature matrix, where N represents the number of points in this motion feature; Track feature processing: Filter four-dimensional information such as timestamp, distance, bearing, and echo intensity to construct a two-dimensional track map.
3. The target classification method based on the fusion of motion features and track features according to claim 2, characterized in that, The process of constructing the two-dimensional track map is as follows: (1) Using each D meter as a pixel, calculate the pixel to which the point belongs in the track map based on the distance and orientation. Using the first timestamp as the starting point and the last timestamp as the ending point, construct a complete two-dimensional track map. (2) After normalizing the echo intensity to [0,1], map it to the pixel value of each track point and record the variation pattern of the echo intensity in the track map; (3) Mark the values 1 to N from the starting point to the ending point in the order of timestamps to record the direction of the target's movement. If the target has not been passed, mark it as 0. If the target is passed repeatedly, only the value of the first pass is retained. (4) Form a two-dimensional track feature matrix of size M×M from the two-dimensional track map. The value of each track point in the array represents the track order and echo intensity.
4. The target classification method based on the fusion of motion features and track features according to claim 1, characterized in that, Step 3, converting the motion feature matrix and track feature matrix into a fixed size, specifically involves: Step 3-1: Convert the motion feature matrix of N×6 into a fixed-size A1×6 input matrix, that is, convert the first N-dimensional data into A1-dimensional data, while keeping the second 6-dimensional data unchanged. Step 3-2: Convert the M×M track feature matrix into a fixed-size A2×A2 input matrix.
5. The target classification method based on the fusion of motion features and track features according to claim 4, characterized in that, The process of transforming the motion feature matrix in step 3-1 is as follows: When N < A1, the matrix is stretched from N dimensions to A1 dimensions by padding with zeros at both ends. When N≥A1, the 1D Spatial Pyramid Network (1D-SPP) is used to stitch spatial features of different scales into a spatial matrix of fixed length A1, including: a) Determine the pooling levels of the 1D spatial pyramid network 1D-SPP: Suppose that 1D-SPP contains K pooling layers, and the target output space length of each layer is s1_1, s1_2, ..., s1_i, ..., s1_k, which must satisfy: ; b) Calculate the window size and stride for each pooling layer: For the i-th pooling layer, the length s1_i of the output after pooling is controlled by the window and stride: Pooling window size: w1_i=ceil(N / s1_i), ceil(a / b) represents the rounding up of a divided by b, ensuring that all inputs are covered; Pooling step size: stride1_i=floor(N / s1_i), floor(a / b) represents the floor function of a divided by b, controlling the shift interval; c) Perform multi-scale 1D For each of the 6 channels of input X1, perform K 1D pooling operations respectively: The expression for element X1_i in the feature map obtained after the i-th pooling is: ; in, `pool` is the pooling function; d) Concatenate to obtain output Y1 Concatenate the spatial dimensions of the k feature maps: ; in, For concatenation functions, because Therefore, the Y1 matrix is an A1×6 dimensional characteristic matrix.
6. The target classification method based on the fusion of motion features and track features according to claim 4, characterized in that, The process for transforming the track feature matrix in step 3-2 is as follows: When M < A2, in order to ensure the integrity of the internal structure, the surrounding zero-padding method is used to stretch the M×M dimension to A2×A2 dimension; When M≥A2, the Spatial Pyramid Network (SPP) is used to transform spatial features M×M at different scales into a fixed-length A2×A2 spatial matrix, including: a) Design SPP pooling layer branches: Select three pooling layer branches and determine the output sizes s2_1, s2_2, and s2_3 respectively; b) Calculate the stride and window size for each pooling layer. For the i-th pooling layer, the length s2_i of the output after pooling needs to be controlled by the window and stride, as shown in the formula: Pooling step size: stride2_i=floor(M / s2_i), where i=1,2,3, and floor(a / b) represents the floor function of a divided by b, controlling the shift interval; Pooling kernel size: w2_i = M - stride2_i × (s2_i - 1), ensuring coverage of all inputs; c) Pooling operations Pooling the input X2 yields the minimum-size feature map: ; in, Represents the max pooling function; d) Fusion output Adding the three pooled results together yields the final fused result Y2: 。 7. The target classification method based on the fusion of motion features and track features according to claim 4, characterized in that, The generation of the initial classification result in step 4 is specifically as follows: The initial classification network includes two independent sub-networks, each of which contains two LSTM layers, one fully connected layer, and one SoftMax layer; Two independent sub-networks extract features from different dimensions and generate initial classification results, which are then adaptively weighted and fused to obtain the final recognition result. One of the sub-networks takes an A1×6 dimensional motion feature matrix as input, extracts temporal features through two layers of LSTM, and outputs a 6-dimensional classification result Z1. The other network takes an A2×A2 dimensional track feature matrix as input, extracts spatiotemporal features through two layers of LSTM, and outputs a 6-dimensional classification result Z2.
8. The target classification method based on the fusion of motion features and track features according to claim 7, characterized in that, Obtaining the accurate classification result in step 5 specifically involves: Based on the adaptive fusion layer, the weights α and β are optimized through end-to-end learning, and the fusion result is output: Z = α*Z1 + β*Z2 Where α+β=1.
9. A target classification system that fuses motion characteristics and track characteristics, characterized in that, Includes the following modules: Data acquisition module: used to automatically acquire the target's trajectory information in batches after detecting a low, slow, and small target; Feature information acquisition module: used to process motion feature information and visualize track information to obtain motion features and track features; Feature information processing module: used to convert motion feature matrices and track feature matrices into fixed-size matrices; The classification module is used to feed fixed-length track features and motion features into an initial classification network based on LSTM to generate initial classification results. An adaptive weighted summation strategy is used to fuse the target classification results to obtain an accurate classification result.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.