A trajectory feature extraction method based on phase alignment and frequency domain gating
Patent Information
- Application Number
- CN202610948452.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-29
AI Technical Summary
纯时间域方法混合建模,难以清晰区分长期趋势与局部动态
[0075]频域显式建模:在轨迹特征提取阶段显式引入频域建模,使轨迹的长期趋势(低频)与局部细节(高频)在频谱结构下被清晰表征。
Smart Images

Figure CN122471026B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to a multi-scale frequency domain attention trajectory feature extraction method based on phase alignment and high-frequency gating for trajectory prediction. Background Technology
[0002] In autonomous driving systems, trajectory prediction aims to infer the future trajectories of traffic participants based on their historical movement patterns. Existing methods typically extract features from historical trajectories and then combine this with information from maps and interactions to predict future trajectories. Current trajectory feature extraction methods are mostly based on time-domain sequence modeling (such as recurrent neural networks, convolutional neural networks, graph neural networks, or Transformer structures). While these methods can learn the evolution of trajectories over time to some extent, they still have the following problems:
[0003] First, historical trajectories are essentially time-series signals superimposed with different frequency components: low frequencies reflect overall trends and long-term intentions, while high frequencies embody local curvature, instantaneous maneuvers, and motion details. Pure time-domain methods, using a hybrid modeling approach, struggle to clearly distinguish between long-term trends and local dynamics.
[0004] Second, in multi-scale convolution, pooling and other hierarchical feature extraction, high-frequency details are easily weakened by smoothing layer by layer. Although noise is suppressed, it also weakens real local maneuver features such as turning and lane changing, reducing the prediction model's ability to perceive fine movements.
[0005] Third, similar movement patterns among different traffic participants may exhibit temporal differences (such as inconsistent lane change starting points), which manifests as local misalignment in the time domain and phase shift in the frequency domain. Existing frequency domain methods tend to focus on amplitude distribution but neglect phase alignment, leading to insufficient pattern matching.
[0006] Fourth, simply performing a Fourier transform on the trajectory and then concatenating it with time-domain features lacks adaptive adjustment for the importance of different frequency bands, and also lacks a mechanism to stably fuse the frequency-domain enhancement results back to the time domain, resulting in unstable training or poor enhancement effects.
[0007] Therefore, there is an urgent need for a new trajectory feature extraction method that explicitly introduces frequency domain modeling into multi-scale feature extraction, corrects phase shift, and performs adaptive gating enhancement on the high-frequency part to obtain a more discriminative and stable trajectory feature representation, providing high-quality input for subsequent prediction. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention provides a trajectory feature extraction method based on phase alignment and frequency domain gating, comprising the following steps:
[0009] Step 1: Obtain the historical trajectory sequences of several traffic participants in the traffic scenario;
[0010] Step 2: Perform local coordinate normalization and tensor quantization on the historical trajectory to form trajectory input features;
[0011] Step 3: Perform multi-scale temporal domain convolution on the trajectory input features to extract temporal features at different time scales. Each scale of temporal features corresponds to a scale layer, thus constructing multi-scale temporal domain initial trajectory features.
[0012] Step 4: Perform phase-aligned frequency domain attention calculations within each scale layer:
[0013] Step 41: Based on the temporal features of the scale layer, generate query features, key features, and value features;
[0014] Step 42: Transform the query features, key features, and value features to the frequency domain to obtain the corresponding frequency domain query, frequency domain key, and frequency domain value;
[0015] Step 43: Apply learnable phase rotations in opposite directions to the frequency domain query and the frequency domain key to achieve phase alignment;
[0016] Step 44: Based on the phase-aligned frequency domain query and frequency domain key, calculate the frequency domain attention weights, and use the frequency domain attention weights to weight the frequency domain values to obtain the frequency domain attention result;
[0017] Step 5: Divide the frequency domain attention result into low-frequency and high-frequency parts, and use learnable high-frequency gating parameters to adaptively enhance or suppress the high-frequency part to obtain the enhanced frequency domain features.
[0018] Step 6: Inversely transform the enhanced frequency domain features back to the time domain, and fuse them with the original time domain features of the scale layer through residual connection to obtain the enhanced time domain features of the scale layer.
[0019] Step 7: Fuse the enhanced temporal features from all scale layers to output the final trajectory feature representation.
[0020] Furthermore, in step 1, the historical trajectory sequence of traffic participants includes a state vector of a certain time window length. The state vector includes the position coordinates of the traffic participants, as well as one or more types of information such as velocity components, heading angle, acceleration, target category encoding, or scene context.
[0021] Furthermore, in step 2, for the position vector at any moment in the historical trajectory sequence, the local coordinate normalization transforms the historical trajectory to a local coordinate system based on the heading angle at the current moment through a rotation matrix. After normalization, the historical trajectories of all targets are organized into tensor form.
[0022] Furthermore, in step 3, the multi-scale temporal convolution includes at least two layers of residual one-dimensional convolution modules. First, the historical trajectory tensor is input into the first layer of residual one-dimensional convolution module to obtain shallow temporal features. Then, after downsampling, it is input into the next layer of residual one-dimensional convolution module to finally obtain deep temporal features.
[0023] Furthermore, step 3, which involves constructing multi-scale temporal domain initial trajectory features, specifically includes:
[0024] Input the historical trajectory tensor into the first layer residual one-dimensional convolution module to obtain shallow temporal features;
[0025] The shallow temporal features are downsampled and input into the second layer residual one-dimensional convolution module to obtain the middle temporal features;
[0026] The mid-level temporal features are downsampled and input into the third-layer residual one-dimensional convolution module to obtain deep temporal features.
[0027] Furthermore, in step 41, for the first... Temporal features of each scale layer The query features are obtained through three independent linear transformations. Key features Sum value characteristics :
[0028]
[0029] in, , , These are the learnable linear mapping parameters for the corresponding features;
[0030] In step 42, the above query features are analyzed along the time dimension. Key features Sum value characteristics Perform a frequency domain transformation to obtain the corresponding frequency domain query. Frequency domain key and frequency domain value express:
[0031]
[0032] in, This represents the Discrete Fourier Transform; after the transform, , as well as All are complex frequency domain tensors, each containing amplitude and phase information of the corresponding feature.
[0033] Furthermore, in step 43, a learnable phase rotation parameter is introduced. :
[0034]
[0035] in, Indicates the first Number of channels per scale layer Indicates the first The time length of each scale layer;
[0036] Based on the phase rotation parameters, a frequency domain query is performed. and frequency domain key Apply phase rotation in the opposite direction:
[0037]
[0038]
[0039] in, , These represent the frequency domain lookup and frequency domain key after phase rotation, respectively. Indicates frequency index, The imaginary unit, Represents the natural constant. This represents element-wise multiplication;
[0040] Step 44: Calculate the frequency domain attention weights based on the frequency domain query and frequency domain key after phase rotation alignment. as follows:
[0041]
[0042] in, This represents the softmax function. To take the real part of a complex number, This indicates the conjugate transpose. To scale the dimensions; , They represent the first Frequency domain query and frequency domain key after phase rotation at each scale level;
[0043] Then, the frequency domain attention weights are used to adjust the frequency domain values. Weighting is performed to obtain the frequency domain attention result. :
[0044] .
[0045] Furthermore, in step 5, let the first... The frequency band boundary corresponding to each scale layer is The low-frequency mask and high-frequency mask are constructed as follows:
[0046]
[0047]
[0048] in, For high-frequency mask, For low-frequency masking, Indicates frequency index;
[0049] Based on low-frequency and high-frequency masks, the frequency domain attention results are... Divided into low frequency part and high frequency part :
[0050]
[0051]
[0052] in, , They represent the first Low-frequency and high-frequency masks for each scale layer This indicates element-wise multiplication; the low-frequency part mainly represents long-term motion trends, while the high-frequency part mainly represents local short-term maneuvers and trajectory detail changes.
[0053] Define high-frequency gating factor for:
[0054]
[0055] in, For learnable parameter vectors, Represents the hyperbolic tangent function; The range of values is (0, 2), when When, the high-frequency component is suppressed; when When, the high-frequency part remains unchanged; when The high-frequency components are enhanced;
[0056] By applying a high-frequency gating factor to the high-frequency portion, the enhanced frequency domain features are obtained. :
[0057] .
[0058] Furthermore, in step 5, the frequency band boundary Based on the adaptive determination of the spectral energy distribution of the current trajectory characteristics, the cumulative normalized spectral energy is calculated as follows:
[0059]
[0060] in, Indicates the first Each scale in the frequency index The cumulative normalized spectral energy at the location, Indicates the first Each scale in the frequency index Normalized spectral energy at the location, Indicates frequency index, This indicates the current frequency index and the upper limit of the cumulative summation; the frequency band boundary is determined based on the cumulative energy.
[0061]
[0062] in, ∈(0,1) represents the preset target low-frequency energy ratio.
[0063] Furthermore, in step 6, the enhanced frequency domain features are... Performing the inverse Fourier transform yields the corresponding... Time-domain reconstruction features at each scale level :
[0064]
[0065] in, This represents the inverse discrete Fourier transform. This indicates taking the real part of a complex number;
[0066] right Perform instance normalization:
[0067]
[0068] in, This represents the time-domain features after instance normalization. Indicates instance normalization;
[0069] By using residual connection method With the Original temporal features of each scale layer By merging, we obtain the first... Enhanced temporal features at each scale layer :
[0070]
[0071] in, This is the residual scaling factor.
[0072] Furthermore, in step 7, the enhanced temporal features of each scale layer are projected onto a unified channel dimension, and then a top-down cross-scale fusion method is used to perform multi-scale fusion.
[0073] After completing multi-scale fusion, the final fused features are pooled or time-selected to obtain trajectory-level feature representations.
[0074] Compared with the prior art, the present invention has the following beneficial effects:
[0075] Explicit frequency domain modeling: Explicitly introducing frequency domain modeling in the trajectory feature extraction stage allows the long-term trend (low frequency) and local details (high frequency) of the trajectory to be clearly represented in the spectral structure.
[0076] Phase alignment: Through a learnable phase alignment rotation mechanism, the timing misalignment between different trajectories is corrected, the comparability of different trajectory patterns in the frequency domain space is improved, and the ability of frequency domain attention to recognize real motion patterns is enhanced.
[0077] Adaptive high-frequency enhancement: Through a frequency-domain gated attention mechanism, it adaptively enhances high-frequency local maneuver information (such as turning and lane changing) that is beneficial to the prediction task, while suppressing noisy high-frequency disturbances.
[0078] Stable back-fusion mechanism: By fusing frequency domain enhancement information back to time domain features through residual connections, the continuity of the time domain is preserved and the discriminative power of the frequency domain is supplemented, making training more stable.
[0079] Multi-scale fusion: It integrates multi-scale features, taking into account both macro trends and micro dynamics, and provides high-quality input for downstream prediction tasks. Attached Figure Description
[0080] Figure 1 This is a schematic diagram of the overall process of Embodiment 1 of the present invention;
[0081] Figure 2 This is a schematic diagram of the overall framework for multi-scale frequency domain attention trajectory feature extraction in Embodiment 1 of the present invention;
[0082] Figure 3 This is a schematic diagram of the phase-aligned frequency domain attention unit structure of a single scale layer in Embodiment 1 of the present invention;
[0083] Figure 4 This is a schematic diagram of high-frequency gated enhancement and time-domain residual fusion in Embodiment 1 of the present invention;
[0084] Figure 5 This is a schematic diagram of multi-scale fusion and final trajectory feature output in Embodiment 1 of the present invention. Detailed Implementation
[0085] Example 1
[0086] See Figure 1-5As shown, this embodiment provides a trajectory feature extraction method based on phase alignment and frequency domain gating. Its core concept is not simply "performing a Fourier transform on the trajectory," but rather embedding frequency domain modeling into the multi-scale trajectory feature extraction process. The following processing links are completed at each scale layer, such as... Figure 1 As shown, it includes the following steps:
[0087] Step 1: Obtain the historical trajectory sequences of several traffic participants in the traffic scenario;
[0088] Step 2: Perform local coordinate normalization and tensor quantization on the historical trajectory to form trajectory input features;
[0089] Step 3: Perform multi-scale temporal domain convolution on the trajectory input features to extract temporal features at different time scales. Each scale of temporal features corresponds to a scale layer, thus constructing multi-scale temporal domain initial trajectory features.
[0090] Step 4: Perform phase-aligned frequency domain attention calculations within each scale layer:
[0091] Step 41: Based on the temporal features of the scale layer, generate query features, key features, and value features;
[0092] Step 42: Transform the query features, key features, and value features to the frequency domain to obtain the corresponding frequency domain query, frequency domain key, and frequency domain value;
[0093] Step 43: Apply learnable phase rotations in opposite directions to the frequency domain query and the frequency domain key to achieve phase alignment;
[0094] Step 44: Based on the phase-aligned frequency domain query and frequency domain key, calculate the frequency domain attention weights, and use the frequency domain attention weights to weight the frequency domain values to obtain the frequency domain attention result;
[0095] Step 5: Divide the frequency domain attention result into low-frequency and high-frequency parts, and use learnable high-frequency gating parameters to adaptively enhance or suppress the high-frequency part to obtain the enhanced frequency domain features.
[0096] Step 6: Inversely transform the enhanced frequency domain features back to the time domain, and fuse them with the original time domain features of the scale layer through residual connection to obtain the enhanced time domain features of the scale layer.
[0097] Step 7: Fuse the enhanced temporal features from all scale layers to output the final trajectory feature representation.
[0098] Furthermore, step 1 of this embodiment includes the following steps:
[0099] Suppose there are a total of The traffic participants to be modeled are denoted as:
[0100]
[0101] in, This represents the set of all traffic participants to be modeled in a traffic scenario. Indicates the first One traffic participant to be modeled;
[0102] For the Each traffic participant, at the current moment Previously, the length was obtained as Historical trajectory sequence:
[0103]
[0104] in, Represents a historical trajectory sequence. They represent , to The state vector at any given time;
[0105] No. The state vector at time 1 Represented as:
[0106]
[0107] in, These are the position coordinates in the x and y directions. These are the velocity components in the x and y directions.
[0108] In other implementations, additional information such as heading angle, acceleration, target category encoding, or scene context may be further incorporated into the state vector.
[0109] Furthermore, step 2 of this embodiment includes the following steps:
[0110] To minimize the impact of global coordinate system changes on trajectory modeling, this invention preferably transforms the historical trajectory to a local coordinate system. For the position vector at any given moment in the historical trajectory sequence... The local coordinate normalization transforms the historical trajectory to a local coordinate system based on the current heading angle using a rotation matrix. The position vector after transformation to the local coordinate system... Represented as:
[0111]
[0112] in, Let the heading angle of the target be at the current moment. It is a two-dimensional rotation matrix. This represents the target's current position vector.
[0113] After normalization, the historical trajectories of all targets are organized into tensor form:
[0114]
[0115] in, Represents the historical trajectory tensor. For the number of traffic participants, The length of the time window. This represents the trajectory feature dimension at each moment. In a preferred embodiment, In another embodiment, corresponding to two-dimensional position coordinates and two-dimensional velocity components, Alternatively, a higher-dimensional trajectory feature dimension can be used. This representation preserves the temporal evolution of the trajectory and facilitates subsequent multi-scale convolution and frequency domain processing.
[0116] Furthermore, step 3 of this embodiment includes the following steps:
[0117] To simultaneously extract trajectory pattern information at different time scales, this invention employs a multi-scale feature extraction structure. First, the historical trajectory tensor... Inputting the first layer residual one-dimensional convolutional module yields shallow temporal features. :
[0118]
[0119] in, This represents the first residual one-dimensional convolutional module, which consists of two one-dimensional convolutional layers, a normalization layer, and an activation function, and retains the original trajectory information through residual connections;
[0120] The shallow temporal features are downsampled and input into the second-layer residual one-dimensional convolution module to obtain the mid-layer temporal features. :
[0121]
[0122] in, This represents the second layer of residual one-dimensional convolutional modules, which has the same structure as the first layer of residual one-dimensional convolutional modules. For downsampling;
[0123] The mid-level temporal features are downsampled and input into the third-layer residual one-dimensional convolution module to obtain the deep temporal features. :
[0124]
[0125] in, This represents the third layer of residual one-dimensional convolutional modules, which has the same structure as the first layer of residual one-dimensional convolutional modules. This is for downsampling.
[0126] In a non-limiting embodiment, the output dimensions of each scale layer are as follows:
[0127]
[0128]
[0129]
[0130] Among them, shallow time-domain features focus more on local short-term dynamics, mid-level time-domain features take into account both local and medium-term relationships, and deep time-domain features are more conducive to encoding long-term motion trends. It should be noted that the present invention is not limited to a three-layer structure, and can also be configured with two, four or more layers as needed.
[0131] Furthermore, step 4 of this embodiment includes the following steps:
[0132] The core of this invention lies in embedding a phase-aligned frequency domain attention feature extraction unit within each scale layer.
[0133] Step 41: Based on the temporal features of the scale layer, generate query features, key features, and value features, as follows:
[0134] For the Temporal features of each scale layer The query features are obtained through three independent linear transformations. Key features Sum value characteristics :
[0135]
[0136] in, , , These are the learnable linear mapping parameters for the corresponding features;
[0137] Step 42: Transform the query features, key features, and value features to the frequency domain to obtain the corresponding frequency domain query, frequency domain key, and frequency domain value, as follows:
[0138] Along the time dimension for the above query features Key features Sum value characteristics Perform a frequency domain transformation to obtain the corresponding frequency domain query. Frequency domain key and frequency domain value express:
[0139]
[0140] in, This represents the Discrete Fourier Transform; after the transform, , as well as All are complex frequency domain tensors, each containing amplitude and phase information of the corresponding feature;
[0141] Step 43: Apply learnable phase rotations in opposite directions to the frequency domain query and the frequency domain key to achieve phase alignment, as follows:
[0142] Considering that different traffic participants may have similar movement patterns, but the timing and location of their local actions are not entirely consistent, this invention introduces a learnable phase rotation parameter. :
[0143]
[0144] in, Indicates the first Number of channels per scale layer This indicates the time length of this scale layer;
[0145] Based on the phase rotation parameters, apply phase rotations in opposite directions to the frequency domain query and the frequency domain key:
[0146]
[0147]
[0148] in, , These represent the frequency domain lookup and frequency domain key after phase rotation, respectively. Indicates frequency index, The imaginary unit, Represents the natural constant. This represents element-wise multiplication;
[0149] The effect of the above processing is that, for trajectory patterns that originally had insufficient frequency domain matching due to time misalignment, by learning appropriate phase rotation parameters, the query and key can be better aligned in the spectral space, thereby enhancing the frequency domain attention's ability to recognize similar trajectory patterns.
[0150] Step 44: Based on the frequency domain query and frequency domain key after phase rotation alignment, calculate the frequency domain attention weights, and use these weights to weight the frequency domain values to obtain the frequency domain attention result, as follows:
[0151] After obtaining the phase-aligned frequency domain query and frequency domain key, calculate the frequency domain attention weights. as follows:
[0152]
[0153] in, This represents the softmax function. To take the real part of a complex number, This indicates the conjugate transpose. To scale the dimensions; , They represent the first Frequency domain query and frequency domain key after phase rotation at each scale level;
[0154] The frequency domain values are then weighted using the aforementioned frequency domain attention weights to obtain the frequency domain attention result. :
[0155]
[0156] This process does not focus attention between time points, but rather establishes correlations between frequency components, achieving dependency modeling in the frequency component space. This allows the model to learn the collaborative relationships between different frequency bands of the trajectory at the spectral level, thus making it more conducive to characterizing the contribution relationship of different frequency bands to the trajectory motion pattern.
[0157] Furthermore, step 5 of this embodiment includes the following steps:
[0158] To prevent high-frequency details from being weakened layer by layer during feature extraction, this invention divides the frequency domain attention results into high and low frequencies, and further introduces learnable high-frequency gating parameters into the high-frequency components of the frequency domain attention results. The steps are as follows:
[0159] Let the first The frequency band boundary corresponding to each scale layer is The low-frequency mask and high-frequency mask are constructed as follows:
[0160]
[0161]
[0162] in, For high-frequency mask, For low-frequency masking, Indicates frequency index;
[0163] Based on low-frequency and high-frequency masks, the frequency domain attention results are... Divided into low frequency part and high frequency part :
[0164]
[0165]
[0166] in, , They represent the first Low-frequency and high-frequency masks for each scale layer This indicates element-wise multiplication; the low-frequency part mainly represents long-term motion trends, while the high-frequency part mainly represents local short-term maneuvers and trajectory detail changes.
[0167] To avoid indiscriminately amplifying high-frequency information in all scenarios, and to enable the model to adaptively enhance or suppress high-frequency information according to task requirements, this invention employs learnable high-frequency gating parameters to adaptively enhance or suppress the high-frequency components. The high-frequency gating factor... Defined as:
[0168]
[0169] in, For learnable parameter vectors, Represents the hyperbolic tangent function; because The range of its value is (-1, 1), therefore The range of is (0, 2), which means that when When, the high-frequency component is suppressed; when When, the high-frequency part remains unchanged; when The high-frequency components are enhanced;
[0170] The advantage of adopting the above form is that it can both suppress and enhance, achieving adaptive adjustment.
[0171] By applying a high-frequency gating factor to the high-frequency portion, the enhanced frequency domain features are obtained. :
[0172]
[0173] Through this mechanism, the model can automatically learn which channels' high-frequency information is more helpful in expressing local actions and which channels' high-frequency components are more likely to originate from noise, thereby achieving targeted enhancement and suppression.
[0174] Furthermore, step 6 of this embodiment includes the following steps:
[0175] To reintroduce frequency domain enhancement information into trajectory time domain features, this invention enhances the frequency domain features. Perform an inverse Fourier transform, then inversely transform back to the time domain to obtain the corresponding... Time-domain reconstruction features at each scale level :
[0176]
[0177] in, This represents the inverse discrete Fourier transform (IDFT). This indicates taking the real part of a complex number;
[0178] To reduce scale differences between different channels and frequency bands, instance normalization is performed:
[0179]
[0180] in, This represents the time-domain features after instance normalization. Indicates instance normalization;
[0181] Then, it is connected to the first through residual connection. Original temporal features of each scale layer By merging, we obtain the first... Enhanced temporal features at each scale layer :
[0182]
[0183] in, This is the residual scaling factor, which can be a fixed hyperparameter or a learnable parameter.
[0184] Through the above design, the enhanced frequency domain features do not directly replace the original time domain features, but are supplemented in an incremental form. Therefore, while preserving the continuous information in the time domain, the trend and detail discrimination capabilities brought about by frequency domain enhancement are introduced.
[0185] Furthermore, step 7 of this embodiment includes the following steps:
[0186] To obtain enhanced temporal features at each scale level , and Then, the enhanced temporal features of all scale layers are fused. First, the enhanced temporal features of each scale layer are projected onto a unified channel dimension through a 1×1 convolution or linear layer. The projected enhanced temporal features are denoted as follows: , , Then, a top-down cross-scale fusion approach is used for multi-scale fusion:
[0187]
[0188]
[0189]
[0190] in, , This indicates the cross-scale fusion characteristics between layer 3 and layer 2. These represent the cross-scale fusion features of all layers, This indicates an upsampling operation along the time dimension. This indicates a structure with two layers of one-dimensional convolutions and residual connections, using the GeLU activation function.
[0191] After completing multi-scale fusion, pooling or time selection is performed on the final fused features to obtain trajectory-level feature representations:
[0192]
[0193] in, This indicates one of the following: average pooling, max pooling, or attention pooling; or it can directly take the features from the last time step. Represents a multilayer perceptron;
[0194] The final result This is the final trajectory feature representation output by the method proposed in this invention; this feature can be used as input to the subsequent trajectory prediction decoder, or it can be used for behavior recognition, risk prediction or other downstream tasks.
[0195] Example 2
[0196] This embodiment provides a trajectory feature extraction method based on phase alignment and frequency domain gating. Building upon Embodiment 1, step 5 involves the use of adaptive frequency band boundaries for high-frequency and low-frequency division. Instead of using a fixed value, it is adaptively determined based on the spectral energy distribution of the current trajectory characteristics. In this embodiment, the cumulative normalized spectral energy is first calculated:
[0197]
[0198] in, Indicates the first Each scale in the frequency index The cumulative normalized spectral energy at the location, Indicates the first Each scale in the frequency index Normalized spectral energy at the location, Indicates frequency index, This indicates the current frequency index and the upper limit of the cumulative summation; further, the frequency band boundary is determined based on the cumulative energy.
[0199]
[0200] in, ∈(0,1) represents the preset target low-frequency energy ratio.
[0201] Since low-frequency components contain the majority of the total spectral energy, the boundary can be determined based on the target low-frequency proportion or scene statistics. This implementation method can improve the adaptability of the present invention to different datasets, different sampling frequencies and different motion modes.
[0202] Example 3
[0203] This embodiment provides a trajectory feature extraction method based on phase alignment and frequency domain gating, which is jointly trained with a downstream prediction module; assuming that the final trajectory features... The prediction head outputs the future trajectory as The true future trajectory is Then a predictive loss can be constructed. for:
[0204]
[0205] or:
[0206]
[0207] in, The L1 norm represents the error between the predicted future trajectory and the actual future trajectory. This represents the smooth L1 norm between the predicted future trajectory and the actual future trajectory. It employs a quadratic penalty when the error is small and a linear penalty when the error is large.
[0208] The prediction loss is then used to jointly optimize the convolution parameters, phase alignment parameters, high-frequency gating parameters, and attention parameters in this invention.
Claims
1. A trajectory feature extraction method based on phase alignment and frequency domain gating, characterized in that: Includes the following steps: Step 1: Obtain the historical trajectory sequences of several traffic participants in the traffic scenario; Step 2: Perform local coordinate normalization and tensor quantization on the historical trajectory to form trajectory input features; Step 3: Perform multi-scale temporal domain convolution on the trajectory input features to extract temporal features at different time scales. Each scale of temporal features corresponds to a scale layer, thus constructing multi-scale temporal domain initial trajectory features. Step 4: Perform phase-aligned frequency domain attention calculations within each scale layer: Step 41: Based on the temporal features of the scale layer, generate query features, key features, and value features; Step 42: Transform the query features, key features, and value features to the frequency domain to obtain the corresponding frequency domain query, frequency domain key, and frequency domain value; Step 43: Apply learnable phase rotations in opposite directions to the frequency domain query and the frequency domain key to achieve phase alignment; Step 44: Based on the phase-aligned frequency domain query and frequency domain key, calculate the frequency domain attention weights, and use the frequency domain attention weights to weight the frequency domain values to obtain the frequency domain attention result; Step 5: Divide the frequency domain attention result into low-frequency and high-frequency parts, and use learnable high-frequency gating parameters to adaptively enhance or suppress the high-frequency part to obtain the enhanced frequency domain features. In step 5, let the first... The frequency band boundary corresponding to each scale layer is The low-frequency mask and high-frequency mask are constructed as follows: in, For high-frequency mask, For low-frequency masking, Indicates frequency index; Based on low-frequency and high-frequency masks, the frequency domain attention results are... Divided into low frequency part and high frequency part : in, , They represent the first Low-frequency and high-frequency masks for each scale layer This indicates element-wise multiplication; the low-frequency part mainly represents long-term motion trends, while the high-frequency part mainly represents local short-term maneuvers and trajectory detail changes. Define high-frequency gating factor for: in, For learnable parameter vectors, Represents the hyperbolic tangent function; The range of values is (0, 2), when When, the high-frequency component is suppressed; when When, the high-frequency part remains unchanged; when The high-frequency components are enhanced; By applying a high-frequency gating factor to the high-frequency portion, the enhanced frequency domain features are obtained. : ; The frequency band boundary Based on the adaptive determination of the spectral energy distribution of the current trajectory characteristics, the cumulative normalized spectral energy is calculated as follows: in, Indicates the first Each scale in the frequency index The cumulative normalized spectral energy at the location, Indicates the first Each scale in the frequency index Normalized spectral energy at the location, Indicates frequency index, This indicates the current frequency index and the upper limit of the cumulative summation; Determine the frequency band boundary based on the cumulative energy: in, ∈(0,1) represents the preset target low-frequency energy ratio; Step 6: Inversely transform the enhanced frequency domain features back to the time domain, and fuse them with the original time domain features of the scale layer through residual connection to obtain the enhanced time domain features of the scale layer. Step 7: Fuse the enhanced temporal features from all scale layers to output the final trajectory feature representation.
2. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 1, the historical trajectory sequence of traffic participants includes a state vector of a certain time window length. The state vector includes the location coordinates of the traffic participants, as well as one or more types of information such as velocity components, heading angle, acceleration, target category encoding, or scene context.
3. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 2, for the position vector at any moment in the historical trajectory sequence, the local coordinate normalization transforms the historical trajectory to a local coordinate system based on the heading angle at the current moment through a rotation matrix. After normalization, the historical trajectories of all targets are organized into tensor form.
4. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 3, the multi-scale temporal convolution includes at least two layers of residual one-dimensional convolution modules. First, the historical trajectory tensor is input into the first layer of residual one-dimensional convolution module to obtain shallow temporal features. Then, after downsampling, it is input into the next layer of residual one-dimensional convolution module to finally obtain deep temporal features.
5. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 41, for the first Temporal features of each scale layer The query features are obtained through three independent linear transformations. Key features Sum value characteristics : in, , , These are the learnable linear mapping parameters for the corresponding features; In step 42, the above query features are analyzed along the time dimension. Key features Sum value characteristics Perform a frequency domain transformation to obtain the corresponding frequency domain query. Frequency domain key and frequency domain value express: in, This represents the Discrete Fourier Transform; after the transform, , as well as All are complex frequency domain tensors, each containing amplitude and phase information of the corresponding feature.
6. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 43, learnable phase rotation parameters are introduced. : in, Indicates the first Number of channels per scale layer Indicates the first The time length of each scale layer; Based on the phase rotation parameters, a frequency domain query is performed. and frequency domain key Apply phase rotation in the opposite direction: in, , These represent the frequency domain lookup and frequency domain key after phase rotation, respectively. Indicates frequency index, The imaginary unit, Represents the natural constant. This represents element-wise multiplication; Step 44: Calculate the frequency domain attention weights based on the frequency domain query and frequency domain key after phase rotation alignment. as follows: in, This represents the softmax function. To take the real part of a complex number, This indicates the conjugate transpose. To scale the dimensions; , They represent the first Frequency domain query and frequency domain key after phase rotation at each scale level; Then, the frequency domain attention weights are used to adjust the frequency domain values. Weighting is performed to obtain the frequency domain attention result. : 。 7. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 6, the enhanced frequency domain features are... Performing the inverse Fourier transform yields the corresponding... Time-domain reconstruction features at each scale level : in, This represents the inverse discrete Fourier transform. This indicates taking the real part of a complex number; right Perform instance normalization: in, This represents the time-domain features after instance normalization. Indicates instance normalization; By using residual connection method With the Original temporal features of each scale layer By merging, we obtain the first... Enhanced temporal features at each scale layer : in, This is the residual scaling factor.
8. The trajectory feature extraction method based on phase alignment and frequency domain gating according to claim 1, characterized in that: In step 7, the enhanced temporal features of each scale layer are projected onto a unified channel dimension, and then multi-scale fusion is performed using a top-down cross-scale fusion method. After completing multi-scale fusion, the final fused features are pooled or time-selected to obtain trajectory-level feature representations.
Citation Information
Patent Citations
Self-adaptive multi-head alignment time-frequency semantic three-mode surrounding vehicle trajectory prediction method
CN121834721A
Retinal vessel segmentation method based on frequency domain enhancement
CN122175997A