Lagrangian trajectory prediction method and system based on u-net and space-time transformer, and medium
Patent Information
- Application Number
- CN202511715644.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-11-21
AI Technical Summary
[0004]虽然CNN擅长提取环境场等网格数据中的局部空间特征,但其固有的局部感受野会限制了其在捕捉大尺度、长距离空间依赖关系的能力
本申请中,核心优势在于卓越的长期预测能力:通过在U-Net网络的每个层级嵌入“预测注意力桥接模块”,使模型在从粗到细的多个尺度上都能进行时空信息推断。通过先预测、后重建的步骤,克服传统模型中的误差累积问题。本申请创新性地实现局部特征与全局依赖的高效协同,解决现有架构的内在矛盾:通过U-Net骨干,继承了卷积神经网络(CNN)在处理网格化数据时高效提取局部空间特征的优势;同时,通过在桥接模块中引入解耦的门控注意力Transformer,又获得了强大的全局长距离时空依赖建模能力。巧妙地解决了背景技术中的核心矛盾:避免了纯CNN的全局感受野不足,也避免了纯Transformer的二次方计算复杂度和空间归纳偏置缺失。因此,本申请在实现高精度的同时,其计算成本可控,能够处理更高分辨率的环境场数据,具备更强的实际应用潜力。
Smart Images

Figure CN121502727B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of spatiotemporal sequence prediction technology, and specifically relates to a method, system and medium for predicting Lagrange trajectories based on U-Net and spatiotemporal Transformer. Background Technology
[0002] Accurate prediction of the trajectories of floating objects (such as buoys, search and rescue targets, and oil spills) is of paramount practical importance in many fields, including marine science, environmental monitoring, and maritime safety. Existing technologies are mainly divided into two categories: physical models and data-driven models. Currently, the technical approaches to trajectory prediction can be categorized into three types: physical-driven numerical models, probability-statistical prediction models, and data-driven models represented by deep learning. The first two are relatively traditional prediction methods, both of which have significant drawbacks.
[0003] With the continuous development of artificial intelligence, data-driven models centered on deep learning have shown enormous potential, capable of directly learning the intrinsic patterns of trajectory evolution from massive amounts of historical observation data. Current mainstream deep learning architectures, when applied to this task, typically use Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs) and their variants, as well as the Transformer self-attention mechanism. However, they still have the following specific technical limitations:
[0004] While CNNs excel at extracting local spatial features from gridded data such as environmental fields, their inherent local receptive field limits their ability to capture large-scale, long-range spatial dependencies. RNNs can handle temporal problems, so many methods combine CNNs and RNNs. However, RNNs generally suffer from the vanishing gradient problem when processing long sequences, making it difficult to learn effective long-range temporal dependencies, and their sequential computation mechanism leads to inefficiency. Leading researchers have then turned their attention to Transformers, using them to capture long-range dependencies. However, the quadratic computational complexity of their self-attention mechanism makes them difficult to handle high-resolution spatiotemporal gridded data, and they lack the built-in inductive biases for spatial locality and hierarchical structure found in CNNs.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] To address or at least alleviate one or more of the above problems, a method, system, and medium for predicting Lagrange trajectories based on U-Net and spatiotemporal Transformer are provided. By embedding prediction attention bridging processing at each layer of the U-Net network, the model can perform spatiotemporal information inference at multiple scales from coarse to fine.
[0007] To achieve the above objectives, according to the first aspect of this application, a Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer is provided, comprising: S1. Input data fusion and encoding: The input initial point coordinates and Eulerian field data are fused and encoded to obtain the fused feature tensor; S2, Layered Spatiotemporal Prediction: Based on the fused feature tensor, a U-Net architecture is used for hierarchical spatiotemporal prediction. The fused feature tensor is processed by an encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The predictive attention bridge is configured as follows: The system receives an input feature map from an encoder layer; converts the input feature map into a spatiotemporal token sequence and incorporates location information; models the spatiotemporal token sequence at least sequentially using a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, and the gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing, and outputs the processed feature map to the corresponding decoder layer; S3, Coordinate Regression: The high-dimensional feature tensor is decoded and regressed into a sequence of trajectory coordinates from a Lagrange perspective.
[0008] To achieve the above objectives, according to a second aspect of this application, a Lagrange trajectory prediction system based on U-Net and spatiotemporal Transformer is provided, the Lagrange trajectory prediction system comprising: Input data fusion and encoding module: The input initial point coordinates and Eulerian field data are fused and encoded to obtain the fused feature tensor; The core module of hierarchical spatiotemporal prediction: hierarchical spatiotemporal prediction is performed based on the fused feature tensor using the U-Net architecture. The fused feature tensor is processed by the encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The hierarchical spatiotemporal prediction core module includes a prediction attention bridging module: receiving an input feature map from an encoder layer; converting the input feature map into a spatiotemporal token sequence and incorporating location information; modeling the spatiotemporal token sequence at least sequentially through a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, and the gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing, and outputs the processed feature map to the corresponding decoder layer; Coordinate Regression Module: Decodes the high-dimensional feature tensor and regresses it into a trajectory coordinate sequence from a Lagrange perspective.
[0009] To achieve the above objectives, according to a third aspect of this application, a computer-readable storage medium is provided storing a computer program, which, when executed by a processor, is used to implement the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer as described above.
[0010] By adopting the above technical solution, this application has the following beneficial effects compared with the prior art: The core advantage of this application lies in its superior long-term predictive capability: by embedding a "prediction attention bridging module" at each layer of the U-Net network, the model can perform spatiotemporal information inference at multiple scales, from coarse to fine. Through a prediction-then-reconstruction step, it overcomes the error accumulation problem in traditional models. This application innovatively achieves efficient collaboration between local features and global dependencies, resolving the inherent contradictions of existing architectures: through the U-Net backbone, it inherits the advantage of Convolutional Neural Networks (CNNs) in efficiently extracting local spatial features when processing gridded data; simultaneously, by introducing a decoupled gated attention Transformer in the bridging module, it gains powerful global long-distance spatiotemporal dependency modeling capabilities. It cleverly solves the core contradictions in the background technology: avoiding the insufficient global receptive field of pure CNNs and the quadratic computational complexity and lack of spatial inductive bias of pure Transformers. Therefore, this application achieves high accuracy while maintaining controllable computational costs, can process higher-resolution environmental field data, and possesses greater potential for practical applications.
[0011] This application effectively enhances the model's generalization ability and environmental adaptability by reformulating the trajectory prediction problem as an end-to-end mapping from the Eulerian environmental field to Lagrange coordinates, independent of the floating object's own historical trajectory data. This design enables it to effectively address the challenge of sparse or missing trajectory observation data in the real world. Furthermore, joint training and independent testing on multiple seasonal datasets with different dynamic characteristics can improve environmental adaptability and model generalization ability.
[0012] In this application, the modular design is clear, exhibiting good scalability and interpretability: the system architecture consists of three clearly defined modules: input fusion, hierarchical prediction, and coordinate regression. This clear structure facilitates implementation and deployment. Future upgrades or replacements of the internal structures of these modules can be made easily without altering the overall framework, demonstrating excellent technical scalability. Furthermore, analyzing the attention weights of the "prediction attention bridging module" at different levels provides insights into how the model performs spatiotemporal reasoning at multiple scales, enhancing the model's interpretability.
[0013] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. Attached Figure Description
[0014] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application. The illustrative embodiments and descriptions of the application are used to explain the application, but do not constitute an undue limitation of the application. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0015] In the attached diagram: Figure 1 This is a flowchart illustrating the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 2 This is a schematic diagram of the model architecture of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 3 This is a schematic diagram of the input data fusion and encoding module of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 4 This is a schematic diagram of the architecture of the hierarchical spatiotemporal prediction core module of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 5This is a schematic diagram of the architecture of the prediction attention bridging module of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 6 This is a schematic diagram of the gating attention architecture of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 7 This is a schematic diagram of the architecture of the pre-trained decoder for the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 8 This is a schematic diagram of the coordinate regression module of the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer in this specific embodiment; Figure 9 This is a schematic diagram of the architecture of the Lagrange trajectory prediction system based on U-Net and spatiotemporal Transformer in this specific embodiment. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0017] Please see Figure 1 and Figure 2 This application provides a Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer, including: S1. Input data fusion and encoding: The input initial point coordinates and Eulerian field data are fused and encoded to obtain the fused feature tensor; S2, Layered Spatiotemporal Prediction: Based on the fused feature tensor, a U-Net architecture is used for hierarchical spatiotemporal prediction. The fused feature tensor is processed by an encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The predictive attention bridge is configured as follows: The system receives an input feature map from an encoder layer; converts the input feature map into a spatiotemporal token sequence and incorporates location information; models the spatiotemporal token sequence at least sequentially using a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, and the gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing, and outputs the processed feature map to the corresponding decoder layer; S3, Coordinate Regression: The high-dimensional feature tensor is decoded and regressed into a sequence of trajectory coordinates from a Lagrange perspective.
[0018] It should be noted that the execution entity of the Lagrange trajectory prediction method based on U-Net and Spatiotemporal Transformer in this embodiment is a Lagrange trajectory prediction system based on U-Net and Spatiotemporal Transformer. This system can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, etc., and non-mobile electronic devices can be servers and personal computers, etc., which are not specifically limited in this application. The following description uses a server as the execution entity to illustrate the Lagrange trajectory prediction method based on U-Net and Spatiotemporal Transformer in this embodiment.
[0019] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0020] In a preferred embodiment, the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer includes fusing and encoding the input initial point coordinates and Eulerian field data to obtain a fused feature tensor.
[0021] Specifically, input data fusion and encoding: S1 is executed by the input data fusion and encoding module, which aims to combine heterogeneous input data, namely the coordinates of discrete initial points and the data of continuous environmental fields, into a unified, high-dimensional feature tensor for subsequent modules to process.
[0022] In a preferred embodiment, the step of fusing and encoding the input initial point coordinates and Eulerian field data to obtain a fused feature tensor includes: linearly normalizing the input initial point coordinates to obtain normalized coordinates; generating a two-dimensional Gaussian graph by passing the normalized coordinates through a Gaussian encoder; copying the two-dimensional Gaussian graph along the time dimension of the Eulerian field to obtain an embedded coordinate tensor with the same spatiotemporal dimension as the Eulerian field; and concatenating the embedded coordinate tensor with the Eulerian field data to obtain a multimodal spatiotemporal tensor, which is the fused feature tensor.
[0023] For example, S1 consists of roughly three steps: 1.1 Initial point coordinate data processing: The input data fusion and encoding module receives one or a batch of initial coordinate points for the trajectory, with the data format (B, 2), where B is the batch size and 2 is the channel dimension (i.e., longitude and latitude). To ensure the consistency of the coordinate system with the environmental field space, the initial coordinates of the buoy are first determined using equations (2-1a), (2-1b), and (2-2). After performing linear normalization, the normalized coordinates are obtained: (2-1a); (2-1b); (2-2); in, Represents the normalized coordinate points, These represent the normalized longitude and latitude values, respectively. , , These represent the original longitude value of the current input, the minimum longitude value, and the maximum longitude value in the training dataset, respectively. , , These represent the original dimension value of the current input, the minimum value of all dimension values in the training dataset, and the maximum value, respectively. Normalized coordinates It will be transformed into a two-dimensional Gaussian probability map through a Gaussian encoder. Its value decreases smoothly outwards from the buoy's position, as shown in the following formula: (2-3); in, This represents a two-dimensional Gaussian probability plot. Let x and y represent the indices of the two-dimensional Gaussian probability plot, respectively. The scale parameter controls the diffusion range of the Gaussian distribution. Finally, to maximize the effectiveness of the initial position information across all prediction steps, the generated 2D Gaussian map is copied along the time dimension T, resulting in an embedding coordinate tensor with the same spatiotemporal dimension as the environment field, and its size is... .
[0024] 1.2 Environmental Field Data Input: Receives multi-channel Eulerian field data over T time steps, the form of which is... Where c represents the number of channels, including the latitude and longitude of the environmental field and the values of various attributes, and H×W represents the spatial resolution. In this example, the spatial resolution is 24×24, and the time step T is 24.
[0025] 1.3 Feature Fusion: By concatenating the embedded coordinate tensor with the Eulerian field data on the channel, a unified and information-rich multimodal spatiotemporal tensor is obtained. Its size is This tensor contains both the object's initial position information and time-varying environmental background information, serving as input for the next stage.
[0026] In a preferred embodiment, the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer includes hierarchical spatiotemporal prediction using the U-Net architecture based on the fused feature tensor, wherein the fused feature tensor is processed by an encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective.
[0027] Specifically, S2 is executed by the hierarchical spatiotemporal prediction core module, which adopts the U-Net architecture to perform multi-scale feature extraction and spatiotemporal evolution prediction on the fused feature tensor. The hierarchical spatiotemporal prediction core module is roughly divided into three sub-modules: encoder, predictive attention bridging, and decoder.
[0028] 2.1 Encoder: like Figure 2 As shown on the left, the data processed by the input data fusion and encoding module flows through multiple cascaded 3D convolutional modules for downsampling, decomposing the input spatiotemporal tensor into a series of multi-scale feature representations. In each transformation layer, the tensor undergoes a combination of "channel transformation followed by strided convolution" to achieve deep feature extraction and spatiotemporal dimension downsampling. Compared to traditional convolutional pooling, the strided convolution method used in this approach can simultaneously complete feature extraction and dimensionality compression in a single operation, better preserving spatial features. Deeper 3D convolutional modules may contain 3D batch normalization layers (BN3D) to stabilize the training process and accelerate model convergence.
[0029] For example, the input to the encoder path is a fused feature tensor of shape (B, C=c+1, T, H, W) from the input data fusion and encoding module. After passing through an initial 3D convolutional module, its size becomes (B, 64, T, H, W), and the stepwise downsampling begins. The aforementioned tensor of (B, 64, T, H, W) first passes through a SiLU activation function, and then through a second 3D convolutional layer, which uses strided convolution to achieve downsampling. After the second 3D convolutional layer, the spatiotemporal resolution of the feature map is halved, and a feature map of shape (B, 64, T / 2, H / 2, W / 2) is output. This feature map is then sent to the next level encoder module and the predictive attention bridging module at the same level.
[0030] The feature map with shape (B, 64, T / 2, H / 2, W / 2) is fed to the next level encoder module. This feature map, with dimensions (B, 64, T / 2, H / 2, W / 2), is then passed through a 3D convolutional layer, increasing its channel count from 64 to 128, resulting in a feature map with shape (B, 128, T / 2, H / 2, W / 2). This resulting feature map is then downsampled by another 3D convolutional layer, outputting a feature map with shape (B, 128, T / 4, H / 4, W / 4). This feature map with shape (B, 128, T / 4, H / 4, W / 4) is also fed to the next level encoder module and the predictive attention bridging module at the same level.
[0031] Finally, the feature map with shape (B, 128, T / 4, H / 4, W / 4) is passed through a composite module containing a 3D batch normalization layer and 3D convolutions to increase the number of channels of the feature map to 256, outputting a feature map with shape (B, 256, T / 4, H / 4, W / 4). This feature map is then downsampled by another composite module containing 3D convolutions and 3D batch normalizations, outputting a feature map with shape (B, 256, T / 8, H / 8, W / 8). This feature map with shape (B, 256, T / 8, H / 8, W / 8) is the output of the deepest layer of the encoder path, and it is sent to the predictive attention bridging module of its corresponding layer. The addition of 3D batch normalization makes the model more stable and more effective in learning the complex mapping relationship from marine environmental data to buoy trajectories.
[0032] After L-layer downsampling, the encoder outputs a series of spatiotemporal feature maps representing different physical scales. Shallow feature maps (such as the feature map with shape (B, 64, T / 2, H / 2, W / 2) in the image) have high resolution and capture small-scale local details such as eddies and fronts; deep feature maps (such as the feature map with shape (B, 256, T / 8, H / 8, W / 8) in the image) have a wide perceptual field and capture macroscopic trends such as large-scale circulation and weather systems. The convolutional coding results generated after each downsampling step... In other words, the output feature maps of this layer will be temporarily stored. The output feature maps of this layer will be used as input to the prediction attention bridging module of this layer for subsequent cross-layer spatiotemporal inference, and finally fused with the corresponding layer of the decoder path.
[0033] In a preferred embodiment, the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer includes receiving an input feature map from an encoder layer; converting the input feature map into a spatiotemporal token sequence and incorporating location information; modeling the spatiotemporal token sequence at least sequentially using a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, wherein the gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to nonlinearly modulate the self-attention-processed spatiotemporal token sequence, and outputs the processed feature map to the corresponding decoder layer.
[0034] 2.2 Predictive Attention Bridging Module: To overcome the inherent limitations of traditional models in handling long-term dependencies, such as limited receptive field and vanishing gradients, this embodiment embeds a prediction attention bridging module in the skip connection portion of the U-Net structure. At each layer of the U-Net, the feature map output from the encoder path is not directly passed to the decoder, but is first fed into this prediction attention bridging module. The detailed structure of the prediction attention bridging module is as follows... Figure 5 As shown.
[0035] The predictive attention bridging module receives the encoder's output at each scale and models its deep spatiotemporal relationships, extrapolating the predicted future information.
[0036] For example, such as Figure 2 and Figure 5 As shown, the input to the prediction attention bridging module is a feature map from layer i of the encoder, and its data shape is... Where B is the batch size. is the number of feature channels, and i represents the layer in the U-Net encoder path, with a value of 1 to 3. T, H, and W are the time, height, and width of the input prediction attention bridging module at layer i, respectively. Since the Transformer architecture natively processes sequence data, not grid-like data, the prediction attention bridging module first performs a rearrange operation to transform the input grid-like feature map data into a grid with the shape of... This is transformed into a token sequence. Spatial dimensions H and W are flattened into a single dimension N = H * W, and the order of the dimensions is adjusted to obtain a shape as follows: The output is as follows: At this point, the data consists of B batches, each batch containing T time steps, and each time step has N tokens at spatial locations.
[0037] To enable the temporal and spatial Transformers in the predictive attention bridging module to capture the absolute or relative position of each token in the original spatiotemporal grid, this embodiment uses position embedding (shape: ...). The tokens are added element by element to the token sequence, which not only preserves the original spatial details but also injects the spatiotemporal position information of each token, thus enhancing the model's spatiotemporal perception capability.
[0038] After a rearrangement operation, the shape of the feature tensor with positional encoding changes from... Transform into Each of the B*N spatial points is considered an independent sample, and each sample has a time series of length T.
[0039] In a preferred embodiment, both the temporal Transformer and the spatial Transformer integrate at least one gated attention. The gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing, and outputs the processed feature map to the corresponding decoder layer.
[0040] Specifically, to efficiently capture global dependencies, this embodiment employs a decoupled spatiotemporal self-attention mechanism, processing them sequentially through a Temporal Transformer, a rearrangement operation, and a Space Transformer. The Temporal and Space Transformers used in this embodiment are not standard Transformers, but rather are stacked from one or more gated attention units proposed in this embodiment. The structure of the gated attention unit is as follows... Figure 6 As shown.
[0041] In a preferred embodiment, the gated attention mechanism uses a gated signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing. This includes: input features sequentially passing through layer normalization, multi-head self-attention, and layer normalization to obtain intermediate features; the intermediate features being processed in parallel through a linear path and a gated path: in the linear path, the intermediate features are processed through a linear layer to obtain linearly transformed features; in the gated path, the intermediate features are processed through a linear layer and a SiLU activation function to generate a dynamic gated signal; the features output from the linear path are multiplied element-wise with the gated signal output from the gated path to achieve nonlinear modulation; the modulated features are sequentially processed through regularization, layer normalization, a linear layer, and regularization to output a feature map processed by the gated attention mechanism.
[0042] For example, in the gated attention unit, the input features first pass through a layer normalization layer, an attention module (containing a standard multi-head self-attention mechanism), and a second normalization layer to capture contextual dependencies within the sequence. The second normalization layer outputs an intermediate feature representation X. Then, the intermediate feature representation X is simultaneously fed into two parallel paths for processing: in the linear path, X passes through a linear layer; in the gated path, X passes through a gated signal generation unit (GVIN). Figure 6The SiLinGate part of the model (the gated signal generation unit) consists of an independent linear layer and a SiLU activation function. The intermediate feature representation X is first processed by the linear layer, then non-linearly activated by the SiLU function, ultimately generating a dynamic, data-dependent gated signal. Subsequently, the representation from the linear path is multiplied element-wise with the gated signal from the gated path. The gating mechanism provided by the gated signal generation unit allows the model to dynamically and non-linearly amplify or suppress specific feature channels in the data stream based on contextual information learned from the standard multi-head self-attention module. This allows the model to adaptively enhance key information and filter out irrelevant noise. The output data from the linear layer is multiplied element-wise with the gated signal from the gated path. The resulting final feature is then regularly regularized through a Dropout regularization layer, a normalization layer, a linear layer, and a final Dropout regularization layer, ultimately forming the output of the gated attention unit.
[0043] After processing by the temporal and spatial Transformers, the final feature sequence is rearranged, layer normalized, and then transformed by a multilayer linear perceptron head (MLP Head) to form a shape of (B*T*N, d i The 2D tensor projection is (B*T*N,C), which is then reshaped and permuteed. After reshaping, the flattened B*T*N dimensions are restored according to the known dimensions B, T, H, W (where N = H*W), restoring the 2D tensor (B*T*N, C) to a 5D tensor with the shape (B, T, H, W, C). A permute operation is then performed to output a tensor with the shape... The enhanced feature map, which is the output map, will be passed as input to the corresponding layer of the decoder path.
[0044] In a preferred embodiment, the hierarchical spatiotemporal prediction based on the fused feature tensor using a U-Net architecture is performed, wherein the fused feature tensor is processed by an encoder and passed to the decoder through a prediction attention bridge to output a high-dimensional feature tensor from an Euler perspective. This includes: inputting the fused feature tensor into an encoder path using a U-Net architecture, performing downsampling and multi-scale feature extraction through the encoder path to obtain an enhanced feature map of at least one level in the encoder path;
[0045] The enhanced feature map output from the deepest level of the encoder path, after being bridged by the corresponding level's predictive attention, is used as the initial input and fed into the decoder path using the U-Net architecture. The decoder path upsamples the input feature map through one or more upsampling operations and outputs a high-dimensional feature tensor processed by the highest level of the decoder path.
[0046] 2.3 Decoder Module: like Figure 2 As shown on the right, the hierarchical spatiotemporal prediction core of this embodiment reconstructs high-resolution prediction results step by step and layer by layer from the highly abstract spatiotemporal features captured by the encoder at the deepest layer, with dimensions (B, 256, T / 8, H / 8, W / 8), through the decoder path. The core idea of the decoder path is to effectively fuse deep macroscopic prediction information with shallow, predictively enhanced detailed information.
[0047] The decoder path proceeds from bottom to top. The initial input to the decoder path is a feature map from the deepest layer of the encoder path, processed by the prediction attention bridging module of the corresponding decoder layer. This feature map first passes through an upsampling module. The deepest upsampling module contains a transposed convolution (TransConv) layer and a SiLU activation function, transforming the tensor of size (B, 256, T / 8, H / 8, W / 8) from the deepest layer output into a tensor of size (B, 256, T / 4, H / 4, W / 4). The resulting (B, 256, T / 4, H / 4, W / 4) feature map is then concatenated along the channel dimension with a feature map (shaped (B, 128, T / 4, H / 4, W / 4)) processed by the prediction attention bridging module from the second layer of the encoder. The concatenated fused feature map becomes (B, 384, T / 4, H / 4, W / 4) in shape. It then passes through a second upsampling module, which includes transposed convolutions, SiLU, and a regularization layer. The feature map's shape is then transformed from (B, 384, T / 4, H / 4, W / 4) to (B, 128, T / 4, H / 4, W / 4). The feature map continues through a third upsampling module consisting of transposed convolutions and SiLU, further improving its spatiotemporal resolution. Its shape becomes (B, 128, T / 2, H / 2, W / 2). This is then concatenated with the feature map processed by the first-level prediction attention bridging module (shape (B, 64, T / 2, H / 2, W / 2)), resulting in a size of (B, 192, T / 2, H / 2, W / 2). Finally, it is processed by a final module containing transposed convolutions and SiLU for final feature extraction and channel adjustment. The final output is a high-dimensional feature tensor (feature map) with shape (B, 64, T / 2, H / 2, W / 2). The transposed convolution is the opposite of convolution; it gradually increases the resolution of the feature map and reduces the number of channels by padding the elements with zeros and then performing convolution.
[0048] In a preferred embodiment, the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer includes: performing feature fusion operations at one or more intermediate layers in the decoder path.
[0049] The upsampled feature map from the next level is concatenated with the enhanced feature map from the same level of the encoder path, which has been bridged by prediction attention, to form a fused feature map. The fused feature map is then input to the upsampled feature map of the current level for processing and then passed to the next level of the decoder path.
[0050] Specifically, for each intermediate layer in the decoder path, a key feature fusion operation is performed once, i.e. Figure 4 The "stitching" circular node on the right, representing an upsampled feature map from the next deeper layer (L+1), is stitched together with an enhanced feature map from the same layer (L) of the encoder path that has already been processed by the prediction attention bridging module. The stitched fused feature map is then processed by the upsampling module of the current layer to further restore the spatiotemporal resolution of the fused feature map and adjust the number of channels before being passed to the next shallower layer. For example... Figure 2 As shown, since performing regularization operations in the deepest layer may lose useful information, while performing regularization operations in the shallow layer may easily interfere with the final output, the intermediate layer upsampling module includes regularization operations, namely the Dropout regularization layer, which is used to randomly drop a portion of neurons during training to prevent the model from overfitting and enhance its generalization ability.
[0051] The process of progressive fusion and upsampling is repeated until the highest layer of the decoder path. At the highest layer, after the final fusion and upsampling operations, a high-dimensional feature tensor is output, which integrates spatial details at all scales and prediction information derived from global spatiotemporal inference. Its size is... The output tensor, i.e., the high-dimensional feature tensor, will then be fed into the final coordinate regression module for final coordinate decoding.
[0052] In a preferred embodiment, decoding and regressing the high-dimensional feature tensor into a trajectory coordinate sequence from a Lagrange perspective includes: upsampling the high-dimensional feature tensor from the top-level output of the decoder path through deconvolution to obtain a full-resolution feature tensor with the same spatiotemporal resolution as the original input Eulerian field data; performing feature integration and channel dimensionality reduction on the full-resolution feature tensor through 3D convolution to output a single-channel spatiotemporal feature tensor, wherein the spatiotemporal feature tensor represents the probability distribution of the object appearing at each grid position within multiple future time steps from an Eulerian perspective; and inputting the spatiotemporal feature tensor into a pre-trained decoder to map and output the trajectory coordinate sequence of the target within the multiple future time steps.
[0053] Specifically, the coordinate regression process is as follows: Coordinate regression module, such as Figure 2 Module c) and Figure 8 As shown, the core task is to decode and regress the high-dimensional feature tensor of the hierarchical spatiotemporal prediction core output, which is in the Euler perspective, into a specific trajectory coordinate sequence in the Lagrange perspective.
[0054] The high-dimensional feature tensor from the top layer of the hierarchical spatiotemporal prediction core decoder path undergoes an upsampling operation in a deconvolutional layer (DeConv) to restore the spatiotemporal resolution of the feature map, i.e., the high-dimensional feature tensor, from (T / 2, H / 2, W / 2) to (T, H, W), resulting in a tensor with shape […]. The full-resolution feature tensor with the exact same spatiotemporal resolution as the original input data.
[0055] The generated full-resolution feature tensor is then subjected to 3D convolution through a 3D convolutional layer. The main function of the 3D convolutional layer is to perform final integration and dimensionality reduction of the multi-channel features, compressing the full-resolution feature tensor from 64 channels into a single-channel spatiotemporal tensor, with an output shape of... The spatiotemporal tensor of this single channel can be understood as the probability distribution of an object appearing at each grid position (H, W) within the next T time steps from Euler's perspective.
[0056] The probability distribution map is not the final trajectory sequence form desired in this embodiment. This probability map must be processed by a pre-trained decoder to obtain a sequence with the following shape: The target trajectory tensor of size is the output trajectory, representing the trajectory points of B buoys at T time steps.
[0057] In a preferred embodiment, the step of inputting the spatiotemporal feature tensor into a pre-trained decoder to map and output the trajectory coordinate sequence of the target within the future multiple time steps includes: transforming the spatiotemporal feature tensor through a reshaping operation; inputting the reshaped spatiotemporal feature tensor into a series of cascaded two-dimensional convolutional layers to extract spatial features step by step and finally outputting a one-dimensional feature vector; mapping the one-dimensional feature vector into two-dimensional coordinate values through one-dimensional convolution, and normalizing the two-dimensional coordinate values to the [0, 1] interval using the Sigmoid activation function to output a normalized coordinate sequence; and reshaping the normalized coordinate sequence to obtain the final trajectory coordinate sequence.
[0058] For example, 3.2 Pre-trained decoder structure: The internal structure of the pre-trained decoder is as follows Figure 7 As shown, the pre-trained decoder receives a sequence of probability heatmaps for T frames and regresses a two-dimensional coordinate point for each frame. To facilitate independent spatial feature extraction of features at each time step by subsequent two-dimensional convolutional layers, After the probability distribution map is fed into the pre-trained decoder, it first undergoes a reshape operation to transform the dimensions of the input probability distribution map. This involves stacking the heatmaps of each of the B samples at T time steps to form a batch containing B*T independent 2D images. The reshaped tensor is then processed through five 2D convolutional blocks and a ReLU activation function to extract features step by step. The tensor size input to the first 2D convolutional layer is [size missing]. ,Right now The tensor is processed through a first 2D convolutional layer (with parameters changing the number of channels C from 1 to 16, the kernel size K to 3, the stride S to 2, and the padding P to 0), and then through a ReLU activation function, resulting in a tensor size of […]. Then, through a second 2D convolutional layer (with parameters changing from 16 channels to 32, kernel size K of 3, stride S of 2, and padding P of 0) and an activation function, the tensor size becomes... After passing through a third 2D convolutional layer (with parameters changing from 32 to 64 channels, kernel size K of 3, stride S of 1, and padding P of 0), the tensor, after passing through this convolutional layer and the ReLU activation function, has a size that becomes... Then, the tensor undergoes a fourth 2D convolution (the number of channels changes from 64 to 64, the kernel size is K = 2, the stride is S = 1, and the padding is P = 0), followed by a ReLU activation function, after which the tensor size becomes... After passing through a final 2D convolutional layer (the number of channels changes from 64 to 64, the kernel size K is 1, the stride S is 1, and the padding P is [value missing]), the tensor size remains unchanged after this 2D convolutional module. The tensor output after five 2D convolutional layers is further flattened by a flattening layer, resulting in 24B feature vectors of length 256, with a data shape of (24B, 256). The flattened tensor is then transformed into a 3D tensor through an unsqueeze operation, resulting in a shape of (24B, 256, 1). A 1D convolutional layer (with the parameters C changing from 256 to 2 channels and K being a kernel size of 1) maps the 256-dimensional abstract features to a 2D coordinate space, yielding (24B, 2, 1). The regressed 2D coordinate values are normalized to the [0, 1] interval using a sigmoid activation function, and then subjected to a squeeze, resulting in a normalized coordinate sequence of shape (24B, 2), where 24 represents T. This final coordinate sequence only requires a simple reshape operation to restore its batch B and time T dimensions, yielding the final output tensor with a shape of (B, T, 2).
[0059] The purpose of pre-training a pre-trained decoder is to enable it to learn the ability to accurately regress the peak values from a two-dimensional heatmap. That is, given a heatmap with a known center point generated by a Gaussian encoder, the pre-trained decoder is required to accurately regress the center coordinates of the heatmap.
[0060] During pre-training, this embodiment randomly samples a large number of coordinate points in a normalized coordinate field as training samples. Then, a Gaussian encoder is used to encode each coordinate point into a two-dimensional Gaussian distribution heatmap. These heatmaps serve as the training input for the pre-trained decoder, while the corresponding original coordinate points serve as training labels. Pre-training employs a composite loss function, which is a weighted sum of the following two parts: Reconstruction loss: The coordinates output by the decoder, after Gaussian encoding, can reconstruct the distribution of the input field, ensuring that the decoder has the ability to correctly recover the coordinates from the distribution. This involves calculating the loss between the distributions. Coordinate constraint loss: Calculate the Euclidean distance between the decoded coordinates and the true coordinates, applying constraints at the output level of the final target. The sum of these two distances is used as the final loss function. (2-4); in, Indicates the losses incurred during reconstruction. This represents the coordinate constraint loss.
[0061] 4. Model Training: This embodiment proposes a model for executing the methods described above. The model needs to learn and optimize its internal parameters through a specific training process. The model training process includes at least three steps: constructing training data, designing a loss function, and iteratively optimizing the model.
[0062] 4.1 Construction of training data: In order for the model to learn the mapping relationship from the ocean environmental field to the Lagrange trajectory, it is necessary to construct a training dataset suitable for the model in this embodiment.
[0063] It is necessary to acquire marine environmental field data (including gridded data of environmental field latitude and longitude, ocean current u / v components, and wind speed u / v components with multiple channels) covering the target sea area and having a certain spatiotemporal resolution, as well as Lagrange trajectory data generated in this environmental field. The trajectory data can be generated by numerical integration on this environmental field (e.g., using the fourth-order Runge-Kutta method).
[0064] A sliding window method can be used to divide large-scale environmental field data into a series of local environmental field data with fixed spatial dimensions. Then, Lagrange trajectories are matched with these spatiotemporal slices to select trajectories whose motion trajectory is completely contained within a specific spatiotemporal slice throughout the entire time period. This forms a training sample pair (local environmental field sequence, internal trajectory sequence) for efficient training. Before training, all constructed training sample pairs should be divided into training, validation, and test sets according to a certain ratio.
[0065] In a preferred embodiment, the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer further includes: a step of using a loss function to guide model training, including: The loss function is a weighted sum of coordinate regression loss and probability distribution loss, expressed as: (2-5); Where λ is a hyperparameter used to balance the weights of the two loss components. This represents the coordinate regression loss. Represents probability distribution loss; Coordinate regression loss is primarily used to measure the deviation between the predicted trajectory and the true trajectory. It is calculated using a weighted Euclidean distance and is expressed as follows: (2-6); in, This represents the trajectory points predicted by the model simulation. These are the actual trajectory points; To represent the importance of data at different times, use The weight coefficient representing the t-th time step is expressed as: (2-7); Where t represents the time step, and T represents the total number of time steps in the predicted sequence. As a decay factor, it ensures that the early prediction points of the model have higher weights, while the later prediction points have relatively lower weights; through normalization, it ensures that the overall mean of the weight sequence is 1, thereby avoiding changing the numerical scale of the loss function. The probability distribution loss is calculated based on the predicted distribution of the model output before the final coordinate regression. Compared with the two-dimensional distribution map generated by the Gaussian encoder from the real trajectory points. The difference between them is measured using the KL divergence to determine the similarity between the two probability distributions, expressed as: (2-8); in, This represents the predicted distribution map. This represents a two-dimensional distribution map.
[0066] Based on the same inventive concept, please see Figure 9 This application also provides a Lagrange trajectory prediction system based on U-Net and spatiotemporal Transformer, the Lagrange trajectory prediction system comprising: Input data fusion and encoding module: The input initial point coordinates and Eulerian field data are fused and encoded to obtain the fused feature tensor; The core module of hierarchical spatiotemporal prediction: hierarchical spatiotemporal prediction is performed based on the fused feature tensor using the U-Net architecture. The fused feature tensor is processed by the encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The hierarchical spatiotemporal prediction core module includes a prediction attention bridging module: receiving an input feature map from an encoder layer; converting the input feature map into a spatiotemporal token sequence and incorporating location information; modeling the spatiotemporal token sequence at least sequentially through a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, and the gated attention uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the spatiotemporal token sequence after self-attention processing, and outputs the processed feature map to the corresponding decoder layer; Coordinate Regression Module: Decodes the high-dimensional feature tensor and regresses it into a trajectory coordinate sequence from a Lagrange perspective.
[0067] Based on the same inventive concept, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer as described above.
[0068] The program product of this application for implementing the above method may employ a portable compact disk read-only memory and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, system, or device.
[0069] It should be noted that a computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, system, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0070] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-mentioned technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. The implementation schemes in the above embodiments can also be further combined or replaced. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of this application.
Claims
1. A Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer, characterized in that, include: S1. Input data fusion and encoding: The initial point coordinates of floating objects in a marine environment and Eulerian field data are fused and encoded to obtain a fused feature tensor; S2, Layered Spatiotemporal Prediction: Based on the fused feature tensor, a U-Net architecture is used for hierarchical spatiotemporal prediction. The fused feature tensor is processed by an encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The predictive attention bridge is configured as follows: The process involves receiving an input feature map from an encoder layer; converting the input feature map into a spatiotemporal token sequence and incorporating location information; modeling the spatiotemporal token sequence at least sequentially using a temporal Transformer and a spatial Transformer; wherein both the temporal Transformer and the spatial Transformer integrate at least one gated attention mechanism, and the gated attention mechanism uses a gate signal dynamically generated from the spatiotemporal token sequence to perform nonlinear modulation on the self-attention processed spatiotemporal token sequence, including: The input features are sequentially processed through layer normalization, multi-head self-attention, and layer normalization to obtain intermediate features. These intermediate features are then processed in parallel through a linear path and a gated path: in the linear path, the intermediate features are processed through a linear layer to obtain linearly transformed features; in the gated path, the intermediate features are processed through a linear layer and a SiLU activation function to generate a dynamic gated signal; the features output from the linear path are multiplied element-wise with the gated signal output from the gated path to achieve nonlinear modulation; the modulated features are sequentially processed through regularization, layer normalization, a linear layer, and regularization to output a feature map processed by the gated attention mechanism. The processed feature map is output to the corresponding decoder layer; S3, Coordinate Regression: The high-dimensional feature tensor is decoded and regressed into a sequence of trajectory coordinates from a Lagrange perspective.
2. The method according to claim 1, characterized in that, The input initial point coordinates and Eulerian field data are fused and encoded to obtain a fused feature tensor, including: The initial point coordinates of the input are linearly normalized to obtain normalized coordinates. The normalized coordinates are then passed through a Gaussian encoder to generate a two-dimensional Gaussian graph. The two-dimensional Gaussian graph is copied along the time dimension of an Eulerian field to obtain an embedded coordinate tensor with the same spatiotemporal dimension as the Eulerian field. The embedded coordinate tensor is then concatenated with the Eulerian field data to obtain a multimodal spatiotemporal tensor, which is the fused feature tensor.
3. The method according to claim 1, characterized in that, The hierarchical spatiotemporal prediction based on the fused feature tensor using a U-Net architecture, wherein the fused feature tensor is processed by an encoder, and through prediction attention bridging, is passed to the decoder to output a high-dimensional feature tensor from an Euler perspective, including: The fused feature tensor is input into the encoder path using the U-Net architecture. Downsampling and multi-scale feature extraction are performed through the encoder path to obtain an enhanced feature map of at least one level in the encoder path. The enhanced feature map output from the deepest level of the encoder path, after being bridged by the corresponding level's predictive attention, is used as the initial input and fed into the decoder path using the U-Net architecture. The decoder path upsamples the input feature map through one or more upsampling operations and outputs a high-dimensional feature tensor processed by the highest level of the decoder path.
4. The method according to claim 3, characterized in that, The hierarchical spatiotemporal prediction based on the fused feature tensor using a U-Net architecture, wherein the fused feature tensor is processed by an encoder, and through prediction attention bridging, is passed to the decoder to output a high-dimensional feature tensor from an Euler perspective, further comprising: One or more decoder layers in the decoder path concatenate the upsampled feature map from the next layer with the enhanced feature map from the same layer of the encoder path that has been bridged by prediction attention to form a fused feature map; the fused feature map is then input to the upsampling of the current layer for processing, and then passed to the layer above the decoder path.
5. The method according to claim 4, characterized in that, Decoding and regressing the high-dimensional feature tensor into a sequence of trajectory coordinates from a Lagrange perspective includes: The high-dimensional feature tensor from the top-level output of the decoder path is upsampled by deconvolution to obtain a full-resolution feature tensor with the same spatiotemporal resolution as the original input Eulerian field data. The full-resolution feature tensor is integrated and channel-wise reduced through 3D convolution to output a single-channel spatiotemporal feature tensor. The spatiotemporal feature tensor represents the probability distribution of the object appearing at each grid position in multiple future time steps from the Euler perspective. The spatiotemporal feature tensor is input into a pre-trained decoder, which maps and outputs the trajectory coordinate sequence of the target within the future multiple time steps.
6. The method according to claim 5, characterized in that, The step of inputting the spatiotemporal feature tensor into a pre-trained decoder, mapping and outputting the trajectory coordinate sequence of the target within the future multiple time steps includes: The spatiotemporal feature tensor is transformed through a reshaping operation. The reshaped spatiotemporal feature tensor is then input into a series of cascaded two-dimensional convolutional layers to extract spatial features step by step, and finally outputs a one-dimensional feature vector. The one-dimensional feature vector is mapped to two-dimensional coordinate values through one-dimensional convolution, and the two-dimensional coordinate values are normalized using the Sigmoid activation function to output a normalized coordinate sequence; the normalized coordinate sequence is then reshaped to obtain the final trajectory coordinate sequence.
7. The method according to claim 1, characterized in that, The method further includes the step of training the model using a loss function, wherein the model is used to perform the method of claim 1, and the step of using a loss function to guide model training includes: The loss function is a weighted sum of coordinate regression loss and probability distribution loss, expressed as: (2-5); Where λ is a hyperparameter used to balance the weights of the two loss components. This represents the coordinate regression loss. Represents probability distribution loss; The coordinate regression loss is expressed as: (2-6); in, This represents the trajectory points predicted by the model simulation. These are the actual trajectory points; The weight coefficient representing the t-th time step is expressed as: (2-7); Where t represents the time step, and T represents the total number of time steps in the predicted sequence. It is the attenuation factor; The probability distribution loss is expressed as: (2-8); in, This represents the predicted distribution map. This represents a two-dimensional distribution plot, and KL represents the KL divergence.
8. A Lagrange trajectory prediction system based on U-Net and spatiotemporal Transformer, characterized in that, The Lagrange trajectory prediction system includes: Input data fusion and encoding module: The initial point coordinates of floating objects in the marine environment and Eulerian field data are fused and encoded to obtain the fused feature tensor; The core module of hierarchical spatiotemporal prediction: hierarchical spatiotemporal prediction is performed based on the fused feature tensor using the U-Net architecture. The fused feature tensor is processed by the encoder and passed to the decoder through prediction attention bridging to output a high-dimensional feature tensor from the Euler perspective. The hierarchical spatiotemporal prediction core module includes a prediction attention bridging module: receiving an input feature map from an encoder layer; converting the input feature map into a spatiotemporal token sequence and incorporating location information; modeling the spatiotemporal token sequence at least sequentially using a temporal Transformer and a spatial Transformer; wherein each of the temporal Transformer and the spatial Transformer integrates at least one gated attention, and the gated attention performs nonlinear modulation on the self-attention-processed spatiotemporal token sequence using a gate signal dynamically generated from the spatiotemporal token sequence, including: The input features are sequentially processed through layer normalization, multi-head self-attention, and layer normalization to obtain intermediate features. These intermediate features are then processed in parallel through a linear path and a gated path: in the linear path, the intermediate features are processed through a linear layer to obtain linearly transformed features; in the gated path, the intermediate features are processed through a linear layer and a SiLU activation function to generate a dynamic gated signal; the features output from the linear path are multiplied element-wise with the gated signal output from the gated path to achieve nonlinear modulation; the modulated features are sequentially processed through regularization, layer normalization, a linear layer, and regularization to output a feature map processed by the gated attention mechanism. The processed feature map is output to the corresponding decoder layer; Coordinate Regression Module: Decodes the high-dimensional feature tensor and regresses it into a trajectory coordinate sequence from a Lagrange perspective.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it is used to implement the Lagrange trajectory prediction method based on U-Net and spatiotemporal Transformer as described in any one of claims 1-7.