Ship movement behavior prediction method based on BiGRU-Seq2Seq network
The improved BiGRU-Seq2Seq network extracts the time characteristics and destination characteristics of the ship trajectory and combines multiple loss functions for model training, which solves the shortcomings of the existing technology in long-term ship trajectory prediction and achieves higher prediction accuracy and stability.
Patent Information
- Application Number
- CN202510269801.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The existing ship trajectory prediction methods perform poorly in long-term prediction and are difficult to meet the needs of maritime traffic estimation, international shipping services and port resource planning.
The ship movement behavior prediction method based on the improved BiGRU-Seq2Seq network is adopted. By extracting the time characteristics and destination characteristics of the trajectory segment, and combining multiple loss functions for model training, the trajectory prediction accuracy is optimized.
It improves the accuracy and stability of ship movement behavior prediction, enhances the accuracy of long-term trajectory prediction, and meets the actual needs of the maritime field.
Smart Images

Figure CN120220467A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of ship trajectory prediction, and particularly relates to a ship movement behavior prediction method based on a BiGRU-Seq2Seq network. Background Art
[0002] Maritime transportation is an important driving force for global economic growth and cooperation, and over 90% of world trade is completed by sea. However, maritime safety is affected by various factors such as adverse weather, sea conditions, traffic density, and human factors, among which human error causes a significant proportion of maritime accidents. Accurate ship trajectory prediction is crucial for tasks such as risk reduction, route planning, and energy conservation, and helps to early warn potential risks and reduce collision accidents.
[0003] With the expansion of global maritime trade, the pressure on traffic control centers to monitor and manage ship trajectories has increased. Traditional manual collision avoidance methods rely on the experience of duty officers and are prone to errors in complex environments. Vessel Traffic Services (VTS) with the support of technologies such as the Automatic Identification System (AIS) enables accurate monitoring and prediction of ship movements. AIS records the real-time data of ships, providing rich spatio-temporal data for ship trajectory prediction research.
[0004] Currently, ship trajectory prediction methods are mainly divided into model-driven, data-driven, and hybrid methods. Model-driven methods are based on physical models with limited accuracy and complexity, and are difficult to handle complex long-term trajectory predictions. Data-driven methods, especially deep learning-based methods, have advantages in dealing with trajectory non-linearity. For example, Recurrent Neural Networks (RNN) and their variants Long Short-Term Memory networks (LSTM) and Gated Recurrent Units (GRU), as well as Sequence-to-Sequence (Seq2Seq) architectures, etc. are widely used in trajectory prediction. However, current methods perform well in short-term prediction, but long-term prediction still faces challenges and is difficult to meet the needs of maritime traffic estimation, international shipping services, and port resource planning, etc. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem of long-term ship trajectory prediction, and propose a solution for ship movement behavior prediction based on an improved BiGRU-Seq2Seq network, using the possible destinations of trajectory segments to improve the estimation accuracy for long time spans.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A ship movement behavior prediction method based on a BiGRU-Seq2Seq network, comprising the following steps:
[0008] 1) Extract the trajectory data of each ship from the collected AIS data, first perform data preprocessing, and then perform whereabouts positioning and annotation;
[0009] 2) Construct a destination-guided trajectory prediction model DGTP based on the BiGRU-Seq2Seq network. The DGTP model includes a destination prediction module and a trajectory prediction module, each of which contains an encoder and a decoder;
[0010] 3) Input the trajectory sequence marked in step 1) into the destination prediction module. The encoder of this module extracts the time features of the trajectory sequence and predicts the destination based on the historical trajectory data through the decoder; Input the predicted destination into the trajectory prediction module. The encoder of this module generates the destination features and concatenates them with the historical trajectory features to generate a unified tensor. The decoder predicts the future trajectory based on the generated unified tensor;
[0011] 4) Design a trajectory alignment loss consisting of mean square error loss, trajectory length loss, direction loss and dynamic time warping loss. Use this trajectory alignment loss to optimize DGTP model training and improve trajectory prediction accuracy.
[0012] 5) Using the trained DGTP model, the future trajectory is predicted based on the input ship historical trajectory data, and the predicted complete trajectory sequence and destination are output.
[0013] Furthermore, the data preprocessing in step 1) includes cleaning the data by filtering out small ship data, as well as data points whose latitude, longitude, ground speed and ground heading values exceed a preset range, and removing incomplete trajectories with less than 20 trajectory points.
[0014] Furthermore, the data preprocessing in step 1) includes repairing the trajectory, and the method is as follows: for single ship trajectory data, if the time interval between adjacent trajectory points is less than 1 minute, the original data is retained; if the time interval is between 1 and 30 minutes, the missing points are filled in with 1-minute intervals using cubic spline interpolation, and the latitude, longitude, ground speed and ground heading data are supplemented.
[0015] Furthermore, the tracking positioning in step 1) includes trajectory merging, and the method is: based on the dynamic time warping algorithm DTW, the pairwise similarity between the trajectories is calculated. If the DTW distance between two trajectories is lower than a set threshold, they are merged, and the trajectory with the longest physical length is selected as the representative trajectory in the similar trajectory group.
[0016] Furthermore, the tracking positioning in step 1) includes trajectory clustering, and the method is: clustering the trajectories using the density-based spatial clustering algorithm DBSCAN, marking the starting point and the end point of the trajectory as the land destination, and clustering the non-starting and non-end points.
[0017] Further, the method for performing whereabouts annotation in step 1) is as follows: a heuristic method is used to perform semantic annotation on the clustering results. If the sailing speed of the cluster is slower than the preset threshold, the cluster is marked as a common stopping point.
[0018] Further, step 1) also includes rearranging the trajectory. The method is as follows: adjust the time interval of the trajectory points to 1 minute, select the trajectories with the number of trajectory points between 20 and 70 as the complete trajectory dataset. If the trajectory length exceeds 70, truncate it to the format including 10 historical data points and 10 to 60 future data points.
[0019] Further, step 1) also includes destination assignment for the trajectory segments. The method is as follows: calculate the included angle of the direction vectors between each point on the trajectory and the potential destinations. If the trajectory passes through a marked destination, assign this destination to the previous trajectory segment; if it does not pass through a marked destination, select the destination that is the closest to the trajectory point and the included angle is less than the preset threshold. If neither of the above two destinations exists, assign the onshore destination that is the closest to the trajectory endpoint.
[0020] Further, the mean squared error loss in step 4): is used to measure the Euclidean distance between the predicted trajectory and the true trajectory at each time step. The formula is:
[0021]
[0022] where, and are the latitude and longitude of the i-th predicted time step, lat i and lon i are the corresponding true values, and T is the total number of time steps considered;
[0023] The trajectory length loss: is used to make the model maintain the consistency in shape between the predicted trajectory and the actual trajectory. The formula is:
[0024]
[0025] where, and p i respectively represent the coordinates of the i-th predicted and true time steps, and ‖·‖2 is the L2 norm of the Euclidean distance between the two coordinates;
[0026] The direction loss: is used to measure the angular difference between two trajectories. The formula is:
[0027]
[0028] where, and θ iThey are the direction angles of the i-th predicted and true points relative to the starting point (lat0, lon0); atan2(,) is used to calculate the angle from the positive x-axis to the trajectory point in the two-dimensional plane;
[0029] Dynamic Time Warping loss: It is used to align the predicted trajectory and the true trajectory, and the formula is:
[0030]
[0031] where and p j are a pair of coordinate points in the predicted and true trajectories, and π represents the corresponding elements in the two trajectories and p j group.
[0032] Furthermore, in step 5), the DGTP model uses a sliding window method for cyclic prediction. The method is as follows: First, combine the first predicted future trajectory point with other points in the original historical trajectory data except the last point to form a new input sequence; then use the new input sequence to perform destination prediction and trajectory prediction again to obtain the next future trajectory point; continuously cycle according to the above steps, and each time add the newly predicted trajectory point to the input sequence, replacing the earliest historical trajectory point, and continue to perform prediction until a complete predicted trajectory sequence with the required time span is generated.
[0033] The beneficial effects achieved by the present invention are as follows:
[0034] 1. Through the improved BiGRU-Seq2Seq network model, the present invention combines a double-layer BiGRU encoder to extract the time features of the trajectory and combines the destination features for trajectory prediction. Compared with traditional methods, it improves the accuracy and stability of ship movement behavior prediction;
[0035] 2. The present invention uses AIS data for trajectory preprocessing, including trajectory sorting, abnormal data cleaning, interpolation and filling, etc., which improves the quality of the input data, reduces noise interference, and thus optimizes the prediction effect of the model;
[0036] 3. The present invention uses the DTW algorithm to calculate the trajectory similarity and combines the DBSCAN clustering algorithm to classify the trajectories and label the berthing points, which improves the interpretability of the trajectory data and reduces the influence of redundant data;
[0037] 4. The present invention uses a heuristic method to assign the destination of the trajectory points and combines the direction vector angle calculation method to optimize the destination matching, which improves the accuracy of destination prediction and makes the predicted trajectory more in line with the actual movement law of the ship;
[0038] 5. The present invention introduces a variety of loss functions, including mean square error loss, trajectory length loss, direction loss, DTW loss, etc., and proposes a trajectory alignment loss to optimize model training, improving the accuracy and generalization ability of trajectory prediction.
[0039] 6. The present invention uses a sliding window method to generate future trajectories, which can dynamically update prediction results, is applicable to different types of ships and navigation environments, and improves the applicability and computational efficiency of the prediction method. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the ship movement behavior prediction method in the embodiment.
[0041] Figure 2 is a diagram of the trajectory annotation result in the embodiment.
[0042] Figure 3 is a structural diagram of the destination-guided trajectory prediction model in the embodiment.
[0043] Figure 4 is an example diagram of the trajectory predicted by the ship movement behavior prediction method in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below through specific embodiments and in combination with relevant principles.
[0045] An embodiment of the present invention discloses a ship movement behavior prediction method based on a BiGRU-Seq2Seq network, and its process is as Figure 1 shown. The specific implementation steps are as follows:
[0046] Step 1: Data processing and movement annotation
[0047] (1) Data preprocessing:
[0048] Single-ship data extraction: From the collected AIS dataset, according to the unique MMSI (Maritime Mobile Service Identity) of each ship, the data belonging to the same ship are completely extracted. At the same time, the extracted trajectory points are strictly sorted according to the chronological order of the data records to ensure the continuity and logic of the data in the time dimension. For example, if a ship has an MMSI of "123456789", then all the trajectory points corresponding to this MMSI are arranged in chronological order to ensure the accuracy of subsequent analysis and processing.
[0049] Data cleaning: To avoid interference from noisy data in subsequent model learning, the following filtering measures are taken. First, data of small ships with a length less than 3m or a width less than 2m are excluded, because the trajectory characteristics of such small ships may be quite different from those of large merchant ships, etc., and may have a negative impact on the generalization ability of the overall model. Second, data points with LAT (latitude), LON (longitude), SOG (speed over ground), and COG (course over ground) values outside the reasonable range are filtered. For example, the latitude range is set to [-90°, 90°], the longitude range is [-180°, 180°], the reasonable range of SOG is [0, 50] knots (set according to the general ship navigation speed range), and the COG range is [0, 360°]. Data points outside this range are regarded as abnormal data points and are excluded. In addition, trajectories with less than 20 trajectory points are also filtered out because their data volume is too small to reflect the complete motion characteristics of the ship.
[0050] Trajectory repair: For single-ship trajectory data, different processing methods are adopted according to the time interval between adjacent trajectory points. If the time interval between adjacent trajectory points is less than 1 minute, it means the data is relatively dense, and the original data can be retained. If the time interval is between 1 and 30 minutes, cubic spline interpolation is used to fill in the missing points in the trajectory. Taking 1 minute as the interval, using the information of known trajectory points, through the cubic spline interpolation formula, the LAT, LON, SOG, and COG information of the missing points is supplemented. For example, given trajectory points P1(x1, y1, t1), P2(x2, y2, t2), and t2 - t1 > 1 minute, let the time interval between t1 and t2 be Δt. At time points such as t1 + 1 minute, t1 + 2 minutes... t2 - 1 minute, the corresponding x and y coordinates (i.e., latitude and longitude) and the corresponding SOG and COG values are calculated according to the cubic spline interpolation formula. If the time interval exceeds 30 minutes, since the time interval is too long and the trajectory may have changed significantly, these points are regarded as parts of different trajectories and no interpolation is performed.
[0051] (2) Location and annotation of whereabouts:
[0052] Trajectory merging: The Dynamic Time Warping (DTW) algorithm is used to calculate the pairwise similarity between trajectories. The DTW algorithm can effectively capture the similarity in the shape of time series data. Its principle is to find an optimal time warping path between two time series so that the sum of the distances between the two series along this path is minimized. The specific calculation process is as follows. For two trajectories T1 = {p1, p2,..., p m}, and T2 = {q1, q2, …, q n} where p i and q j are points on trajectories T1 and T2 respectively, calculate the distance matrix D between them, where D(i, j) represents the distance between point p i and q j (e.g., Euclidean distance). Then, find an optimal path π = {(i1, j1), (i2, j2), …, (i k , j k )} from D(1, 1) to D(m, n) through dynamic programming, such that is minimized, and this minimum value is the DTW distance between the two trajectories. If the DTW distance between the two trajectories is lower than a pre-set threshold (e.g., an empirical value 0.5 determined through multiple experiments), they are considered highly similar and then merged. In each group of similar trajectories, by calculating the cumulative Euclidean distance between consecutive trajectory points, select the trajectory with the longest physical length as the representative of the merged result. For example, for trajectories T a , T b and T c in a group of similar trajectories, calculate their cumulative Euclidean distances (for the T a trajectory), L b and L c similarly, compare the magnitudes of L a , L b and L c , and select the trajectory with the longest length as the merged representative.
[0053] Trajectory Clustering: Use DBSCAN (Density-Based Spatial Clustering of Applications with Noise) to cluster the merged trajectory set. The core idea of the DBSCAN algorithm is based on the density of data points. If the density of data points in a region exceeds a certain threshold (i.e., the number of data points in the neighborhood is greater than the set minimum number of points, such as 10 points), then these points are divided into a cluster. At the same time, the algorithm can automatically identify clusters of any shape and size, and does not require prior specification of the number of clusters. It also has good handling ability for noise and outliers, which is very suitable for dealing with irregular points caused by signal interference or data loss in AIS data. During the clustering process, not only the starting and ending points of the trajectory (usually ports / terminals) are clustered by DBSCAN and marked as onshore destinations, but also the non-starting and non-ending points in the trajectory set are clustered to dig out more meaningful whereabouts information, such as intermediate stop points and route intersection points during the ship's voyage.
[0054] Figure 2 It is a graph showing the results of trajectory merging and trajectory clustering in the embodiment.
[0055] Semantic Whereabouts Annotation: Use a heuristic method to annotate the whereabouts obtained by clustering. For example, for a certain maritime cluster, if the sailing speeds of most ships in this cluster are slow, that is, after statistics, it is found that this cluster contains a certain proportion (such as more than 50%) of points with SOG less than a specific threshold (such as 5 knots, set according to the actual situation), then this cluster is regarded as a common stop point, indicating that ships may have carried out stop operations in this area. If the COG of ships in a certain cluster shows a high degree of diversity, specifically, assuming that the distribution of COG in this cluster conforms to a normal distribution, when the cluster contains a sufficient proportion (such as more than 30%) of points with COG deviating from the range SOG±Δd (Δd is the COG diversity threshold, for example 30°), then this cluster is regarded as a route point hub, meaning that this area may be the intersection of multiple routes or an area where ships turn frequently.
[0056] Trajectory Rearrangement: This invention focuses on long-term trajectory prediction with a time span of 10 minutes to 1 hour. Therefore, the time interval between consecutive trajectory points in the processed data is set to 1 minute. Select trajectories with the number of points m + n ∈ [20, 70] from the dataset to form a complete trajectory dataset. If m + n > 70, then truncate the trajectory so that each trajectory contains n = 10 historical data points X and m ∈ [10, 60] future data points X ′ .
[0057] Trajectory Segment Destination Allocation: First, calculate the included angle between the direction vectors from each point on the trajectory to all potential destinations. Let the trajectory point be P(x0, y0) and the potential destination be D(x d , y d ), then the direction vector Calculate the included angle through the vector included angle formula (where is a reference direction vector, such as the due east direction (1, 0)). Then, if the trajectory passes through a marked destination, assign this destination to the previous trajectory segment. For example, if the trajectory passes through a location marked as "Port A" at a certain moment, then for the segment of the trajectory from the starting point to this passing point, its destination is "Port A". Otherwise, if there is a nearest destination that is close enough to the trajectory point (such as the distance is less than a certain threshold, for example, 1 nautical mile, set according to the actual situation) and the included angle of the direction vector is small enough (such as less than 30°, set according to the actual situation), then select this destination. If there is no such suitable destination, use the onshore destination closest to the trajectory endpoint as the destination of the current point. After rearrangement, each trajectory information not only includes the LAT, LON, SOG, and COG of each point, but also the LAT and LON of the corresponding destination. At this time, the trajectory prediction task can be expressed as f: X → (X ′ , Y ′ ), where Y ′ is the predicted destination of the future trajectory X ′ , and accurately estimate the LAT and LON of Y ′ through the model. To achieve predictions for any time span, a sliding window method is used for cyclic prediction, that is, use the currently predicted trajectory point and the previous points to form a new input for subsequent predictions. For example, first use the historical trajectory points X = {x1, x2,..., x 10} to predict the first future point Then use as the new input to continue predicting the next future point And so on.
[0058] Step 2: Construction of the Destination-Guided Trajectory Prediction Model (DGTP)
[0059] BiGRU-Seq2Seq Structure: In the present invention, the encoder part of the BiGRU-Seq2Seq architecture uses BiGRU to extract temporal features from the input trajectory sequence. Specifically, the input trajectory sequence X = {x1, x2,..., x n}, where x i is the trajectory point information at each time step. BiGRU receives input information from both the forward and backward directions at each time step. Through the gating mechanism (reset gate r t and update gate zt ) to control the flow and memory of information. For example, at time step t, the reset gate r t = σ(W r ·[h t-1 , x t + b r ), the update gate z t = σ(W z ·[h t-1 , x t + b z ), where σ is the Sigmoid function, W r , W z are weight matrices, b r , b z are bias terms, and h t-1 is the hidden state of the previous time step. Then, through the candidate hidden state (where ⊙ is element-wise multiplication, W is the weight matrix, and b is the bias term) and the updated hidden state , the processing of input information and feature extraction are completed. After being processed by BiGRU, the encoded hidden states contain rich trajectory time feature information. These hidden states are then processed by a fully connected layer for dimensionality reduction to match the input requirements of the decoder. The decoder also uses BiGRU to predict future trajectories based on the output of the encoder. Its hidden state is processed by two linear layers. The first linear layer uses the hyperbolic tangent (tanh) activation function for non-linear transformation to map the hidden state to a new feature space, and the second linear layer maps the transformed representation to the final predicted trajectory points. For example, after the first linear layer y1 = tanh(W1·h + b1), and then after the second linear layer y2 = W2·y1 + b2, the final predicted trajectory points
[0060] DGTP Model Architecture: The DGTP model adopts a cascaded Seq2Seq architecture combined with BiGRU. This unique architecture design enables the model to simultaneously predict the destination and future trajectories. Specifically, it first predicts the destination and then uses it as a key guiding factor for long-term trajectory prediction. The model mainly consists of two core parts: the Destination Prediction Module (DestPM) and the Trajectory Prediction Module (TrajPM), as Figure 3 shown.
[0061] Destination Prediction Module: In the first stage of the DGTP model, the destination prediction module starts to work. The encoder Encoder1 based on bidirectional GRU is used to process the historical trajectory data. With the historical trajectory data X as the input, Encoder1 encodes the historical motion sequence from both the forward and backward directions through the bidirectional GRU structure, captures the forward and backward dependencies therein, and transforms it into a hidden representation. These hidden representations are passed to the decoder Decoder1 which is also based on bidirectional GRU and a linear layer. Decoder1 further refines the hidden representations, performs a dimensionality reduction operation through the tanh activation function, maps the hidden representations to a lower-dimensional space, and at the same time introduces a non-linear transformation to enhance the expressive power of the model. Finally, after being processed by Decoder1, the destination prediction result is generated and output in the form of longitude and latitude coordinates of.
[0062] Trajectory Prediction Module: In the second stage, the trajectory prediction module comes into play. The encoder Encoder2 of the TrajPM module receives the destination information predicted by the DestPM module, encodes it to generate the encoded destination features. At the same time, the historical trajectory features generated by Encoder1 are also prepared. Then, the encoded destination features are concatenated with the historical trajectory features generated by Encoder1 to form a unified tensor. This tensor combines the past motion information and the predicted future destination information, providing rich context information for subsequent trajectory prediction. Next, this combined representation is input into the decoder Decoder2. Decoder2 is also based on bidirectional GRU and linear transformation. Combining the historical and future destination context information, it predicts the future trajectory. The structure of Decoder2 is similar to that of Decoder1. It processes the input combined features through bidirectional GRU, captures the time series features and dependencies therein, and then maps them to the predicted values of future trajectory points through linear transformation, thereby generating the complete future trajectory prediction X ′ .
[0063] Training Process: The training of the DGTP model follows a two-stage framework. In the first stage, historical trajectory data is input into the model, and the focus is on training the DestPM module to optimize the accuracy of destination prediction. During each backpropagation process, according to the difference between the predicted destination and the actual destination, the weights of Encoder1 and Decoder1 are updated accordingly through the gradient descent algorithm, enabling the model to continuously adjust the parameters and improve the performance of destination prediction. At the same time, during the same iteration, although Encoder2 and Decoder2 are not directly updated based on the destination prediction error, they are also synchronously updated to ensure the consistency and coordination of the overall model parameters. In the second stage, the destination predicted by the DestPM module is input into Encoder2 to obtain a high-dimensional destination feature representation. After these features are merged with the historical trajectory features generated by Encoder1, Decoder2 predicts the future trajectory. In this stage, the proposed TAL loss function is used to better optimize the parameters of Encoder2 and Decoder2. The TAL loss function comprehensively considers multiple factors such as the position difference, trajectory length difference, direction difference, and dynamic time warping (DTW) difference between the predicted trajectory and the true trajectory. By minimizing the TAL loss function, the model can more accurately predict the future trajectory considering multiple factors.
[0064] Step 3: Trajectory Alignment Loss (TAL) Calculation and Application
[0065] Loss Function Requirement Analysis: When training the DestPM module, the mean squared error (MSE) loss function is used to measure the quality of destination prediction. This is because the MSE loss function can directly quantify the difference between the predicted destination and the actual destination. By minimizing the MSE loss, the model can be prompted to predict the position of the destination as accurately as possible. However, in the TrajPM module, for the destination-guided prediction method, especially in the long-term prediction scenario, it is far from enough to only consider the position difference between the predicted trajectory and the true trajectory. Factors such as the motion trend of ship navigation, such as the sailing direction and trajectory shape, are crucial for accurate trajectory prediction. For example, even if the end position of the predicted trajectory is close to the end position of the true trajectory, if the sailing direction or trajectory shape is significantly different from the actual situation, such a prediction result is unreliable in practical applications. Therefore, in order to better align the predicted trajectory with the destination and ensure that the predicted trajectory has a high degree of consistency with the true trajectory in terms of direction, shape, etc., it is necessary to comprehensively consider the correlation between these factors when optimizing the trajectory prediction. Based on this requirement, the Trajectory Alignment Loss (TAL) function is proposed.
[0066] Calculation of TAL Components:
[0067] Mean Squared Error Loss: This loss is used to measure the Euclidean distance between the predicted trajectory and the true trajectory at each time step, ensuring that the model can accurately predict the spatial position of each point in the trajectory. Its formula is:
[0068]
[0069] where, and are the latitude and longitude values predicted by the model at the i-th time step respectively, lat i and lon i are the corresponding true latitude and longitude values, and T is the total number of time steps considered. For example, if T = 10, at the 3rd time step, the predicted latitude the true latitude lat3 = 30.3°, the predicted longitude the true longitude lon3 = 120.1°, then the contribution of this time step to is (30.5 - 30.3) 2 + (120.2 - 120.1) 2 = 0.04 + 0.01 = 0.05. After calculating the contributions of all T time steps, the average value is taken to obtain
[0070] Trajectory Length Loss: Since trajectories going to the destination over a long time interval often have different shapes, in order to encourage the model to maintain the consistency of the predicted trajectory and the actual trajectory in shape and ensure that the lengths of the two trajectories are similar, the trajectory length loss is introduced. The calculation formula is:
[0071]
[0072] where, and p i represent the coordinates of the i-th time step of the prediction and the true respectively, and ‖·‖2 is the Euclidean distance (L2 norm) between the two coordinates. For example, assume that the coordinates of three consecutive points of the predicted trajectory are the coordinates of the corresponding points of the true trajectory are p1 = (1,1), p2 = (2,2), p3 = (3,4). Then the length of the predicted trajectory is the length of the true trajectory is
[0073] Direction Loss: Similar to the length loss, the direction loss measures the angular difference between the two trajectories by quantifying the cumulative vector angles formed by adjacent points along the trajectory, ensuring that the predicted trajectory goes towards the correct direction to the destination and maintains a smooth transition between adjacent points. Its formula is:
[0074]
[0075]
[0076] θ i = atan2(lat i - lat0, lon i - lon0)
[0077] where and θ i are the direction angles of the i-th predicted and true points relative to the starting point (lat0, lon0), respectively. atan2(y, x) is a mathematical function used to calculate the angle from the positive x-axis to the point (x, y) in a two-dimensional plane and takes into account the correct quadrant. For example, if the starting point coordinates are (0, 0), the coordinates of the second predicted point are (1, 1), and the coordinates of the second true point are (1, -1), then the direction loss at this time step is After calculating the direction losses for all T time steps, the average value is taken to obtain
[0078] Dynamic Time Warping (DTW) loss: As an important metric for measuring the similarity of time series, DTW is widely used in tasks such as trajectory similarity comparison. By minimizing the DTW loss, the trained DGTP model can better align the predicted trajectory and the true trajectory, even in the presence of local time warping. Its formula is:
[0079]
[0080] where and p j are a pair of coordinate points in the predicted and true trajectories, and π is the set of alignment paths, representing the corresponding elements in the two trajectories and p j group. DTW minimizes the total distance by selecting the optimal path in π. Specifically, when calculating, a distance matrix is constructed, and the matrix elements represent the Euclidean distance between the predicted trajectory points and the true trajectory points. Then, the optimal path is found through a dynamic programming algorithm, such that the sum of the distances along this path is minimized, and this minimum value is
[0081] TAL formula: Combining the above different losses, the TAL loss function adds them together in a weighted manner, and the formula is as follows:
[0082]
[0083] Among them, λ is the weight parameter, which is selected as 0.7 after hyperparameter tuning in the experiment. The selection of this weight is achieved by experimenting with different λ values on the validation set, comparing the performance of the model on the validation set, and choosing the λ value that optimizes the model's performance.
[0084] For example, when λ = 0.7, if then During the model training process, by continuously adjusting the model parameters, gradually decreases, thereby improving the prediction accuracy of the model.
[0086] Step 4: Use the trained model for trajectory prediction
[0087] Input preparation: Obtain the historical trajectory data X = {x t-n , …, x t-1 , x t} of the ship to be predicted, where x i represents the position at the i-th time step, i ∈ [t - n, t], and here n = 10. These data need to undergo the same data preprocessing steps as the training data, including extracting single-ship data according to the ship's MMSI, data cleaning, and possible trajectory repair operations, etc., to ensure that the data format and quality meet the model input requirements. For example, if the MMSI of the ship to be predicted is "987654321", extract the relevant trajectory point data of this ship from the AIS data, sort them in chronological order, and then check and process the outliers, missing values, etc. in the data to make it meet the model input format.
[0088] Destination prediction: Input the historical trajectory data X into Encoder1 in the trained Destination Prediction Module (DestPM). Encoder1 processes the input historical trajectory sequence based on bidirectional GRU, captures the forward and backward dependencies therein, and encodes it into a hidden representation. The specific process is as described above, and the input at each time step is processed through mechanisms such as the reset gate and update gate. The hidden representation is passed to Decoder1, and Decoder1 also refines the hidden representation based on bidirectional GRU and a linear layer. Decoder1 reduces the dimension of the processed hidden representation through the tanh activation function and finally generates the predicted destination Y ′ , which is represented in the form of longitude and latitude coordinates .
[0089] Trajectory prediction: Input the predicted destination Y ′ into Encoder2 in the Trajectory Prediction Module (TrajPM). At the same time, prepare the hidden state (i.e., the historical trajectory feature) obtained by Encoder1 processing the historical trajectory data X.
[0090] Encoder2 processes the predicted destination Y ′ to generate the encoded destination features. The encoded destination features are concatenated with the historical trajectory features generated by Encoder1 to form a unified tensor, which fuses the past motion information and the predicted future destination information. This combined representation is input into Decoder2, and Decoder2 predicts the future trajectory X ′ ={x t+1 , x t+2 , …, x t+m} through bidirectional GRU and linear transformation, combining historical and future destination contexts. Here, m represents the number of future time steps predicted, and m ∈ [10, 60].
[0091] Recursive prediction (sliding window method): To achieve the prediction of trajectories with a longer time span, a sliding window method is adopted for recursive prediction. First, the first predicted future trajectory point x t+1 is combined with other points in the original historical trajectory data X except the last point x t to form a new input sequence X new1 ={x t-n , …, x t-1 , x t+1}. Repeat the steps, and use the new input sequence X new1 to perform destination prediction and trajectory prediction again to obtain the next future trajectory point x t+2 . Keep looping in the above way, each time adding the newly predicted trajectory point to the input sequence and replacing the earliest historical trajectory point, and continuously perform prediction until a complete predicted trajectory sequence with the required time span is generated.
[0092] Result output: The finally output predicted result is the complete future trajectory X ′ and the predicted destination Y ′ , which can be used in related application scenarios such as ship traffic management and risk warning to provide support for the navigation decision-making of ships. For example, in a ship traffic management system, the predicted future trajectory and destination information are presented to the management personnel to help them plan the shipping lanes in advance, allocate resources, and prevent potential collision accidents.
[0093] Figure 4 is an example diagram of the trajectory predicted by the ship movement behavior prediction method in the embodiment.
[0094] Through the above detailed and comprehensive specific implementation manners, the method of the present invention can effectively achieve long-term accurate prediction of ship trajectories and meet the actual needs of the maritime field for ship trajectory prediction.
[0095] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall all be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to what is defined by the claims.
Claims
1. A ship movement behavior prediction method based on BiGRU-Seq2Seq network, characterized in that: The following steps are involved: 1) Extract the track data of each ship from the collected AIS data, perform data preprocessing, and then locate and mark the track; 2) Construct a destination-guided trajectory prediction model DGTP based on the BiGRU-Seq2Seq network. The DGTP model includes a destination prediction module and a trajectory prediction module, each of which contains an encoder and a decoder; 3) Input the trajectory sequence marked in step 1) into the destination prediction module. The encoder of this module extracts the time features of the trajectory sequence and predicts the destination based on the historical trajectory data through the decoder; Input the predicted destination into the trajectory prediction module. The encoder of this module generates the destination features and concatenates them with the historical trajectory features to generate a unified tensor. The decoder predicts the future trajectory based on the generated unified tensor; 4) Design a trajectory alignment loss consisting of mean square error loss, trajectory length loss, direction loss and dynamic time warping loss. Use this trajectory alignment loss to optimize DGTP model training and improve trajectory prediction accuracy. 5) Using the trained DGTP model, the future trajectory is predicted based on the input ship historical trajectory data, and the predicted complete trajectory sequence and destination are output.
2. The method according to claim 1, characterized in that The data preprocessing in step 1) includes cleaning the data by filtering out small ship data, data points whose latitude, longitude, ground speed and ground heading values are beyond the preset range, and removing incomplete tracks with less than 20 track points.
3. The method according to claim 1, characterized in that The data preprocessing in step 1) includes repairing the trajectory. The method is as follows: for single-ship trajectory data, if the time interval between adjacent trajectory points is less than 1 minute, the original data is retained; if the time interval is between 1 and 30 minutes, the missing points are filled in with 1-minute intervals using cubic spline interpolation to supplement the latitude, longitude, ground speed and ground heading data.
4. The method according to claim 1, characterized in that In step 1), the tracking positioning includes trajectory merging, and the method is: based on the dynamic time warping algorithm DTW, the pairwise similarity between the trajectories is calculated. If the DTW distance between two trajectories is lower than the set threshold, they are merged, and the trajectory with the longest physical length is selected as the representative trajectory in the similar trajectory group.
5. The method according to claim 1, characterized in that The tracking location in step 1) includes trajectory clustering, which is performed by clustering the trajectories using the density-based spatial clustering algorithm DBSCAN, marking the starting point and the end point of the trajectory as the land destination, and clustering the non-starting and non-end points.
6. The method according to claim 1, characterized in that The method for marking the whereabouts in step 1) is: using a heuristic method to semantically mark the clustering results, if the navigation speed of the cluster is slower than a preset threshold, the cluster is marked as a common stop point.
7. The method according to claim 1, characterized in that Step 1) also includes rearranging the trajectory by adjusting the time interval of the trajectory points to 1 minute, selecting trajectories with a number of trajectory points between 20 and 70 as the complete trajectory data set, and truncating the trajectory to a format containing 10 historical data points and 10 to 60 future data points if the trajectory length exceeds 70.
8. The method according to claim 1, characterized in that Step 1) also includes the assignment of destinations to trajectory segments, the method being: calculating the angle between the direction vectors of each point on the trajectory and the potential destination; if the trajectory passes through a marked destination, assigning the destination to the previous trajectory segment; if it does not pass through a marked destination, selecting the destination closest to the trajectory point and with a direction angle less than a preset threshold; if the above two destinations do not exist, assigning the land destination closest to the trajectory endpoint.
9. The method according to claim 1, characterized in that Mean square error loss in step 4): used to measure the Euclidean distance between the predicted trajectory and the true trajectory at each time step, the formula is: in, and is the latitude and longitude of the predicted i-th time step, lat i and lon i is the corresponding true value, T is the total number of time steps considered; Trajectory length loss: It is used to make the model maintain the consistency of the shape of the predicted trajectory and the actual trajectory. The formula is: in, and p i denote the predicted and true coordinates of the ith time step, respectively, ‖·‖2 is the L2 norm of the Euclidean distance between the two coordinates; Directional loss: used to measure the angular difference between two trajectories. The formula is: in, and θ i are the direction angles of the ith prediction and true point relative to the starting point (lat0, lon0), respectively; atan2(,) is used to calculate the angle from the positive x-axis to the trajectory point in the two-dimensional plane; Dynamic time warping loss: used to align the predicted trajectory and the true trajectory, the formula is: in, and p j is a pair of coordinate points in the predicted and true trajectories, and π represents the corresponding elements in the two trajectories and p j group.
10. The method according to claim 1, characterized in that In step 5), the DGTP model uses a sliding window method for cyclic prediction. The method is as follows: first, the predicted first future trajectory point is combined with other points in the original historical trajectory data except the last point to form a new input sequence; then the new input sequence is used to perform destination prediction and trajectory prediction again to obtain the next future trajectory point; the above steps are continuously cycled, each time the newly predicted trajectory point is added to the input sequence to replace the earliest historical trajectory point, and the prediction is continued until a complete predicted trajectory sequence of the required time span is generated.