Ship behavior prediction method, system and device fusing trajectory semantics and environment encoding

CN122799441APending Publication Date: 2026-09-22WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611267924.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

上述现有技术将轨迹坐标回归与行为意图分析割裂为两个独立阶段,先预测输出轨迹坐标再事后解释,缺乏在同一框架内协同输出预测轨迹与行为意图描述的端到端机制,既降低了推理效率又破坏了输出一致性

Benefits of technology

[0073]1、本发明所述船舶行为预测方法将海域地理环境约束、船舶运动状态语义与轨迹坐标信息统一融合于同一端到端框架内,利用GPT2大模型主干网络对组合特征进行推理,通过轨迹输出分支输出预测轨迹坐标,通过语言预测分支同步输出船舶运动趋势与朝向的文本描述,同时实现了轨迹坐标预测与行为语义解释的协同输出。因此,本发明解决了现有方法多模态融合不足和预测与解释脱节的问题,能够在复杂海域场景下实现高精度的、语义可解释的端到端船舶行为预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799441A_ABST
    Figure CN122799441A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of ship behavior prediction, and particularly relates to a ship behavior prediction method, system and device fusing trajectory semantics and environment coding, the method comprising: constructing a training sample set comprising historical trajectory data and future trajectory data; constructing multi-modal input information comprising historical trajectory data, three-channel sparse obstacle image and trajectory semantic intention based on the historical trajectory data; extracting trajectory local detail feature, obstacle environment feature and sentence feature vector based on the three kinds of multi-modal input information respectively and fusing to obtain combined features; outputting predicted trajectory coordinates by a trajectory output branch after the combined features are inferred by a GPT2 large model backbone network, and outputting text description of ship movement trend and direction by a language prediction branch; the model is trained to realize ship behavior prediction. The present application can realize end-to-end collaborative output of ship trajectory coordinate prediction and behavior semantic explanation at the same time, and is suitable for ship behavior prediction in complex sea environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ship behavior prediction technology, specifically relating to a ship behavior prediction method, system, device and medium that integrates trajectory semantics and environmental coding. Background Technology

[0002] Ship behavior prediction plays a crucial role in ship early warning, abnormal trajectory identification, and encounter situation differentiation. The Automatic Identification System (AIS) provides a large amount of real-time operational data, including the latitude and longitude coordinates of trajectory points and ship speed. By deeply mining AIS data to extract ship motion patterns and achieve accurate prediction of future trajectories, it has significant research significance and engineering application value.

[0003] Large Language Models (LLMs) such as GPT2, with their powerful temporal context modeling capabilities and cross-modal semantic understanding, can capture long-range spatiotemporal dependencies in trajectory sequences, bringing new technological possibilities for improving trajectory prediction accuracy. Patent application No. 202411288440.1 discloses a method, apparatus, and device for ship trajectory prediction based on scene semantic information. This method combines scene semantic information, a self-attention mechanism network, and a large language model, incorporating the ship's environmental location information into the network computation. It considers the impact of the environment on the ship during prediction, making the predicted trajectory more reasonable and accurate. Patent application No. 202411196913.5 discloses a method, apparatus, device, and storage medium for ship trajectory prediction. This method inputs the current trajectory point's location information, semantic location vector, and historical trajectory feature vector into a ship guidance decoder, outputting a sequence of predicted trajectory points corresponding to the current trajectory point where the ship to be predicted is located. However, the above-mentioned existing ship trajectory prediction methods still face the following technical problems:

[0004] First, the methods for utilizing scene information are limited, and multimodal information fusion is insufficient. The existing technologies either extract general, dense scene semantic features from scene images or use coarse-grained marine grid semantic vectors, which cannot accurately reflect the continuous spatial distribution of obstacles and navigable areas. Small-scale obstacles (such as individual reefs) may be lost due to grid size limitations, restricting the model's predictive robustness in complex marine scenes. Furthermore, navigation status is not encoded as semantic information for inference, making it difficult to leverage the alignment advantages of natural language forms in the LLM pre-training space.

[0005] Second, the end-to-end predicted trajectory is disconnected from the behavioral intent description. The existing technologies mentioned above separate trajectory coordinate regression and behavioral intent analysis into two independent stages, predicting the output trajectory coordinates first and then interpreting them afterward. They lack an end-to-end mechanism that coordinates the output of the predicted trajectory and the behavioral intent description within the same framework, which reduces inference efficiency and undermines output consistency.

[0006] Third, there is a lack of innovative applications at the algorithm model level. How to efficiently transfer large models such as GPT2 to the specific task of ship trajectory prediction, and accurately inject trajectory and environmental features while freezing the pre-trained language capabilities, remains a key engineering problem to be solved.

[0007] In summary, there is an urgent need for a ship behavior prediction method that can efficiently integrate trajectory semantics and environmental coding to achieve end-to-end trajectory prediction and collaborative output of behavioral intentions. Summary of the Invention

[0008] The purpose of this invention is to address the aforementioned problems in the existing technology by providing a method, system, device, and medium for ship behavior prediction that integrates trajectory semantics and environmental coding, enabling high-precision, semantically interpretable end-to-end ship behavior prediction in complex marine scenarios.

[0009] To achieve the above objectives, the technical solution of the present invention is as follows:

[0010] In a first aspect, the present invention provides a method for predicting ship behavior by fusing trajectory semantics and environmental coding, the method comprising:

[0011] S1. Construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data;

[0012] S2. Construct multimodal input information based on historical trajectory data. The multimodal input information includes historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent. The three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension. The trajectory semantic intent refers to the prompt text obtained by converting the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation, and the motion trend.

[0013] S3. Construct a multimodal data fusion feature extraction model. The multimodal data fusion feature extraction model is based on historical trajectory data, three-channel sparse obstacle images and trajectory semantic intent extraction to obtain trajectory local detail features, obstacle environment features and sentence feature vectors, and fuses the three features to obtain combined features.

[0014] S4. Construct a ship behavior prediction model, which includes a GPT2 large model backbone network and an output head. The combined features are input into the GPT2 large model backbone network to obtain fused features, and then the fused features are input into the output head. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch is used to output the latitude and longitude coordinate vectors of multiple trajectory points predicted for the future. The language prediction branch is used to output a text description containing the predicted ship motion trend and ship orientation for the future.

[0015] S5. Train the ship behavior prediction model based on future trajectory data, and use the trained ship behavior prediction model to realize ship behavior prediction.

[0016] In S3, the specific steps for extracting local trajectory detail features based on historical trajectory data include:

[0017] First, the latitude and longitude coordinates of all trajectory points in the historical trajectory are extracted and normalized to obtain trajectory features. Then, a first linear layer is used to linearly map the latitude and longitude coordinates of all trajectory points, resulting in a high-dimensional spatial feature vector. The self-attention mechanism in the Transformer encoder is used to model the relationship between these high-dimensional spatial feature vectors. The high-dimensional spatial features at each time step in the high-dimensional spatial feature vector are used as queries, and the high-dimensional spatial features at all time steps in the high-dimensional spatial feature vector are used as keys and values ​​for self-attention calculation. Finally, a second linear layer is used to reduce the dimensionality of the attention calculation result, and a residual connection is performed with the trajectory features to obtain a globally enhanced trajectory. The expression for the globally enhanced trajectory is:

[0018] ;

[0019] In the above formula, express Global feature enhancement trajectory at any given time; Represents the normalized result Latitude and longitude coordinates of the time trajectory point; , These represent the first linear layer and the second linear layer, respectively. This indicates a self-attention computation operation;

[0020] Then, the global feature enhancement trajectory is sliced ​​to obtain multiple sub-trajectory segments. All sub-trajectory segments are processed through a linear layer to obtain high-dimensional feature vectors. These high-dimensional feature vectors are then processed using a first-layer one-dimensional convolution and a ReLU activation function to obtain local trajectory encoding features. A second-layer one-dimensional convolution extracts higher-level semantic features from these local trajectory encoding features to obtain local trajectory detail features. The expression for these local trajectory detail features is as follows:

[0021] ;

[0022] In the above formula, Represents local detailed features of the trajectory; , These represent the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. Indicates a linear layer; This indicates a slice operation.

[0023] In S3, the specific steps for extracting obstacle environment features based on three-channel sparse obstacle images include:

[0024] The three-channel sparse obstacle image is subjected to a first-layer two-dimensional convolution operation. The resulting convolution result is then processed by the ReLU activation function and max pooling to obtain the first feature map.

[0025] ;

[0026] In the above formula, Represents the first feature map; Represents a three-channel sparse obstacle image; Indicates the max pooling layer; This represents the first two-dimensional convolutional layer;

[0027] Perform a second two-dimensional convolution operation on the obtained first feature map, and then apply the convolution result to the ReLU activation function to obtain the deep spatial feature map:

[0028] ;

[0029] In the above formula, Represents deep spatial feature maps; This represents the second two-dimensional convolutional layer;

[0030] An adaptive max-pooling operation is performed on the obtained deep spatial feature map to obtain an adaptive feature map:

[0031] ;

[0032] In the above formula, For adaptive feature maps; This indicates an adaptive max-pooling layer;

[0033] The adaptive feature map is flattened in both spatial and feature channel dimensions to obtain a one-dimensional representation of the high-dimensional environment feature vector:

[0034] ;

[0035] In the above formula, A high-dimensional environmental feature vector represented in one dimension; Indicates flattening;

[0036] The obtained high-dimensional environmental feature vectors are linearly projected and aligned using layer normalization to make them consistent with the embedding dimension required by the GPT2 large model, thus obtaining the obstacle environment features. .

[0037] In S3, the specific steps for obtaining sentence feature vectors based on trajectory semantic intent extraction include:

[0038] First, the GPT2 large model's word segmenter is used to segment the prompt text, resulting in a segmented sequence. ,in, Indicates the prompt text, A word segmenter representing the GPT2 large model;

[0039] Then, a lookup table is performed on the segmented sequence to obtain the word feature sequence of all words in the segmented sequence Token. ,in, This indicates a table lookup mapping operation;

[0040] Finally, the obtained word feature sequence The mean is calculated along the sequence dimension to obtain the sentence feature vector. ,in, For sequence dimensions.

[0041] The formula for calculating the average velocity of the trajectory is:

[0042] ;

[0043] In the above formula, The average velocity of the trajectory; , They are respectively time, The latitude and longitude coordinates of the trajectory point at any given moment; This represents the total number of all trajectory points in the historical trajectory data. The time slot is between two adjacent moments;

[0044] The formula for calculating the average acceleration of the trajectory is:

[0045] ;

[0046] In the above formula, The average acceleration of the trajectory; , They are respectively time, The velocity of the trajectory point at any given time;

[0047] The method for obtaining the ship's orientation is as follows:

[0048] First, calculate the bearing angle using the following formula, and then determine the ship's orientation based on the bearing angle:

[0049] ;

[0050] In the above formula, This refers to the azimuth angle; , These are the latitude and longitude coordinates of the first and last trajectory points in the historical trajectory data, respectively.

[0051] The method for obtaining the movement trend is as follows:

[0052] Determining motion trends based on trajectory average acceleration:

[0053] ;

[0054] In the above formula, Indicates the trend of movement.

[0055] The steps for constructing the three-channel sparse obstacle image include:

[0056] First, satellite nautical charts of the area where the ship's trajectory is located are obtained and processed into grayscale to obtain a grayscale map. Based on the grayscale values ​​of the grayscale map, obstacle areas and passable areas are divided.

[0057] Then, the pixel positions of the obstacle region are encoded as 1, and the rest are encoded as 0, to obtain the binary spatial mask matrix of the obstacle. Encode the pixel positions of the passable area as 1 and the rest as 0 to obtain the binary spatial mask matrix of the passable area. The historical trajectory is plotted on the grayscale image, and the pixel position of the trajectory location is encoded as 1, while the rest are encoded as 0, thus obtaining the binary spatial mask matrix of the trajectory. ;

[0058] Finally, the three binary spatial mask matrices are concatenated along the feature channel dimension to construct a three-channel sparse obstacle image. .

[0059] In S5, the ship behavior prediction model is trained based on a multi-task joint loss function, the expression of which is:

[0060] ;

[0061] ;

[0062] ;

[0063] In the above formula, For multi-task joint loss function; The mean squared error loss function used for the trajectory output branch measures the latitude and longitude coordinates of the predicted trajectory points. Latitude and longitude coordinates of the actual trajectory point Geometric deviations between them; The number of predicted trajectory points; The cross-entropy loss function used for the language prediction branch is used to evaluate the accuracy of the text output; The probability of correctly predicting the true label for the text output; The effective sequence length for the text output.

[0064] Secondly, the present invention provides a ship behavior prediction system that integrates trajectory semantics and environmental coding, the ship behavior prediction system comprising:

[0065] The data acquisition module is used to construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data;

[0066] A multimodal input construction module is used to construct multimodal input information based on historical trajectory data. The multimodal input information includes historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent. The three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension. The trajectory semantic intent refers to the prompt text obtained by converting the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation, and the motion trend.

[0067] The feature extraction and fusion module is used to construct a multimodal data fusion feature extraction model. The multimodal data fusion feature extraction model extracts local trajectory details, obstacle environment features, and sentence feature vectors based on historical trajectory data, three-channel sparse obstacle images, and trajectory semantic intent extraction, and then fuses the three features to obtain combined features.

[0068] The prediction model construction module is used to construct a ship behavior prediction model. The ship behavior prediction model includes a GPT2 large model backbone network and an output head. The combined features are input into the GPT2 large model backbone network to obtain fused features, and then the fused features are input into the output head. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch outputs the latitude and longitude coordinate vectors of multiple trajectory points predicted for the future; the language prediction branch outputs a text description containing the predicted ship motion trend and ship orientation.

[0069] The model training and output module is used to train the ship behavior prediction model based on future trajectory data, and to use the trained ship behavior prediction model to realize ship behavior prediction.

[0070] Thirdly, the present invention provides a ship behavior prediction device that integrates trajectory semantics and environmental coding. The ship behavior prediction device includes a memory and a processor. The memory is used to store computer program code and transmit the computer program code to the processor. The processor is used to execute the aforementioned ship behavior prediction method according to the instructions in the computer program code.

[0071] Fourthly, the present invention provides a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, implements the aforementioned ship behavior prediction method.

[0072] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0073] 1. The ship behavior prediction method of this invention integrates marine geographical environment constraints, ship motion state semantics, and trajectory coordinate information into a single end-to-end framework. It utilizes the GPT2 large-scale model backbone network to infer the combined features, outputs predicted trajectory coordinates through the trajectory output branch, and simultaneously outputs textual descriptions of ship motion trends and orientations through the language prediction branch. This achieves coordinated output of trajectory coordinate prediction and behavioral semantic interpretation. Therefore, this invention solves the problems of insufficient multimodal fusion and the disconnect between prediction and interpretation in existing methods, enabling high-precision, semantically interpretable end-to-end ship behavior prediction in complex marine scenarios.

[0074] 2. The ship behavior prediction method of this invention, when extracting local detail features of the trajectory, first utilizes the self-attention mechanism of the Transformer encoder to perform global relational modeling of the high-dimensional spatial features of all time steps of the entire trajectory, so that the trajectory point features at each moment are integrated into the contextual information of the entire trajectory; then, by connecting with the residuals of the normalized trajectory features, the accurate position reference is preserved, resulting in a globally enhanced trajectory that contains both global temporal dependencies and maintains positional accuracy; finally, through slicing and two layers of one-dimensional convolution, trajectory motion details within local time periods are further extracted from the globally enhanced trajectory. The above-mentioned serial structure from global modeling to residual preservation to local refinement enables the trajectory features to simultaneously possess the ability to perceive global motion trends and the ability to depict details such as local turning and speed changes, suppressing the interference of position and scale differences between different coordinates on feature extraction. Therefore, this invention can extract richer and more accurate motion feature representations from historical trajectories.

[0075] 3. The ship behavior prediction method of this invention, when performing obstacle environment feature extraction, first extracts features layer by layer from the three-channel sparse obstacle image through a two-layer concatenated two-dimensional convolutional network. Specifically, the first convolutional layer increases the feature dimension and enlarges the receptive field through max pooling; the second convolutional layer further extracts the spatial association pattern between obstacles, passable areas, and trajectories in the three channels. Then, adaptive max pooling compresses the obtained feature map into a fixed spatial grid. Finally, flattening and linear projection normalization precisely align the environmental feature dimension to the embedding dimension required by the GPT2 large model, ensuring that environmental information can be fused with trajectory features and semantic features in the same embedding space. Therefore, this invention can inject marine geographical environment constraints into the large model inference process, providing prior spatial perception of the marine environment for behavior prediction.

[0076] 4. In the ship behavior prediction method of this invention, when extracting semantic features, the method first segments and maps the trajectory semantic intent prompt text using the built-in word segmenter and word embedding layer of the GPT2 large model. This ensures that the word vector space of the semantic features is naturally consistent with the pre-trained language space of the GPT2 large model, eliminating the need for additional training of a separate text encoder and reducing model complexity. Then, the semantics of the entire sentence are aggregated by calculating the mean across the word feature sequence dimension. This preserves core semantic information such as trajectory movement trends and ship orientation while avoiding the inconsistency in feature dimensions caused by variable-length prompt text. Therefore, this invention can inject ship motion state semantic information into a large model in a lightweight manner, providing interpretable semantic support for behavior prediction. Attached Figure Description

[0077] Figure 1 This is a flowchart of the method described in this invention.

[0078] Figure 2 This is the ship orientation determination rule described in this invention.

[0079] Figure 3 This is a schematic diagram of the structure of the multimodal data fusion feature extraction model described in this invention.

[0080] Figure 4 for Figure 3 A schematic diagram of global and local information embedding in the trajectory feature encoding module.

[0081] Figure 5 This is a schematic diagram of the ship behavior prediction model described in this invention.

[0082] Figure 6 This is a structural block diagram of the system described in this invention.

[0083] Figure 7 This is a structural block diagram of the device described in this invention. Detailed Implementation

[0084] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0085] Example 1:

[0086] See Figure 1 A method for predicting ship behavior that integrates trajectory semantics and environmental coding is proposed, and is carried out in the following steps:

[0087] S1. Construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data.

[0088] Specifically, ship trajectory data from the Automatic Identification System (AIS) is collected and preprocessed. A training sample set is then constructed based on the preprocessed ship trajectory data. A complete segment of ship trajectory data refers to a time-series sequence composed of trajectory data from multiple trajectory points, including the latitude and longitude coordinates, speed, and timestamps of the trajectory points. The preprocessing steps include deleting abnormal drift points, deleting missing points, and deleting stationary trajectories. A single training sample in the training sample set is formed by dividing a complete segment of ship trajectory data into historical trajectory data and future trajectory data. Subsequently, the historical trajectory data will be used as input data to construct a ship behavior prediction model, predicting the future trajectory point latitude and longitude coordinates, ship movement trends, and ship orientation. Simultaneously, the future trajectory data is used as ground truth and input into a pre-constructed multi-task joint loss function. The ship behavior prediction model is trained by minimizing the multi-task joint loss function.

[0089] S2. Construct multimodal input information based on historical trajectory data in the training sample set. The multimodal input information includes historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent. The three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension. The trajectory semantic intent includes the following information: the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation (e.g., eight directions such as east, south, west, north, northeast, northwest, southeast, and southwest), and the motion trend (e.g., acceleration, deceleration, constant speed). Convert the above information into prompt text, which is the trajectory semantic intent.

[0090] Specifically, this invention uses a three-channel sparse obstacle image to decompose complex marine scene information into three non-overlapping feature channels through sparse coding. This sparse coding method preserves the independence of each information source and achieves spatial alignment through the channel dimension, enabling the transformation of marine spatial constraints into standardized image tensor inputs with low-complexity coding. The construction steps of the three-channel sparse obstacle image include:

[0091] First, the satellite nautical chart is processed into a grayscale image. Taking advantage of the different grayscale values ​​of obstacle areas and passable areas (land or islands have low grayscale values ​​close to 0-80, while passable waters have high grayscale values ​​of 180-255), obstacle areas and passable areas are clearly delineated by the difference in grayscale values.

[0092] Then, the pixel positions of the obstacle region are encoded as 1, and the rest as 0, to obtain the binary spatial mask matrix of the obstacle. The pixel positions of the passable area are encoded as 1, and the rest as 0, to obtain the binary spatial mask matrix of the passable area. The historical trajectory is drawn on the grayscale image, and the pixel position of the trajectory is encoded as 1, while the rest are 0, to obtain the binary spatial mask matrix of the trajectory. For example, the actual latitude and longitude of the trajectory points are first encoded into image pixel coordinates according to the following formula:

[0093] ;

[0094] The binary space mask matrix is ​​obtained using the following formula. :

[0095] ;

[0096] In the above formula, for The latitude and longitude coordinates of the trajectory point in the image pixels at any given time are determined by the actual latitude and longitude coordinates of the trajectory point. Encoded; , These represent the maximum and minimum longitudes of the area where the satellite chart is located, respectively. , These represent the maximum and minimum latitudes of the area where the satellite chart is located, respectively. , These are pixel height and pixel width, respectively.

[0097] Finally, the three binary spatial mask matrices are concatenated along the feature channel dimension to construct a three-channel sparse obstacle image with a dimension of 3×64×64. .

[0098] Specifically, in the trajectory semantic intent, the formula for calculating the average velocity of the trajectory is:

[0099] ;

[0100] In the above formula, The average velocity of the trajectory; for The latitude and longitude coordinates of the trajectory point at any given moment; This represents the total number of all trajectory points in the historical trajectory data. The time slot is between two adjacent moments;

[0101] The formula for calculating the average acceleration of the trajectory is:

[0102] ;

[0103] In the above formula, The average acceleration of the trajectory; , They are respectively time, The velocity of the trajectory point at any given time;

[0104] First, calculate the azimuth angle using the following formula, then combine the azimuth angle with... Figure 2 The shown judgment rule yields the ship's orientation:

[0105] ;

[0106] In the above formula, It is the azimuth angle; , These are the latitude and longitude coordinates of the first and last trajectory points in the historical trajectory data, respectively. It is a two-parameter arctangent function;

[0107] Determine the movement trend using the following formula:

[0108] ;

[0109] In the above formula, Indicates the trend of movement.

[0110] For future trajectory data, the semantic intent of the trajectory and the trajectory data of each trajectory point are extracted and used as the real values ​​for training the subsequent ship behavior prediction model.

[0111] S3, Construction as follows Figure 3 , Figure 4The multimodal data fusion feature extraction model shown includes a trajectory feature encoding module, an environment feature encoding module, and a semantic prompt word feature encoding module. The trajectory feature encoding module, the environment feature encoding module, and the semantic prompt word feature encoding module obtain trajectory local detail features, obstacle environment features, and sentence feature vectors based on historical trajectory data, three-channel sparse obstacle images, and trajectory semantic intent extraction, respectively. The above three features are fused to obtain combined features.

[0112] Specifically, the trajectory feature encoding module is a Transformer encoder containing residual connections, used to extract and enhance global trajectory information, and to perform local time-series modeling through slicing and convolutional networks to enhance local information of trajectory coordinates. Specific steps include:

[0113] First, the latitude and longitude coordinates of all trajectory points in the historical trajectory are extracted and normalized to obtain trajectory features. Global information embedding is then performed on these features: a first linear layer linearly maps the latitude and longitude coordinates of all trajectory points to obtain high-dimensional spatial feature vectors. The self-attention mechanism in the Transformer encoder is used to model the relationships between these high-dimensional spatial feature vectors. Self-attention is calculated using the time-series high-dimensional spatial features in the trajectory point high-dimensional spatial feature vector as the query and the high-dimensional spatial features at all time steps in the trajectory point high-dimensional spatial feature vector as the key and value. A second linear layer reduces the dimensionality of the attention calculation results. Finally, the dimensionality-reduced results are residually concatenated with the trajectory features to obtain the corresponding globally enhanced trajectory.

[0114] ;

[0115] In the above formula, express Global feature enhancement trajectory at any given time; Represents the normalized result The latitude and longitude coordinates of the time trajectory points are normalized using Z-score standardization. , These represent the first linear layer and the second linear layer, respectively. This indicates a self-attention computation operation;

[0116] Then, local information embedding is performed: first, the global feature enhancement trajectory is processed. The process involves slicing the trajectory to obtain multiple sub-trajectory segments. These segments are then processed through a linear layer to obtain high-dimensional feature vectors. Next, a first-layer one-dimensional convolution and ReLU activation function are used to further enhance local trajectory details, resulting in local trajectory encoding features. Finally, a second-layer one-dimensional convolution extracts higher-level semantic features from these local encoding features, yielding the local trajectory detail features.

[0117] ;

[0118] In the above formula, Represents local detailed features of the trajectory; , These represent the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. Indicates a linear layer; This indicates a slice operation.

[0119] Specifically, the environmental feature encoding module is a concatenated network structure containing multiple convolutional layers, which encodes the input three-channel sparse obstacle image. By employing two layers of 2D convolution and adaptive max pooling, features with a spatial resolution of 4×4 are locked. These features are then flattened along the channel dimension and aligned to the embedding dimension required by the GPT2 large model to obtain obstacle environment features near the trajectory. Specific steps include:

[0120] For the three-channel sparse obstacle image The first layer of two-dimensional convolution operation is performed to transform the original three-channel sparse obstacle image. The image is transformed into a 16-dimensional feature image. High-dimensional features are extracted, and the resulting convolutional result is spatially downsampled using the ReLU activation function and max pooling to obtain the first feature map. :

[0121] ;

[0122] In the above formula, Indicates the max pooling layer; This represents the first two-dimensional convolutional layer;

[0123] The obtained first feature map A second layer of two-dimensional convolution is performed, and the resulting convolution is processed by the ReLU activation function to obtain the deep spatial feature map. The purpose of this step is to transform the 16-dimensional features after the first layer of 2D convolution into 32-dimensional features to display deeper layers of features.

[0124] ;

[0125] In the above formula, This represents the second two-dimensional convolutional layer;

[0126] The obtained deep spatial feature map An adaptive max pooling operation is performed to force the result to be spatially locked into a 4×4 feature grid while preserving spatial orientation information, resulting in an adaptive feature map. :

[0127] ;

[0128] In the above formula, This indicates an adaptive max-pooling layer;

[0129] For adaptive feature maps Flattening is performed in the spatial and feature channel dimensions to obtain a one-dimensional high-dimensional environmental feature vector. The image dimensions before flattening are 32×4×4, and the image dimensions after flattening are 1×512.

[0130] ;

[0131] In the above formula, Indicates flattening;

[0132] The obtained high-dimensional environmental feature vector Linear projection processing is performed, and layer normalization alignment is applied to make it consistent with the embedding dimension required by the GPT2 large model (dimension 768), thus obtaining the final obstacle environment features. .

[0133] Specifically, the semantic prompt word feature encoding module uses the GPT2 large model's word segmenter to segment the prompt word text, obtaining a segmented sequence. ,in, Indicates the prompt text, This represents the tokenizer of the GPT2 large model. Subsequently, a lookup table is performed on the segmented sequence to obtain the word feature sequence of all words in the token segmentation sequence. ,in, This represents a table lookup mapping operation; it applies to the resulting word feature sequence. The mean is calculated along the sequence dimension to obtain the sentence feature vector. ,in, The sequence dimension is 256. The sequence dimension of a prompt word is typically fixed at 256.

[0134] Sentence feature vector Environmental characteristics of obstacles Local details of the trajectory By splicing them together, we can obtain the combined features. .

[0135] S4, construct as follows Figure 5 The ship behavior prediction model shown includes a GPT2 large model backbone network and an output head. Combined features are input into the GPT2 large model backbone network to obtain fused features. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch outputs the latitude and longitude coordinate vectors of multiple trajectory points predicted in the future through a linear mapping layer based on the fused features. The language prediction branch outputs a text description containing the predicted ship motion trend and ship orientation based on the fused features using the language model head.

[0136] S5. Train the ship behavior prediction model based on future trajectory data, and use the trained ship behavior prediction model to realize ship behavior prediction.

[0137] Specifically, a partial fine-tuning strategy was employed to train the ship behavior prediction model. Only the first three transformer layers of the GPT2 main model backbone were used as the inference core. The weight parameters of the self-attention layer and feedforward network within the backbone were frozen, and only the layer normalization and position encoding layers were opened for weight updates and fine-tuning. Simultaneously, the trajectory feature encoding module, environmental feature encoding module, trajectory output branch, and language prediction branch outside the GPT2 main model were all updated as trainable parameters through backpropagation. This guided the frozen main model backbone to adapt to the multimodal joint distribution of trajectory coordinates and text prediction. Since the semantic prompt word feature encoding module calls the word embedding layer built into the GPT2 main model for feature encoding, and the word embedding layer is frozen during training, the semantic prompt word feature encoding module does not participate in training.

[0138] Specifically, the multi-task joint loss function used for model training consists of trajectory regression loss and language classification loss, allowing the model to naturally balance trajectory accuracy and semantic accuracy during training, effectively achieving joint optimization. The expression for the multi-task joint loss function is:

[0139] ;

[0140] ;

[0141] ;

[0142] In the above formula, For multi-task joint loss function; The mean squared error loss function used for the trajectory output branch measures the latitude and longitude coordinates of the predicted trajectory points. Latitude and longitude coordinates of the actual trajectory point Geometric deviations between them; The number of predicted trajectory points; The cross-entropy loss function used for the language prediction branch is used to evaluate the accuracy of the text output; The probability of correctly predicting the true label for the text output; The effective sequence length for the text output.

[0143] The ship behavior prediction results include the future ship movement trend, ship orientation, and the latitude and longitude coordinates of each trajectory point.

[0144] Performance verification:

[0145] 1. To verify the effectiveness of the method described in this invention, a training dataset of nearly 20,000 ship trajectories from a certain waterway was created, and comparative experiments were conducted with other models. Other models included the baseline models LSTM and Transformer. The experimental environment was a Windows 11 system, configured with Python 3.9.21, PyTorch 2.7.0, and CUDA 13.1. The training parameters were set as follows: the initial learning rate was 1×10⁻⁶. -4 The training run consisted of 100 rounds, with a batch size of 128. The comparison results are shown in Table 1.

[0146] Table 1 shows the comparison results between the method described in this invention and other models.

[0147] ;

[0148] As can be seen, the method described in this invention outperforms existing models in both the average displacement error (ADE) and final displacement error (FDE).

[0149] 2. To further demonstrate the advantages of the method described in this invention in trajectory behavior semantic understanding, a typical ship trajectory is selected as a specific example for illustration. Historical trajectory data of a typical ship trajectory stored in .npy format is input into the multimodal data fusion feature extraction model and the ship behavior prediction model. This historical trajectory data contains a sequence of latitude and longitude coordinates for 60 consecutive time steps (a total of 600 seconds), and a satellite image of the waterway containing the historical trajectory is also input. Finally, the ship behavior prediction model outputs a sequence of predicted trajectory point coordinates for the next 60 time steps, along with textual descriptions explaining the ship's movement trend and orientation.

[0150] The final predicted text description is "the ship is heading east and its movement is decelerating," which is consistent with the true meaning of this typical ship trajectory. Figure 1This indicates that the method described in this invention achieves accurate semantic prediction of behavioral direction and speed trend. The ADE of the ship trajectory point coordinates is 0.0912, and the FDE is 0.1351. This further illustrates that the method described in this invention simultaneously achieves high-precision trajectory coordinate prediction and interpretable behavioral intent text output within the same framework. It not only outperforms existing models in prediction accuracy but also possesses textual descriptions of the ship's future orientation and motion trend predictions that traditional prediction models cannot provide. This verifies the advancement and practicality of the method described in this invention.

[0151] Example 2:

[0152] See Figure 6A ship behavior prediction system integrating trajectory semantics and environmental coding is disclosed, comprising a data acquisition module and a multimodal input construction module. The data acquisition module is used to construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data; specifically, the data acquisition module is used to execute S1 in Embodiment 1, which will not be elaborated here. The multimodal input construction module is used to construct multimodal input information based on historical trajectory data, the multimodal input information including historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent; the three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension; the trajectory semantic intent refers to prompt text obtained from the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation, and the transformation of motion trend; specifically, the multimodal input construction module is used to execute S2 in Embodiment 1, which will not be elaborated here. The feature extraction and fusion module is used to construct a multimodal data fusion feature extraction model, the multimodal data fusion feature extraction model being divided into... Based on historical trajectory data, three-channel sparse obstacle images, and trajectory semantic intent extraction, local trajectory detail features, obstacle environment features, and sentence feature vectors are obtained, and the three features are fused to obtain combined features. Specifically, the feature extraction and fusion module is used to execute S3 in Example 1, which will not be elaborated here. The prediction model construction module is used to construct a ship behavior prediction model, which includes a GPT2 large model backbone network and an output head. The combined features are input into the GPT2 large model backbone network to obtain fused features, and then the fused features are input into the output head. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch outputs the latitude and longitude coordinate vectors of multiple trajectory points predicted in the future. The language prediction branch outputs a text description containing the predicted ship movement trend and ship orientation in the future. Specifically, the prediction model construction module is used to execute S4 in Example 1, which will not be elaborated here. The model training and output module is used to train the ship behavior prediction model based on future trajectory data, and uses the trained ship behavior prediction model to realize ship behavior prediction. Specifically, the model training and output module is used to execute S5 in Example 1, which will not be elaborated here.

[0153] Example 3:

[0154] See Figure 7 A ship behavior prediction device integrating trajectory semantics and environmental coding includes a memory and a processor; the memory is used to store computer program code and transmit the computer program code to the processor; the processor is used to execute the ship behavior prediction method described in Embodiment 1 according to the instructions in the computer program code.

[0155] Example 4:

[0156] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the ship behavior prediction method described in Embodiment 1.

[0157] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program goods. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0158] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0159] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0160] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for predicting ship behavior by fusing trajectory semantics and environmental coding, characterized in that: The ship behavior prediction method includes: S1. Construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data; S2. Construct multimodal input information based on historical trajectory data. The multimodal input information includes historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent. The three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension. The trajectory semantic intent refers to the prompt text obtained by converting the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation, and the motion trend. S3. Construct a multimodal data fusion feature extraction model. The multimodal data fusion feature extraction model is based on historical trajectory data, three-channel sparse obstacle images and trajectory semantic intent extraction to obtain trajectory local detail features, obstacle environment features and sentence feature vectors, and fuses the three features to obtain combined features. S4. Construct a ship behavior prediction model, which includes a GPT2 large model backbone network and an output head. The combined features are input into the GPT2 large model backbone network to obtain fused features, and then the fused features are input into the output head. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch is used to output the latitude and longitude coordinate vectors of multiple trajectory points predicted for the future. The language prediction branch is used to output a text description containing the predicted ship motion trend and ship orientation for the future. S5. Train the ship behavior prediction model based on future trajectory data, and use the trained ship behavior prediction model to realize ship behavior prediction.

2. The ship behavior prediction method fusing trajectory semantics and environmental coding according to claim 1, characterized in that: In S3, the specific steps for extracting local trajectory detail features based on historical trajectory data include: First, the latitude and longitude coordinates of all trajectory points in the historical trajectory data are extracted and normalized to obtain trajectory features. Then, a first linear layer is used to linearly map the latitude and longitude coordinates of all trajectory points, resulting in a high-dimensional spatial feature vector. The self-attention mechanism in the Transformer encoder is used to model the relationship between these high-dimensional spatial feature vectors. The high-dimensional spatial features at each time step in the trajectory point's high-dimensional spatial feature vector are used as queries, and the high-dimensional spatial features at all time steps in the trajectory point's high-dimensional spatial feature vector are used as keys and values ​​for self-attention calculation. Finally, a second linear layer is used to reduce the dimensionality of the attention calculation results, and a residual connection is performed with the trajectory features to obtain a globally enhanced trajectory. The expression for the globally enhanced trajectory is: ; In the above formula, express Global feature enhancement trajectory at any given time; Represents the normalized result Latitude and longitude coordinates of the time trajectory point; , These represent the first linear layer and the second linear layer, respectively. This indicates a self-attention computation operation; Then, the global feature enhancement trajectory is sliced ​​to obtain multiple sub-trajectory segments. All sub-trajectory segments are processed through a linear layer to obtain high-dimensional feature vectors. These high-dimensional feature vectors are then processed using a first-layer one-dimensional convolution and a ReLU activation function to obtain local trajectory encoding features. A second-layer one-dimensional convolution extracts higher-level semantic features from these local trajectory encoding features to obtain local trajectory detail features. The expression for these local trajectory detail features is as follows: ; In the above formula, Represents local detailed features of the trajectory; , These represent the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. Indicates a linear layer; This indicates a slice operation.

3. The ship behavior prediction method integrating trajectory semantics and environmental coding according to claim 1, characterized in that: In S3, the specific steps for extracting obstacle environment features based on three-channel sparse obstacle images include: The three-channel sparse obstacle image is subjected to a first-layer two-dimensional convolution operation. The resulting convolution result is then processed by the ReLU activation function and max pooling to obtain the first feature map. ; In the above formula, Represents the first feature map; Represents a three-channel sparse obstacle image; Indicates the max pooling layer; This represents the first two-dimensional convolutional layer; Perform a second two-dimensional convolution operation on the obtained first feature map, and then apply the convolution result to the ReLU activation function to obtain the deep spatial feature map: ; In the above formula, Represents deep spatial feature maps; This represents the second two-dimensional convolutional layer; An adaptive max-pooling operation is performed on the obtained deep spatial feature map to obtain an adaptive feature map: ; In the above formula, For adaptive feature maps; This indicates an adaptive max-pooling layer; The adaptive feature map is flattened in both spatial and feature channel dimensions to obtain a one-dimensional representation of the high-dimensional environment feature vector: ; In the above formula, A high-dimensional environmental feature vector represented in one dimension; This indicates a flattening process; The obtained high-dimensional environmental feature vectors are linearly projected and aligned using layer normalization to make them consistent with the embedding dimension required by the GPT2 large model, thus obtaining the obstacle environment features. .

4. The ship behavior prediction method integrating trajectory semantics and environmental coding according to claim 1, characterized in that: In S3, the specific steps for obtaining sentence feature vectors based on trajectory semantic intent extraction include: First, the GPT2 large model's word segmenter is used to segment the prompt text, resulting in a segmented sequence. ,in, Indicates the prompt text, A word segmenter representing the GPT2 large model; Then, a lookup table is performed on the segmented sequence to obtain the word feature sequence of all words in the segmented sequence Token. ,in, This indicates a table lookup mapping operation; Finally, the obtained word feature sequence The mean is calculated along the sequence dimension to obtain the sentence feature vector. ,in, For sequence dimensions.

5. The ship behavior prediction method fusing trajectory semantics and environmental coding according to claim 1, characterized in that: The formula for calculating the average velocity of the trajectory is: ; In the above formula, The average velocity of the trajectory; , They are respectively time, The latitude and longitude coordinates of the trajectory point at any given moment; This represents the total number of all trajectory points in the historical trajectory data. The time slot is between two adjacent moments; The formula for calculating the average acceleration of the trajectory is: ; In the above formula, The average acceleration of the trajectory; , They are respectively time, The velocity of the trajectory point at any given time; The method for obtaining the ship's orientation is as follows: First, calculate the azimuth angle using the following formula, and then determine the ship's heading based on the azimuth angle: ; In the above formula, This refers to the azimuth angle. , These are the latitude and longitude coordinates of the first and last trajectory points in the historical trajectory data, respectively. The method for obtaining the movement trend is as follows: Determining motion trends based on trajectory average acceleration: ; In the above formula, Indicates the trend of movement.

6. The ship behavior prediction method fusing trajectory semantics and environmental coding according to claim 1, characterized in that: The steps for constructing the three-channel sparse obstacle image include: First, satellite nautical charts of the area where the ship's trajectory is located are obtained and processed into grayscale to obtain a grayscale map. Based on the grayscale values ​​of the grayscale map, obstacle areas and passable areas are divided. Then, the pixel positions of the obstacle region are encoded as 1, and the rest are encoded as 0, to obtain the binary spatial mask matrix of the obstacle. Encode the pixel positions of the passable area as 1 and the rest as 0 to obtain the binary spatial mask matrix of the passable area. The historical trajectory is plotted on the grayscale image, and the pixel position of the trajectory location is encoded as 1, while the rest are encoded as 0, thus obtaining the binary spatial mask matrix of the trajectory. ; Finally, the three binary spatial mask matrices are concatenated along the feature channel dimension to construct a three-channel sparse obstacle image. .

7. The ship behavior prediction method fusing trajectory semantics and environmental coding according to claim 1, characterized in that: In S5, the ship behavior prediction model is trained based on a multi-task joint loss function, the expression of which is: ; ; ; In the above formula, For multi-task joint loss function; The mean squared error loss function used for the trajectory output branch measures the latitude and longitude coordinates of the predicted trajectory points. Latitude and longitude coordinates of the actual trajectory point Geometric deviations between them; The number of predicted trajectory points; The cross-entropy loss function used for the language prediction branch is used to evaluate the accuracy of the text output; The probability of correctly predicting the true label for the text output; The effective sequence length for the text output.

8. A ship behavior prediction system integrating trajectory semantics and environmental coding, characterized in that: The ship behavior prediction system includes: The data acquisition module is used to construct a training sample set based on ship trajectory data, the training sample set including historical trajectory data and future trajectory data; A multimodal input construction module is used to construct multimodal input information based on historical trajectory data. The multimodal input information includes historical trajectory data, a three-channel sparse obstacle image, and trajectory semantic intent. The three-channel sparse obstacle image refers to a sparse image obtained by stitching together obstacle masks, passable area masks, and ship trajectory masks in the channel dimension. The trajectory semantic intent refers to the prompt text obtained by converting the minimum and maximum values ​​of the latitude and longitude coordinates of all trajectory points in the historical trajectory data in each coordinate dimension, the average trajectory velocity, the average trajectory acceleration, the ship orientation, and the motion trend. The feature extraction and fusion module is used to construct a multimodal data fusion feature extraction model. The multimodal data fusion feature extraction model extracts local trajectory details, obstacle environment features, and sentence feature vectors based on historical trajectory data, three-channel sparse obstacle images, and trajectory semantic intent extraction, and then fuses the three features to obtain combined features. The prediction model construction module is used to construct a ship behavior prediction model. The ship behavior prediction model includes a GPT2 large model backbone network and an output head. The combined features are input into the GPT2 large model backbone network to obtain fused features, and then the fused features are input into the output head. The output head consists of a trajectory output branch and a language prediction branch. The trajectory output branch outputs the latitude and longitude coordinate vectors of multiple trajectory points predicted for the future; the language prediction branch outputs a text description containing the predicted ship motion trend and ship orientation. The model training and output module is used to train the ship behavior prediction model based on future trajectory data, and to use the trained ship behavior prediction model to realize ship behavior prediction.

9. A ship behavior prediction device that integrates trajectory semantics and environmental coding, characterized in that: The ship behavior prediction device includes a memory and a processor; the memory is used to store computer program code and transmit the computer program code to the processor; the processor is used to execute the ship behavior prediction method as described in claim 1 according to the instructions in the computer program code.

10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the ship behavior prediction method as described in claim 1.

Citation Information

Patent Citations

  • Ship trajectory prediction method, device and equipment based on scene semantic information

    CN118799830A

  • Ship trajectory prediction method, device and equipment and storage medium

    CN119066363A