A method for predicting the drifting trajectory of a person falling into the sea

By improving the Transformer model, adding an adaptive multi-head attention weight mechanism and LSTM layer, the problem of low prediction accuracy of drift trajectory of people who fell into the water on the sea in the existing technology is solved, and more efficient and reliable prediction results are achieved, and more effective search and rescue operations are supported.

CN119538755BActive Publication Date: 2025-06-20FUJIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510106186.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing method of drift trajectory prediction for people who fell into the sea lacks generalization when processing different locations and meteorological data, and cannot effectively capture the timing relationship between the data, resulting in low prediction accuracy.

Method used

Deep learning method is adopted to improve the Transformer model, and the adaptive multi-head attention weight mechanism (AMWM) and LSTM layers are added to better capture the timing relationship and the differences between different influencing factors.

Benefits of technology

It improves the accuracy and reliability of the drift trajectory prediction of sea-falling personnel, enhances the efficiency of search and rescue operations, reduces resource waste, and increases the survival chance of people falling into the water.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538755B_ABST
    Figure CN119538755B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of maritime search and rescue, and specifically discloses a method for predicting the drifting trajectory of maritime personnel falling into the water, including the following steps: S1, collecting relevant data of maritime personnel falling into the water; S2, performing normalization processing on the data; S3, making data Loaders for training and testing according to the processed data set; S4, iterating multiple times in a loop, feeding the data in each batch into the AMWM-Transformer model, obtaining the output of the model and calculating the loss value with the original label data, and updating the model parameters according to the loss value to complete the model training; S5, according to the trained model, feeding the input features of the test to obtain the output, calculating the cumulative loss value with the true label, obtaining the total loss value of the final test data, and using ADE and FDE to evaluate the experimental results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maritime search and rescue, and in particular to a method for predicting the drifting trajectory of a person falling into the sea. Background Art

[0002] In maritime search and rescue operations, the rescue of a person falling into the sea is a very crucial and urgent task. However, due to the uncertainty of meteorology and the complexity of drifting, the operation becomes particularly difficult. First of all, the search and rescue of a person falling into the sea requires accurate positioning. A drifting object moves under the action of the net balance force of the wind, water flow and waves acting on it. If the trajectory of the drifting object or the most likely search area cannot be accurately predicted, it is very difficult to find the drifting object. Therefore, in the past two decades, how to model the movement of drifting objects has been a research hotspot. And the prediction of the drifting trajectory can enable rescue personnel to more effectively and conveniently understand the position and drifting movement of the person falling into the sea, and then formulate targeted search and rescue operations. In addition, searching in such a vast area at sea is very time-consuming and laborious. Through the prediction of the drifting trajectory, the search area can be reduced, the search time can be minimized to the greatest extent, and the rescue success rate can be improved. Therefore, accurately predicting the drifting trajectory of a person falling into the sea to improve the rescue success rate has become an important research direction, and how to achieve efficient and accurate trajectory prediction of the falling target is particularly important.

[0003] In the category of trajectory prediction, the prediction of the drifting path of a person falling into the sea also belongs to a type of particle prediction. In particle prediction, traditional numerical model methods are usually used to simulate and predict the movement of particles. Some researchers use the analytical method to predict the drifting trajectory. Since the drifting movement of a person falling into the sea is a complex synthetic movement under the action of external forces such as wind, waves and currents, this method first obtains some data factors that will directly affect the drifting trajectory of a person falling into the sea, so as to determine the motion equation model (a relatively simple flow field model). Then, according to the initial conditions of the current person falling into the sea (initial velocity, weight, falling position, direction, etc.) as the input of the motion equation model for solution, the displacement of the drifting trajectory obtained according to different times is obtained, so as to determine the position at a certain moment. For example, taking u, v (drifting decomposition velocity), x, y (drifting coordinate position) as an example, to predict the drifting position and decomposition velocity at a future time t, through this numerical simulation means, the time t is divided into several blocks , after predicting the drifting position and decomposition velocity of each block, the Runge-Kutta method is used for prediction, and the specific formula is as follows:

[0004] ;

[0005] Due to the deficiencies of the analytical method, such as ignoring randomness (the analytical method describes changes with a deterministic formula, ignoring the occurrence of sudden situations) and being limited to simplified models, some researchers have proposed the Monte Carlo simulation method. Its basic idea is that when the problem to be solved is the probability of a certain random event occurring or the expected value of a certain random variable, through a certain experimental method, the probability of this random event is estimated by the frequency of the occurrence of this event, or the numerical characteristics of this random variable are obtained and used as the solution to the problem. This method adds randomly generated data to the numerical model to simulate the random influence during the drifting process. Subsequently, a large number of samples are taken, multiple simulations are carried out by adding different random values, and multiple corresponding drifting prediction trajectories are obtained, and then comprehensive evaluation and analysis are carried out. This method can consider various situations in the simulation, making the finally obtained predicted trajectory more accurate and reliable.

[0006] Secondly, some researchers have also proposed the Lagrangian particle tracking method. Its basic idea is to track the movement of fluid particles and integrate the movements of all particles to construct the entire hydrodynamics. Considering the characteristics of the research object (the lost ship or aircraft) and the environment, the constraint conditions and the solution process have been studied. Simply put, the Lagrangian particle tracking method is to simulate a large number of particles and estimate the drifting trajectory of real particles based on the simulated path. First, the flow field domain to be studied needs to be divided into discrete networks, that is, several sub-networks. Multiple particles can be set in each sub-network, and these particles are given initial positions, velocities, and other necessary attributes. Subsequently, according to the flow field data or the flow field motion equation, the motion of each particle is calculated to obtain the dynamic behavior of the entire flow field domain and help understand the flow field motion law. However, this method also has certain limitations. The attribute data such as the initial position of real particles, such as people falling into the water at sea, may be different from the attributes of the simulated particles, resulting in a poor prediction effect.

[0007] However, due to the complexity and uncertainty of the ocean environment, the traditional method is based on certain assumptions, and the predicted drifting position is not only related to the information at the previous moment. The traditional numerical simulation method cannot obtain a relatively accurate prediction path. In addition, the task of drift trajectory prediction is usually defined as predicting the trajectory within a future time period based on a series of historical drift data. Therefore, researchers often use statistical and machine learning-based methods to convert the trajectory prediction problem into an ordered time series prediction problem to improve the prediction accuracy. Summary of the Invention

[0008] The object of the present invention is to provide a method for predicting the drifting trajectory of a person falling into the sea, which solves some technical problems faced in the existing prediction of the drifting trajectory of a person falling into the sea. For example, there is no strong generalization ability for different position and meteorological data, and it is unable to capture the temporal relationship between data well, resulting in low prediction accuracy. The present invention directly uses the method of deep learning for experiments and comparisons. By improving the original Transformer, a novel and effective model method is provided, and relatively reliable prediction results are obtained in the field of search and rescue for people falling into the sea, which is expected to provide new tools and methods for the research in this field.

[0009] To achieve the above object, the present invention provides a method for predicting the drifting trajectory of a person falling into the sea, including the following steps:

[0010] S1. Collect data related to a person falling into the sea;

[0011] S2. Normalize the data;

[0012] S3. Make data Loaders for training and testing according to the processed data set;

[0013] S4. Iterate multiple times in a loop, feed the data in each batch into the AMWM-Transformer model, calculate the loss value by obtaining the output of the model and the original label data, and update the model parameters according to the loss value to complete the model training;

[0014] S5. According to the trained model, feed the input features of the test into it to obtain the output, calculate the cumulative sum of the loss value with the true label, obtain the total loss value of the final test data, and use ADE and FDE to evaluate the experimental results.

[0015] Preferably, in S1, the data related to a person falling into the sea includes the drifting trajectory information data of a certain simulation dummy in the offshore area, including the dummy id number, date, specific time, longitude, latitude, and the decomposed speed information of sea breeze and sea current. Extract the data except the id number, date, and specific time as the experimental data.

[0016] Preferably, in S2, the maximum-minimum normalization method is used to process the original data, and the formula is as follows:

[0017] ;

[0018] Among them, X represents the current data, X min represents the minimum value in the current data, X max represents the maximum value in the current data, Xnorm It represents the result after the current data is normalized.

[0019] Preferably, the specific process of S3 is as follows:

[0020] First, divide the training set and the test set according to the set ratio. Subsequently, obtain the input features and labels based on the set historical sliding window size and prediction window size, and name them X train , Y train , X test , Y test . Add the input features and labels of the training set to a new list in tuple form, and then split them by batch_size to form train_loader and test_loader.

[0021] Preferably, in S4, the loss value is calculated using MSE , and the formula is as follows:

[0022] ;

[0023] Among them, represents the data sample size, represents the label value, represents the model prediction value.

[0024] Preferably, in S4, a multi-head adaptive attention weight mechanism AMWM is added to the Transformer. By changing the corresponding weights through the differential activation degrees of the attention heads, the mechanism that only uses the same weight allocation in the original attention matrix is replaced, and the feed-forward network in the Transformer structure that only contains fully connected layers is replaced with an LSTM layer.

[0025] Preferably, in S4, the adaptive attention weight mechanism AMWM is based on the obtained attention matrix. In the part of obtaining the attention matrix of the original Transformer, the input data is a time-series data sequence, that is, the data stacked by the drift trajectory position data and meteorological data in the past period of time:

[0026] Assume its data shape is [batch_size, seq_len, features_num]. Among them, batch_size represents the number of samples in each batch, seq_len represents the length of the time-series data, and features_num represents the number of features of each data point;

[0027] Based on the input data, three vectors of the attention matrix are constructed: the query Q vector, the key K vector, and the value V vector. Using matrix multiplication and softmax operation, the result of Q and K is calculated to obtain the attention scores, and then weighted average calculation is performed with the V vector to obtain the attention matrix. The obtained result shape is [batch_size, seq_len, seq_len], representing the relationship between each time-series data point and other time-series data points in each sample data;

[0028] In the multi-head Transformer, the data shape of the attention matrix is [batch_size, n_heads, seq_len, seq_len], which represents the relationship matrix between each time-series data point and other data points in different attention head environments. Each attention head focuses on different context features to capture different dependencies in the data, where n_heads represents the number of attention heads in the multi-head Transformer.

[0029] Preferably, in S4, variance, standard deviation, and mean square error are used to obtain the attribute data in the attention matrix. The sum of different attributes is calculated and then the standard deviation is obtained to get the degree of dispersion of the overall data. Finally, a mean operation is performed on the attribute data to obtain a specific index for comprehensively evaluating the activity degree of the attention heads. The specific formula is as follows:

[0030] ;

[0031] Among them, is expressed as the th attention head, variance , std_deviation and mean_square_ error respectively represent the variance, standard deviation, and mean square error of the attention matrix corresponding to the attention head, represents the current analysis of the th batch;

[0032] After multiplying the result obtained according to the softmax function by the attention matrix, the final calculation result is obtained. After a full-batch data analysis is completed, the attention head weight list under this loop iteration is obtained. The next time, this weight list is used to multiply with the attention matrix to obtain the attention matrix under different influence degrees, and then the model parameters are updated.

[0033] Preferably, in S4, the LSTM neurons in the LSTM layer are used to add, maintain, or delete the information transmitted through the feedback loop. The gate consists of a sigmoid neural network layer, which selects to let part of the information pass according to the input value. This sigmoid layer outputs a value between 0 and 1 to indicate how much information is allowed to pass, and controls how much information is retained according to the new input and the previous iteration value. The specific formula is as follows:

[0034] ;

[0035] Among them, W and b respectively refer to the weights and bias terms of different gates. The subscript represents the current time t, represents the input at the current time t, usually a vector or matrix, represents the activation function.

[0036] Preferably, in S5, the difference unit on the longitude and latitude is converted from degrees to kilometers, and then evaluated through the ADE_custom and FDE_custom formulas as shown below:

[0037] ;

[0038] ;

[0039] Among them, represents the th sample, represents the subscript of the last data, represents the th value in the predicted data, represents the th value in the real data.

[0040] Therefore, the present invention adopts the above-mentioned method for predicting the drifting trajectory of a person falling into the sea, and the beneficial effects are as follows:

[0041] (1) Traditional drifting trajectories often only consider conventional influencing factors and mostly use numerical model simulation methods for prediction. This results in the inability of models based on stable influencing factors and deterministic laws to capture the non-linear characteristics of drifting behavior, posing a great challenge to providing accurate and reliable drifting trajectories and causing unnecessary losses. However, the drifting trajectory prediction model based on deep learning in the present invention can capture the influencing factors at different time periods and different positions.

[0042] (2) Under the basic framework of the deep learning model, the present invention first adds an LSTM recurrent network layer to the original model, which can better model the complex relationships between time-series data points and improve spatio-temporal continuity.

[0043] (3) The present invention adopts a new mechanism of adaptive multi-head attention weights, which assigns different numerical differential attention head weights to different influencing factors. The drift trajectory prediction after adding this mechanism helps to improve the efficiency of search and rescue operations, reduce the waste of human and material resources, and can respond quickly and take more precise measures to increase the survival probability of the drowning person.

[0044] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0045] Figure 1 is the overall flowchart of an embodiment of the method for predicting the drift trajectory of a person falling into the sea of the present invention;

[0046] Figure 2 is the flowchart of the adaptive multi-head attention weight update mechanism of an embodiment of the method for predicting the drift trajectory of a person falling into the sea of the present invention;

[0047] Figure 3 is the block diagram of the LSTM structure of an embodiment of the method for predicting the drift trajectory of a person falling into the sea of the present invention;

[0048] Figure 4 is the experimental result diagram of the single-step prediction multi-model comparison of the dummy No. 582401 of an embodiment of the method for predicting the drift trajectory of a person falling into the sea of the present invention. Specific Embodiments

[0049] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0050] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the technical field to which the present invention belongs.

[0051] Compared with the traditional numerical simulation method for trajectory prediction, the deep learning method can achieve a faster prediction speed, reduce the computational time complexity and time cost; in addition, it can learn the complex non-linear relationship between the input and output to improve its accuracy. Since the prediction of the drifting trajectory of a person falling into the water is itself a time series prediction, recursive neural networks (RNNs), long short-term memory networks (LSTMs), gated recurrent units (GRUs), Transformers and their variants can be used to predict the drifting trajectory. Similarly, using an RNN neural network for vehicle path prediction has good results and can learn the relationship between moments from the data, thus providing more accurate prediction results.

[0052] As Figure 1 shown, a method for predicting the drifting trajectory of a person falling into the water at sea includes the following steps:

[0053] S1. Collect data related to a person falling into the water at sea;

[0054] The dataset formed by the data related to a person falling into the water at sea contains the drifting trajectory information data of a certain simulation dummy in the offshore sea area, with a total of 2,880 pieces of data, including the dummy id number, date, specific time, longitude, latitude, and the decomposed speed information of sea breeze and sea current, a total of nine attributes. Extract the data except the id number, date, and specific time as the experimental data.

[0055] S2. Normalize the data;

[0056] Taking the sea breeze data as an example, obtain the maximum and minimum values in the sea breeze data at all time points, and use the maximum-minimum normalization method to process the original data. The formula is as follows:

[0057] ;

[0058] where X represents the current data, X min represents the minimum value in the current data, X max represents the maximum value in the current data, X norm represents the result of the current data after being standardized.

[0059] S3. Make training and test data Loaders according to the processed dataset. The specific process is as follows:

[0060] First, divide the training set and the test set according to the set ratio. In this embodiment, the training set and the test set are divided in a ratio of 8:2. Subsequently, obtain the input features and labels according to the set historical sliding window size and prediction window size. In this embodiment, a historical data window of 30 and a prediction data window of 1 are used to make a single-step prediction data set, which are respectively named X train , Y train , X test , Y test . Add the input features and labels of the training set to a new list in tuple form; then divide by batch_size to form train_loader and test_loader.

[0061] S4. Iterate multiple times in a loop, feed the data in each batch into the AMWM-Transformer model, obtain the output of the model and calculate the loss value with the original label data. The loss value is calculated using MSE . The formula is as follows:

[0062] ;

[0063] Among them, represents the data sample size, represents the label value, represents the model prediction value. Update the model parameters according to the loss value to complete the model training.

[0064] This embodiment mainly uses the original Transformer structure as the basic framework. First, add a multi-head adaptive attention weight mechanism AMWM to the Transformer. By changing the corresponding weights through the differential activation degrees of the attention heads, the mechanism that only uses the same weight distribution in the original attention matrix is replaced, avoiding the inability to fully consider the influence degree of different attention heads on its prediction results. Secondly, replace the feed-forward network containing only fully connected layers in the Transformer structure with an LSTM layer to capture the temporal relationship between data that cannot be obtained only by using fully connected layers and strengthen the spatio-temporal continuity of the prediction.

[0065] Regarding the adaptive attention weight mechanism AMWM, it is based on the obtained attention matrix. In the part of obtaining the attention matrix of the original Transformer, the input data is a time-series data sequence, that is, the data stacked by the drift trajectory position data and meteorological data in the past period of time:

[0066] Assume its data shape is [batch_size, seq_len, features_num], where batch_size represents the number of samples in each batch, seq_len represents the length of the time-series data, and features_num represents the number of features in each data point.

[0067] Based on this input data, three vectors for constructing the attention matrix are generated: the query Q vector, the key K vector, and the value V vector. Using matrix multiplication and the softmax operation, the result of calculating Q and K is used to obtain the attention scores. Finally, a weighted average calculation is performed with the V vector to obtain the attention matrix. The resulting shape is [batch_size, seq_len, seq_len], representing the relationship between each time-series data point and other time-series data points in each sample data.

[0068] In the multi-head Transformer, the data shape of the attention matrix is [batch_size, n_heads, seq_len, seq_len], representing the relationship matrix between each time-series data point and other data points in different attention head environments. Each attention head focuses on different context features to capture different dependencies in the data, where n_heads represents the number of attention heads in the multi-head Transformer.

[0069] However, since the relationship matrices obtained by each attention head are different, and at the same time, their influence on the final result also varies. After obtaining the multi-head attention matrix of all sample quantities in a batch, different data samples are randomly sampled to calculate the attribute data of the attention matrix under all attention head environments based on these sampled samples. Based on these data, the weights of the attention heads are determined.

[0070] Therefore, in this embodiment, variance, standard deviation, and mean square error are used to obtain the attribute data in the attention matrix. These three selected in this embodiment are common indicators for measuring data distribution and dispersion. These three attribute data are all used to measure the degree of dispersion or deviation. A larger value means a larger degree of dispersion of the element values in the attention matrix, indicating that the attention head may be paying attention to more positions and information; conversely, it means that the attention distribution is relatively concentrated, perhaps only focusing on the information of a few positions. The larger the value of these indicators, the stronger the activity degree corresponding to each attention head. In each iteration process, when the attributes of all data are analyzed, a set of three attributes of all attention heads is obtained.

[0071] Subsequently, the different attributes are added together and then the standard deviation is calculated to obtain the degree of dispersion of the overall data. Finally, a mean operation is performed on these three attribute data to obtain a specific index for comprehensively evaluating the activity degree of the attention heads. The specific formula is as follows:

[0072] ;

[0073] Among them, is denoted as the th attention head, variance, std_deviation and mean_square_ error respectively represent the variance, standard deviation, and mean squared error of the attention matrix corresponding to the attention head. represents the current analysis of the th batch. This index lays the foundation for subsequently obtaining the attention head weight list and thus updating the influence degree of each attention head on the final result.

[0074] Finally, after multiplying the result obtained from the softmax function by the attention matrix, the final calculation result is obtained. After a full-batch data analysis is completed, the attention head weight list under this loop iteration is obtained. The next time this weight list is used to multiply with the attention matrix, attention matrices under different influence degrees are obtained, and then the model parameters are updated.

[0075] The multi-attention head weight optimization algorithm aims to dynamically adjust the weights of each attention head by analyzing the performance of multi-attention heads in different batches (batches) to optimize the performance of the model.

[0076] Define n to represent the number of attention heads, all_batch_list is used to store the list of feature analysis results for all batches in each loop iteration, power is the attention head weight list at the current loop, var, std, and mse respectively represent variance (Variance), standard deviation (StandardDeviation), and mean squared error (MeanSquaredError), which are statistics used to measure the characteristics of the attention matrix, and selected_num is the number of random samplings.

[0077] As Figure 2 shown, the adaptive multi-head attention weight update mechanism includes the following steps:

[0078] 1. Initialize weights: Initialize the power list so that the initial weight of each attention head is 1, that is, Initialpower = [1] * n;

[0079] 2. Start loop iteration: Enter the outer loop, and this loop will continue until a specific termination condition is met (such as reaching the maximum number of iterations or the convergence criterion);

[0080] 3. Calculate the product of the attention matrix and the weights: For each iteration, calculate the result of multiplying the attention matrix by the current power list, i.e., attention ← attention matrix * Initialpower;

[0081] 4. Process the data of each batch: Start the inner loop and iterate through the attention data of the current batch : Extract the attention data of the current batch from attention, current_attention_data ← attention[i]; Perform feature analysis on the extracted attention data to obtain its variance var, standard deviation std, and mean squared error mse; Add these feature data [var, std, mse] to the batch_features list; Add the batch_features list to the all_batch_list for subsequent processing;

[0082] 5. End batch processing: After the inner loop ends, the feature analysis results of all batches have been collected in the all_batch_list;

[0083] 6. Random sampling and feature aggregation: Randomly sample the all_batch_list, and select using the number of samples specified by selected_num; For the sampled data, calculate the standard deviation for each of the three attributes var, std, and mse respectively, and then calculate the mean of these standard deviations to obtain the activity level of each attention head in the current iteration; Update the power list to be equal to the mean calculated above and perform normalization on it, i.e., Power ← mean(all_batch) & normalization, and perform reverse weighting on the weight list, i.e., use 1 minus the obtained weight list to update the model parameters;

[0084] 7. End the loop iteration: When the termination condition is met, end the outer loop;

[0085] 8. Output the result: Output the final optimized attention head weight list power;

[0086] Through the above steps, this algorithm can dynamically adjust the weights of each attention head, thereby optimizing the performance of the model on different batches. In the final test stage, there is no need to additionally add a weight list to increase the model complexity. Just input the test data directly to obtain the final prediction result.

[0087] In S4, regarding the LSTM structure layer, the original purpose of this recurrent neural network structure layer was to solve the long-term dependencies of the typical recursive RNN model with gradient vanishing or explosion. The specific structure diagram is shown inFigure 3 。

[0088] The LSTM neurons in the LSTM layer can be used to add, maintain, or delete information transmitted through the feedback loop. The gate consists of a sigmoid neural network layer, which can select to let part of the information pass according to the input value. This sigmoid layer outputs a value between 0 and 1 to indicate how much information it can allow to pass, so as to control how much information is retained according to the new input and the previous iteration value. The specific formula is as follows:

[0089] ;

[0090] Among them, W and b respectively refer to the weights and bias terms of different gates. The subscript represents the current time step t, represents the input at the current time step t, usually a vector or matrix, represents the activation function.

[0091] S5. According to the trained model, feed the input features of the test into it to obtain the output, calculate the cumulative sum of the loss value with the true label, obtain the total loss value of the final test data (the loss value calculation formula is the same as above), and visualize its effect. Use ADE_custom and FDE_custom to evaluate the experimental results.

[0092] In terms of evaluation metrics, first convert the unit of the difference in longitude and latitude from degrees to kilometers, and then evaluate through the ADE_custom and FDE_custom formulas shown below, named ADE_custom and FDE_custom:

[0093] ;

[0094] ;

[0095] Among them, represents the th sample, represents the subscript of the last data, represents the th value in the predicted data, represents the th value in the true data.

[0096] Finally, as Figure 4 shown, conduct a comparative experiment visualization and experimental result evaluation with the currently more common time series prediction models. The detailed results are shown in Table 1.

[0097] Table 1 Evaluation of Drifting Trajectories of Each Model, Dummy Number 582401, Unit: km

[0098] ;

[0099] It can be found that the ADE and FDE errors of the AMWM-Transformer proposed in the present invention are 1.2645 km and 0.7931 km respectively. Compared with the best comparison model mentioned, the results are reduced by 0.7854 km and 0.7788 km respectively.

[0100] In summary, on the basis of the original Transformer model, the present invention adds an adaptive multi-head attention weight mechanism to calculate different weights for the attention matrices corresponding to each attention head to determine their influence on the final result. The core idea of this mechanism is to change the weights of each attention head, setting large weights for those with greater influence on the result; otherwise, setting small weights to increase the influence of positive elements and decrease the influence of negative elements. Additionally, in order not to increase the model complexity in the final test stage, reverse weighting is performed to better update the model parameters.

[0101] Meanwhile, in the multi-head attention mechanism, it is carried out through multiple parallel attention heads, and each head can focus on different local context information. However, in the added adaptive multi-head attention weight mechanism, each attention head learns a specific feature weight pattern, and the final result is obtained by weighted summation of different features, thereby improving the prediction accuracy of the drifting trajectories of maritime drowning persons.

[0102] The present invention modifies the feed-forward fully connected neural network layer in the original Transformer into an LSTM structure layer: First, since the LSTM structure is a recurrent neural network specifically used to process time-series data, compared with the fully connected layer, LSTM can capture the long-term dependence relationship of time-series data. And because the task of predicting the drifting trajectories of maritime drowning persons often has a certain temporal nature (the current drifting position is related to the historical drifting position), using this structure can better perform modeling. Additionally, since the past drifting position and meteorological data information are very important when predicting the current drifting position, the LSTM network is needed to retain and utilize the previous data information, thereby better predicting the current position and improving the accuracy and stability of the prediction.

[0103] Therefore, the present invention adopts the above-mentioned method for predicting the drifting trajectories of maritime drowning persons, conducts experiments and comparisons using deep learning methods, and provides a novel and effective model method by improving the original Transformer, obtaining relatively reliable prediction results in the field of search and rescue for drowning persons, and is expected to provide new tools and methods for the research in this field.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for predicting the drifting trajectory of a person falling into the water at sea, characterized in that: The following steps are involved: S1. Collect data related to people falling into the water at sea; S2, normalize the data; S3, create data Loader for training and testing based on the processed data set; S4, iterate the loop multiple times, feed the data in each batch into the AMWM-Transformer model, obtain the model output and the original label data to calculate the loss value, update the model parameters according to the loss value, and complete the model training; S5. According to the trained model, the input features of the test are fed into the output, and the loss value is accumulated with the real label to obtain the total loss value of the final test data. ADE and FDE are used to evaluate the experimental results. In S4, the AMWM-Transformer model adds a multi-head adaptive attention weight mechanism AMWM to the Transformer; The input data is a time series data sequence formed by stacking the drift trajectory position data and meteorological data over a period of time; The multi-head adaptive attention weight mechanism AMWM includes the following steps: Initialize weights: Initialize the weight list so that the initial weight of each attention head is 1; Start loop iteration: enter the outer loop; Calculate the product of the attention matrix and the weights: For each iteration, calculate the result of multiplying the attention matrix by the current weight list; Process the data of each batch: start the inner loop, traverse the attention matrix of the current batch i, extract the attention matrix of the current batch; obtain the variance var, standard deviation std and mean square error mse of the extracted attention matrix; add [var, std, mse] to the batch_features list; add the batch_features list to the all_batch_list; End batch processing: After the inner loop ends, the feature analysis results of all batches are collected into all_batch_list; Random sampling and feature aggregation: Randomly sample all_batch_list, calculate the standard deviation of the three attributes of var, std and mse for the sampled data, and then average these standard deviations to obtain the activity level of each attention head in the current iteration; update the weight list to make it equal to the mean calculated above, and normalize it, and reverse weight the weight list, specifically using 1 to subtract the normalized result; update the model parameters; End loop iteration: When the termination condition is met, end the outer loop; Output result: Output the final optimized attention head weight list.

2. A method for predicting the drifting trajectory of a person falling into the water at sea according to claim 1, characterized in that: In S1, the data related to people falling into the sea include the drifting trajectory information data of a simulated dummy in the offshore waters, including the dummy ID number, date, specific time, longitude, latitude, and decomposition speed information of sea breeze and current. Data other than the ID number, date and specific time are extracted as experimental data.

3. A method for predicting the drifting trajectory of a person falling into the water at sea according to claim 2, characterized in that: In S2, the maximum-minimum normalization method is used to process the original data. The formula is as follows: Among them, X represents the current data, X min Indicates the minimum value in the current data, X max Indicates the maximum value in the current data, X norm Indicates the result after the current data is standardized.

4. A method for predicting the drifting trajectory of a person falling into the water at sea according to claim 3, characterized in that: In S4, the loss value is calculated using MSE, and the formula is as follows: Among them, n represents the data sample size, y i Indicates the label value, y pred Represents the model prediction value.

5. A method for predicting the drifting trajectory of a person falling into the water at sea according to claim 4, characterized in that: In S4, the LSTM neurons in the LSTM layer are used to add, maintain, or delete information transmitted through the feedback loop. The gate consists of a sigmoid neural network layer that selects to let part of the information pass according to the input value. This sigmoid layer outputs a value between 0 and 1 to indicate how much information is allowed to pass, and controls how much information is retained according to the new input and the previous iteration value.

Citation Information

Patent Citations

  • Short-term photovoltaic power generation power prediction method based on L-Transform

    CN117034055A

  • DBASformer model-based AUV drift trajectory prediction method

    CN118735030A