Transform model tide level prediction method and system based on multi-source heterogeneous data fusion
By using a multimodal fusion model based on the Transformer architecture to integrate tide level, meteorological, and remote sensing data for tide level prediction, the problem of insufficient integration of multi-source data in existing technologies is solved, high-precision multi-step tide level prediction is achieved, and the prediction accuracy under extreme weather conditions is improved.
Patent Information
- Application Number
- CN202511795300.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-02-24
AI Technical Summary
Existing tide level prediction technologies cannot effectively integrate multi-source heterogeneous data, especially under extreme weather events, resulting in low prediction accuracy and an inability to respond to nonlinear interference factors, leading to insufficient accuracy and generalization ability in tide level prediction.
A multi-modal fusion model based on the Transformer architecture is adopted. Through multi-source observation data preprocessing, multi-modal coding sub-network and cross-modal fusion mechanism, a multi-step time series prediction model is constructed to predict tide level by fusing tide level, meteorological and remote sensing data.
It enables high-precision multi-step tide level prediction under extreme weather conditions, improving the accuracy and timeliness of tide level prediction and adapting to complex nonlinear ocean processes.
Smart Images

Figure CN121562918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coastal engineering, and in particular to a method and system for predicting tide levels using a transformer model based on the fusion of multi-source heterogeneous data. Background Technology
[0002] Ocean tide levels are a crucial parameter reflecting Earth's ocean dynamics, playing a vital role in national maritime security, port logistics operations, coastal zone management, disaster prevention and mitigation, and climate change monitoring. With the intensifying trend of global warming, extreme weather events (such as strong typhoons and torrential rains) are becoming increasingly frequent, triggering storm surges and floods that cause significant casualties and economic losses. Therefore, improving the accuracy and timeliness of tide level forecasts has become a key requirement for emergency management in coastal cities and for coastal engineering design.
[0003] Existing tide prediction technologies have significant limitations in terms of modeling capabilities and data utilization: astronomical harmonic methods cannot respond to non-astronomical disturbances such as wind pressure; statistical regression models are limited by linear assumptions and struggle to adapt to complex nonlinear ocean processes; and existing deep learning methods are mostly single-modal modeling, failing to effectively integrate multi-source heterogeneous data such as tide levels, meteorological data, and remote sensing data, resulting in low prediction accuracy and poor generalization ability in the face of sudden events such as extreme weather. These existing tide prediction technologies cannot effectively consider the nonlinear interference of meteorological elements such as wind, air pressure, and precipitation, and their prediction performance deteriorates significantly, especially during sudden events such as storm surges, making it impossible to ensure the accuracy of tide predictions under extreme weather conditions. Summary of the Invention
[0004] This invention provides a method and system for tide prediction based on a transformer model using multi-source heterogeneous data fusion, to address the problems existing in related technologies. The technical solution is as follows: In a first aspect, embodiments of the present invention provide a tide prediction method based on a transformer model using multi-source heterogeneous data fusion, comprising: Acquire multi-source observation data and preprocess the multi-source observation data to obtain model samples; the multi-source observation data includes tide level observation data, meteorological station data and satellite remote sensing image data; A multimodal fusion model is constructed based on the Transformer architecture. The multimodal fusion model is trained by combining model samples and a multi-step time series prediction loss function to obtain a tide level prediction model. The multimodal fusion model includes multiple modality coding sub-networks and a cross-modal fusion mechanism. Based on the trained tide prediction model, the latest multi-source observation data acquired in real time are used to predict the tide level and generate future multi-time step tide prediction results.
[0005] In one implementation, preprocessing multi-source observation data to obtain model samples includes: Multi-source observation data are preprocessed to obtain standardized data; Based on the preset historical sequence length and future prediction steps, a single-step sliding window strategy is used to segment standardized data into time series, constructing continuous and partially overlapping training sample sequences to obtain corresponding model samples.
[0006] In one implementation, preprocessing of multi-source observation data includes time alignment, missing value imputation, normalization, and image grayscale processing.
[0007] In one implementation, the multimodal fusion module includes: Multiple modal coding subnetworks are used to encode the features of tide observation data, meteorological station data and satellite remote sensing image data respectively, and output tide feature sequences, meteorological feature sequences and remote sensing image feature sequences; The cross-modal attention fusion module connects the outputs of each modality coding sub-network and is used to fuse tidal feature sequences, meteorological feature sequences, and remote sensing image feature sequences based on the cross-modal fusion mechanism to generate multimodal fusion features. The time-series decoding module is used to receive multimodal fusion features, perform time-series decoding on the multimodal fusion features, and generate a future multi-time-step tide prediction sequence.
[0008] In one implementation, a multimodal fusion feature is generated by fusing tidal level feature sequences, meteorological feature sequences, and remote sensing image feature sequences based on a cross-modal fusion mechanism, including: Meteorological feature sequences and remote sensing image feature sequences are spliced together to obtain meteorological and satellite image features; Based on the attention mechanism, the tidal feature sequence is mapped into a query vector through linear mapping, and meteorological and satellite image features are mapped into key vectors and value vectors; Attention weights are obtained by calculating the similarity between the query vector and the key vector. The value vectors are then weighted and summed based on the attention weights to output the fused multimodal fusion features.
[0009] In one implementation, the multi-step time series prediction loss function is:
[0010] in, The number of samples in the model; Predict the number of steps for the future; Let be the predicted value of the i-th sample at step t; Let be the actual value of the i-th sample at step t.
[0011] In one implementation, it further includes: The model outputs a multi-time step tide prediction sequence based on the tide level prediction model, and obtains the actual tide level observation values provided by the tide gauge at the same time. The prediction error index is calculated based on the predicted tide level sequence and the actual tide level observation. The prediction error index includes the mean absolute error, root mean square error and mean absolute percentage error. Each prediction error index value is compared with a preset accuracy threshold range, and the model accuracy evaluation result is output.
[0012] Secondly, embodiments of the present invention provide a tide prediction system based on a transformer model using multi-source heterogeneous data fusion, comprising: The data acquisition module is used to acquire multi-source observation data, including tide level observation data, meteorological station data, and satellite remote sensing image data. The data preprocessing module is used to preprocess multi-source observation data to obtain model samples; The model training module is used to train a pre-built multimodal fusion model by combining model samples and a multi-step time series prediction loss function to obtain a tide prediction model; the multimodal fusion model includes multiple modality coding sub-networks and a cross-modal fusion mechanism; The model prediction module is used to predict the latest multi-source observation data acquired in real time using a trained tide prediction model, and generate tide prediction results for future multiple time steps.
[0013] Thirdly, embodiments of the present invention provide an electronic device comprising a memory and a processor. The memory and the processor communicate with each other via an internal connection path. The memory stores instructions, and the processor executes the instructions stored in the memory. When the processor executes the instructions stored in the memory, it causes the processor to perform the method described in any of the above embodiments.
[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium that stores a computer program, wherein when the computer program is run on a computer, the methods in any of the embodiments described above are executed.
[0015] The advantages or beneficial effects of the above technical solutions include at least the following: This invention obtains a multimodal fusion model through collaborative modeling using multimodal coding and attention mechanisms. The multimodal fusion model is then used to fuse tide level observation data, meteorological station data, and satellite remote sensing image data to achieve high-precision, multi-step prediction of future tide levels, thereby improving the accuracy of tide level prediction under extreme weather conditions.
[0016] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0017] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in the invention and should not be construed as limiting the scope of the invention.
[0018] Figure 1 This is a flowchart illustrating the tide prediction method based on the transformer model using multi-source heterogeneous data fusion according to the present invention. Figure 2 This is a schematic diagram of the functional structure of the transformer model tide prediction system based on multi-source heterogeneous data fusion according to the present invention. Figure 3 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0019] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0020] Currently, existing tidal level prediction techniques can be broadly categorized into three types: astronomical harmonic methods, statistical regression models, and deep learning models, as detailed below: (1) Astronomical harmonic analysis method This type of method is based on the periodicity of tidal forces and uses the Fourier decomposition principle to represent historical tide level sequences as a linear combination of harmonic components, such as the classic Doodson decomposition method. This method is suitable for modeling long-term stable tide level fields, but it cannot be dynamically adjusted under strong disturbance conditions (such as storm surges and rainfall peaks), resulting in large prediction errors.
[0021] (2) Statistical regression method These methods include linear models such as ARIMA (Autoregressive Integral Moving Average) and SARIMA (Seasonal ARIMA), which mainly utilize the autocorrelation of time series data to predict future values. Although these methods have good linear modeling capabilities, they are limited to historical tidal data, cannot incorporate external factors, and have insufficient predictive ability under complex sea conditions with strong coupling of multiple factors.
[0022] (3) Single-modal deep learning method Deep learning architectures such as LSTM, GRU, CNN-LSTM, and Transformer are increasingly being used for tide prediction to model tide sequence data. While these methods offer significant advantages in modeling nonlinear dynamics, they generally suffer from the following problems: they use only univariate tide or wind speed sequences, ignoring meteorological and image information; they lack multimodal collaborative modeling capabilities and cannot handle spatial information such as remote sensing imagery; and while the models have strong fitting ability for specific regions, their generalization ability is weak, making them prone to overfitting.
[0023] The fundamental reason for the shortcomings of existing tide prediction technologies lies in the lack of a unified feature fusion mechanism and multimodal information modeling structure, making it impossible to capture the overall tide evolution pattern driven by multiple factors. Therefore, to address this issue, this embodiment provides a tide prediction method based on a transformer model using multi-source heterogeneous data fusion. This method constructs a unified model that effectively integrates multi-source data and enhances nonlinear modeling capabilities, enabling high-precision multi-step prediction of future tide levels, thereby overcoming the limitations of existing technologies.
[0024] refer to Figure 1 As shown in the figure, the tide level prediction method based on multi-source heterogeneous data fusion using a transformer model specifically includes the following steps: Step S1: Acquire multi-source observation data and preprocess the multi-source observation data to obtain model samples; wherein, the multi-source observation data includes tide level observation data, meteorological station data and satellite remote sensing image data.
[0025] It should be noted that tide level observation data refers to historical sampling intervals obtained from coastal tide level stations. The tide level height is denoted as ,in Indicates time The tide level. Weather station data refers to time-series data such as wind speed, air pressure, and precipitation obtained from weather stations, denoted as... ,in ( (Number of meteorological elements), sampling frequency aligned with tide level. Received satellite remote sensing images, recorded as satellite remote sensing image data. ,in Indicates time Satellite images are used to extract spatial features.
[0026] To achieve unified input and efficient training of multi-source heterogeneous data, standardized preprocessing is required for tidal observation data, meteorological station data, and satellite remote sensing image data. This embodiment's preprocessing process includes five sub-steps: time alignment, missing value imputation, normalization, image grayscale processing, and sliding window construction, as detailed below: The sampling frequencies of multi-source observation data are different. To ensure the synchronization of time series, a unified time is used. Interpolate all data from the multi-source observation data to Time granularity: ; in, The unified target time series set; This is the start timestamp; For time step; This represents the i-th standardized time point; The number of steps in the time series depends on the overall duration of the data and the sampling frequency.
[0027] For all low-frequency data (such as weather station data and remote sensing image data) with relatively low frequency of change on the time scale, since the time frequency of low-frequency data differs from other data, linear interpolation is used to upsample the low-frequency data to the high-frequency time axis and align it with high-frequency observation data such as tide level observation data and weather station data to complete subsequent feature fusion. The linear interpolation expression is as follows:
[0028] in, It is the value of the target independent variable, that is, the time point or spatial location where interpolation is to be performed at a certain moment; It is the value of the dependent variable obtained by interpolation; and They are the closest The time points before and after; and They are the observation times that are closest The data points before and after.
[0029] If linear interpolation cannot be completed (e.g., due to missing endpoints), then mean filler method is used:
[0030] in, For at a certain point in time The variable value; For time points The actual observed value at the location; This represents the number of samples within the local time window, i.e., the number of neighborhood points involved in the mean calculation. For time points Symmetrical A set of neighboring time points, for example: hour, for .
[0031] It should be noted that outliers exceeding 3 standard deviations are removed, and then the remaining outliers are added according to the above rules.
[0032] To avoid interference from differences in the dimensions of different physical quantities during model training, a uniform min-max normalization method is adopted:
[0033] in, These are the normalized sample values; , Let be the minimum and maximum values of the variable in the training set samples, respectively.
[0034] It should be noted that the normalization operation in this embodiment is performed independently on each variable in the tide level observation data and the meteorological station data.
[0035] Satellite remote sensing image data, on the other hand, needs to be standardized in size and number of channels, and undergo grayscale normalization.
[0036] in, The pixel value (usually an integer between 0 and 255) in the i-th row and j-th column of the original image. These are the normalized pixel values; This represents the coordinates of all pixels in the image. If the image is missing, it can be filled using the image from the previous observation (last-observation carry-forward).
[0037] After the above time alignment, missing value imputation and normalization processes, and image preprocessing, the resulting data is called standardized data. Before using standardized data as model samples, data segmentation is required, as follows: Construct a sliding window and use it to extract continuous input sequences from standardized data as sample inputs, specifically as follows: Input the length of the historical sequence, and set the historical sequence length as the sample length. , that is, the former The data at each time point constitutes a training sample:
[0038] The number of future prediction steps is set to That is, the model objective is the future continuous Tidal range:
[0039] in, For the first The input vector at any given time contains standardized data such as tide level, wind speed, air pressure, and precipitation. For length is Multidimensional time series, The feature dimension; For the first Output vector at time step; For predicting goals, for the future The tidal level value of the step.
[0040] Based on the preset historical sequence length and future prediction steps, a single-step sliding window strategy is used to segment the standardized data into time series. Each time, one time step is slid to construct a continuous and partially overlapping training sample sequence, thereby obtaining the corresponding model samples. The model samples are then divided into training set and prediction set to meet the model training requirements.
[0041] Step S2: Construct a multimodal fusion model based on the Transformer architecture, and train the multimodal fusion model by combining model samples and multi-step time series prediction loss function to obtain the tide level prediction model.
[0042] This embodiment uses a Transformer structure for multi-source data fusion to model the tide level, resulting in a multimodal fusion model. The model input includes three types of data: tide gauge sequences, meteorological sequences, and satellite images. The multimodal fusion model as a whole includes multiple modality coding sub-networks and a cross-modality fusion mechanism. Finally, the model is input into a unified prediction head to output the future tide level.
[0043] In this embodiment, the modal coding subnetwork includes a tide level sequence encoder, a meteorological sequence encoder, and a satellite image feature extractor.
[0044] The tide level sequence encoder is used to encode the features of tide level observation data and output a tide level feature sequence. Specifically: The historical tide levels recorded by the tide gauge are a continuous one-dimensional time series:
[0045] in, The input data represents the tidal level sequence, containing One time step; This represents the total length of the time series (time window size).
[0046] The tide level sequence encoder uses an LSTM network to extract tide level features for each time step. LSTM will update the hidden state. and memory unit This includes the following processes: a. The forget gate determines whether to retain the memory information from the previous moment based on the current input and the previous hidden state; b. Input gates control the degree of influence of the current input information on the memory unit; c. Candidate memory units generate potential memory update content at the current time step; d. Memory state updates are completed through the combined action of the forgetting gate and the input gate; e. The output gate combines the current memory state to generate the hidden state at that time step as a tidal feature representation.
[0047]
[0048] in, Forgotten Gate; For input gates; For output gate; Candidate memories; This is the current memory unit; The hidden state of the LSTM at each time step; Use the Sigmoid activation function; For Hadamard (element-by-element) multiplication; This is the gated weight matrix; This is the gating bias term.
[0049] The tide level sequence encoder processes the data through an LSTM network, and the final output is the hidden state sequence features:
[0050] in, The encoded tide position feature sequence is the output of the entire LSTM encoder; These are the tidal feature vectors at different time steps; This indicates that the final output is a shape of The feature matrix is the hidden state vector that is preserved at each time step.
[0051] The meteorological sequence encoder encodes features from meteorological station data, outputting a meteorological feature sequence. Meteorological station data is a multivariate time series.
[0052] in, This represents a multivariate time series input of weather station data.
[0053] After adding location encoding, the weather station data is input into the weather sequence encoder. In this embodiment, the weather sequence encoder is a Transformer encoder.
[0054] in, The sequence of meteorological features encoded at each time step; It is a multi-layered self-attention structure used to model the interactions between meteorological elements and between time steps; This is used for position encoding, which provides timing information to the Transformer.
[0055] The satellite image feature extractor encodes features from satellite remote sensing image data, outputting a sequence of remote sensing image features. Satellite remote sensing image data is defined as follows:
[0056] in, Indicates the duration as A sequence of satellite remote sensing images; Image height; Image width; This represents the number of image channels.
[0057] Image features from satellite remote sensing data are extracted using CNN:
[0058] in, For the first Spatial features of a frame image; It is a convolutional neural network used to extract features from an input image.
[0059] After image spatial features are stitched together, they are further modeled using a Transformer to obtain a remote sensing image feature sequence. : .
[0060] To effectively integrate multi-source information from tidal level modalities and meteorological and remote sensing image modalities, this embodiment provides a cross-modal fusion mechanism, also known as a cross-attention feature fusion method. This method can model the semantic alignment relationships between different modalities, thereby improving the model's ability to understand complex tidal level evolution patterns.
[0061] Cross-modal fusion mechanisms fuse the outputs of various modality coding sub-networks to generate multimodal fusion features. Specifically: The tidal level feature sequence, meteorological feature sequence, and remote sensing image feature sequence are used as inputs, with the tidal level feature sequence as the main modal feature and the meteorological feature sequence and remote sensing image feature sequence as auxiliary modal features.
[0062] The auxiliary modal features are concatenated, that is, the meteorological feature sequence and the remote sensing image feature sequence are concatenated to obtain a context information matrix, which contains meteorological and satellite image features:
[0063] In the attention mechanism, input features are mapped to three sets of vectors—query, key, and value—through a trainable linear mapping, which are used to calculate the correlation between modalities and the fusion result. In this embodiment, based on the attention mechanism, the tidal feature sequence is mapped to a query vector, and meteorological and satellite image features are mapped to key and value vectors through linear mapping.
[0064]
[0065]
[0066] Where Q (Query) represents what information the current tide level feature wants to focus on; K (Key) represents what information each position in the auxiliary modality contains; and V (Value) is the semantic representation actually passed to the fusion result. This is a learnable weight matrix.
[0067] Subsequently, attention weights are obtained by calculating the similarity between the query vector and the key vector. Based on these attention weights, the value vectors are weighted and summed to output the fused multimodal fusion features. Specifically: The attention weight matrix A is calculated as follows:
[0068] in: Represents the similarity scoring matrix; This is a scaling factor to prevent gradient explosion; This ensures that the sum of the weights of each row is 1.
[0069] The fusion weight is determined by the similarity between the query and the key, and the calculation formula is as follows:
[0070] The fused output is calculated using a weighted vector:
[0071] Weight This indicates that the intensity of the auxiliary modality at the j-th position should be considered at the tidal level at the i-th time step. The higher the weight, the more important the auxiliary feature is in the prediction. This mechanism realizes dynamic matching and semantic alignment of information between modalities.
[0072] This embodiment ultimately yields the fused feature sequence:
[0073] in, The fused multimodal features; To indicate with For query Attention fusion is performed on features from other modalities.
[0074] Ultimately, it is achieved through the decoder of the Transformer architecture. Output the future Tide level prediction:
[0075] in, The output of the model is the tide level prediction sequence, with dimension . ; This is the decoder module, which receives the fused features and outputs the prediction results.
[0076] After the above construction is completed, the tide level prediction model is trained using the model samples obtained in step S1, and the actual observed tide levels are used during the training process. To monitor the signal, a multi-step regression loss function is used to train the entire network of the tide prediction model end-to-end.
[0077] In this embodiment, the multi-step time series prediction loss function is:
[0078] in, The number of samples in the model; Predict the number of steps for the future; Let be the predicted value of the i-th sample at step t; Let be the actual value of the i-th sample at step t.
[0079] It should be noted that the predicted value in the loss function refers to the value output by the tide level prediction model, while the actual value is the real tide level observation value collected by the tide level station.
[0080] Model samples provide the foundational material for the multimodal fusion model's learning, the supervision signal defines the learning objective, and the loss function quantifies the gap between predicted and true values to provide optimization direction. These three elements work together to drive the model to learn how to effectively fuse multimodal information to approximate the complex mapping relationships of the real world. In this embodiment, the tide prediction model is trained using the Adam optimizer, dynamically adjusting the learning rate, and incorporating an early stopping mechanism to prevent overfitting.
[0081] During or after model training, the prediction accuracy of the multimodal fusion model is evaluated using metrics such as mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE). This method is compared with traditional ARIMA and single LSTM models to verify the performance advantages of the fusion model. Details are as follows: The model outputs a multi-time step tide prediction sequence based on the tide level prediction model, and obtains the actual tide level observation values provided by the tide gauge at the same time. Based on the predicted tide level sequence and the actual tide level observation, multiple prediction error index values are calculated, including mean absolute error, root mean square error and mean absolute percentage error. Each prediction error index value is compared with a preset accuracy threshold range, and the model accuracy evaluation result is output.
[0082] This involves calculating multiple prediction error metrics, including but not limited to: Mean Absolute Error (MAE):
[0083] Root Mean Square Error (RMSE):
[0084] Mean Absolute Percentage Error (MAPE):
[0085] in, This represents the actual tide level observation at the i-th time point; These are the model predictions for the corresponding time points; This represents the number of samples.
[0086] The calculated MAE, RMSE, and MAPE values are compared with preset accuracy threshold ranges, or with the error indices of at least one benchmark model on the same test set. The benchmark model includes a historical mean model, an ARIMA model, or a single-modal neural network model. If all error indices of the multimodal fusion model are lower than the preset thresholds or significantly better than the benchmark model, the multimodal fusion model is determined to have high prediction accuracy and can be used for practical deployment. Otherwise, the model is returned to the model optimization stage for structural adjustment or retraining.
[0087] Step S3: Based on the trained tide prediction model, predict the latest multi-source observation data acquired in real time to generate tide prediction results for future multi-time steps.
[0088] Obtain the latest multi-source observation data, including: (1) Tide level sequence:
[0089] (2) Meteorological sequence:
[0090] (3) Satellite image sequence: ; Predicting future tide levels using a trained tide prediction model Tide position:
[0091] in, These are the model's predicted values; This is the model mapping function; it inputs the latest multi-source observation data into the tide level prediction model to predict tide levels for multiple future time steps and outputs the tide level prediction results. These tide level prediction results can be used in scenarios such as flood control scheduling, port operation management, and early warning of construction near the coast, improving the intelligence level of the forecasting system.
[0092] In another embodiment, a tide prediction system based on a transformer model using multi-source heterogeneous data fusion is also provided. This system executes the tide prediction method based on a transformer model using multi-source heterogeneous data fusion as described above. Figure 2 As shown, the system includes a data acquisition module, a data preprocessing module, a model training module, and a model prediction module. Details are as follows: The data acquisition module is connected to the tide gauge, weather station and satellite signals to acquire multi-source observation data, including tide observation data, weather station data and satellite remote sensing image data. The data preprocessing module is used to preprocess multi-source observation data to obtain model samples; The model building module is used to build multimodal fusion models; The model training module is used to train a pre-built multimodal fusion model by combining model samples and a multi-step time series prediction loss function to obtain a tide prediction model. The model evaluation module is used to evaluate the prediction accuracy of the multimodal fusion model using metrics such as mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE). The model prediction module is used to predict the latest multi-source observation data acquired in real time using a trained tide prediction model, and generate tide prediction results for future multiple time steps.
[0093] The multimodal fusion model in this embodiment includes: Multiple modal coding subnetworks are used to encode the features of tide observation data, meteorological station data and satellite remote sensing image data respectively, and output tide feature sequences, meteorological feature sequences and remote sensing image feature sequences; The cross-modal attention fusion module connects the outputs of each modality coding sub-network and is used to fuse tidal feature sequences, meteorological feature sequences, and remote sensing image feature sequences based on the cross-modal fusion mechanism to generate multimodal fusion features. The time-series decoding module is used to receive multimodal fusion features, perform time-series decoding on the multimodal fusion features, and generate a future multi-time-step tide prediction sequence.
[0094] It should be noted that the functions of each module in this embodiment system can be found in the corresponding descriptions in the above methods, and will not be repeated here.
[0095] In another embodiment, an electronic device is also provided. Figure 3 A structural block diagram of an electronic device according to an embodiment of the present invention is shown. Figure 3 As shown, the electronic device includes a memory 100 and a processor 200. The memory 100 stores a computer program that can run on the processor 200. When the processor 200 executes the computer program, it implements the tide prediction method based on a transformer model using multi-source heterogeneous data fusion as described in the above embodiments. The number of memories 100 and processors 200 can be one or more.
[0096] The electronic device also includes: The communication interface 300 is used to communicate with external devices and perform data exchange and transmission.
[0097] If the memory 100, processor 200, and communication interface 300 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc.
[0098] Optionally, in a specific implementation, if the memory 100, processor 200, and communication interface 300 are integrated on a single chip, then the memory 100, processor 200, and communication interface 300 can communicate with each other through an internal interface.
[0099] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this invention.
[0100] This invention also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the method provided in this invention.
[0101] This invention also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in this invention.
[0102] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (RISC) machine (ARM) architecture.
[0103] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0104] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0105] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0106] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0107] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A tide prediction method based on a transformer model using multi-source heterogeneous data fusion, characterized in that, include: Acquire multi-source observation data and preprocess the multi-source observation data to obtain model samples; wherein, the multi-source observation data includes tide level observation data, meteorological station data and satellite remote sensing image data; A multimodal fusion model is constructed based on the Transformer architecture. The multimodal fusion model is trained by combining the model samples and a multi-step time series prediction loss function to obtain a tide level prediction model. The multimodal fusion model includes multiple modality coding sub-networks and a cross-modal fusion mechanism. Based on the trained tide prediction model, the latest multi-source observation data acquired in real time are used to predict the tide level and generate future multi-time step tide prediction results.
2. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 1, characterized in that, The process of preprocessing the multi-source observation data to obtain model samples includes: The multi-source observation data is preprocessed to obtain standardized data; Based on the preset historical sequence length and future prediction steps, the standardized data is segmented into time series using a single-step sliding window strategy to construct a continuous and partially overlapping training sample sequence, thereby obtaining the corresponding model samples.
3. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 2, characterized in that, The preprocessing of the multi-source observation data includes time alignment, missing value filling, normalization, and image grayscale processing.
4. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 1, characterized in that, The multimodal fusion module includes: Multiple modality coding subnetworks are used to perform feature coding on the tide level observation data, the meteorological station data and the satellite remote sensing image data respectively, and output tide level feature sequence, meteorological feature sequence and remote sensing image feature sequence; A cross-modal attention fusion module, connected to the outputs of each modality coding sub-network, is used to fuse the tide level feature sequence, the meteorological feature sequence, and the remote sensing image feature sequence based on the cross-modal fusion mechanism to generate multimodal fusion features; The time-series decoding module is used to receive the multimodal fusion features, perform time-series decoding on the multimodal fusion features, and generate a future multi-time-step tide prediction sequence.
5. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 4, characterized in that, Based on the cross-modal fusion mechanism, the tidal feature sequence, the meteorological feature sequence, and the remote sensing image feature sequence are fused to generate multimodal fused features, including: The meteorological feature sequence and the remote sensing image feature sequence are spliced together to obtain meteorological and satellite image features; Based on the attention mechanism, the tide level feature sequence is mapped into a query vector through linear mapping, and the meteorological and satellite image features are mapped into key vectors and value vectors; Attention weights are obtained by calculating the similarity between the query vector and the key vector. The value vectors are then weighted and summed based on the attention weights to output the fused multimodal fusion features.
6. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 1, characterized in that, The multi-step time series prediction loss function is: in, The number of samples in the model; Predict the number of steps for the future; Let be the predicted value of the i-th sample at step t; Let be the actual value of the i-th sample at step t.
7. The tide prediction method based on multi-source heterogeneous data fusion using a transformer model according to claim 1, characterized in that, Also includes: Based on the tide prediction model, a multi-time step tide prediction sequence is output, and the actual tide observation values provided by the tide gauge at the same time are obtained. The prediction error index value is calculated based on the tidal level prediction sequence and the actual tidal level observation value. The prediction error index value includes the mean absolute error, the root mean square error and the mean absolute percentage error. Each prediction error index value is compared with a preset accuracy threshold range, and the model accuracy evaluation result is output.
8. A tide prediction system based on a transformer model using multi-source heterogeneous data fusion, characterized in that, include: The data acquisition module is used to acquire multi-source observation data, including tide level observation data, meteorological station data, and satellite remote sensing image data. The data preprocessing module is used to preprocess the multi-source observation data to obtain model samples; The model training module is used to train the pre-constructed multimodal fusion model by combining the model samples and the multi-step time series prediction loss function to obtain the tide level prediction model; wherein, the multimodal fusion model includes multiple modality coding sub-networks and a cross-modal fusion mechanism; The model prediction module is used to predict the latest multi-source observation data acquired in real time using the trained tide prediction model, and generate tide prediction results for future multiple time steps.
9. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores instructions that are loaded and executed by the processor to implement the tide prediction method based on multi-source heterogeneous data fusion of the transformer model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the tide prediction method based on the transformer model of multi-source heterogeneous data fusion as described in any one of claims 1 to 7.