Ocean buoy missing data filling method and system based on pre-training language model
Through the ocean interpolation model based on the pre-trained language model, the problems of spatiotemporal relationships and nonlinear features in the processing of missing value of ocean buoy data are solved, and efficient filling and analysis of ocean buoy data is achieved, which improves the completeness and accuracy of the data.
Patent Information
- Application Number
- CN202510305393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
Existing methods for missing value processing of marine buoy data are difficult to effectively deal with complex spatiotemporal relationships and nonlinear features in marine data, especially the insufficient spatial dependence between multiple buoys or sensors, which affects the completeness and accuracy of data analysis.
Using the ocean interpolation model based on the pre-trained language model, the space-time feature extraction module and the spatiotemporal tokenization module, combined with partial freezing attention strategy and LoRA technology, an interpolation method specifically targets the spatiotemporal data of the ocean buoy is designed, and the powerful semantic understanding and cross-modal knowledge transfer capabilities of the pre-trained language model are used to fill missing values.
The effective filling of missing data of ocean in situ buoys is achieved, the data continuity and analysis accuracy are improved, and the feasibility and effectiveness of pre-trained language models are shown in the analysis of ocean spatiotemporal data, with high interpolation accuracy and robustness.
Smart Images

Figure CN120234543A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for filling missing data of ocean buoys based on a pre-trained language model. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] Ocean data has extremely important application values in many fields such as climate research, environmental monitoring, and ocean resource development. The ocean buoy monitoring system, as a core tool for obtaining in-situ observation data, plays a key role in research fields such as marine meteorology, physical oceanography, and marine ecology with its advantages of high frequency, real-time, and economy. The buoy can monitor parameters such as temperature, salinity, ocean current, and wave in the ocean in real time, providing valuable ocean data for researchers, and thus providing necessary theoretical basis and data support for ocean-related research such as climate change and ocean environmental change. However, during long-term use, due to the influence of various factors such as harsh ocean environment, equipment failure, or human error, the missing of ocean buoy data is inevitable. The missing values of ocean buoy data will hinder the continuity of real-time monitoring, and further affect the integrity of data analysis, the reliability of model construction, and the accuracy of ocean environmental monitoring results.
[0004] To address the problem of missing values in ocean buoy data, the simplest and most direct method is to delete all incomplete data and only analyze the complete data, but this approach will lead to serious data bias, especially when the missing rate is relatively large. Another common method is to use some interpolation techniques to fill the missing values through a certain estimation method. Traditional interpolation methods often rely on specific rules or statistical assumptions and are difficult to effectively handle the complex spatio-temporal relationships and non-linear characteristics in ocean data. Therefore, more and more machine learning and deep learning-based methods have been applied to ocean data processing, such as gated recurrent unit (GRU), long short-term memory network (LSTM), and Transformer, etc., but most of these methods focus on the field of ocean time series data prediction, and there is relatively little research on ocean missing value interpolation. And the existing methods for imputing missing values in ocean data pay more attention to the variable relationship and time series dependence between single buoys, and rarely consider the spatial dependence between multiple buoys or sensors.
[0005] Meanwhile, pre-trained language models (PLMs) have achieved remarkable success in the fields of natural language processing and computer vision. Their powerful semantic understanding and cross-modal knowledge transfer and reasoning capabilities have gained wide recognition. This progress has inspired researchers to combine PLMs with spatio-temporal data, especially in the field of traffic spatio-temporal data prediction, where significant results have been achieved. Recently, methods such as STG-LLM and TPLLM have used pre-trained large language models to predict traffic data, explored the effectiveness of pre-trained language models in understanding and predicting spatio-temporal data, and compared them with deep learning methods of non-language models. Compared with the methods of non-language models for processing spatio-temporal data, the language model-based methods have stronger generalization capabilities and can capture both the temporal and spatial characteristics of the data simultaneously, thus enabling a comprehensive understanding and processing of spatio-temporal data. Summary of the Invention
[0006] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for filling missing data of ocean buoys based on a pre-trained language model, applies the pre-trained language model (PLM) to the interpolation task of in-situ ocean buoy observation data, and designs an ocean interpolation model specifically for spatio-temporal data interpolation of ocean buoys.
[0007] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0008] In a first aspect, the present invention provides a method for filling missing data of ocean buoys based on a pre-trained language model, including:
[0009] Obtain the ocean buoy data to be filled and the graph structure of the buoy station;
[0010] Input the ocean buoy data and the graph structure into the ocean interpolation model for interpolation filling to obtain complete ocean buoy data;
[0011] The ocean interpolation model includes a spatio-temporal feature extraction module, a spatio-temporal tokenization module, a pre-trained language model, and an output layer connected in sequence. The spatio-temporal feature extraction module extracts temporal feature representations from the ocean buoy data and spatial feature representations from the graph structure. Input the temporal feature representations and spatial feature representations into the spatio-temporal tokenization module for feature extraction and conversion to obtain temporal tokens and spatial tokens and connect them to obtain a token sequence. Input the token sequence into the pre-trained language model for learning to obtain high-dimensional hidden layer representations, and the output layer receives the high-dimensional hidden layer representations and converts them into interpolation values to obtain complete ocean buoy data.
[0012] A further technical solution is that specifically extracting the temporal feature representation from the ocean buoy data is:
[0013] Extract the hourly feature vector and the monthly feature vector from the ocean buoy data respectively, and concatenate the two to obtain the time feature vector;
[0014] Concatenate the time feature vectors of all time steps to obtain the time feature representation.
[0015] A further technical solution is that the spatial feature representation extracted from the graph structure is specifically:
[0016] Establish a graph Laplacian matrix according to the graph adjacency matrix in the graph structure;
[0017] Select the largest multiple eigenvectors from the eigenvectors of the graph Laplacian matrix, and perform a linear transformation on the eigenvectors to obtain the spatial feature vector;
[0018] Concatenate the spatial feature vectors of all buoy stations to obtain the spatial feature representation.
[0019] A further technical solution is that the time Token is specifically:
[0020] Obtain the overall state according to the average value of the ocean buoy data, and calculate the first-order difference of the overall state to obtain the overall trend;
[0021] Use a multi-layer perceptron to concatenate the overall state and the time feature representation to obtain the state Token;
[0022] Use a multi-layer perceptron to concatenate the overall trend and the time feature representation to obtain the trend Token;
[0023] Connect and normalize the state Token and the trend Token to obtain the time Token.
[0024] A further technical solution is that the spatial Token is specifically:
[0025] Use a multi-layer perceptron to concatenate the time feature representation and the spatial feature representation to obtain the static essence Token;
[0026] Use a multi-layer perceptron to extract and transform the dynamic features from the historical observation data to obtain the dynamic change Token;
[0027] Use a multi-layer perceptron to extract the missing features from the mask matrix to obtain the missing pattern Token;
[0028] Connect the static essence Token, the dynamic change Token and the missing pattern Token to obtain the spatial Token.
[0029] In a further technical solution, the pre-trained language model introduces a partial freezing attention strategy, freezing the multi-head attention layer and the feed-forward network of the first F layers, and unfreezing and fine-tuning the multi-head attention layer of the last U layers.
[0030] In a further technical solution, the output layer uses a decoder to convert the high-dimensional hidden layer representation generated by the pre-trained language model into an interpolation value.
[0031] In a second aspect, the present invention provides an ocean buoy missing data filling system based on a pre-trained language model, including:
[0032] A data acquisition module, which is configured to: acquire the ocean buoy data to be filled and the graph structure of the buoy site;
[0033] A model filling module, which is configured to: input the ocean buoy data and the graph structure into an ocean interpolation model for interpolation filling to obtain complete ocean buoy data;
[0034] The ocean interpolation model includes a spatio-temporal feature extraction module, a spatio-temporal Tokenization module, a pre-trained language model, and an output layer connected in sequence; the spatio-temporal feature extraction module extracts a time feature representation from the ocean buoy data and a space feature representation from the graph structure; input the time feature representation and the space feature representation into the spatio-temporal Tokenization module for feature extraction and conversion to obtain a time Token and a space Token and connect them to obtain a Token sequence; input the Token sequence into the pre-trained language model for learning to obtain a high-dimensional hidden layer representation, and the output layer receives the high-dimensional hidden layer representation and converts it into an interpolation value to obtain complete ocean buoy data.
[0035] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the method for filling missing data of ocean buoys based on a pre-trained language model as described in the first aspect.
[0036] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for filling missing data of ocean buoys based on a pre-trained language model as described in the first aspect.
[0037] The above one or more technical solutions have the following beneficial effects:
[0038] The present invention proposes a method for filling missing values in spatio-temporal data among multiple ocean buoys using a pre-trained language model. By leveraging the powerful semantic understanding and cross-modal knowledge transfer and reasoning capabilities of the pre-trained language model, it understands and processes spatio-temporal data, and combines the PLM with ocean spatio-temporal data through two steps of Token transformation and fine-tuning to effectively fill the missing data of in-situ ocean buoys.
[0039] The present invention designs a pre-trained language model framework for filling missing values in ocean buoy spatio-temporal data, namely the ocean interpolation model. Through the design of the spatio-temporal feature extraction module and the spatio-temporal Tokenization module, it effectively extracts the time dependence of ocean time series data and the spatial correlation among multiple buoys, thereby helping the model better understand ocean data.
[0040] The present invention successfully introduces the PFA strategy and the LoRA technique. During the fine-tuning of the PLM, it adopts the PFA strategy different from the traditional fine-tuning strategy, freezes the multi-head attention of the first F layers to retain the rich knowledge obtained by the PLM during the pre-training stage, and unfreezes and fine-tunes the last U layers to enable the PLM to better learn and adapt to ocean spatio-temporal data. The LoRA technique is added during the fine-tuning stage, which not only significantly reduces the amount of parameter adjustment of the model and effectively reduces the computational cost, but also maintains the good performance of the model in ocean interpolation.
[0041] The present invention compares the ocean interpolation model with other interpolation models. The results not only prove the spatial nature of ocean buoy data, but also prove the feasibility and effectiveness of using pre-trained language models for ocean spatio-temporal data analysis, thus providing a new idea for ocean data research. Compared with the baseline, the OSTI-PLM model achieves the optimal interpolation under different missing rates and different missing types, proving the effectiveness of the OSTI-PLM model for interpolating ocean multi-buoy spatio-temporal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0043] Figure 1 is the architecture diagram of the ocean interpolation model in the embodiment of the present invention;
[0044] Figure 2 is the comparison diagram between the traditional frozen pre-trained transformer and the partial frozen attention strategy in the embodiment of the present invention;
[0045] Figure 3 is the optimal parameter curve diagram under the PFA strategy in the embodiment of the present invention;
[0046] Figure 4It is the distribution map of multiple buoy stations in the Mediterranean Sea in the embodiments of the present invention;
[0047] Figure 5 It is the map of missing data types in the experiment in the embodiments of the present invention;
[0048] Figure 6 It is the example diagram of the setting of parameters of the ocean interpolation model in the experiment in the embodiments of the present invention;
[0049] Figure 7 It is the comparison of the model interpolation performance of different missing types under different missing rates in the Mediterranean Sea dataset in the embodiments of the present invention;
[0050] Figure 8 It is the ablation experiment result under different missing types with a missing rate of 30% in the embodiments of the present invention;
[0051] Figure 9 It is the ablation experiment result under different missing types with a missing rate of 50% in the embodiments of the present invention. Detailed implementation manners
[0052] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0053] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0054] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0055] Inspired by the method of using a pre-trained large language model to predict traffic data, the present invention proposes an ocean interpolation model (OSTI-PLM) specifically designed for interpolating spatio-temporal data of ocean buoys to solve the problem of a large number of missing values in ocean buoys. This method combines ocean multi-buoy spatio-temporal data with a pre-trained language model PLM, and uses the powerful knowledge base and reasoning ability of the pre-trained language model to understand the ocean buoy spatio-temporal data, and then effectively fills in the in-situ missing data. Experiments conducted on the data of 20 buoys of in-situ observation data in the Mediterranean Sea show that the method of using PLM to fill ocean data is feasible and effective, which provides new ideas and solutions for ocean research.
[0056] Example 1
[0057] This example discloses a method for filling missing data of ocean buoys based on a pre-trained language model, and the method includes the following steps:
[0058] S1: Obtain the ocean buoy data to be filled and the graph structure of the buoy stations;
[0059] In this example, in-situ buoy observation data, i.e., ocean buoy data, is obtained using an ocean buoy monitoring system.
[0060] Considering the spatial interdependence between buoys, it is necessary to obtain the spatial relationship between buoy stations and construct a graph structure of the buoy stations, that is, a graph network with each buoy station as a node and the distance relationship between buoys as an edge. The graph structure is defined as G=(V, E, A) to describe the data used, where V is a finite set composed of each buoy node, E is a finite set composed of edges between connected nodes, and A∈R N×N is the adjacency matrix, indicating the connectivity between nodes; for the adjacency matrix A, if the distance between two nodes is less than or equal to the distance threshold, it is considered that there is an edge between the two nodes, and the corresponding value is 1, otherwise it is 0.
[0061] The specific construction process of the graph structure is as follows: Each buoy station is used as a node in the graph structure, and its distance is calculated through the longitude and latitude between stations; a distance threshold is set (to avoid interference caused by data between stations with too large distances). If the distance between two nodes is less than or equal to this threshold, it is considered that there is an edge between the two nodes in the graph, otherwise there is no edge; the reciprocal of the distance is used as the weight of the edge between the two nodes in the graph. The smaller the distance, the greater the weight, that is, the closer the space between the two nodes and the higher the spatial similarity. The spatial correlation between buoy stations can be more intuitively reflected through the graph structure.
[0062] At the same time, a binary mask matrix M is also defined to represent the observation situation of the data. M t,n =1 indicates that the observed value of the nth node exists at time t. Then, M t,n =0 indicates that the observed value of the nth node is missing at time t. Through the mask matrix, the ocean interpolation model can distinguish which data should be interpolated.
[0063] S2: Input the marine buoy data and the graph structure into a marine interpolation model for interpolation and filling to obtain complete marine buoy data. The marine interpolation model includes a spatio-temporal feature extraction module, a spatio-temporal tokenization module, a pre-trained language model, and an output layer connected in sequence. The spatio-temporal feature extraction module extracts a time feature representation from the marine buoy data and a spatial feature representation from the graph structure. Input the time feature representation and the spatial feature representation into the spatio-temporal tokenization module for feature extraction and transformation to obtain time tokens and spatial tokens and connect them to obtain a token sequence. Input the token sequence into the pre-trained language model for learning to obtain a high-dimensional hidden representation. The output layer receives the high-dimensional hidden representation and converts it into an interpolation value to obtain complete marine buoy data.
[0064] The architecture of the marine interpolation model is as Figure 1 shown. Specifically, first, a spatio-temporal feature extraction module for marine data is designed to capture the periodic and seasonal changes of the data and the spatial relationships between buoys to help the model effectively understand the marine spatio-temporal dependence relationships. Then, a spatio-temporal tokenization module is designed to convert the spatio-temporal feature vectors into token embeddings acceptable to the PLM. In addition, to optimize the training process of the model and reduce the computational overhead, during the fine-tuning stage of the PLM, the present invention also adopts a partial freezing attention (PFA) strategy and a Low-Rank Adaptation (LoRA) technique, enabling the PLM to more effectively adapt to the missing value filling task of marine spatio-temporal data while retaining the knowledge in the pre-training stage, improving the interpolation accuracy and the adaptability of the model. Finally, the output layer is a decoder that maps the high-dimensional hidden representation output by the PLM into actual filling values to reconstruct the complete marine spatio-temporal data.
[0065] (I) Spatio-temporal feature extraction module
[0066] In this embodiment, in the processing of marine spatio-temporal data, how to effectively capture and express the heterogeneity of the data at different timestamps and different stations is an important challenge faced by the present invention. To fully exploit the spatio-temporal features in marine data, a spatio-temporal feature extraction module is designed in the pre-trained language model framework specifically for interpolating marine buoy spatio-temporal data. This module includes two parts: time feature extraction and spatial feature extraction, aiming to comprehensively reveal the time dynamics and spatial relationships of marine buoy data.
[0067] (1) Time feature extraction: Time feature extraction aims to capture the diurnal and seasonal changes in marine buoy data. Specifically, hour-based features and month-based features are extracted for modeling diurnal and seasonal patterns respectively.
[0068] Hour-based feature extraction aims to capture daily cycle features, such as the daily fluctuations of ocean temperature driven by sunlight changes. Let the timestamp be t, the hour of the timestamp be h(t), and the corresponding hour feature vector be H(t) ∈ R d , where d is the feature dimension. Month-based feature extraction is used to capture the seasonal changes of ocean parameters. Let the timestamp be t, the month of the timestamp be m(t), and the corresponding month feature vector be M(t) ∈ R d , where d is the feature dimension.
[0069] Finally, the hour feature vector H(t) and the month feature vector M(t) are concatenated to obtain a comprehensive time feature vector T(t), which can comprehensively reflect the time dimension features in the data, especially the features in capturing periodic and seasonal changes. The time feature vector is specifically represented as:
[0070]
[0071] where T(t) represents the time feature vector corresponding to the timestamp t, represents the concatenation operation.
[0072] Therefore, the time features corresponding to all time steps are represented as T T ∈ R N×T×2d , where the superscript T represents the number of time steps, and N represents the number of buoy stations. That is, the time feature vectors of all time steps are concatenated to form the overall time feature representation, so as to better capture the overall time series features.
[0073] (2) Spatial feature extraction: Spatial feature extraction is mainly used to capture the spatial correlation between different buoy stations.
[0074] To effectively reveal the mutual dependence of spatial data, a graph Laplacian matrix L is constructed according to the adjacency matrix A. This matrix not only considers the degree information of each station but also combines the connection relationships between adjacent stations. In this way, the topological structure and spatial dependence between stations can be effectively captured, helping the pre-trained language model to express spatial features more accurately. The process of constructing the graph Laplacian matrix is as follows: First, obtain the degree corresponding to each node from the adjacency matrix to get the corresponding degree matrix; subtract the adjacency matrix from the degree matrix to obtain the corresponding graph Laplacian matrix.
[0075] To enhance the expression ability of the pre-trained language model for local spatial relationships and dynamic change patterns, the largest K eigenvectors corresponding to the eigenvalues are selected from the eigenvectors of the graph Laplacian matrix V(n), and a linear transformation is performed on them to generate the spatial feature vector S(n) of each buoy station, which is expressed as:
[0076] S(n) = WV(n) + b
[0077] Among them, W represents the weight matrix and b represents the bias vector.
[0078] Therefore, the spatial features corresponding to all buoy stations are represented as S N ∈R N×T×d , where T represents the number of time steps and N represents the number of buoy stations. That is, the spatial feature vectors of all stations are concatenated into an overall spatial feature representation.
[0079] In summary, by deeply mining the features in the time and space dimensions, the spatio-temporal feature extraction module can effectively capture the spatio-temporal variation laws in ocean buoy data, providing a richer and more accurate feature representation for subsequent data filling and analysis.
[0080] (2) Spatio-temporal Tokenization Module
[0081] Since pre-trained language models were initially designed to process natural language, when combining a language model with ocean buoy spatio-temporal data, first, the ocean spatio-temporal data needs to be converted into a Token embedding form that can be received by the PLM, and then the adaptability of the model can be improved through the fine-tuning process of the PLM.
[0082] After spatio-temporal feature extraction of ocean buoy data, in order to better input the ocean buoy data into the pre-trained language model PLM, a spatio-temporal Tokenization module is designed in the model framework. This module is a key step in converting ocean buoy data into Tokens that the PLM can understand. Spatio-temporal Tokenization not only needs to capture the static and dynamic features of buoy stations but also be able to represent the spatio-temporal state and change trend of the entire system. Therefore, the spatio-temporal Tokenization module includes two parts: time Token conversion and space Token conversion, which respectively process further feature extraction and conversion in the time dimension and the space dimension to better model and analyze spatio-temporal data.
[0083] (1) Time Tokenization
[0084] Given that ocean conditions are highly dynamic and spatially dependent, a single time step cannot fully capture the local patterns and long-term trends in buoy data. Therefore, in the present invention, all time steps are combined into one "time block". Therefore, the core goal of time Tokenization is to encapsulate the state T state and the change trend T trend of the entire system by aggregating the information of all buoy stations in each time step.
[0085] First, calculate the average value of all buoy stations (the average value of ocean observation data) to obtain an overall state μ∈R 1×T, this value reflects the overall state of the system. Then, the first-order difference of μ is calculated, denoted as Δμ ∈ R 1×(T-1) , to represent the overall trend. Then, μ and Δμ are respectively concatenated with the time feature representation T T ∈ R 1×2d , and a multi-layer perceptron (MLP) is used to perform feature transformation on these combinations to obtain the overall state Token (T state ) and the overall trend Token (T trend ). Finally, these two Tokens are concatenated and normalized to obtain the final time Token, and the time Token is denoted as The specific process is expressed as:
[0086]
[0087] T token = LayerNorm(T state + T trend )
[0088] where represents the concatenation operation of features, and the torch.concat function is used to concatenate different features in the feature dimension; + represents the addition operation between vectors.
[0089] (2) Spatial Tokenization
[0090] The purpose of spatial Tokenization is to capture and transform the static and dynamic features of each buoy station. The static features mainly reflect the geographical location and periodic features of the buoy station, and the dynamic features capture the change patterns of the station over time. In addition, the missing value pattern is also important information in the spatial features. Therefore, the present invention also considers the influence of the missing pattern on the spatial Token.
[0091] The specific steps to obtain the spatial Token are as follows: The time feature representation and the spatial feature are concatenated, and after concatenation, a non-linear transformation is performed through a multi-layer perceptron to obtain the static essence Token; a multi-layer perceptron is used to perform a non-linear transformation on the historical observation data of the buoy station to obtain the dynamic change Token; a multi-layer perceptron is used to perform a non-linear transformation on the mask matrix to obtain the missing pattern Token; the static essence Token, the dynamic change Token, and the missing pattern Token are combined to obtain the spatial Token.
[0092] The multi-layer perceptron MLP is essentially composed of two linear layers, with a ReLU activation function in the middle. First, for the time feature representation T T and the spatial feature representation S NPerform a splicing operation, and splice the two in the feature dimension to obtain the combined total feature representation. Then, perform a non-linear transformation on it through an MLP to construct the static essential Token, namely S inherent ∈R N×D Then, extract the dynamic features from the historical observation data X through an MLP and convert them into the dynamic change Token, namely S dynamic ∈R N×D Construct the missing pattern Token, namely S, based on the mask matrix M through an MLP missing ∈R N×D Finally, add the static essential Token, the dynamic change Token, and the missing pattern Token to form the final spatial Token, namely S token ∈R N×D The specific process is expressed as:
[0093]
[0094] S dynamic = MLP(X)
[0095] S missing = MLP(M)
[0096] S token = LayerNorm(S inherent + S dynamic + S missing )
[0097] Among them, represents the splicing operation of features
[0098] Through the above two tokenization processes, not only can the buoy data be converted into a form that the PLM can understand, but also the variation laws of the time and space dimensions are effectively combined, which helps the PLM better understand and process the marine spatio-temporal data and achieve higher performance in the task of filling missing values
[0099] (III) Pre-trained language model
[0100] To improve the accuracy of filling missing values in marine spatio-temporal data, a partially frozen attention (PFA) strategy is introduced in the pre-trained language model PLM. As Figure 2 shown, different from the traditional frozen pre-trained transformer (FPT) method, the PFA strategy only freezes the multi-head attention layers and feed-forward networks of the first few layers, and unfreezes the multi-head attention of the last few layers for fine-tuning to better capture the spatio-temporal dependence of the data. In the pre-trained language model of the present invention, the multi-head attention layers and feed-forward networks of the first F layers are kept frozen, retaining the basic knowledge learned in the pre-training stage, and the multi-head attention layers of the last U layers are unfrozen for fine-tuning, enabling the model to better adapt to the task of filling missing values in marine spatio-temporal data
[0101] The fine-tuning mechanism mainly refers to training on the basis of a pre-trained language model, using data for a specific task and freezing a part of the parameters to optimize the model's performance on that task. Its purpose is to enable the PLM to better adapt to the needs of a specific domain or task, thereby improving its accuracy and generalization ability.
[0102] To better set the U parameter, the present invention conducted an optimization experiment on the U parameter in the PFA strategy under the condition of 50% randomly consecutive block missing data. As Figure 3 shown, when the number of GPT layers is 5, F is set to 3, and U is set to 2, the model performs best. Therefore, in the final model framework of the present invention, the U parameter is set to 2.
[0103] In addition, the present invention combines the LoRA (Low-Rank Adaptation) technology in the fine-tuning stage, and further optimizes the model's performance on specific tasks by performing low-rank adaptation on the unfrozen part of the parameters. The LoRA technology enables the model to more efficiently learn task-related features by adding adjustments of low-rank matrices during training, while reducing parameter redundancy and enhancing the model's training efficiency and generalization ability.
[0104] By introducing the PFA strategy and LoRA technology, the PLM of the present invention can, while retaining the knowledge in the pre-training stage, more effectively adapt to the task of filling missing values in ocean spatio-temporal data, improving the interpolation accuracy and the adaptability of the model.
[0105] (IV) Output layer
[0106] The spatio-temporal Tokenization module outputs time Tokens and space Tokens, connects the time Tokens and space Tokens to form a Token sequence. Then, these sequences are input into the PLM. The PLM processes these sequences through its self-attention mechanism to learn the complex relationships between Tokens. This process enables the PLM to generate a high-dimensional hidden layer representation containing spatio-temporal information for each time step and all stations. Next, the decoder, as the output layer, is responsible for converting the high-dimensional hidden layer representation generated by the PLM into actual imputation values. Specifically, it maps the high-dimensional features to the target output dimension through a fully connected neural network (including linear transformation and ReLU activation function) to generate imputation values. The model fills the imputation values into the positions of the missing data and outputs the complete data after filling. Specifically, it replaces the missing data corresponding to the positions with a value of 0 in the mask matrix with the imputation values, and a value of 1 indicates that the observed value already exists. Therefore, the final output is a complete data sequence.
[0107] By mapping these high-dimensional hidden layer representations to the target output dimension, it effectively generates imputation values for all stations at each time step, filling in the missing parts in the original sequence to form a complete data sequence.
[0108] The following is a specific introduction to the experiment:
[0109] In this embodiment, the real data used in the experiment comes from the official website of the Copernicus Marine Environment Monitoring Service. As Figure 4 shown, in-situ buoy observation data of 20 stations in the Mediterranean Sea are selected, covering hourly ocean temperature data throughout 2023.
[0110] In the experiment, to better understand and process ocean data, as Figure 5 shown, the ocean missing data is divided into two categories: random point missing and random continuous block missing. Random point missing refers to the missing of single buoy observation data at a specific time point, while random continuous block missing refers to the continuous time-series missing of single buoy data within a period of time or the simultaneous missing of multiple adjacent buoy observation data in the same period.
[0111] The parameter settings of this experiment are as Figure 6 shown. To prove the effectiveness of the method of the present invention, the present invention conducts experiments on the Mediterranean Sea data under the conditions of missing rates of 10%, 30%, 50% and missing types of random point missing and random continuous block missing, and takes the mean absolute error (MAE) and root mean square error (RMSE) as evaluation indicators to measure the performance of the filling model. The pre-trained language model framework specifically for ocean buoy data interpolation of the present invention is compared with the baseline models (MEAN, SAITS, BRITS, SPIN and Imputeformer). Among the baseline models, MEAN is the mean interpolation, SAITS and BRITS models are time-series interpolation models, and Imputeformer and SPIN are spatio-temporal data interpolation models.
[0112] The experimental results are as Figure 7 shown. It can be seen from the experimental results that the performance of all models in dealing with random point missing is generally better than that of random continuous block missing, because random point missing does not destroy the overall spatio-temporal structure of the data, while random continuous block missing requires the model to have stronger ability to extract time dependence and spatial correlation modeling ability.
[0113] Compared with all baseline models, the OSTI-PLM model achieves the best results in the cases of random point missing and random block missing at different missing rates, which further proves the feasibility and effectiveness of using the pre-trained language model for ocean buoy data interpolation.
[0114] In the case of the increasing missing rate, compared with other baseline models, the performance of the ocean interpolation model of the present invention decreases minimally, indicating that the model of the present invention has high interpolation accuracy and strong robustness.
[0115] From the comparison results, the spatio-temporal data interpolation model has better effects than the pure time-series interpolation model, which indicates that there is spatial correlation among ocean buoy data, and it is difficult to obtain accurate interpolation results by only considering time dependence. The OSTI-PLM model has better interpolation effects compared with other spatio-temporal interpolation models, which depends on the spatio-temporal feature extraction module and spatio-temporal Tokenization module specifically designed for ocean spatio-temporal data of the present invention, which can effectively capture the time-series features and spatial correlation among ocean buoy data.
[0116] To better prove the effectiveness of each component of the OSTI-PLM model in filling ocean buoy data, the present invention conducted a series of ablation experiments, removing time feature extraction, spatial feature extraction, PLM, and PFA strategies respectively. Experiments were carried out on ocean buoy data with different missing types at missing rates of 30% and 50%, and the results are as Figure 8 and Figure 9 shown.
[0117] The experimental results show that the complete OSTI-PLM model has achieved the best MAE and RMSE performance in all missing types of the two missing rates, which proves that time feature extraction, spatial feature extraction, PLM, and PFA strategies have a significant positive impact on the filling performance of the model. From the contributions of each module, the performance degradation is the most obvious after removing PLM, especially in the case of continuous block missing, which indicates that PLM plays a key role in modeling global spatio-temporal dependencies. This also further proves the feasibility and effectiveness of using the pre-trained language model of the present invention for ocean buoy data interpolation.
[0118] The present invention aims at the problems of missing and incomplete ocean buoy data, proposes an ocean data processing method combining a pre-trained language model with ocean spatio-temporal data, and designs an ocean interpolation model specifically constructed for ocean buoy spatio-temporal data interpolation.
[0119] The present invention proposes a method of using a pre-trained language model for the task of filling ocean spatio-temporal missing data. This method extracts time and space features and Token embeddings from ocean spatio-temporal data to generate the input form received by the PLM, and uses the powerful knowledge base and reasoning ability of the large language model to analyze ocean spatio-temporal data and interpolate missing values, so as to achieve efficient interpolation performance.
[0120] The present invention proposes a brand-new ocean spatio-temporal data missing value filling pre-trained language model, namely the ocean interpolation model, called the OSTI-PLM model. This model can effectively handle the missing value problem in ocean multi-buoy in-situ observation data to form a complete ocean observation data set.
[0121] The OSTI-PLM model designs a spatio-temporal feature extraction module and a spatio-temporal Tokenization module, enabling the model to better understand and extract the spatio-temporal characteristics of ocean data. During the fine-tuning process of the pre-trained language model PLM, a partial freezing attention (PFA) strategy and LoRA technology are introduced to optimize the training process of PLM and improve the interpolation accuracy of the model. Experiments are also carried out on real Mediterranean data sets to verify the effectiveness and superiority of this method and model.
[0122] Embodiment 2
[0123] This embodiment discloses an ocean buoy missing data filling system based on a pre-trained language model, including:
[0124] A data acquisition module, which is configured to: acquire ocean buoy data to be filled and the graph structure of the buoy site;
[0125] A model filling module, which is configured to: input the ocean buoy data and the graph structure into the ocean interpolation model for interpolation filling to obtain complete ocean buoy data;
[0126] The ocean interpolation model includes a spatio-temporal feature extraction module, a spatio-temporal Tokenization module, a pre-trained language model, and an output layer connected in sequence. The spatio-temporal feature extraction module extracts a time feature representation from the ocean buoy data and a spatial feature representation from the graph structure. The time feature representation and the spatial feature representation are input into the spatio-temporal Tokenization module for feature extraction and conversion to obtain time Tokens and spatial Tokens, which are then connected to form a Token sequence. The Token sequence is input into the pre-trained language model for learning to obtain a high-dimensional hidden layer representation, and the output layer receives the high-dimensional hidden layer representation and converts it into an interpolation value to obtain complete ocean buoy data.
[0127] Embodiment 3
[0128] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment 1 are implemented.
[0129] Embodiment 4
[0130] The purpose of this embodiment is to provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it performs the steps of the method in Embodiment 1.
[0131] The steps involved in the devices in the above Embodiments 3 and 4 correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0132] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0133] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0134] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A method for filling missing data of ocean buoys based on a pre-trained language model, characterized in that: include: Obtain the ocean buoy data to be filled and the graph structure of the buoy sites; Inputting the ocean buoy data and the graph structure into an ocean interpolation model for interpolation and filling to obtain complete ocean buoy data; The ocean interpolation model includes a spatiotemporal feature extraction module, a spatiotemporal Tokenization module, a pre-trained language model and an output layer that are connected in sequence; the spatiotemporal feature extraction module extracts the time feature representation from the ocean buoy data and extracts the space feature representation from the graph structure; the time feature representation and the space feature representation are input into the spatiotemporal Tokenization module for feature extraction and conversion to obtain the time Token and the space Token and are connected to obtain a Token sequence; the Token sequence is input into the pre-trained language model for learning to obtain a high-dimensional hidden layer representation, the output layer receives the high-dimensional hidden layer representation and converts it into an interpolation value to obtain complete ocean buoy data.
2. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The time feature extracted from the ocean buoy data is specifically expressed as: Extract hourly feature vectors and monthly feature vectors from ocean buoy data respectively, and connect them to get time feature vectors; The temporal feature vectors of all time steps are concatenated to obtain the temporal feature representation.
3. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The spatial feature representation extracted from the graph structure is specifically: Establish a graph Laplacian matrix based on the graph adjacency matrix in the graph structure; Selecting the largest multiple eigenvectors from the eigenvectors of the graph Laplacian matrix, and performing linear transformation on the eigenvectors to obtain spatial eigenvectors; The spatial feature vectors of all buoy sites are concatenated to obtain the spatial feature representation.
4. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The specific time Token is obtained as follows: Obtaining an overall state according to an average value of the ocean buoy data, and calculating a first-order difference of the overall state to obtain an overall trend; A multi-layer perceptron is used to concatenate the overall state and the time feature representation to obtain the state Token; A multi-layer perceptron is used to combine the overall trend with the time feature representation to obtain the trend token; The state Token and the trend Token are connected and normalized to obtain a time Token.
5. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The specific method to get the space token is: A multi-layer perceptron is used to concatenate the temporal feature representation and the spatial feature representation to obtain a static essential token. A multi-layer perceptron is used to extract dynamic features from historical observation data and convert them to obtain dynamic change tokens; A multi-layer perceptron is used to extract missing features from the mask matrix and obtain missing pattern tokens; The static essence Token, the dynamic change Token and the missing pattern Token are connected to obtain a space Token.
6. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The pre-trained language model introduces a partial frozen attention strategy, keeping the multi-head attention layer and feedforward network of the first F layers frozen, and unfreezing and fine-tuning the multi-head attention layer of the last U layer.
7. The method for filling missing data of ocean buoys based on a pre-trained language model according to claim 1, characterized in that: The output layer uses a decoder to convert the high-dimensional hidden layer representation generated by the pre-trained language model into interpolated values.
8. Ocean buoy missing data filling system based on pre-trained language model, characterized by: include: A data acquisition module is configured to: acquire ocean buoy data to be filled and a graph structure of buoy sites; A model filling module is configured to: input the ocean buoy data and the graph structure into the ocean interpolation model for interpolation filling to obtain complete ocean buoy data; The ocean interpolation model includes a spatiotemporal feature extraction module, a spatiotemporal Tokenization module, a pre-trained language model and an output layer that are connected in sequence; the spatiotemporal feature extraction module extracts the time feature representation from the ocean buoy data and extracts the space feature representation from the graph structure; the time feature representation and the space feature representation are input into the spatiotemporal Tokenization module for feature extraction and conversion to obtain the time Token and the space Token and are connected to obtain a Token sequence; the Token sequence is input into the pre-trained language model for learning to obtain a high-dimensional hidden layer representation, the output layer receives the high-dimensional hidden layer representation and converts it into an interpolation value to obtain complete ocean buoy data.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for filling missing data of an ocean buoy based on a pre-trained language model as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for filling missing data of an ocean buoy based on a pre-trained language model as described in any one of claims 1-7 are implemented.