Buoy meteorological data restoration method based on space-time double-attention neural network
By combining a spatiotemporal dual attention neural network with Transformer and an improved GAT model, the problem of combining temporal and spatial features of meteorological data was solved, achieving efficient data repair and improving prediction accuracy and computational efficiency.
Patent Information
- Application Number
- CN202511331639.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing technologies struggle to effectively combine the temporal and spatial characteristics of meteorological data, leading to information loss and low prediction accuracy during data repair. Furthermore, traditional neural network models suffer from wasted computational resources and noise interference when processing meteorological data.
A spatiotemporal dual attention neural network-based approach is adopted, combining the Transformer model with an improved GAT graph attention network. Feature extraction and fusion are performed through temporal and spatial encoding links. An adaptive physical constraint matrix is used to optimize the correlation between meteorological elements, and a bidirectional prediction method is used to handle continuous missing values.
It enables high-precision prediction and replacement of outliers and missing values in meteorological data, improves the accuracy and computational efficiency of data repair, reduces the waste of computing resources, and enhances the interpretability and prediction accuracy of the model.
Smart Images

Figure CN120832480A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data quality control of ocean data buoys, and particularly relates to a buoy meteorological data repairing method based on a space-time dual attention neural network. BACKGROUND
[0002] Meteorological research needs to use long-term observation data in the research area. Through these complete data, local meteorological characteristics can be analyzed to form inferences. Among them, the ocean buoy is an important platform for obtaining meteorological and hydrological data, with characteristics of long-term, continuity and stability, and can continuously measure various ocean and atmospheric parameters such as sea surface temperature, wind speed, wind direction and air pressure. Real-time data provided by the buoy can be used for verification and improvement of numerical weather prediction models, and is crucial for studying the interaction between the ocean and the atmosphere.
[0003] Ocean buoy meteorological sensors often have missing data and abnormalities due to external environmental factors (such as damage to components / sensors caused by exposure to sunlight, strong winds, high waves, and high salt mist) and internal problems (such as missing calibration and aging of components). These problems distort data characteristics and change statistical quantities, making analysis more difficult, so it is necessary to identify and repair abnormal data to provide reliable data.
[0004] For statistical methods, mean interpolation, median interpolation, and mode interpolation are extremely simple, but they severely underestimate the variability and correlation between variables, and are obviously not suitable for meteorological data, which contains time series characteristics.
[0005] In deep learning, RNN and its derivative models GRU and LSTM are effective and widely used for time series data, but have defects: RNN has gradient explosion / vanishing problem, and improved RNN still cannot get rid of the gradient problem; GRU alleviates this problem, but does not explore the importance of information, making it difficult to learn long sequences; LSTM is better at processing long sequences, but has high computational cost, and it is difficult to capture long-term dependencies when processing extremely long sequences, and it is prone to overfitting (especially when training data is scarce).
[0006] The Transformer's global self-attention mechanism demonstrates unique advantages in capturing complex temporal dependencies and long-term trends within time series, making it possible to capture dependencies over longer distances. Furthermore, meteorological factors are not independent but rather interrelated. This correlation structure can be represented in a computer via a graph structure, but this graph structure is not a regular Euclidean space model and cannot be processed using traditional CNNs. GNNs are neural network models specifically designed for learning and predicting non-Euclidean spatial data. GNNs are primarily categorized as GCN, GAT, and GAE. However, GNN models generally suffer from limitations, such as weak generalization, high computational overhead, and limited expressive power. Furthermore, in meteorological data applications, a single meteorological factor is often only partially influenced by multiple related factors. Relying solely on inter-factor correlations for predictions fails to fully capture the complex patterns of meteorological factor variation.
[0007] Furthermore, the core attention mechanisms of both the Transformer and GAT suffer from a common flaw: they allocate unnecessary attention to irrelevant or weakly correlated connections or sequence units. This redundant attention not only wastes computational resources but also introduces noise into the model, reducing prediction accuracy. For example, in meteorological data, wave period and temperature are variables with very weak physical correlation. If the model assigns significant attention weight to them, it may interfere with the model's ability to capture true causal relationships, thereby affecting the accuracy of its judgments. In summary, while existing methods can repair and interpolate buoy meteorological data from either the temporal or spatial dimensions, meteorological data itself possesses heterogeneous characteristics in both temporal and spatial dimensions. Relying solely on a single property for interpolation often results in the loss of other types of feature information, failing to truly achieve deep integration and coordinated prediction of spatiotemporal features. Furthermore, accurately capturing the physical connections between different meteorological elements and constructing interpretable, strongly correlated influence relationship models are key to improving data interpolation quality. Summary of the Invention
[0008] In order to solve the above technical problems, the present invention provides a buoy meteorological data repair method based on a spatiotemporal dual attention neural network, which effectively combines time and space constraints, realizes the prediction and replacement of outliers and missing values in meteorological data, and thus achieves the purpose of repairing meteorological data.
[0009] To achieve the above object, the technical solution of the present invention is as follows: A buoy meteorological data repair method based on spatiotemporal dual attention neural network includes the following steps: Step 1, data collection and preprocessing: The historical multi-dimensional meteorological element data obtained by the ocean data buoy is collected to divide a training set and a test set, wherein the data in the test set is preprocessed to obtain a test set containing missing values and abnormal values; Step 2, establishing a spatio-temporal double-attention neural network model: The model comprises a time encoding link, a space encoding link and a feature fusion layer; The time encoding link comprises a linear embedding layer, a position encoder and a multi-layer Transformer encoder, each layer of the Transformer encoder comprising a multi-head cross-attention mechanism layer, a residual link and a layer normalization layer and a feedforward neural network layer; The space encoding link comprises an adaptive physical constraint matrix, a graph convolution operation layer and a feature processing layer; The feature fusion layer outputs after fusing the features output by the time encoding link and the space encoding link; Step 3, model training and evaluation: The data in the training set is used to train the model, first, the data is divided into time features containing a target element to be predicted and meteorological elements, the time features containing the target element to be predicted are input into the time encoding link, the meteorological elements are input into the space encoding link, and the model outputs the target element at the time to be predicted; after the model is trained, the test set is used to evaluate the model, the predicted value of the model is compared with the true value, and a qualified model is obtained; Step 4, data repair: The missing values or abnormal values in the real-time collected multi-dimensional meteorological element data are predicted by using the qualified model to realize data repair.
[0010] In the above scheme, the time feature sequence containing the target element to be predicted is input into the time encoding link First, the linear embedding layer is used to map the discrete input symbols to dense vector representations, and then the position encoding is added to obtain the fused sequence ; the sequence enters the multi-head cross-attention mechanism layer, and the calculated attention score is combined with the sequence to perform a residual link and layer normalization, and output the layer normalization result x; then x enters the feedforward neural network layer, which performs nonlinear transformation and deep feature abstraction on each position vector that has already contained global context information, and outputs ; then and x are combined again to perform a residual link and layer normalization, and finally the time series feature matrix of the target element is obtained and output.
[0011] In a further technical solution, the position encoding is an encoding of time information by a position encoder, and the position encoder generates a fixed position encoding vector using a cosine function: ; ; wherein, represents the position encoding of even dimension index, represents the position encoding of odd dimension index, pos is the time step position, and i is the dimension index; represents the total dimension of the input, including the dimension of the position vector and the dimension of the word embedding vector.
[0012] In a further technical solution, the sequence After entering the multi-head cross-attention mechanism layer, first will be multiplied by the attention weight matrix to obtain Q, K, and V, which represent the query, key, and value matrices, respectively; then the attention score is calculated: ; wherein, is the attention score, N is the length of the input data, is the dimension of ; After obtaining the attention score , it will be linked with the input sequence once and layer normalized, and the calculation process is as follows: ; wherein, x is the layer normalization result, and LayerNorm is the layer normalization function.
[0013] In a further technical solution, after x enters the feedforward neural network layer, the calculation formula is as follows: ; wherein, represents the abstraction of the features calculated by the feedforward fully connected layer after nonlinear transformation, is the function; , is the weight matrix of the feedforward neural network, , is the bias of the feedforward neural network; Subsequently, and x are linked again once and layer normalized to finally obtain the time series feature matrix of the target element and output, and the calculation process is as follows: ; wherein, represents a layer normalization function.
[0014] In the above scheme, the meteorological element input sequence After inputting the spatial encoding link, the adaptive physical constraint matrix The graph convolution calculation formula is as follows: ; wherein, represents the meteorological element obtained after graph convolution calculation; Then enter the feature processing layer, in the feature processing layer, the meteorological element After the ReLU function, it is nonlinearly transformed and projected to a high bit space, and finally the spatial feature of the target element is obtained and output, and the calculation process is as follows: ; wherein, , is a weight matrix of the spatial encoding link, , is a bias vector of the spatial encoding link.
[0015] In a further technical scheme, the adaptive physical constraint matrix is obtained after adaptive adjustment during the model training process, and the calculation formula is as follows: ; wherein, represents a Sigmoid function, represents a Hadamard product; is a physical prior matrix, which is a n*n Boolean matrix, and the row and column correspond to the meteorological element set , wherein n is the number of meteorological elements; the row index in the matrix corresponds to the target element, i.e. the affected element, and the column index corresponds to the original meteorological element, i.e. the influencing element; for the row where the target element is located, the weight value of the column position corresponding to the influencing original meteorological element is set to 0 or 1: wherein 0 represents that the original meteorological element and the target element are irrelevant elements, and 1 represents that the original meteorological element and the target element are relevant elements; is a physical learning matrix, which is a learnable parameter matrix, and the initial value is the same as , the weight value in the matrix will be automatically adjusted according to the gradient descent process during training, and the range is kept in [0, 2].
[0016] In the above scheme, in the feature fusion layer, the output fusion feature is calculated as follows: ; wherein, a spatial feature output for a spatial encoding link, a temporal feature output for a temporal encoding link, denotes a Hadamard product, is a gating vector: ; wherein, denotes a Sigmoid function, , is a weight matrix of the feature fusion layer.
[0017] In the above scheme, in step 4, for the consecutive missing values, a bidirectional prediction method is used in the prediction stage, that is, the input sequence is first predicted in the time direction, and then predicted in the reverse time direction, and the output value of the prediction is: ; wherein, is a prediction output value, is a forward output fusion feature, is an output fusion feature, is a position weight.
[0018] In the further technical scheme, in step 4, according to the relative position of the missing value to the prediction direction, a dynamic weight is used for prediction, and the position closer to the prediction direction is given a higher weight, and the weight range is limited to [0.1, 0.9]; the function expression is: ; wherein, is the relative position of the missing value to the consecutive missing values at this position, is the consecutive length of the missing value.
[0019] Through the above technical scheme, the buoy meteorological data repairing method based on the spatio-temporal double attention neural network has the following beneficial effects: 1. The spatio-temporal double attention neural network model of the present application combines the Transformer model and the GAT graph attention network improved by the attention mechanism, constructs a spatio-temporal double link reasoning structure, and realizes the separation and fusion prediction of time sequence and spatial features; 2. In the time sequence link, the time feature is first extracted and combined with the target meteorological element to form a time sequence input, so as to avoid information leakage and pollution to attention calculation; then, the Transformer encoder structure is used to model the sequence to effectively capture the long-range dependence relationship; 3. In the space link, an adaptive physical constraint matrix GAT model is innovatively designed to capture the features between the space dependence, that is, the spatial dependence, and the features between the space dependence are captured by using the adaptive physical constraint matrix to replace the attention mechanism of the original GAT, which effectively suppresses the link between irrelevant features, strengthens the correlation between related features, and dynamically adjusts the weight relationship with training, reduces the calculation amount, and improves the prediction accuracy; 4. In order to effectively process the continuous abnormal values and missing values in the data, a bidirectional prediction method is adopted according to the error accumulation rule and the time series data characteristics. In order to ensure the rationality of the prediction result, a dynamic weight mechanism is introduced according to the position of the missing data relative to the prediction direction: the closer the position along the prediction direction, the higher the weight. The weight range is limited to [0.1, 0.9]. The experimental results show that this method significantly improves the prediction accuracy.
[0020] Overall, the model of the application effectively combines time and space constraints, realizes the prediction and replacement of abnormal values and missing values in meteorological data, and realizes the repair of meteorological data. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiment or prior art description will be briefly introduced below.
[0022] Figure 1 A space-time dual attention neural network ST-DAN structure disclosed by the embodiment of the application is shown in the figure; Figure 2 A time encoding link structure diagram is shown in the figure; Figure 3 An attention mechanism calculation flowchart is shown in the figure; Figure 4 An adaptive physical mask matrix structure diagram is shown in the figure; Figure 5 A temperature data repair effect diagram is shown in the figure, (a) is a Transformer, and (b) is a ST-DAN; Figure 6 A Figure 5 An A area enlargement comparison diagram in the figure, (a) is a Transformer, and (b) is a ST-DAN; Figure 7 A Figure 5 A B area enlargement comparison diagram in the figure, (a) is a Transformer, and (b) is a ST-DAN; Figure 8 A wind speed data repair effect diagram is shown in the figure, (a) is a Transformer, and (b) is a ST-DAN; Figure 9 A Figure 8Zoom-in comparison of region A, (a) is Transformer, (b) is ST-DAN; Figure 10 For Figure 8 Zoom-in comparison of region B, (a) is Transformer, (b) is ST-DAN. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application.
[0024] The present application provides a kind of based on spatio-temporal double attention neural network's buoy weather data repair method, comprising the following steps:
[0025] Step 1, data acquisition and pretreatment: Collect the historical multi-dimensional meteorological element data obtained by the ocean data buoy, including but not limited to air pressure, air temperature, wave period, wind speed, humidity and other multi-dimensional meteorological elements. The collected data has good time continuity, and the selected meteorological variables can be flexibly adjusted according to the actual available data type and content. In this embodiment, the data set uses the intact data collected by the ocean data buoy located in Qingdao Xiaomai Island area for experimental verification, the data set has a time resolution of ten minutes, contains five meteorological elements of air temperature, wind speed, air pressure, wave period and humidity, and the training set: test set is divided into training set and test set-true value .
[0026] The data in the test set is pretreated to obtain a test set containing missing values and abnormal values.
[0027] Detect the target element data, find out the abnormal value points and missing value points by gradient detection, peak detection and continuity detection method, and mark them. To verify the effectiveness of the model proposed in the present application, this embodiment generates a test set Y` by inserting abnormal values into Y, and then repairs Y` through the model. First, replace 15% of the data in the test set-true value Y with randomly distributed missing values, with a maximum continuous time of not more than 4, to simulate the missing values in the actual collected data, and then mark these data to obtain . Then, artificial noise is inserted, and 5% of the data is converted into abnormal values. For this purpose, an adaptive Gaussian noise injection method based on stratified sampling is designed. The noise injection process includes the following three steps: First, randomly sort the candidate set , where V is the non-empty value point in the data , and then divide it into H layers, denoted as The number of abnormal values m allocated to each layer is calculated by the following formula.
[0028] ; wherein, is the total number of noise points to be inserted. For the remainder , the first R layers are each allocated 1 sample to ensure that the sample allocation is as balanced as possible. After determining the number of abnormal values for each layer, m original data points are independently extracted from each layer of data.
[0029] The second step is the generation of adaptive Gaussian noise. Noise intensity is crucial to model prediction results, and the present application generates Gaussian noise adaptively according to the standard deviation of the original data. The noise generation function is wherein, is the standard deviation of the original data. By scaling the standard deviation by 0.5, the noise intensity is controlled within a reasonable range, ensuring that the noise can introduce data changes without excessively damaging the characteristics of the original data.
[0030] The third step is noise injection. The generated noise is injected into to obtain the data points after noise injection , and the calculation formula is: ; This process precisely adds noise to selected data points through index matching, and uniformly injects a specific proportion of noise in the time dimension, which can better simulate the real situation while preserving the distribution characteristics of the original data. After the above steps, the test set Y` containing 15% missing values and 5% abnormal values is obtained.
[0031] Step 2, establish a spatio-temporal dual attention neural network model: The present application proposes a spatio-temporal dual attention neural network (ST-DAN) combining a Transformer model and a Graph Attention Network (GAT) deep learning model to achieve high-precision time series data repair. The model has a time and space dual-link structure. The time link uses a Transformer encoder layer to capture the time series characteristics of a single element, while the space link captures the correlation between elements through GAT, and optimizes the algorithm by introducing a new physical adjacency matrix to dynamically adjust the weight of the influence between elements to obtain more realistic prediction results.
[0032] The overall architecture of ST-DAN is as follows: Figure 1As shown, it consists of a time encoding link, a space encoding link and a feature fusion layer. The original input data is calculated through the two links of the time encoding link and the space encoding link in parallel, and then the feature fusion is performed to finally output the prediction data. Before data input, the data is first standardized using the z-score method, converting the input features into a distribution with a mean of 0 and a standard deviation of 1, to improve the stability of model training. The input data format is m-dimensional time elements, such as year, month, day, hour, minute, etc.; it also includes n-dimensional meteorological elements, such as temperature, pressure, wind speed, wave period, etc. Then the data will be split into m-dimensional time features plus 1-dimensional target meteorological elements into the time encoding link; and all n-dimensional meteorological elements into the space encoding link.
[0033] 1. Time encoding link
[0034] As shown in Figure 2 , the time encoding link includes a linear embedding layer, a position encoder and a multi-layer Transformer encoder, each of which includes a multi-head cross-attention mechanism layer, a residual link and a layer normalization layer, and a feedforward neural network layer.
[0035] The self-attention mechanism included in the traditional Transformer encoder can capture long-distance time sequence correlation in the data, and can more effectively fit the time evolution trend of the data, thereby obtaining time features. However, this mechanism performs global calculation on the attention scores between all elements in the input sequence, which not only produces a huge amount of calculation when processing complex data, but also may produce attention scores that do not conform to objective reality. To this end, the present application first imposes strict constraints on the input of the Transformer encoder, so that it only receives time elements and target elements, thereby avoiding redundant associations between irrelevant elements, reducing computational complexity, and improving training efficiency and prediction accuracy. At the same time, the use of multi-layer Transformer encoder further enhances the expression ability of the model through a hierarchical feature extraction mechanism: the bottom layer encoder can focus on capturing local time sequence patterns and fine-grained features, and the high layer encoder can realize deep integration of global context on the basis of low layer features, forming a progressive feature abstraction process from specific to abstract. This multi-layer stacking structure not only allows the model to capture multi-scale time sequence dependencies - both preserving local dynamics within short periods and being able to depict associated patterns in long periods, but also alleviates the gradient vanishing problem in deep networks through residual connections and layer normalization mechanisms, allowing the model to still be stable when stacking more encoding layers and continuously improving feature representation capabilities, ultimately providing stronger fitting ability and generalization performance for modeling complex time series data.
[0036] The time encoder first encodes the observation time through the position encoding function, and then establishes a multi-layer Transformer encoder to learn and train the m+1-dimensional data by using the multi-head attention mechanism thereof, so as to fully capture the change characteristics of the target element in the time law, such as the diurnal and seasonal changes of the temperature, while the input of the m-dimensional time element and the 1-dimensional target element ensures that only the characteristics of the target element are captured and learned in the model learning process, thereby avoiding learning pollution caused by irrelevant links and reducing the learning accuracy. After the learning and training of the multi-layer Transformer encoder, the time encoding link outputs the time characteristics.
[0037] Input sequence First, the discrete input symbols are mapped to dense vector representations through a linear embedding layer, and then added to the position encoding. The position encoding is the encoding of the time information by the position encoder. Since the Transformer model only focuses on the content similarity between elements, it lacks the ability to perceive the order of element arrangement, so the position encoding is essential. The position encoder generates a fixed position encoding vector by using the sine / cosine function: ; ; wherein, represents the position encoding of the even-dimensional index, represents the position encoding of the odd-dimensional index, pos is the time step position, and i is the dimension index. represents the total dimension of the input, including the dimension of the position vector and the dimension of the word embedding vector.
[0038] Input sequence after position encoding fusion enter the multi-head cross-attention mechanism layer, the structure and calculation process of which are shown in Figure 3 . First, will be multiplied by the attention weight matrix respectively to obtain Q, K, and V, which represent the query, key, and value matrices, respectively, and then the attention score is calculated. The attention score calculation method is: ; wherein, is the attention score, N is the length of the input data, is the dimension of the input data , and is the dimension of . Taking the experiment as an example, contains day, hour, and temperature, so is 3, at this time is an N*3 matrix, and is set to 256, that is is 256, The shape is , which is 3*256. and Perform matrix multiplication to obtain the Q, K, and V matrices, all of which have a shape of N*256.
[0039] Getting attention score After that, it will be combined with the input sequence Perform a residual connection and layer normalization, and the calculation process is: ; Among them, x is the layer normalization result, and LayerNorm is the layer normalization function.
[0040] Then x enters the feed-forward neural network (FFN). The feed-forward neural network acts independently on the representation vector of each position in the sequence output by the self-attention layer. Its main function is to perform nonlinear transformation and deep feature abstraction on each position vector that already contains global context information. Its structure adopts the expansion-nonlinear-contraction design. First, the input vector x is passed through a weight matrix and bias Linear projection to higher dimension d ff , ,here Take 256. Then apply a nonlinear activation function to the result of the dimension increase. This embodiment uses the ReLU function, and then pass the activated high-dimensional vector through the weight matrix and bias Linear projection back to the original Dimension, its mathematical formula is: ; in, Represents the abstraction of features calculated after nonlinear changes in the feedforward fully connected layer. 、 is the weight matrix of the feedforward neural network, 、 is the bias of the feedforward neural network; the parameters of FFN (W1, W2, b1, b2) are shared for all sequence positions in the same layer, but the calculation is position-independent.
[0041] This layer significantly enhances the model's representational capabilities by introducing powerful nonlinear activations and high-dimensional mappings. It is a key component in the model's ability to learn complex features and contributes significantly to the number of parameters. The FFN complements the self-attention mechanism, which focuses on learning dependencies between positions, and together they contribute to the Transformer encoder layer's powerful feature learning capabilities.
[0042] Subsequently, Residual connection and layer normalization will be performed again with x, and finally the time sequence feature matrix of the target element is obtained And output, the calculation process is as follows: ; Wherein, Indicates a layer normalization function.
[0043] 2. Spatial encoding link
[0044] The spatial encoding path is composed of the physical constraint GAT (Physics Aware GAT) designed by us as a spatial encoder, as shown in Figure 3 The spatial encoding link includes an adaptive physical constraint matrix, a graph convolution operation layer and a feature processing layer.
[0045] Input sequence of spatial encoding link After entering the spatial encoding path, an adaptive physical constraint matrix will be obtained first. The attention mechanism in the traditional graph attention network relies on data-driven implicit learning relationship, which can automatically capture the relationship between features and dynamically adjust the weight. It is extremely convenient and effective for data that cannot directly know the correlation between features, but there may be false correlation captured and lack of physical interpretation during the learning process. For meteorological data, we can directly know the relationship between features, such as temperature, wind speed and pressure, which are related to precipitation. The adaptive physical mask matrix mechanism proposed in the present application replaces the traditional attention mechanism, which can forcibly prohibit the connection that violates the domain knowledge through the predefined physical constraint matrix, ensuring that the elements with correlation can be correctly connected, enhancing the interpretability of the model, avoiding unnecessary calculation, and improving the model precision and training efficiency.
[0046] As shown in Figure 4 First, a physical prior matrix is preset , which is an n*n boolean matrix, and the row and column correspond to the meteorological element set , where n is the number of meteorological elements; the row index of the matrix corresponds to the target element, i.e. the affected element, and the column index corresponds to the original meteorological element, i.e. the influencing element; for the row of the target element, the weight value of the column position corresponding to the original meteorological element influencing the element is set to 0 or 1: where 0 indicates that the original meteorological element and the target element are irrelevant elements, and 1 indicates that the original meteorological element and the target element are related elements; taking the of the present embodiment as an example, If the wind speed is to be predicted, the weight of the second row can be set as , so as to ensure the correlation between irrelevant elements.
[0047] However, the Boolean matrix cannot quantify the influence weight between meteorological elements, for this reason, the application designs an adaptive physical constraint matrix module, in which, will be copied and marked as a physical learning matrix , which is a learnable parameter matrix, its initial value is the same as , the weight in the matrix will be automatically adjusted according to the gradient descent process during training, and the range is kept in [0, 2]. After the gradient descent update, will take Hadamard product with the Boolean matrix , during the training process, is fixed and unchanged, so as to effectively ensure that the weight of 0 in the matrix is always 0, that is, to forcibly prohibit the link between irrelevant elements, after obtaining the result, it will be regularized by the Sigmoid function, the range is controlled in [0, 1], and then multiplied by 2 to restore the weight range in [0, 2], which ensures that the weight in the matrix will not exceed the limit due to cyclic accumulation during the update process, and keeps the weight in [0, 2], enlarges the range interval, which can increase the weight optimization gradient, so that the change is more smooth. After this calculation, the updated matrix is obtained as the output of the adaptive physical constraint matrix, the calculation formula is as follows: ; Among them, represents the Sigmoid function, represents the Hadamard product.
[0048] This process realizes the organic combination of physical priori and data-driven learning, which not only preserves the physical constraint of feature relationship, but also allows the model to adaptively adjust the influence weight. In the actual calculation process, the weight range is between [0, 2], 0 is irrelevant, and 2 is strongly related. After obtaining the adaptive physical constraint matrix , as a mask and the input sequence will start the graph convolution operation, the graph convolution calculation formula is as follows: ; Among them, represents the meteorological element obtained after graph convolution calculation, at this time the meteorological element is not the spatial feature needed, which needs to be further converted through the feature processing layer.
[0049] In the feature processing layer, the derived meteorological features will be projected to a high-dimensional space through a ReLU function for nonlinear transformation, and finally the spatial features of the target element will be derived and output, and the calculation process is as follows: ; wherein, , is a weight matrix of the spatial encoding link, , is a bias vector of the spatial encoding link.
[0050] The spatial encoding link also learns and trains the input n-dimensional meteorological elements while the time encoding link is learning and training, wherein the n-dimensional meteorological elements already contain the target element, and the GAT neural network is used to focus on the feature correlation between meteorological elements. The GAT here is not the traditional graph attention neural network, but the GAT neural network optimized by the attention mechanism is replaced by the learnable physical constraint matrix proposed in the application. The matrix can strengthen the correlation between related meteorological elements and prohibit the connection between irrelevant variables, for example, when the air temperature data is the repair object, the related variables are air pressure and wind speed, and the irrelevant variables are wave periods, which avoids the poor prediction effect caused by the traditional attention mechanism capturing the correlation between irrelevant elements. At the same time, the parameters in the matrix are not fixed once set, but the changed matrix is continuously taken Hadamard product with the original set matrix to ensure that the parameters of related elements are continuously optimized and the parameters between irrelevant elements are always 0. After learning and training through the spatial link, the spatial features are output.
[0051] 3. Feature fusion layer
[0052] The feature fusion layer outputs the features output by the time encoding link and the spatial encoding link after fusion.
[0053] Subsequently, after obtaining the time features and the space features in the feature fusion module, a decoupling-fusion strategy is used to dynamically adjust the space-time weight by using linear transformation and sigmoid generation gating mechanism. The design follows: physical consistency, which ensures that the spatial features carry clear meteorological physical relationship; time sequence dynamics, which captures the periodicity and trend pattern in the time features; adaptive interaction, which dynamically adjusts the contribution of the space-time features through the gating mechanism. Finally, the fusion features are generated, and the final prediction result is output based on the features.
[0054] Specifically, the spatial encoding link outputs a feature matrix through the improved graph attention network; and the time encoding link generates a time sequence feature matrix through the Transformer encoder after splicing the target meteorological features and the time features Both are spliced in the hidden layer dimension, and adaptive weight allocation is realized through the gated fusion network: first, linear calculation and Sigmoid activation are used to generate a gating vector: ; Wherein, , is the learnable weight matrix of the linear fusion layer, and then: ; Feature fusion is realized, and the fused features are output The design constrains the physical consistency of spatial features through a physical mask matrix, and dynamically adjusts the contribution of spatio-temporal features using a gating mechanism.
[0055] In the prediction output stage, the sliding window method is used for prediction output, and experimental tests show that the sliding window parameters set by the model are the data amount of one day, if the minimum time unit of data collection is hour, then set to 24, if the minimum time unit of data collection is ten minutes, then the input window size is set to 144, and so on, the output window size is 1, and the step is 1. That is, the sliding window slides one unit of data each time, and one day of data is used to predict one unit of data. This method allows the model to capture the mapping relationship from the past one day to the future one time step.
[0056] Step 3, model training and evaluation:
[0057] The data in the training set is used to train the model, first, the data is divided into time characteristics containing the target elements to be predicted and meteorological elements, the time characteristics containing the target elements to be predicted are input into the time encoding link, the meteorological elements are input into the space encoding link, and the model outputs the target elements to be predicted at the time.
[0058] The data set used for training does not contain abnormal values and missing values, and all data is known to the model. The model will learn the trend of the target meteorological elements in the data through unsupervised learning. After training is completed, the model will be packaged and saved for prediction and repair tasks of the to-be-repaired data. After the model is trained, the packaged model is used to start predicting the test set Y`, and the data marked as abnormal values in Y` is predicted and replaced in the prediction output stage.
[0059] In order to effectively model the time dependence and feature interaction, the sliding window technique is used to convert the original sequence into an input-output pair set. Given a time series of length , d is the feature dimension, and the training sample is generated through the sliding window, where the input sequence contains L time step features, and the target value is the target variable of the next time step, and this method allows the model to learn the mapping relationship from the past L time steps to the future P time steps. In the data set collected in this embodiment, since the time resolution of the data is 1 hour, L = 24 and P = 1 are set, that is, when data labeled as an abnormal value or a missing value is identified, the data 24 hours before the data point is used to predict the data of the point and replace it, realizing data repair.
[0060] During the experiment, the set hyperparameters include model structure parameters and training optimization parameters. In the model structure parameters, the size of the hidden layer is set to 256, the number of attention heads in the multi-head cross-attention mechanism is 8, and the random inactivation rate is set to 0.1. The model structure parameters control the complexity of the model, and when the prediction result is overfitting or underfitting, the model structure parameters can be adjusted. In the training optimization parameters, the learning rate is set to 1e-4, the batch size is set to 64, and the gradient clipping size is set to 1.0. The training optimization parameters control the convergence speed and training speed of the model, and when the model training convergence process appears large amplitude oscillation or training speed is too slow, the training optimization parameters can be adjusted.
[0061] After the model is trained, the test set is used to evaluate the model, and the predicted value of the model is compared with the true value to obtain the qualified model.
[0062] The loss function uses the mean absolute error (MAE) as the training target, which is calculated as follows: ; where M is the number of samples, is the data in the test set-true value Y, is the predicted value. MAE has physical abnormal robustness, and when facing extreme values in meteorological data, MAE is less affected, which meets the fault tolerance demand of meteorological prediction on abnormal points. Moreover, MAE has the advantage of interpretability, and the physical unit of MAE is consistent with the prediction target, which is convenient for direct evaluation of model performance.
[0063] The gradient clipping mechanism is introduced in the training process to prevent gradient explosion, and the gradient clipping is calculated as shown in the following formula: ; where θ is the gradient threshold, which is preset to 1, G is the current gradient, is the L2 norm of G, is the updated gradient. During the training process, if the current gradient reaches the threshold θ, the gradient clipping is triggered, and then the updated gradient is used to continue the gradient descent for prediction.
[0064] In addition to the loss function using MAE, this embodiment also uses mean square error MSE, root mean square error RMSE and coefficient of determination R2 as evaluation indexes respectively, and the calculation formulas are as follows: ; ; ; The four data can reflect the difference between the predicted data and the true value of the data. The four data are usually between 0 and 1, wherein MAE, MSE and RMSE are more effective as close to 0, and R 2 is more effective as close to 1.
[0065] Step 4, data repair:
[0066] The evaluation qualified model is used to predict the missing values or abnormal values in the real-time collected multi-dimensional meteorological element data, so as to realize the repair of the data.
[0067] For continuous missing values, a bidirectional prediction method is used in the prediction stage, that is, the input sequence is first predicted in the time direction, and then the time direction is predicted in the reverse direction. The output value of the prediction is: ; Wherein, is the prediction output value, is the fusion feature of the forward output, is the fusion feature of the output, is the position weight.
[0068] According to the relative position of the missing value to the prediction direction, dynamic weight is used for prediction. The closer the position is to the prediction direction, the higher the weight is, and the weight range is limited to [0.1, 0.9]. The function expression is: ; Wherein, is the relative position of the missing value to the continuous missing value at this position, is the continuous length of the missing value.
[0069] Experimental verification
[0070] 1. Experimental data source: The ocean data buoy measured data arranged by wheat island is selected for experimental test, five meteorological elements of air temperature, wind speed, pressure, wave period and humidity are selected, and the data set is divided for training and test.
[0071] 2. Model parameter setting: In this embodiment, the measured data of air temperature and wind speed are predicted and repaired respectively, so different parameter settings are made according to different target elements, especially in the physical constraint matrix part, for air temperature, wind speed, pressure and humidity are selected as related elements, and for wind speed, air temperature and wave period are selected as related elements.
[0072] 3. Comparison of experimental results: In this experiment, the existing Transformer is selected for result comparison, the repair effect of air temperature data is as shown in Figure 5 Figure 5 The horizontal coordinate of each point is 00:00 of each date, the green area and the red area in the figure are two parts with obvious differences, and the local enlarged detail comparison is shown in Figure 6 and Figure 7 , Figure 6 The time of each point in the horizontal coordinate is each time of December 22, Figure 7 The time of each point in the horizontal coordinate is each time of December 28. The repair effect of wind speed data is as shown in Figure 8 Figure 8 The horizontal coordinate of each point is 00:00 of each date, the green area and the red area in the figure are also two parts with obvious differences, and the local enlarged comparison chart is as shown in Figure 9 and Figure 10 , Figure 9 The time of each point in the horizontal coordinate is each time of March 22, Figure 10 The time of each point in the horizontal coordinate is each time of March 23. The experimental results show that the model proposed in the present application has good prediction accuracy and stronger robustness for meteorological data repair.
[0073] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for repairing buoy meteorological data based on a spatiotemporal dual-attention neural network, characterized in that, The method comprises the following steps: Step 1, data acquisition and pretreatment: Collecting historical multi-dimensional meteorological element data obtained by the ocean data buoy, and dividing the data into a training set and a test set, wherein the data in the test set is pretreated to obtain a test set containing missing values and abnormal values; Step 2, establishing a spatio-temporal double-attention neural network model: The model comprises a time encoding link, a space encoding link and a feature fusion layer; The time encoding link comprises a linear embedding layer, a position encoder and a multi-layer Transformer encoder, each layer of the Transformer encoder comprising a multi-head cross-attention mechanism layer, a residual link, a layer normalization layer and a feedforward neural network layer; The space encoding link comprises an adaptive physical constraint matrix, a graph convolution operation layer and a feature processing layer; The feature fusion layer outputs after fusing the features output by the time encoding link and the space encoding link; Step 3, model training and evaluation: The data in the training set is used to train the model, first, the data is divided into time features containing a target element to be predicted and meteorological elements, the time features containing the target element to be predicted are input into the time encoding link, the meteorological elements are input into the space encoding link, and the model outputs the target element at the time to be predicted; after the model is trained, the test set is used to evaluate the model, the predicted value of the model is compared with the true value, and a qualified model is obtained; Step 4, data repair: The missing values or abnormal values in the real-time collected multi-dimensional meteorological element data are predicted by using the qualified model to realize data repair.
2. The method of claim 1, wherein, A time feature input sequence comprising a target element to be predicted First, the discrete input symbols are mapped to dense vector representations by a linear embedding layer, then added with position encodings to get the fused sequence ; sequence Enter the multi-head cross-attention mechanism layer, calculate the attention score and sequence Residual link and layer normalization are performed, and the layer normalization result x is output Subsequently, x enters the feedforward neural network layer, and each position vector that has contained global context information is subjected to nonlinear transformation and deep feature abstraction, outputting ; subsequently and the residual link and layer normalization are performed again with x, and finally the time sequence feature matrix of the target element is obtained and output.
3. The method of claim 2, wherein, The position encoding is the encoding of the time information by the position encoder, and the position encoder generates a fixed position encoding vector by using a sine / cosine function: ; ; wherein, pos encodes the position of even dimension index, pos encodes the position of odd dimension index, pos is the position of time step, i is the index of dimension; represents the total dimension of input, including the dimension of position vector and the dimension of word embedding vector.
4. The method of claim 2, wherein, Sequence After entering the multi-head cross attention mechanism layer, first will be multiplied by the attention weight matrix to get Q, K, V, which represent the query, key, and value matrices, respectively; then the attention score is calculated: ; wherein, is an attention score, N is the length of the input data, is dimension. After obtaining the attention score The input sequence A residual link and layer normalization are then performed, which are calculated as follows: ; Wherein, x is the layer normalization result, and LayerNorm is a layer normalization function.
5. The method of claim 2, wherein, After x enters the feedforward neural network layer, the calculation formula is: ; wherein, represents an abstraction of the features after a non-linear transformation by a fully connected layer in feed-forward, is a function; , is a weight matrix of the feed-forward neural network, , is a bias of the feed-forward neural network; Subsequently, With x again one residual link and layer normalization, eventually get the target element of the timing characteristics matrix And output, the calculation process is as follows: ; wherein represents a layer normalization function.
6. The method of claim 1, wherein, Meteorological element input sequence After inputting the spatial encoding link, the adaptive physical constraint matrix The graph convolution operation is performed, and the graph convolution calculation formula is as follows: ; wherein, represents the meteorological element obtained after the graph convolution calculation; Then enter the feature processing layer, in the feature processing layer, meteorological elements After nonlinear transformation projection by ReLU function to high bit space, the spatial features of target elements are finally obtained And output, the calculation process is as follows: ; wherein , is a weight matrix for the spatially encoded links, , is a bias vector for the spatially encoded links.
7. The method of claim 6, wherein, The adaptive physical constraint matrix After the adaptive adjustment during the model training process, the calculation formula is as follows: ; wherein, represents a Sigmoid function, represents a Hadamard product; is a physical prior matrix, which is a Boolean matrix of n*n, and both the row and column correspond to the set of meteorological elements , wherein n is the number of meteorological elements; the row index in the matrix corresponds to the target element, i.e. the affected element, and the column index corresponds to the original meteorological element, i.e. the influencing element; for the row where the target element is located, the weight value of the column position corresponding to the influencing original meteorological element is set to 0 or 1: wherein 0 indicates that the original meteorological element and the target element are irrelevant elements, and 1 indicates that the original meteorological element and the target element are related elements; is a physical learning matrix, which is a learnable parameter matrix, and the initial value is the same as , the weight value in the matrix will be automatically adjusted according to the gradient descent process during training, and the range is kept in [0, 2].
8. The method of claim 1, wherein, At the feature fusion layer, the output fused feature is calculated as follows: ; wherein, a spatial signature for a spatially encoded link, a temporal signature for a temporally encoded link, denotes a Hadamard product, is a gating vector: ; wherein, denotes a Sigmoid function, , is a weight matrix of the feature fusion layer.
9. The method of claim 1, wherein, In step 4, for the continuous missing values, a bidirectional prediction method is used in the prediction stage, that is, the input sequence is first predicted in the time direction, and then predicted in the reverse time direction, and the output value of the prediction is: ; wherein, is the predicted output value, is the fused feature for the forward output, is the fused feature for the output, is the position weight.
10. The method of claim 9, wherein, In step 4, according to the relative position of the missing value to the prediction direction, a dynamic weight is used for prediction, and the weight of the position closer to the prediction direction is higher, and the weight range is limited to [0.1, 0.9]; The function expression is: ; wherein, for a missing value, the relative position of the missing value to the consecutive missing values, for a missing value, the consecutive length of the missing value.
Citation Information
Patent Citations
Device for in-situ regulation and restoration of organic contaminated soil and underground water
CN103191912A
Low-complexity MIMO radar SR STAP sea clutter suppression method
CN114609607A
Pure electric ship power battery SOC estimation method based on Extrol-Transform deep learning neural network
CN118169572A
Space-time prediction method and device based on space-time mask auto-encoder, and electronic equipment
CN118171773A
Multi-mode ocean temperature and salinity remote sensing prediction method and device
CN118968330A
Cited By
Photovoltaic time sequence dynamic prediction method and system based on attention mechanism
CN122118688A
Photovoltaic time series dynamic prediction method and system based on attention mechanism
CN122118688B
Numerical control machining center monitoring data processing system based on cloud analysis
CN122286603A
A cloud-based analytics-based monitoring data processing system for CNC machining centers
CN122286603B