Method and System for Sea Surface Temperature Prediction Based on Distributed Cross-Scale Joint Learning
By adopting distributed cross-scale joint learning method in sea surface temperature prediction, the problem of failure to effectively consider the cross-scale characteristics of SST data and data leakage risks in the existing technology is solved, and an efficient and safe SST prediction model is realized, which significantly improves the space-time modeling capabilities.
Patent Information
- Application Number
- CN202510307300.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The prior art fails to effectively consider the cross-scale characteristics of SST data in sea surface temperature prediction, and deep learning models based on centralized training have the problem of data leakage risk and inability to directly apply to SST image prediction.
Using a distributed cross-scale joint learning method, a distributed system architecture is built through local feature learning modules and cross-scale joint learning modules to realize cross-node joint feature learning, avoid data leakage, and dynamically optimize the communication weight matrix to integrate different scale features.
On the premise of ensuring data security, joint feature learning across nodes is realized, applicable scenarios of prediction models are expanded, and the space-time modeling capabilities of SST prediction are significantly improved.
Smart Images

Figure CN119832392B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ocean temperature prediction, and particularly to a sea surface temperature prediction method and system based on distributed cross-scale joint learning. Background Art
[0002] Sea surface temperature prediction is an important research content in the field of ocean science. Sea surface temperature data mostly comes from satellite remote sensing, buoys and other sensors, with large amounts and diversity of data, and involves sensitive information. The prediction of sea surface temperature spatio-temporal data field based on deep neural network is an important attempt to apply artificial intelligence to the field of ocean science. Common methods include physical model-based and data-driven machine learning methods. Physical model-based methods usually cannot accurately model some complex non-linear relationships, while data-driven machine learning can more accurately predict sea surface temperature by learning patterns and features in a large amount of historical ocean data, so as to make more accurate forecast predictions.
[0003] Currently, deep learning-based sea surface temperature prediction uses a centralized training method, and mostly uses recurrent neural networks (RNNs) to process time-series sea surface temperature data, including variants of long short-term memory (LSTM) and gated recurrent unit (GRU), to model the temporal relationships between SST time series data, and applies fully connected layers to map to the final prediction result, achieving good prediction results.
[0004] However, the current methods have several problems: First, the cross-scale characteristics of sea temperature data are not considered. For example, the current methods only train for single-scale images and do not consider letting the model learn the cross-scale characteristics of sea temperature. Second, the deep learning model uses centralized training, that is, all data is first transmitted to the processor and then centrally processed, occupying network transmission resources and resulting in a risk of information leakage. For example, sea temperature image data is transmitted to the central server for model training at the central node. During the data upload process, neither the experimental data is encrypted nor means to ensure the security of the transmission process are added. During the data transmission process, it occupies the network bandwidth resources and may also lead to the leakage of real data. Third, the existing federated learning architectures cannot be directly applied to sea temperature image prediction. Federated learning methods allow data to be parallelly trained on multiple nodes, significantly reducing the training time, being able to more effectively utilize existing computing resources, and ensuring data privacy and security. However, ocean science data has a complex structure with requirements in both processing and application, and there is no distributed learning architecture that can be directly applied to sea temperature prediction. Summary of the Invention
[0005] The technical problem to be solved by this application is to overcome the deficiencies of the prior art and provide a sea surface temperature prediction method and system based on distributed cross-scale joint learning. On the premise of ensuring the local storage of original data, cross-node joint feature learning is achieved through distributed model parameter interaction, avoiding security problems such as data leakage, and expanding the applicable scenarios of the prediction model.
[0006] To achieve the above object, the first aspect of this application provides a sea surface temperature prediction method based on distributed cross-scale joint learning, constructing a sea surface temperature prediction model based on distributed cross-scale joint learning. The prediction model adopts a local system and a distributed system architecture. The local system includes several local nodes, and the communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether communication can occur between local nodes. Different sources of sea surface temperature image sequences with the same data structure are distributed on each local node as the input of the prediction model. The prediction model includes:
[0007] A local feature learning module, deployed on each local node in the local system, for performing local feature learning on the sea surface temperature image sequence stored on the local node and outputting the sea surface temperature prediction result of the local node;
[0008] A cross-scale joint learning module, deployed in the distributed system, learning the different-scale sea surface temperature image features on different local nodes through a graph learning module, dynamically optimizing the communication weights between each local node, achieving gradient aggregation of different local nodes, obtaining global gradient updates, and distributing the updated gradient values to each local node to guide cross-scale joint learning.
[0009] Further, the local feature learning module performing local feature learning on the sea surface temperature image sequence stored on the local node includes: constructing the sampling points of the sea surface temperature image as graph nodes, each sampling point containing longitude and latitude coordinates and temperature values, generating an adjacency matrix based on the Euclidean distance between the sampling points. The adjacency matrix is expressed as:
[0010] ;
[0011] where, represents the Euclidean distance between sampling points i and j, represents the threshold distance. When the distance between two sampling points is less than or equal to the threshold distance, the edge is connected. When the distance between two sampling points is greater than the threshold distance, the edge is not connected.
[0012] Further, the local feature learning module performing local feature learning on the sea surface temperature image sequence stored on the local node specifically includes:
[0013] Feature extraction is performed through a distributed graph convolution module, including, on one of the local nodes, the sea surface temperature image features at time t and the adjacency matrix A are used as the inputs of the distributed graph convolution module, and the output of the local node i is obtained through GCN , define the network weight as W, and the GCN is expressed as:
[0014] ;
[0015] Among them, the initial input , represents the feature representation of the sea surface temperature image at time t on the local node i, define , is the identity matrix, D is the degree matrix, , is the activation function, and the model gradient of one local node i is defined as ;
[0016] The temporal feature learning module is used to perform operations, and the output of the distributed graph convolution module is learned through the LSTM network;
[0017] The output of the temporal feature learning module is predicted by using a fully connected layer, and the prediction result of the local node is output.
[0018] Furthermore, the cross-scale joint learning module learns the sea surface temperature image features at different scales on different local nodes through the graph learning module, including: the output of the distributed graph convolution module learns the weights during distributed gradient aggregation through the graph learning module, and cross-scale joint learning is performed according to the value returned by the graph learning module to optimize the network weights. The graph learning module is expressed as:
[0019] ;
[0020] Among them, represents the output of the graph convolution network from the local node i, represents the Frobenius norm of the matrix.
[0021] Furthermore, the cross-scale joint learning module dynamically optimizes the communication weights between each local node, including: optimizing the communication weight matrix through the graph learning module, and the communication weight matrix is used to guide the aggregation parameters of each local node during gradient aggregation. The aggregation parameters include the sea surface temperature image features at different scales. The distributed gradient aggregation is expressed as:
[0022] ;
[0023] Among them, Denote the set of nodes that can communicate with the local node i
[0024] Furthermore, the output features obtained by the distributed graph convolution module will be used as the input of the temporal feature learning module. The temporal feature learning module uses an LSTM network to learn the temporal information of the input features. The operations performed by the LSTM network include forget gate calculation, input gate calculation, candidate cell state generation, cell state update, output gate calculation, and hidden state update.
[0025] Furthermore, the output of the temporal feature learning module is predicted through a fully connected layer. The composition of the fully connected layer is a multi-layer perceptron containing two hidden layers, which maps the hidden state output by the LSTM network to the sea surface temperature prediction at future times .
[0026] Furthermore, the loss function is expressed as:
[0027] ;
[0028] where is the prediction loss, is the loss of the cross-scale distributed joint learning module, and the parameter is a trade-off parameter;
[0029] The prediction loss is expressed as:
[0030] ;
[0031] where represents the prediction result on the local node i, represents the true value;
[0032] The loss of the cross-scale distributed joint learning module is expressed as:
[0033] ;
[0034] where represents the output of the graph convolution network on the local node i, represents the result of the graph learning module, is a trade-off parameter.
[0035] To achieve the above object, the second aspect of the present application provides a sea surface temperature prediction system based on distributed cross-scale joint learning. The prediction system includes a sea surface temperature prediction model based on distributed cross-scale joint learning. The prediction model adopts a local system and a distributed system architecture. The local system includes several local nodes. The communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether communication is possible between local nodes. Different but identically structured sea surface temperature image sequences from different sources are distributed on each local node and used as the input of the prediction model. The prediction model includes:
[0036] A local feature learning module, deployed on each local node in the local system, for performing local feature learning on the sea surface temperature image sequence stored on the local node and outputting the sea surface temperature prediction result of the local node;
[0037] A cross-scale joint learning module, deployed in the distributed system, which learns the different-scale sea surface temperature image features on different local nodes through a graph learning module, dynamically optimizes the communication weights between each local node, realizes the gradient aggregation of different local nodes, obtains the global gradient update, and distributes the updated gradient values to each local node to guide the cross-scale joint learning.
[0038] After adopting the above technical solution, the present application has the following beneficial effects compared with the prior art:
[0039] In the present application, the prediction model adopts a distributed federated learning framework, which effectively ensures the security of image data. On the premise of ensuring the local storage of the original data, cross-node joint feature learning is realized through the interaction of distributed model parameters. Compared with the traditional centralized training method, this framework uses an innovative privacy protection mechanism to avoid directly transmitting image data during the training process, thus avoiding security problems such as data leakage. While avoiding the transmission of the original data, a globally optimized deep learning model is constructed, effectively solving the security bottleneck problem in marine data sharing and expanding the applicable scenarios of the prediction model.
[0040] In the present application, aiming at the multi-scale spatio-temporal characteristics of the ocean temperature field, a feature fusion mechanism driven by a dynamic communication weight matrix is constructed, and the communication weight matrix is used to guide the joint learning of distributed multi-scale features. By fusing the multi-scale feature information between different nodes and establishing a hierarchical feature interaction criterion, it can adaptively fuse local refined features and global macroscopic features. Joint learning of the feature information in different scale spaces is carried out to achieve precise capture and collaborative optimization of the key scale features during the evolution process of the sea temperature field, significantly improving the spatio-temporal modeling ability of cross-regional sea temperature prediction.
[0041] The following further describes in detail the specific implementation manners of the present application with reference to the accompanying drawings. Brief Description of the Drawings
[0042] The accompanying drawings, as part of the present application, are used to provide a further understanding of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, but do not constitute an improper limitation to the present application. Obviously, the accompanying drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0043] In the drawings of the specification:
[0044] Figure 1 is a schematic diagram of the architectures of the local system and the distributed system in this specific embodiment;
[0045] Figure 2 is a schematic diagram of the architecture of the local feature learning module in the local system in this specific embodiment;
[0046] Figure 3 is a schematic diagram of the architectures of the distributed graph convolution module in the local system and the cross-scale joint learning module in the distributed system in this specific embodiment;
[0047] Figure 4 is a schematic diagram of the architecture of the temporal feature learning module in this specific embodiment;
[0048] Figure 5 is a logical diagram of the graph learning module in this specific embodiment. Specific Embodiments
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but do not limit the scope of the present application.
[0050] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0051] Please refer to Figure 1, this application provides a sea surface temperature prediction method based on distributed cross-scale joint learning, constructs a sea surface temperature prediction model based on distributed cross-scale joint learning. The prediction model adopts a local system and a distributed system architecture. The local system includes several local nodes, and the communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether communication can occur between local nodes. Different but identically structured sea surface temperature image sequences from different sources are distributed on each local node as the input of the prediction model. The prediction model includes:
[0052] A local feature learning module, deployed on each local node in the local system, is used to perform local feature learning on the sea surface temperature image sequence stored on the local node and output the sea surface temperature prediction result of the local node;
[0053] A cross-scale joint learning module, deployed in the distributed system, learns the sea surface temperature image features at different scales on different local nodes through a graph learning module, dynamically optimizes the communication weights between each local node, realizes the gradient aggregation of different local nodes, obtains the global gradient update, distributes the updated gradient values to each local node, and guides the cross-scale joint learning.
[0054] It should be noted that the execution subject of the prediction method in this embodiment is a sea surface temperature prediction device based on distributed cross-scale joint learning. This device can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, etc., and the non-mobile electronic device can be a server and a personal computer, etc. This application does not make specific limitations. Hereinafter, taking the execution subject as a server as an example, the sea surface temperature prediction method based on distributed cross-scale joint learning in this embodiment is described.
[0055] In a realizable implementation manner, the local feature learning module performing local feature learning on the sea surface temperature image sequence stored on the local node includes: constructing the sampling points of the sea surface temperature image into graph nodes, each sampling point containing longitude and latitude coordinates and temperature values, generating an adjacency matrix based on the Euclidean distance between the sampling points. The adjacency matrix is expressed as:
[0056] ;
[0057] where represents the Euclidean distance between sampling points i and j, represents the threshold distance. When the distance between two sampling points is less than or equal to the threshold distance, the edge is connected. When the distance between two sampling points is greater than the threshold distance, the edge is not connected.
[0058] Specifically, for any local node i, the SST image sequence it owns is represented as , for any one of the images , the resolution is k k pixels, and each pixel contains longitude and latitude information and sea surface temperature information. On the image , sampling is performed every s kilometers in the horizontal direction and every m kilometers in the vertical direction, then there are a total of sampling points. The sampling points are regarded as the nodes of the graph . represents the Euclidean distance between sampling points i and j. Set the threshold distance . When the Euclidean distance is greater than or equal to the threshold distance, it is considered that the two points are connected; otherwise, they are not connected. Therefore, the adjacency matrix is defined.
[0059] Please refer to Figure 2 and Figure 3 . In an implementable embodiment, the local feature learning module performs local feature learning on the sea surface temperature image sequence stored on the local node, specifically including:
[0060] Feature extraction is performed through the distributed graph convolution module, including on one of the local nodes, the sea surface temperature image features at time t and the adjacency matrix A are used as the input of the distributed graph convolution module, and the output of local node i is obtained through GCN. Define the network weight as W, and GCN is expressed as:
[0061] ;
[0062] Among them, the initial input , represents the feature representation of the sea surface temperature image at time t on local node i. Define , is the identity matrix, D is the degree matrix, , is the activation function. The model gradient of one local node i is defined as ;
[0063] The temporal feature learning module is used to perform operations, and the output of the distributed graph convolution module is learned through the LSTM network;
[0064] The output of the temporal feature learning module is predicted using the fully connected layer, and the prediction result of the local node is output.
[0065] It should be noted that is the sequence of feature representations of the temporal sea surface temperature image on local node i, Represents the feature representation of the SST image at time t on the local node i (where i is an integer greater than or equal to 1). Represents the number of the graph node on the SST image, and A is the adjacency matrix of the graph.
[0066] Please refer to Figure 3 and Figure 5 , in a realizable embodiment, the cross-scale joint learning module learns the SST image features of different scales on different local nodes through the graph learning module, including: the output of the distributed graph convolution module learns the weights during distributed gradient aggregation through the graph learning module, and cross-scale joint learning is performed according to the value returned by the graph learning module to optimize the network weights. The graph learning module is expressed as:
[0067] ;
[0068] Wherein, Represents the output of the graph convolution network from the local node i. Represents the F-norm of the matrix.
[0069] In another realizable embodiment, the cross-scale joint learning module dynamically optimizes the communication weights between each local node, including: optimizing the communication weight matrix through the graph learning module. The communication weight matrix is used to guide the aggregation parameters of each local node during gradient aggregation. The aggregation parameters include the SST image features of different scales. The distributed gradient aggregation is expressed as:
[0070] ;
[0071] Wherein, Represents the set of nodes that can communicate with the local node i. Represents the gradient of the node j that communicates with the local node i. Represents the communication weight between the local nodes i and j. Represents the gradient of the local node i after distributed aggregation.
[0072] It should be noted that in this embodiment, by optimizing the communication weight matrix, the feature information of different scales of each node is utilized, the information in different scale spaces is jointly learned, the optimized communication weight matrix is used to guide the cross-scale joint learning, and by calculating the similarity of different scale feature spaces, the parameters before the gradient in the distributed gradient update process are calculated, rather than simply taking the average directly, so as to obtain better model performance.
[0073] Specifically, in the local system, during the model initialization phase, the initial input feature representation of the local node i at time step t is where is the original feature vector of the node, and l represents the number of layers of the neural network. The initial feature representation and the adjacency matrix A are input into the graph convolutional layer GCN to obtain the output of the local node i . The output feature of the local node i will pass through the distributed system; in the distributed system, the graph learning module (GL) receives the output features from different local nodes , and calculates the optimized communication weight matrix according to this value , and guides the distributed gradient aggregation according to the values in the communication weight matrix ; then, according to the distributed graph convolution (D-GCNAvg) strategy, the gradient aggregation of the local nodes is performed, and the optimized gradient value is sent back to the local system, and the local system optimizes the network weights according to the updated gradient value.
[0074] After that, the output feature after passing through the double-layer GCN will enter the temporal feature learning module. There are multiple cascaded LSTM structures in the temporal feature learning module. The output features at different time steps will enter different LSTMs in sequence, and the final output feature will be obtained after the layer-by-layer propagation of the LSTM network , and the final prediction result can be obtained after passing this value through the fully connected layer .
[0075] Please refer to Figure 2 , Figure 3 and Figure 4 , in practical applications, the output features obtained through the distributed graph convolution module will be used as the input of the temporal feature learning module. The temporal feature learning module uses an LSTM network to learn the temporal information of the input features. The operations performed by the LSTM network include forget gate calculation, input gate calculation, candidate cell state generation, cell state update, output gate calculation, and hidden state update.
[0076] Specifically, when the output of the distributed graph convolution module enters the LSTM network, the input data first passes through the forget gate. The forget gate determines which information should be forgotten or retained. At each time step t, the calculation formula of the forget gate is as follows:
[0077] ;
[0078] where is the sigmod function, where is the weight matrix of the forget gate, is the bias term, represents the input at the current time step, is the hidden state at the previous time step, and the output of the forget gate is a value between 0 and 1, indicating how much past information should be forgotten. When is close to 1, past information is completely retained. When is close to 0, past information is completely forgotten.
[0079] When new input enters the LSTM network, the input gate determines which information should be retained and updates the cell state. At each time step t, the calculation formula for the input gate is as follows:
[0080] ;
[0081] where is the weight matrix of the input gate, is the bias term, represents the input at the current time step, is the hidden state of the previous time step. Next, the LSTM calculates the candidate cell state , which indicates how much the new input at the current time step can affect the cell state. The calculation formula for the candidate cell state is as follows:
[0082] where is the weight matrix of the input gate, is the bias term, represents the input at the current time step, is the hidden state of the previous time step, is the hyperbolic tangent function, used to map the input value to between -1 and 1. The role of the input gate is to control the weight of the new input at the current time step. Through the input gate, the LSTM can better process long sequence data, avoid the problems of gradient vanishing and gradient explosion, and thus improve the effect and stability of the model.
[0083] The cell state of the LSTM will be updated and passed to the next time step. At each time step t, the update formula for the cell state is as follows:
[0084] ;
[0085] where is the forget gate, indicating the weight for forgetting the cell state, is the cell state of the previous time step, is the input gate, indicating the weight for updating the cell state, is the candidate cell state at the current time step, indicating how much the new input at the current time step can affect the cell state. During the training process, the LSTM network can adaptively update the cell state through the learned weights, retaining and passing important information.
[0086] When it is necessary to transfer the information of the current time step to the next layer or the output layer, an output gate is required to control which information should be output. The calculation formula of the output gate is as follows:
[0087] ;
[0088] where, is the weight matrix of the output gate, is the bias term, represents the input of the current time step, is the hidden state of the previous time step. After that, the LSTM will process the cell state through a function to obtain the hidden state of the current time step. The specific calculation formula is as follows:
[0089] ;
[0090] Through the output gate, the LSTM can adaptively control the output of information according to the cell state and the hidden state of the current time step. In this way, in the LSTM network, important information can be automatically screened out and transferred to the next layer or the output layer.
[0091] In an implementable embodiment, the output of the temporal feature learning module is predicted through a fully connected layer. The fully connected layer consists of a multi-layer perceptron with two hidden layers, which maps the hidden state output by the LSTM network to the predicted sea surface temperature at future times.
[0092] In an implementable embodiment, the loss function is expressed as:
[0093] ;
[0094] where, is the prediction loss, is the loss of the cross-scale distributed joint learning module, and the parameter is a trade-off parameter;
[0095] The prediction loss is expressed as:
[0096] ;
[0097] where, represents the prediction result on the local node i, represents the true value;
[0098] The loss of the cross-scale distributed joint learning module is expressed as:
[0099] ;
[0100] Among them, represents the output of the graph convolutional network on the local node i, represents the result of the graph learning module, is a trade-off parameter.
[0101] Based on the same inventive concept, the present application further provides a sea surface temperature prediction system based on distributed cross-scale joint learning. The prediction system includes a sea surface temperature prediction model based on distributed cross-scale joint learning. The prediction model adopts a local system and a distributed system architecture. The local system includes several local nodes. The communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether communication is possible between local nodes. Different but identically structured sea surface temperature image sequences from different sources are distributed on each local node as the input of the prediction model. The prediction model includes:
[0102] A local feature learning module, deployed on each local node in the local system, for performing local feature learning on the sea surface temperature image sequences stored on the local node and outputting the sea surface temperature prediction result of the local node;
[0103] A cross-scale joint learning module, deployed in the distributed system, learns the different-scale sea surface temperature image features on different local nodes through a graph learning module, dynamically optimizes the communication weights between each local node, realizes the gradient aggregation of different local nodes, obtains the global gradient update, distributes the updated gradient values to each local node, and guides the cross-scale joint learning.
[0104] The above are only the preferred embodiments of the present application, and do not impose any form of limitation on the present application. Although the present application has been disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content within the scope of the technical solution of the present application to obtain equivalent embodiments with equivalent changes. The implementation schemes in the above embodiments can also be further combined or replaced. However, as long as the content does not depart from the technical solution of the present application, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application still belong to the scope of the present application's solution.
Claims
1. A sea surface temperature prediction method based on distributed cross-scale joint learning, characterized in that: A sea surface temperature prediction model based on distributed cross-scale joint learning is constructed. The prediction model adopts a local system and a distributed system architecture. The local system includes a number of local nodes. The communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether the local nodes can communicate. Each local node is distributed with a sea surface temperature image sequence from different sources but with the same data structure as the input of the prediction model. The prediction model includes: A local feature learning module, deployed on each local node in the local system, for performing local feature learning on a sea surface temperature image sequence stored on the local node and outputting a sea surface temperature prediction result of the local node; The cross-scale joint learning module is deployed in the distributed system, and learns the sea surface temperature image features of different scales on different local nodes through the graph learning module, dynamically optimizes the communication weight between each local node, realizes the gradient aggregation of different local nodes, obtains the global gradient update, and distributes the updated gradient value to each local node to guide the cross-scale joint learning; the cross-scale joint learning module learns the sea surface temperature image features of different scales on different local nodes through the graph learning module, including: the output of the distributed graph convolution module learns the communication weight during the distributed gradient aggregation through the graph learning module, and performs cross-scale joint learning according to the value returned by the graph learning module to optimize the network weight. The graph learning module is expressed as: in, represents the graph convolutional network output from local node i, l represents the number of layers of the neural network, ||.|| F represents the F norm of the matrix, ReLU represents the activation function, Θ represents the communication weight matrix, G ij Represents the results of the graph learning module; The cross-scale joint learning module dynamically optimizes the communication weight between each local node, including: optimizing the communication weight matrix θ through the graph learning module ij , the communication weight matrix Θ ij Aggregation parameters used to guide each local node in performing gradient aggregation, the aggregation parameters include sea surface temperature image features at different scales, and distributed gradient aggregation is expressed as: in, represents the set of nodes that can communicate with the local node i, represents the gradient of node j that communicates with local node i, Θ ij represents the communication weight of local nodes i and j, Represents the gradient of the local node i after distributed aggregation.
2. The method according to claim 1, characterized in that The local feature learning module performs local feature learning on the sea surface temperature image sequence stored on the local node, including: constructing the sampling points of the sea surface temperature image as graph nodes, each sampling point includes latitude and longitude coordinates and temperature values, and generating an adjacency matrix based on the Euclidean distance between the sampling points. The adjacency matrix is expressed as: Among them, d ij represents the Euclidean distance between sampling points i and j, d min It is expressed as a threshold distance. When the distance between two sampling points is less than or equal to the threshold distance, the edge is connected. When the distance between two sampling points is greater than the threshold distance, the edge is not connected.
3. The method according to claim 2, characterized in that The local feature learning module performs local feature learning on the sea surface temperature image sequence stored on the local node, specifically including: Feature extraction is performed through a distributed graph convolution module, including the sea surface temperature image feature X at time t on one of the local nodes. i,t and the adjacency matrix A as the input of the distributed graph convolution module, and the output of the local node i is obtained through GCN. Define the network weight as W, and the GCN is expressed as: Among them, the initial input X i,t Represents the feature representation of the sea surface temperature image at time t on the local node i, and defines I is the identity matrix, D is the degree matrix, the degree matrix is a diagonal matrix, and the diagonal elements D ii =∑ j A ij ,∑ j A ij is the sum of the elements in the i-th row or i-th column of the adjacency matrix A, σ is the activation function, and the model gradient of a local node i is defined as The operation is performed using a time series feature learning module, and the output of the distributed graph convolution module is learned through an LSTM network; the output of the time series feature learning module is predicted using a fully connected layer, and the prediction result of the local node is output.
4. The method according to claim 3, characterized in that Output features obtained by the distributed graph convolution module It will be used as the input of the temporal feature learning module, and the temporal feature learning module adopts the LSTM network to learn the temporal information of the input features. The LSTM network performs operations including forget gate calculation, input gate calculation, candidate cell state generation, cell state update, output gate calculation and hidden state update.
5. The method according to claim 4, characterized in that The output of the time series feature learning module is predicted by a fully connected layer, which is composed of a multi-layer perceptron with two hidden layers. The hidden state of the LSTM network output is mapped to the sea surface temperature prediction y at the future moment. i,p .
6. The method according to claim 1, characterized in that The loss function is expressed as: THE 总 =L P +λL G ; Among them, L P To predict the loss, L G is the loss of the cross-scale distributed joint learning module, and the parameter λ ≥ 0 is a trade-off parameter; the prediction loss L P It is expressed as: Among them, y i,p Represents the prediction result on the local node i, y t represents the truth value; The loss L of the cross-scale distributed joint learning module G It is expressed as: Among them, H i represents the graph convolutional network output on the local node i, G ij represents the result of the graph learning module, and γ is a trade-off parameter.
7. The sea surface temperature prediction system based on distributed cross-scale joint learning is characterized by: The prediction system includes a sea surface temperature prediction model based on distributed cross-scale joint learning. The prediction model adopts a local system and a distributed system architecture. The local system includes a plurality of local nodes. The communication relationship between each local node is defined by a connection matrix. The matrix elements contain 0 or 1 to indicate whether the local nodes can communicate. Each local node is distributed with a sea surface temperature image sequence from different sources but with the same data structure as the input of the prediction model. The prediction model includes: A local feature learning module, deployed on each local node in the local system, for performing local feature learning on a sea surface temperature image sequence stored on the local node and outputting a sea surface temperature prediction result of the local node; The cross-scale joint learning module is deployed in the distributed system, and learns the sea surface temperature image features of different scales on different local nodes through the graph learning module, dynamically optimizes the communication weight between each local node, realizes the gradient aggregation of different local nodes, obtains the global gradient update, and distributes the updated gradient value to each local node to guide the cross-scale joint learning; the cross-scale joint learning module learns the sea surface temperature image features of different scales on different local nodes through the graph learning module, including: the output of the distributed graph convolution module learns the communication weight during the distributed gradient aggregation through the graph learning module, and performs cross-scale joint learning according to the value returned by the graph learning module to optimize the network weight. The graph learning module is expressed as: in, represents the graph convolutional network output from local node i, l represents the number of layers of the neural network, ||.|| F represents the F norm of the matrix, ReLU represents the activation function, Θ represents the communication weight matrix, G ij Represents the results of the graph learning module; The cross-scale joint learning module dynamically optimizes the communication weight between each local node, including: optimizing the communication weight matrix θ through the graph learning module ij , the communication weight matrix Θ ij Aggregation parameters used to guide each local node in performing gradient aggregation, the aggregation parameters include sea surface temperature image features at different scales, and distributed gradient aggregation is expressed as: in, represents the set of nodes that can communicate with the local node i, represents the gradient of node j that communicates with local node i, Θ ij represents the communication weight of local nodes i and j, Represents the gradient of the local node i after distributed aggregation.
Citation Information
Patent Citations
Regional sea surface temperature forecasting method and system
CN118333091A