Meteorological service recommendation method and system of edge side equipment small recurrent neural network based on super network
By using hypernetwork to generate small recurrent neural network models on edge computing devices, the problem of high model update cost on edge computing devices is solved, and efficient and low-cost model updates and customized meteorological service recommendations are achieved.
Patent Information
- Application Number
- CN202510284095.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
On edge computing devices, how to trade off between strict computing resource constraints and high model deployment costs to achieve efficient and low-cost model updates, especially in scenarios where data distribution changes rapidly.
The meteorological service recommendation method of a small recurrent neural network of edge-side devices based on hypernetwork is adopted. By deploying a common pre-trained model and hypernetwork in the cloud, the representation of historical data is extracted, and the unique representation of user's local recent meteorological data is learned by fine-tuning hypernetwork, and the optimal parameters are generated and passed to the edge-side devices are generated to realize customized meteorological service recommendations.
It realizes efficient and low-cost model iteration under the premise of ensuring prediction accuracy, and is suitable for scenarios where data distribution changes rapidly, reducing the cost of model deployment and update, and improving the robustness and adaptability of the model.
Smart Images

Figure CN120216765A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of meteorological service recommendation, and relates to a meteorological service recommendation method and system based on a small recurrent neural network of edge devices with a hypernetwork. Background Art
[0002] With the rapid development of the Internet of Things (IoT) and 5G communication technologies, the penetration of artificial intelligence (AI) application scenarios into various industries has gradually deepened. As an emerging computing mode, edge computing has become an important development trend. Edge computing distributes computing tasks to devices closer to the data source. For example, lightweight models are deployed on portable devices such as watches and tablets to complete real-time tasks, rather than sending tasks to the cloud for centralized processing by cloud computing. This greatly reduces data transmission latency and improves real-time performance and system response capabilities. In addition, compared with traditional cloud computing, edge computing has advantages such as lower latency, higher bandwidth utilization, and better privacy protection. However, this also brings some problems: edge computing devices often have limited memory, and the model deployment cost is also higher.
[0003] In current deep learning, models usually adopt an end-to-end training strategy, that is, training on historical data, validating on a part of the data, and then applying it to actual data. This relies on an assumption that historical data and application data follow the same or similar distributions. However, in practical applications, data distributions often change over time. Especially in time series prediction scenarios, due to the non-stationarity of the data, it often changes over time, resulting in a significant distribution shift. This distribution shift phenomenon will lead to a significant difference between the distributions of the training data and application data of the model, and it becomes more obvious as time goes by, thus reducing the practical value of the model. This distribution shift also exists in the input sequence of the training data, making it challenging to train a model that can generalize well to unknown data. Figure 1This paper describes the problem of distribution shift in time series prediction. The probability distributions of time series may vary across different periods and are likely to be different from those in unseen prediction windows. To maintain the practicality of the model, traditional solutions require the model to be updated and adjusted frequently to adapt to new data distributions. This high-frequency update requirement leads to high retraining costs. Or, by increasing the number of model parameters to accommodate more historical knowledge, however, the models trained by this approach often require a large amount of memory and consume more computing resources during inference. Domain generalization or domain adaptation aims to use small models to ensure stable performance when distributions change. The idea is to learn the common knowledge that can be transferred between different domains, despite the differences in their distributions. This method enhances the flexibility of the model by automatically adjusting the learning rate, structure, or parameters of the model according to the changes in data distributions, enabling it to quickly adapt to new data patterns. However, since the actual data is not visible to the training data, the model often lacks sufficient prior knowledge to characterize these distributions and thus learn the general knowledge of different distributions. At the same time, only retaining the identity between different distributions while discarding the differences in these distributions will lead to insufficient utilization of the information of a specific distribution, resulting in a decrease in prediction accuracy.
[0004] Therefore, how to balance the severe computational resource constraints and high model deployment costs on edge computing devices to achieve efficient and low-cost model updates has become a key challenge. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a meteorological service recommendation method based on a small recurrent neural network for edge devices with a hypernetwork;
[0006] The present invention also provides a meteorological service recommendation system based on a small recurrent neural network for edge devices with a hypernetwork.
[0007] Specifically, considering the meteorological monitoring field (such as the Air Quality dataset, Weather dataset), the present invention deploys a general pre-trained model and a hypernetwork in the cloud. When in use, the general pre-trained model is used to extract the representations of historical data, such as the data collected by meteorological monitoring sensors near individual users, to provide stable prior knowledge support for the hypernetwork. On this basis, the hypernetwork can be fine-tuned to further learn the unique representations of the recent local meteorological data of users by combining the prior knowledge provided by the pre-trained model. Based on these unique representations of the data, the hypernetwork can generate optimal parameters for a small main network (recurrent neural network) based on the current data distribution, and transfer the optimal parameters to the edge device, thereby providing customized meteorological advice services for individual users.
[0008] The advantages of the present invention lie in comprehensively utilizing the general features learned by the large model during the pre-training period and the specific data distribution representations learned through the dynamic adjustment of the hypernetwork, thereby realizing the deployment strategy for specific scenarios. This can be regarded as distilling the prior knowledge of the large model into a small model for specific scenarios, achieving model compression. The application of the hypernetwork provides a parameter adjustment strategy without training for the models of edge devices in scenarios where the data distribution changes rapidly: updating the parameters of the small recurrent neural network on the cloud as needed and sending them to the edge device, thereby realizing efficient and low-cost model iteration while ensuring accuracy.
[0009] Term Explanation:
[0010] 1. Pre-trained model: It refers to a deep learning model that is initially trained on a large-scale dataset to learn the general information of the field and can then be further optimized on specific tasks through fine-tuning or transfer learning. This method can significantly reduce the training time, improve the generalization ability of the model, and avoid overfitting. Pre-trained models often have rich prior knowledge and the ability to perform well on multiple tasks. Large models are typical pre-trained models.
[0011] 2. Hypernetwork, HyperNetwork: It is a neural network used to generate or adjust the weights of the target network. Essentially, it is a function that learns the weight distribution of the model. Different from traditional neural networks, the hypernetwork does not directly perform task prediction but serves as a "weight generator" for dynamically adjusting or generating the parameters of another neural network.
[0012] 3. Visualization tool torchviz: It is a computational graph visualization tool based on PyTorch. It can generate the PyTorch computational graph through the automatic differentiation mechanism of Pytorch and graphically display the computational dependency relationships of tensors, thereby helping researchers understand complex deep learning models.
[0013] 4. Pydot: It is mainly used to generate and operate directed and undirected graphs. It supports the.dot file format, can convert structured data such as computational graphs and flowcharts into graphical representations, and can be exported in formats such as PNG, PDF, and SVG. Pydot is often used in combination with deep learning visualization tools (such as Torchviz and Keras visualization tools) to intuitively display the computational graph structure and dependency relationships.
[0014] 5. Networkx: A Python library for complex network analysis, providing various data structures (such as directed graphs, undirected graphs, multigraphs) and graph theory algorithms (such as shortest paths, connected components, community detection, etc.). In deep learning and graph neural networks (GNNs), NetworkX can be used to construct graph datasets, analyze graph structures, and preprocess node and edge features.
[0015] 6. Standard RNN: A neural network used to process sequential data, such as time series, speech, text, etc. It is characterized by having memory capabilities and can pass past information into the current calculation, so it is particularly suitable for handling context-related tasks. RNN stores historical information through "recurrent connections", but is prone to the problem of vanishing gradients in long-sequence tasks, making it difficult for the model to learn long-distance information.
[0016] 7. Standard LSTM: An improvement over RNN, which adds a gating mechanism to more effectively control the storage and forgetting of information, thus better capturing long-term dependencies. LSTM mainly consists of a forget gate, an input gate, and an output gate, which can intelligently select which information needs to be remembered and which can be discarded. Compared with the standard RNN, LSTM is better at handling data with long-term dependencies such as long texts and video frame sequences.
[0017] 8. Standard GRU: A simplified version of LSTM that removes some additional control gates, making the calculation more efficient while still having good memory capabilities. GRU controls the transfer of information through an update gate and a reset gate. Similar to LSTM, but with fewer parameters and faster calculation speed. It usually performs better in scenarios with limited computing resources (such as mobile devices, edge computing), and can achieve results close to LSTM in many tasks.
[0018] 9. Residual connection: In a neural network, usually the input of a certain previous layer is directly skipped over some intermediate layers and connected to the subsequent layer. This can make the network easier to train, solve problems such as vanishing gradients that occur as the number of network layers increases, enable the network to learn more complex features, and also converge faster.
[0019] 10. Layer normalization: Normalize the input data of a certain layer in a neural network. It adjusts the distribution of the input data of this layer to a standard distribution with a mean of 0 and a variance of 1. This method can accelerate the training speed of the network, reduce the problems of vanishing or exploding gradients, and also make the model have better adaptability to different input data, improving the generalization ability of the model.
[0020] The technical solution of the present invention is as follows:
[0021] A weather service recommendation method for a small recurrent neural network on edge devices based on a hypernetwork, including:
[0022] Deploy a general pre-trained model and a hypernetwork in the cloud;
[0023] During use, use the general pre-trained model to extract the representation of historical data. The hypernetwork, through fine-tuning and combining the representation of historical data provided by the pre-trained model, further learns the unique representation of the user's local recent meteorological data;
[0024] Based on the unique representation of the user's local recent meteorological data, the hypernetwork generates optimal parameters for the small main network based on the current data distribution, and transfers the optimal parameters to the edge device to provide customized weather services for individual users.
[0025] Preferably according to the present invention, use a general pre-trained model to extract the representation of historical data; including:
[0026] 1) Preprocessing: Preprocess the historical meteorological data of the i-th cycle Perform preprocessing, where L h represents the input length of the pre-trained model, D represents the number of different meteorological information recorded by different meteorological sensors at each time point, and the multivariate time series data is regarded as multiple independent univariate time series for processing. Then X h i will be represented as D univariate time series x, X h i =[x j j=1,…,D where x j represents the data recorded by the j-th sensor;
[0027] 2) Normalization: Fill in missing values and normalize each x j which is expressed as: where μ and σ are the mean and standard deviation of each univariate sequence respectively;
[0028] 3) Blocking: Divide the input of the j-th channel along the time dimension, and convert it into a series of time blocks patch, expressed as x p j , and choose the overlapping or non-overlapping method; assume the patch length is Q, and the non-overlapping area between two adjacent patches, that is, the step size is S, then the number of generated time blocks is
[0029] 4) Mapping: The blocked input is mapped to a dimension of d through a trainable linear projection model in a high-dimensional vector space, and apply learnable additive positional encoding to introduce the sequential information of the time series and obtain will pass through multiple stacked Transformer encoding layers;
[0030] 5) For each Transformer encoding layer;
[0031] First, through three transformation matrices W Q 、W K 、W V to be transformed into query Q, key K, and value V, Generally, H Q = H K = H V , M represents the number of heads applied in the multi-head attention mechanism, and H Q is generally an integer multiple of M.
[0032] Secondly, query Q, key K, and value V will be sliced into M attention heads with the same shape. The i-th attention head contains different queries Q i 、keys K i and values V i ; For the i-th attention head, the calculation process is as follows:
[0033]
[0034] Again, after the attention mechanism, residual connection and layer normalization operations are performed;
[0035] Finally, after passing through the feed-forward neural network, residual connection and layer normalization are performed again;
[0036] 6) After passing through multiple stacked Transformer encoding layers, the outputs of each channel are concatenated in the feature dimension, and the result is represented as where D represents the feature dimension of the data, P represents the size of the sequence time dimension, and d model represents the dimension of the latent space where the representation is located; take the information of the last time block in the time dimension as the input of the super network;
[0037] Based on the unique representation of the user's local recent meteorological data, the super network generates optimal parameters for the small main network based on the current data distribution
[0038] According to the preference of the present invention, the super network generates optimal parameters for the small main network based on the current data distribution; including:
[0039] The hypernetwork will adaptively generate optimal parameters for the main network based on When constructing the main network, performing a forward pass using the main network to generate a dynamic computational graph for the corresponding model means: using the dynamic computational graph visualization tool torchviz to obtain a computational graph in DOT format that contains all the operation nodes, tensor nodes, and their connections of the main network;
[0040] Using pydot and networkx to convert the computational graph in DOT format of the operation nodes, tensor nodes, and their connections into an operable graph object, relabeling the nodes, filtering and deleting unnecessary intermediate computational nodes, retaining important dependencies, integrating the dependencies of the nodes, and obtaining the parameter matrix names and corresponding parameter matrix shapes that the main network needs to generate;
[0041] Assume the main network is of GRU structure, and the weight generation process is as follows:
[0042] Through reading the computational graph, the GRU structure includes the following 6 nodes: GRU.weight_ih_l0 represents the weight matrix of the GRU kernel from input to hidden state; GRU.weight_hh_l0 represents the weight matrix of the GRU kernel from hidden state to hidden state; GRU.bias_ih_l0 is the bias vector of the GRU kernel from input to hidden state; GRU.bias_hh_l0 is the bias vector of the GRU kernel from hidden state to hidden state; fc.weight is the weight matrix of the feedforward layer; fc.bias is the bias vector of the feedforward layer. According to the shapes of the parameter matrices of different nodes, the hypernetwork generates C groups of different preselected parameter key-value pairs for the 6 nodes of the GRU structure during initialization
[0043] Each group of preselected parameters will optimize the specific representation of the data from different perspectives, and C can be selected according to the specific application scenario:
[0044] Using a fully connected matrix to perform mapping, and generate queries after layer normalization and activation functions which are used as keys and values respectively, and calculate using the standard attention mechanism, and the output is the most adaptable model parameters of the main network B in the i-th cycle i The optimal parameters refer to generating the most matching or adaptable parameters for the data of the current cycle, and the implementation process is described as:
[0045]
[0046]
[0047] Among them, FFN represents the feedforward network layer.
[0048] Preferably according to the present invention, a dual loss function is adopted as follows:
[0049]
[0050] Use configured B i to calculate the loss At the same time, select the candidate parameter group with the largest attention score configured B i to calculate the additional loss Among them, the maximum attention score is obtained from calculation, c * represents the index corresponding to the maximum attention score, and the candidate parameter group selected according to this index is Take as the parameters of the main network for forward propagation; λ represents the weighted parameter.
[0051] Further preferably, use and configured B i to calculate the losses and respectively, including:
[0052] the inputs of all windows in the current period share the same configured B i for forward propagation to obtain the predicted values for future time series data Use the MSE loss to measure the difference between the predicted values of each window and the standard value Y w in the i-th period That is:
[0053]
[0054] Further preferably, the architecture of the main network is optional and includes a standard RNN, LSTM, GRU.
[0055] A computer device includes a memory and a processor. When the processor executes the computer program, the steps of the meteorological service recommendation method based on the small recurrent neural network of the edge device based on the hypernetwork are implemented.
[0056] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of a meteorological service recommendation method based on a small recurrent neural network of edge devices with a hypernetwork.
[0057] A meteorological service recommendation system based on a small recurrent neural network of edge devices with a hypernetwork includes:
[0058] A model deployment and training module is configured to: deploy a general pre-trained model and a hypernetwork in the cloud; during use, use the general pre-trained model to extract the representation of historical data, and the hypernetwork, through fine-tuning, combines the representation of historical data provided by the pre-trained model to further learn the specific representation of the user's local recent meteorological data;
[0059] A service recommendation module is configured to: based on the specific representation of the user's local recent meteorological data, the hypernetwork generates optimal parameters for a small main network based on the current data distribution, and transmits the optimal parameters to the edge device to provide customized meteorological services for individual users.
[0060] The beneficial effects of the present invention are:
[0061] By combining the rich prior knowledge provided by a large time series prediction pre-trained model and the scenario-specific expertise provided by the hypernetwork, the present invention effectively improves the prediction accuracy of the generated main network model, and is particularly suitable for complex and real-world application scenarios that require quick responses.
[0062] 1. The hypernetwork generation scheme of the present invention can be applied to various different pre-trained models and main network models, and has strong universality. It can update a model with higher accuracy at a small cost for edge devices without frequent training (even without training).
[0063] 2. Through the interaction between the cloud and edge devices, the present invention can determine on demand the time to customize and update the model, rather than at a fixed frequency.
[0064] 3. The applicability of the present invention in multiple datasets and multiple domains ensures its wide application in complex real-world time series tasks, improves the robustness of the model, reduces the training cost and deployment cost of the model, can provide assistance for actual life and production scenarios, and has good economic value and social value. Description of the Drawings
[0065] Figure 1 It is a schematic diagram of the distribution shift problem in time series prediction;
[0066] Figure 2 It is an implementation framework diagram of a meteorological service recommendation method based on a small recurrent neural network of edge devices with a hypernetwork;
[0067] Figure 3 Schematic diagram of loss change trained for the present invention;
[0068] Figure 4(a) Schematic diagram of the visualization result of the present invention on the subset Dingling of the air quality dataset;
[0069] Figure 4(b) Schematic diagram of the visualization result of the present invention on the subset Nongzhanguan of the air quality dataset;
[0070] Figure 5 Schematic diagram of the connection architecture of the cloud and edge devices in the present invention;
[0071] Figure 6 Is the architecture diagram of the Transformer encoding layer. Detailed implementation manners
[0072] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.
[0073] Embodiment 1
[0074] A meteorological service recommendation method for a small recurrent neural network of edge devices based on a hypernetwork, as Figure 2 shown, includes:
[0075] Considering in the field of meteorological monitoring (such as Air Quality dataset, Weather dataset),
[0076] Deploy a general pre-trained model and a hypernetwork in the cloud;
[0077] When in use, a general pre-trained model is used to extract the representation of historical data, such as the data collected by meteorological monitoring sensors near individual users, to provide stable prior knowledge support for the hypernetwork. On this basis, the hypernetwork further learns the unique representation of the user's local recent meteorological data by fine-tuning in combination with the representation of historical data provided by the pre-trained model;
[0078] Based on the unique representation of the user's local recent meteorological data, the hypernetwork generates optimal parameters for a small main network (recurrent neural network) based on the current data distribution, and transmits the optimal parameters to the edge device to provide customized meteorological services for individual users.
[0079] such as Figure 5As shown, a large amount of user-related data collected from different data sources (including sensors, etc.) is uploaded to the cloud for storage. For example, information such as temperature and humidity recorded by sensors near the user. These data are integrated according to time points and fed into a frozen pre-trained model to extract features. The hypernetwork uses these features as input to generate optimal model parameters for the edge device, and sends the parameters to the user's edge device. The model with a fixed architecture built into the edge device automatically deploys to provide data analysis and suggestions for the user in the next period of time.
[0080] Considering the actual application in meteorological prediction, the deployment frequency of edge devices should not be too high, that is, the hypernetwork generates a model that can be used in the next period of time. The available time of the model is called a cycle. For the i-th model B i , it should perform well on the T time steps of data contained in the i-th cycle. Meeting the application scenario of real-time time series prediction, when processing complete time series data , it is divided into M non-overlapping subsequences of length T in a non-overlapping manner, where represents the ceiling operation, and N is the length of the original time series. Each subsequence is denoted as representing the data in the i-th cycle, where represents the time series data of the i-th non-overlapping cycle. It will be organized in a sliding window manner. Set the window size to W, then there is where X w is the input data in the w-th window of the i-th cycle, Y w is the corresponding target output or label data in the window. The main network of the edge device accurately predicts Y w based on X w ;
[0081] Represent the historical meteorological data as X h i ; As Figure 2 shown, the present invention uses a frozen large Transformer as a feature extractor, referring to self-supervised training schemes such as PatchTST. This model has been pre-trained on other data. Receive the historical input X h i to extract feature representations. The pre-trained model only participates in inference and does not train it or update its parameters. This is because sufficient prior knowledge has been injected into the pre-trained model during the pre-training stage. In the present invention, the pre-trained model is only used as a pluggable tool. The hypernetwork will generate main network parameters specific to based on these feature representations. The main network is applicable to all windows in the i-th cycle
[0082] Example 2
[0083] The meteorological service recommendation method of the small recurrent neural network for edge devices based on a hypernetwork according to Example 1 is characterized in that:
[0084] A general pre-trained model is used to extract the representation of historical data, including:
[0085] 1) Preprocessing: Preprocess the historical meteorological data of the i-th cycle where L h represents the input length of the pre-trained model, and D represents the number of different meteorological information recorded by different meteorological sensors at each time point, such as carbon dioxide, temperature, humidity, etc. Treat the multivariate time series data as multiple independent univariate time series for processing, then X h i will be represented as D univariate time series x, X h i =[x j j=1,…,D where x j represents the data recorded by the j-th sensor;
[0086] 2) Normalization: Fill in the missing values and normalize each x j which is expressed as: where μ and σ are the mean and standard deviation of each univariate sequence respectively; these statistics will be retained for restoring the output value of the model to the original scale.
[0087] 3) Blocking: Split the input of the j-th channel in the time dimension and convert it into a series of time blocks patch, denoted as x p j , and an overlapping or non-overlapping method is selected; assuming the patch length is Q and the non-overlapping region between two adjacent patches, i.e., the stride is S, then the number of generated time blocks is Through the blocking mechanism, the input sequence length is reduced from the original sequence length L to approximately L / S, and the time and space complexity are reduced from O(L h 2 ) to O((L h / S) 2 ), thus significantly reducing the computational complexity of the attention mechanism. Under the premise of limited computing resources, the model will have a longer effective context length.
[0088] 4) Mapping: The blocked input passes through a trainable linear projection Map to a high-dimensional vector space of dimension d model and apply learnable additive positional encoding to introduce the sequential information of the time series and obtain Pass through multiple stacked Transformer encoding layers;
[0089] 5) For each Transformer encoding layer; Figure 6 Is the architecture diagram of the Transformer encoding layer;
[0090] First of all, Convert to query Q, key K and value V through three transformation matrices W Q , W K , W V , and generally, H =H Q =H K =H V , M represents the number of heads applied in the multi-head attention mechanism, and H Q is generally an integer multiple of M.
[0091] Secondly, query Q, key K and value V will be sliced into M attention heads of the same shape. The i-th attention head contains different queries Q i , keys K i and values V i ; In the study of Vaswani et al. (2017), the standard multi-head self-attention mechanism is defined based on these tuple inputs. For the i-th attention head, the calculation process is as follows:
[0092]
[0093] Thirdly, after the attention mechanism, perform residual connection and layer normalization operations;
[0094] Finally, after passing through the feed-forward neural network, perform residual connection and layer normalization again;
[0095] Overall, the whole process can be described by the following formula, MHA P (x d j ) represents applying the multi-head attention mechanism to x d j , x d j +MHA P (x d j) represents the residual connection operation. The output of the residual connection is subjected to layer normalization (LayerNorm) to obtain the output X of the attention part. j , respectively represent two feed-forward neural networks. d ff is the dimension of the middle layer of the feed-forward network. ReLU represents the activation function used to introduce non-linearity. By using two feed-forward neural networks and the activation function, the model's ability to fit data is improved, thereby outputting richer representation information. The output Z of the k-th layer j will be used as the input of the (k + 1)-th layer:
[0096] X j = LayerNorm(x d j + MHA P (x d j ));
[0097] Z j = LayerNorm(X j + FFN2(ReLU(FFN1(X j ))));
[0098] 6) After passing through multiple stacked Transformer encoding layers, the outputs of each channel are concatenated in the feature dimension, and the result is represented as where D represents the feature dimension of the data, P represents the size of the sequence time dimension, and d model represents the dimension of the latent space where the representation is located; it reflects the dynamic change trend and potential features of the time series, providing a basis for generating appropriate target model parameters subsequently. The self-attention mechanism allows representations of all lengths in the time dimension to share information. Therefore, the information of the last time block in the time dimension is taken as the input of the hypernetwork;
[0099] Based on the unique representation of the user's local recent meteorological data, the hypernetwork generates optimal parameters for a small main network (recurrent neural network) based on the current data distribution
[0100] The hypernetwork generates optimal parameters for a small main network based on the current data distribution; including:
[0101] The hypernetwork will be based on adaptively generate optimal parameters for the main network;
[0102] To make the framework more general, when constructing the main network, performing a forward pass using the main network to generate the dynamic computation graph of the corresponding model means: using the dynamic computation graph visualization tool torchviz to obtain the computation graph in DOT format that contains all the operation nodes, tensor nodes of the main network, and the connections between them; in PyTorch, the computation graph of the model is dynamically generated. Therefore, when constructing the main network, first perform a forward pass using the specified network architecture to let PyTorch automatically record this computation process, thereby establishing the computation graph. Each node in the computation graph represents an operation (such as matrix multiplication, ReLU, convolution, etc.) or a tensor (such as input, output, intermediate features). The edges in the computation graph represent the data flow relationships between operations. Through torchviz, the PyTorch computation graph can be exported to DOT format, which is a text representation used to describe the graph structure. Use the torchviz.make_dot method to generate this computation graph. The output result is in the form of:
[0103] digraph G{
[0104] node0[label="x"];# Input tensor x
[0105] node1[label="GRU_Ih_weight"];# First fully connected layer
[0106] node2[label="ReLU"];
[0107] node3[label="fc1"];
[0108] node4[label="y"];# Output tensor y
[0109] node0->node1;
[0110] node1->node2;
[0111] node2->node3;
[0112] node3->node4;
[0113] } respectively describe the names of each node and the corresponding data flow relationships.
[0114] The DOT file is essentially text and can be parsed using the pydot library provided by Python and converted into a graph structure processed by the networkx library provided by Python.
[0115] The node names of the original computational graph are usually automatically generated by PyTorch. In order to facilitate understanding and code reuse, the nodes are relabeled so that the node numbers are clearer and easier to filter and analyze. In the computational graph, many nodes are intermediate calculation steps, such as activation function nodes (such as ReLU, Sigmoid), broadcast nodes (such as Expand, Reshape), and gradient calculation nodes (such as AccumulateGrad). These nodes are not important for model parameter extraction and can be deleted. Only key calculation relationships (such as matrix multiplication, convolution, etc.) are retained. Taking GRU as an example, the final generated node names and the corresponding parameter matrix shapes are {'GRU.weight_ih_l0':(48,6),'GRU.weight_hh_l0':(48,16),'GRU.bias_ih_l0':(48,),'GRU.bias_hh_l0':(48,),'fc.weight':(1,16),'fc.bias':(1,)}
[0116] The dimension of the weight matrix from input to hidden state is (48,6), the dimension of the weight matrix from hidden state to hidden state is (48,16), and the dimension of the bias vector from input to hidden state and from hidden state to hidden state is (48,). The dimension of the weight matrix is (1,16), and the dimension of the bias vector is (1,), mapping the 16-dimensional hidden state output by the GRU layer to a 1-dimensional output.
[0117] Use pydot and networkx to convert the DOT-formatted computational graphs of operation nodes, tensor nodes, and connections between them into operable graph objects, relabel the nodes, filter and delete unnecessary intermediate computational nodes, retain important dependencies, integrate node dependencies, and obtain the parameter matrix name and corresponding parameter matrix shape that needs to be generated for the main network;
[0118] Assuming that the main network is a GRU structure, the weight generation process is as follows:
[0119] By reading the calculation graph, the GRU structure includes the following 6 nodes: GRU.weight_lh_l0, GRU.weight_hh_l0, GRU.bias_ih_l0, GRU.bias_hh_l0, fc.weight, fc.bias, which represent the weight matrix and bias of the GRU loop kernel and the weight matrix and bias of the fully connected matrix respectively;
[0120] According to the shape of different node parameter matrices, the hypernetwork generates C groups of different pre-selected parameter key-value pairs for the six nodes of the GRU structure during initialization. Each set of preselected parameters will optimize the unique characteristics of the data from different perspectives:
[0121] Use a fully connected matrix to perform mapping, and generate queries after layer normalization and activation functions which are used as keys and values respectively, and calculate using the standard attention mechanism. The output is the optimal model parameters of the main network B in the i-th cycle i The optimal parameters refer to the parameters that are the most matched or adapted to the data of the current cycle. The implementation process is described as:
[0122]
[0123] Among them, FFN represents the feed-forward network layer.
[0124] Through the attention mechanism, weights can be dynamically assigned according to the importance of different candidate parameters, so as to select the parameters that are most suitable for the main network to work in the future.
[0125] Transfer the optimal parameters to the edge device, such as pushing them to the mobile APP. The APP is responsible for putting the optimal parameters into the corresponding model that has been initialized in the mobile APP, making predictions based on the data of the sensors near the user, generating future weather information reports and analyses for the user, providing customized weather services for individual users, or in the industrial equipment monitoring scenario, deploying the model to the factory equipment for real-time monitoring, and making prompts when abnormal values appear in the predictions, etc.
[0126] In addition, in order to improve the representativeness of each set of parameter candidates, a dual loss function is adopted as follows:
[0127]
[0128] Use configured B i to calculate the loss so as to measure the performance of this parameter set in the current task. At the same time, select the candidate parameter set with the largest attention score configured B i to calculate the additional loss Among them, the largest attention score is calculated by c * represents the index corresponding to the largest attention score. The candidate parameter set selected according to this index is Use as the parameters of the main network for forward propagation; use the same method to calculate the loss to optimize the model. By weighted combination of the two losses, that is While focusing on the current task, it can ensure that the hypernetwork learns more representative and robust parameter configurations globally. λ represents the weighted parameters, which can be adjusted in different application scenarios to select the optimal value (usually 0.01 - 0.1). The core advantage of this design scheme lies in balancing the learning of local and global optimal weight parameters, which can not only improve the performance of the main network on specific tasks but also enhance the generality and transfer ability of the hypernetwork.
[0129] Use and configured B i Calculate the losses respectively and including:
[0130] The inputs of all windows in the current period share the same configured B i Perform forward propagation to obtain the predicted values for future time series data Use the MSE loss to measure the difference between the predicted values of each window and the standard value Y w in the i - th period That is:
[0131]
[0132] Use the same steps to calculate configured B i Calculate the loss Apply the gradient descent algorithm to optimize the parameters of the hypernetwork. This process can be expressed as where Θ represents the parameters of the hypernetwork, and η represents the learning rate of the hypernetwork. The hypernetwork can be fine - tuned to further learn the specific representations of the current data distribution by combining the prior knowledge provided by the pre - trained model. Based on these specific representations, the hypernetwork can generate network structures or parameters adapted to the current data distribution, thus better coping with data changes and concept drift. The advantage of this approach is that it can utilize the general features learned by large pre - trained models in multiple domains and can be customized and optimized for the particularity of the current data through the dynamic adjustment of the hypernetwork. The generation strategy learned by the hypernetwork enables the model to flexibly adjust its structure or parameters when concept drift occurs, providing an efficient alternative to avoid repeated training. This method can not only save computing resources but also improve the adaptability and generalization ability of the model. H where Θ represents the parameters of the hypernetwork, and η represents the learning rate of the hypernetwork. The hypernetwork can be fine - tuned to further learn the specific representations of the current data distribution by combining the prior knowledge provided by the pre - trained model. Based on these specific representations, the hypernetwork can generate network structures or parameters adapted to the current data distribution, thus better coping with data changes and concept drift. The advantage of this approach is that it can utilize the general features learned by large pre - trained models in multiple domains and can be customized and optimized for the particularity of the current data through the dynamic adjustment of the hypernetwork. The generation strategy learned by the hypernetwork enables the model to flexibly adjust its structure or parameters when concept drift occurs, providing an efficient alternative to avoid repeated training. This method can not only save computing resources but also improve the adaptability and generalization ability of the model.
[0133] In addition, to further reduce the risk of overfitting, a regularization strategy is introduced. By applying weight decay to the model weights, the excessive update of the hypernetwork parameters is suppressed. This multi-dimensional optimization design effectively improves the generalization ability of the hypernetwork, making its performance more robust in complex tasks, and thus more effectively coping with the concept drift phenomenon.
[0134] To enhance the general ability of the present invention, it is further assumed that the same main network framework is used for prediction in all cycles. This means that the main networks in each cycle have the same computational graph structure, differing only in model parameters. Specifically, all cycles share the same network topology and hierarchical structure, but the models in each cycle are individually adjusted through independent parameters to adapt to the specific patterns and trends of the data within different cycles. This design ensures cross-cycle structural consistency, while enabling optimization for the specific data characteristics of each cycle by adjusting the parameters of the main network.
[0135] Since the architecture of the main network is fixed, the system does not need to frequently switch the underlying hardware configurations required for different structures, thereby reducing system overhead and improving the utilization efficiency of computing resources. This advantage is particularly significant in edge device computing scenarios. The fixed architecture also ensures more convenient model deployment. When migrating from the laboratory environment to the actual production environment, it avoids the need to re-adapt to complex software and hardware ecosystems for different cycles, reduces the error probability in the deployment process, and is conducive to the application of technical achievements. In addition, the same computational graph structure ensures smooth cross-cycle knowledge transfer. The knowledge and feature representations accumulated in the previous cycle can be directly passed to subsequent cycles, helping the models in subsequent cycles to be quickly optimized based on the existing foundation, reducing the cost of repeated learning, and improving the training effect and speed.
[0136] In actual application scenarios, a pre-trained model and a hypernetwork are deployed in the cloud. When the edge device synchronizes data with the cloud, it is determined whether to update the model parameters to the edge device based on whether the prediction accuracy of the main network on the latest data generated according to history is lower than the threshold.
[0137] The architecture of the main network is optional and includes standard RNN, LSTM, and GRU.
[0138] To demonstrate the superiority of the present invention, it is trained and tested on an open-source air quality dataset to verify its performance. This dataset contains hourly air quality data collected from 12 stations in Beijing, China, from March 2013 to February 2017. We use data from the same four stations (Dongsi, Tiantan, Nongzhanguan, Dingling) and the six features included in the data (PM2.5, PM10, SO2, NO2, CO, O3). The data from March 1, 2013 to June 30, 2016 is used to train the hypernetwork to learn the patterns of specific data, the data from July 1, 2016 to October 31, 2016 is used to verify the effect, and the data from November 2, 2016 to February 28, 2017 is used for testing. Only pre-trained models and hypernetworks with relatively small numbers of parameters are used in this experiment to verify the performance of the present invention. Specifically, a pre-trained model with the PatchTST architecture is used. The number of parameters of the pre-trained model is approximately 400,000. The input sequence length of the pre-trained model is set to 28 days (including 672 time points), the input length of the main network is 24, the prediction length is 1, the number of parameters of the hypernetwork is only 1.2MB, and the GRU is set as the main network with a hidden dimension of 16. Then the number of parameters of the main network is only 1169. In Figure 3 the change of training loss is shown, and it can be seen that the proposed scheme of the present invention has strong robustness. Figure 3 In it, the horizontal axis represents the training rounds, and the full amount of training data is used for training in each round. The vertical axis is RMSE, that is, Root Mean Square Error, which is a numerical analysis index widely used in multiple fields such as statistics, mathematics, and engineering, and is used to measure the degree of difference between the observed value and the true value. From Figure 3 it can be clearly seen that as the training rounds increase, the RMSE of the model drops rapidly until it reaches a stable region, and there are few severe oscillations during this period, indicating that:
[0139] (1) The model has good learning ability and convergence on the air quality dataset, can effectively learn relevant patterns and features from the full amount of training data, continuously reduce the error between the predicted value and the true value, and thus the RMSE continues to decrease.
[0140] (2) The training process of the model is relatively stable, without serious fluctuations or abnormal situations such as overfitting or underfitting. The few severe oscillations indicate that the optimization algorithm of the model (such as stochastic gradient descent, etc.) performs well and can smoothly adjust the model parameters and gradually approach the optimal solution.
[0141] (3) The fact that the RMSE finally stabilizes means that the model has approached the convergence state. Further increasing the number of training epochs may not significantly reduce the error. At this time, the performance of the model has basically stabilized at a good level (RMSE is 0.01), showing good fitting effect and generalization ability for this air quality dataset, and is expected to make relatively accurate predictions for air quality-related data in practical applications.
[0142] To verify the performance difference between the main network of this configuration and the original model, the control variable method was used to construct LSTM / GRU architectures with exactly the same initial parameter space under the same data distribution, and a quantitative comparative analysis was carried out with the dynamically parameterized model generated by the present invention. During the experiment, the two groups of models were monitored synchronously, and the average MAE loss in each cycle was recorded for comparison to explore the stability of the test accuracy of the generated model and the end-to-end trained model on the test set. Figures 4(a) and 4(b) respectively show the visualization results of the present invention on two subsets, Dingling and Nongzhanguan, of the air quality dataset. The horizontal axis represents the original LSTM and GRU architectures respectively, and PTH-LSTM and PTH-GRU represent the models configured by the present invention. MAE, that is, the Mean Absolute Error, is a commonly used indicator for measuring the difference between the predicted value and the true value of data. The violin plot is a visualization tool for showing the data distribution, combining the characteristics of the box plot and the kernel density plot. Its contour is generated by kernel density estimation, showing the distribution density of data in each value interval. The wider the contour, the denser the data points in that interval; the narrower the contour, the sparser the data points. The horizontal lines in the figure show the statistical information of the distribution, usually including the median (red solid line), mean (green solid line), etc., which can intuitively present the central tendency and dispersion degree of the data. The MAE distribution of the dynamically parameterized model generated by PTH-Gen has a lower mean on each dataset, which indicates that, on average, the model generated by PTH-Gen has a smaller prediction error and can make more accurate predictions. In addition, the variance of the MAE distribution of these models is also smaller, that is, the contour of the violin plot is narrower. This means that the prediction error of the model generated by PTH-Gen fluctuates less on different samples, the performance is more stable, and there will be no large deviation.
[0143] To verify the actual performance of the present invention, two relatively advanced schemes were considered for comparison with the present invention:
[0144] ADARNN (Adversarial Domain Adaptation RNN, 2021): ADARNN divides the data into multiple data segments with large distribution differences based on the principle of maximum entropy, and trains the RNN model to learn the commonalities between different distributions, thereby improving the generalization ability and domain adaptation ability of the model.
[0145] HTSF (Hyper TimeSeries Forecasting, 2023): The concepts of the hypernetwork and the main network introduced by HTSF. The hypernetwork layer uses bidirectional GRUs to encode randomly selected historical data, generating a feature sequence that can represent the distribution, and then generating model parameters for the main layer. The main network layer, based on the parameters generated by the hyperlayer and its own internal state, adaptively adjusts the weights through the attention mechanism for accurate prediction. The comparison results are shown in Table 1 as follows:
[0146] Table 1
[0147]
[0148] By comparing with the latest advanced solutions, it can be seen that the accuracy of the present invention has been improved compared with them.
[0149] In addition, in order to highlight the advantages of the hypernetwork generation solution compared with the end-to-end trained RNN model, the control variable method is used to construct LSTM / GRU architectures with strictly equivalent parameter space initializations under the same data distribution, and a quantitative comparative analysis is carried out with the dynamic parameterized model generated by the hypernetwork of the present invention. During the experiment, the two groups of models are monitored synchronously, and the average MAE loss in each cycle is recorded for comparison to explore the stability of the test accuracy of the generation model and the end-to-end trained model on the test set. To visually display the performance comparison of the models, the results are plotted as violin plots. A violin plot is a visualization tool for showing data distributions, combining the characteristics of box plots and kernel density plots. Its contour is generated by kernel density estimation, showing the distribution density of data in each value interval. The wider the contour, the denser the data points in that interval; the narrower the contour, the sparser the data points. The MAE distribution of the dynamic parameterized model of the present invention has a lower mean on each dataset, which indicates that, on average, the model generated by the present invention has a smaller prediction error and can make more accurate predictions. In addition, the variance of the MAE distribution of these models is also smaller, that is, the contour of the violin plot is narrower. This means that the prediction error fluctuations of the generation model on different samples are smaller, the performance is more stable, and there will be no large deviations.
[0150] Example 3
[0151] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the meteorological service recommendation method of the edge-side device small recurrent neural network based on the hypernetwork described in Example 1 or 2.
[0152] Example 4
[0153] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of the meteorological service recommendation method of the small recurrent neural network of the edge device based on the hypernetwork described in Embodiment 1 or 2.
[0154] Embodiment 5
[0155] A meteorological service recommendation system of a small recurrent neural network of an edge device based on a hypernetwork, comprising:
[0156] A model deployment and training module, configured to: deploy a general pre-trained model and a hypernetwork in the cloud; when in use, use the general pre-trained model to extract the representation of historical data, and the hypernetwork, through fine-tuning, combines the representation of historical data provided by the pre-trained model to further learn the unique representation of the user's local recent meteorological data;
[0157] A service recommendation module, configured to: based on the unique representation of the user's local recent meteorological data, the hypernetwork generates optimal parameters for a small main network based on the current data distribution, and transmits the optimal parameters to the edge device to provide customized meteorological services for individual users.
Claims
1. A weather service recommendation method based on a small recurrent neural network of edge devices on a hypernetwork, characterized in that: include: Deploy common pre-trained models and hypernetworks in the cloud; When in use, a general pre-trained model is used to extract the representation of historical data. The hypernetwork is fine-tuned and combined with the representation of historical data provided by the pre-trained model to further learn the unique representation of the user's local recent meteorological data; Based on the unique representation of the user's local recent meteorological data, the super network generates optimal parameters for the small main network based on the current data distribution, passes the optimal parameters to the edge side device, and provides customized meteorological services for individual users.
2. The method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork according to claim 1 is characterized in that: Use a general pre-trained model to extract representations of historical data; including: 1) Preprocessing: historical meteorological data of the i-th period Pre-processing, L h represents the input length of the pre-trained model, D represents the number of different meteorological information recorded by different meteorological sensors at each time point, and the multivariate time series data is treated as multiple independent univariate time series. Then X h i will be represented as D univariate time series x, where x j Represents the data recorded by the jth sensor; 2) Normalization: For each x j Perform missing value filling and data normalization, expressed as: Where μ and σ are the mean and standard deviation of each univariate series, respectively; 3) Block: Input to the jth channel Split in the time dimension and convert into a series of time blocks, represented as x p j , choose overlapping or non-overlapping mode; let the patch length be Q, the non-overlapping area between two adjacent patches, that is, the step length is S, then the number of generated time blocks is 4) Mapping: Blocked input Through a trainable linear projection Mapped to dimension d model , and apply a learnable additive positional encoding To introduce the order information of the time series, we can obtain Will pass through multiple stacked Transformer encoding layers; 5) For each Transformer encoding layer; first, Through three transformation matrices W Q , W K , W V Transformed into query Q, key K and value V, H Q =H K =H V =M×D, where M represents the number of heads used in the multi-head attention mechanism; Secondly, the query Q, key K and value V will be split into M attention heads of the same shape, and the i-th attention head contains different query Q i , key K i Sum value V i ; For the i-th attention head, the calculation process is as follows: Again, after the attention mechanism, residual connections and layer normalization operations are performed; Finally, after the feedforward neural network, residual connection and layer normalization are performed again; 6) After multiple stacked Transformer encoding layers, the output of each channel Concatenate on the feature dimension and express the result as Among them, D represents the feature dimension contained in the data, P represents the size of the sequence time dimension, and d model Represents the dimension of the latent space where the representation is located; takes the information of the last time block in the time dimension As input to the hypernetwork; Based on the unique representation of the user's local recent meteorological data, the super network generates optimal parameters for the small main network based on the current data distribution.
3. The method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork according to claim 1, characterized in that: The hypernetwork generates optimal parameters for the small main network based on the current data distribution; including: The hypernetwork will be based on Adaptively generate optimal parameters for the main network; When building the main network, use the main network to perform a forward propagation to generate the dynamic calculation graph of the corresponding model, which means: using the dynamic calculation graph visualization tool torchviz to obtain a calculation graph in DOT format containing all operation nodes, tensor nodes and connections between the main network; Use pydot and networkx to convert the DOT-formatted computational graphs of operation nodes, tensor nodes, and connections between them into operable graph objects, relabel the nodes, filter and delete unnecessary intermediate computational nodes, retain important dependencies, integrate node dependencies, and obtain the parameter matrix name and corresponding parameter matrix shape that needs to be generated for the main network; Assuming that the main network is a GRU structure, the weight generation process is as follows: By reading the calculation graph, the GRU structure includes the following 6 nodes: GRU.weight_lh_l0, GRU.weight_hh_l0, GRU.bias_ih_l0, GRU.bias_hh_l0, fc.weight, fc.bias, which represent the weight matrix and bias of the GRU loop kernel and the weight matrix and bias of the fully connected matrix respectively; According to the shape of different node parameter matrices, the hypernetwork generates C groups of different pre-selected parameter key-value pairs for the six nodes of the GRU structure during initialization. Each set of pre-selected parameters will optimize the unique representation of the data from different perspectives: Using the fully connected matrix Mapping is performed to generate queries after layer normalization and activation function As keys and values, respectively, the standard attention mechanism is used for calculation, and the output is the main network B of the i-th cycle. i The most suitable model parameters The optimal parameters refer to the parameters that best match or adapt to the data of the current period. The implementation process is described as follows: in, FFN stands for Feed-Forward Network.
4. The method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork according to claim 1, characterized in that: Using dual loss function As shown below: use Configuration B i Calculating Losses At the same time, select the candidate parameter group with the largest attention score Configuration B i Calculating additional losses Among them, the maximum attention score is given by It is calculated that c * Represents the index corresponding to the maximum attention score, and the candidate parameter group selected according to this index is Will As the parameters of the main network, forward propagation is performed; λ represents the weighted parameter.
5. The method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork according to any one of claims 1 to 4, characterized in that: use and Configuration B i Calculate the loss separately and include: Input of all windows in the current period Share the same Configuration B i Perform forward propagation to obtain the predicted value for future time series data Use MSE loss to measure the predicted value of each window in the i-th period and the standard value Y w The difference between Right now:
6. The method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork according to claim 1, characterized in that: The architecture of the main network is optional, including standard RNN, LSTM, and GRU.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method for recommending weather services based on a small recurrent neural network of edge devices on a hypernetwork as described in any one of claims 1-6 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for recommending meteorological services based on a small recurrent neural network of edge devices on a hypernetwork as described in any one of claims 1 to 6 are implemented.
9. A weather service recommendation system based on a small recurrent neural network on edge devices of a hypernetwork, including: The model deployment and training module is configured to: deploy common pre-trained models and hypernetworks in the cloud; When in use, a general pre-trained model is used to extract the representation of historical data. The hypernetwork is fine-tuned and combined with the representation of historical data provided by the pre-trained model to further learn the unique representation of the user's local recent meteorological data; The service recommendation module is configured as follows: based on the unique representation of the user's local recent meteorological data, the super network generates optimal parameters for the small main network based on the current data distribution, and transmits the optimal parameters to the edge side device to provide customized meteorological services for individual users.
Citation Information
Cited By
Online distributed parameter optimization method and system for intelligent modeling of complex system
CN121902102A