Inter-bank network time sequence prediction method based on generative adversarial hierarchy fusion model
Through the minimum density method and data enhancement, the interbank network feature data is generated, and combined with the hierarchical fusion network and WGAN architecture, the problems of data missing and model training efficiency in interbank network timing prediction are solved, achieving more efficient feature capture and training stability.
Patent Information
- Application Number
- CN202510660707.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing interbank network timing prediction methods have problems such as missing data, insufficient model feature capture capability and low model training efficiency.
The minimum density method is used to generate interbank network characteristic data, and a diversified data set is constructed through data augmentation. The design hierarchical fusion network HFN is combined with GCN, GAT and Transformer feature extraction information to improve the model's feature capture capability. Introduce WGAN architecture and SAEDC modules to improve the stability and efficiency of model training.
It effectively alleviates the problem of missing interbank transaction data, improves the model's ability to capture interbank network structure and asset characteristics, and improves the stability and convergence speed of model training.
Smart Images

Figure CN120180049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and network time series prediction, and in particular to an inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model. Background Art
[0002] As a core component of the modern financial system, the inter-bank market's dynamic changes directly affect the stability and efficiency of the financial system, especially playing a key role in aspects such as systemic risk identification, liquidity management, and evaluation of the effectiveness of monetary policies. In recent years, with the rapid development of the global financial market, the complexity of the inter-bank market network has increased significantly. How to effectively predict the inter-bank time series network, explore the bank systemic risks under the evolution of the time series network, and thus enhance the early warning capabilities of banks and regulatory authorities for potential financial crises has become the focus of current research.
[0003] The time series prediction of the inter-bank network plays a crucial role in the study of bank systemic risks. However, in the existing inter-bank network time series prediction methods, there are still the following three deficiencies: First, for inter-bank network data, inter-bank transaction information is regarded as a commercial secret and is not publicly available. This lack of inter-bank transaction data caused by commercial competition has directly hindered the research on inter-bank network prediction.
[0004] Second, in terms of the feature capture ability of the model, there is a complex interaction relationship between the inter-bank network and bank assets, that is, changes in bank assets will affect the bank network structure and transaction amount, and changes in the bank network structure and transaction amount will also in turn affect bank assets. However, traditional time series machine learning methods such as RNN, LSTM, etc. cannot effectively capture this complex association between the bank network structure and bank asset features. In addition, when the historical data span is large (such as cross-year transaction patterns), the network structure changes significantly. Traditional graph convolutional networks GNN usually average when dealing with node neighbor relationships, unable to distinguish the importance and contribution degree of different inter-bank relationships, and lacking the ability to identify key connections and important nodes in the network.
[0005] Third, in terms of the efficiency of model training, in the initial training process of a model based on traditional GAN, when there is a large deviation between the real data distribution and the generated data distribution, it will lead to unstable training and slow convergence speed. The main manifestations are: in the initial stage, the networks generated by the generator are all judged as fake, resulting in the disappearance of the generator gradient and unstable model training; the initial edge screening threshold is relatively high, resulting in a small number of edges generated by the generator, insufficient positive feedback of the model, and slow convergence speed of model training.
[0006] Therefore, how to provide an effective interbank network time series prediction method to solve the defects of missing interbank network data, insufficient feature capture ability of the model, and low model training efficiency in existing work has become an urgent technical problem for those skilled in the art. Summary of the Invention
[0007] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to propose a solution based on the Hierarchical Fusion Wasserstein Generative Adversarial Network for Interbank Network Temporal Forecasting (HFAN-ITF).
[0008] First, the present invention uses the minimum density method to generate data with interbank network characteristics based on interbank balance sheet data, and expands the data through data augmentation to ensure data diversity, thereby constructing a diversified dataset that conforms to the basic characteristics of the interbank network.
[0009] Secondly, the present invention designs a hierarchical fusion network HFN in the generator part of the WGAN (Wasserstein Generative Adversarial Network) architecture. This network effectively extracts features by sequentially fusing the feature extraction information of GCN, GAT, and Transformer. Among them, the GCN network aggregates global information to extract the structural features of the network. GAT assigns different weights to each node to aggregate neighbor node information, thereby mining the connection relationships of key nodes; Transformer is used to model the time series network. By assigning different weights to the feature vectors of different time steps, the global time series features are fused to predict the network of the next time step, so as to better understand the relationships between different time steps and multiple variables.
[0010] In addition, the present invention uses the Wasserstein distance in the discriminator to measure the distribution deviation between the generated network and the real network, making the training of the model smoother and improving the overall training effect of the model; SAEDC (Self-adaptive Edge Density Control) adopts an edge density control strategy SAEDC with an adaptive threshold, increases the number of edges during initial training, and improves the convergence speed of the model.
[0011] Technical Solution An interbank network time series prediction method based on a hierarchical fusion generative adversarial model, comprising the following steps: S1: Generate the inter-bank network time series data and perform data augmentation; Adopt the minimum density method to generate data with inter-bank network characteristics based on the inter-bank balance sheet data, and expand the data through data augmentation to ensure data diversity, and then construct a diversified dataset that conforms to the basic characteristics of the inter-bank network.
[0012] S2: Construct a generative adversarial hierarchical fusion model; S3: Divide the data-augmented inter-bank network time series data into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model; S4: Conduct error analysis on the predicted results and the original test set time series data, and perform model comparison and ablation experiments.
[0013] Beneficial effects: (1) The present invention generates data that conforms to the inter-bank network characteristics based on the inter-bank balance sheet data through the minimum density method, and expands the dataset in combination with the data augmentation technology, effectively alleviating the data missing problem caused by the non-disclosure of inter-bank transaction data due to commercial secrets.
[0014] (2) The present invention significantly improves the model's ability to capture the complex interaction relationship between the inter-bank network structure and bank assets through the hierarchical fusion network HFN, thereby improving the accuracy of inter-bank network prediction.
[0015] (3) By introducing the WGAN architecture and the SAEDC module, the present invention effectively alleviates the problems of gradient disappearance and training instability caused by excessive distribution deviation in the initial training stage of the traditional GAN model and the problem of sparse edges under the traditional edge density control strategy, thereby accelerating model convergence and improving the overall training efficiency. Description of the drawings
[0016] Figure 1 is the overall flowchart of the present invention; Figure 2 is the division diagram of the HFAN-ITF module of the generative adversarial hierarchical fusion model of the present invention; Figure 3 is the detailed design diagram of the HFAN-ITF of the generative adversarial hierarchical fusion model of the present invention; Figure 4 is the experimental result comparison curve diagram of the model training convergence speed of the embodiment of the present invention and the benchmark model; Figure 5 is the experimental result comparison diagram of the embodiment of the present invention and other models in terms of graph prediction indicators; Figure 6 is the experimental result comparison diagram of the embodiment of the present invention and other models in terms of node prediction indicators; Figure 7 It is a graph showing the ablation experiment results of an embodiment of the present invention. Specific Embodiments
[0017] The following uses specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0018] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0019] A method for inter-bank network time series prediction based on a generative adversarial hierarchical fusion model includes the following steps (as Figure 1 ): S1: Generate inter-bank network time series data and perform data augmentation The present invention adopts the minimum density method to generate data with inter-bank network characteristics based on inter-bank balance sheet data, and expands the data through data augmentation to ensure data diversity, thereby constructing a diverse dataset that conforms to the basic characteristics of the inter-bank network.
[0020] Inter-bank network data includes: the adjacency matrix of the lending network between banks and the node feature vectors in the network (basic features of banks: equity funds, total liabilities, total assets, inter-bank liabilities, inter-bank assets).
[0021] The inter-bank network time series data is represented as , where is the total length of the time series, and the time unit is year.
[0022] is the inter-bank network at time t, represented as , where is the adjacency matrix of the lending network between N banks. , where represents the amount of bank i borrowing from bank j . , representing the node feature vectors in the network (basic features of banks: equity funds, total liabilities, total assets, inter-bank liabilities, inter-bank assets).
[0023] Specifically as follows: S11: Generation of Interbank Network Based on the Minimum Density Method The present invention generates interbank network data through the minimum density method, and the network has the characteristics of a real interbank network. The interbank balance sheet data is used as the characteristic data of banks to conduct research on the time series prediction of the interbank network based on data driving.
[0024] The minimum density method generates an interbank network according to the characteristics of "sparsity" (i.e., the interbank network is a sparse matrix) and "assortativity" (i.e., there is a large difference in the asset sizes of the borrowing and lending banks) of interbank lending.
[0025] The minimum density method can be expressed as a constrained optimization problem with P(X) and X as the solution targets, where P(X) represents the interbank lending matrix and X represents the interbank lending amount.
[0026] The optimization objective of "sparsity" is : where represents the amount borrowed by bank i from bank j, represents the number of banks, represents the fixed cost of establishing a connection between banks, is the total interbank lending amount of bank , is the total interbank borrowing amount of bank , where , .
[0027] In addition, the borrowing tendency that meets the "assortativity" between banks is quantified by the entropy function to quantify this "assortativity", where .
[0028] Finally, by integrating these two parts of "sparsity" and "assortativity", the total optimization objective function of the minimum density method is designed as , where represents the weight parameter of the objective of "assortativity".
[0029] The solution method of the minimum density method is prior art.
[0030] The P(X) and X obtained by solving the minimum density method respectively correspond to those at time t .
[0031] S12: Data Augmentation The present invention adopts a dual data augmentation strategy to expand the interbank network data. Specifically, the data augmentation strategy includes parameter perturbation augmentation and data perturbation augmentation, and through these two augmentations, network samples with diverse features are obtained.
[0032] S121: Parameter perturbation augmentation Parameter perturbation augmentation is performed on the key parameters in the minimum density network generation method and for random perturbation. In the real interbank market, the transaction costs and transaction propensities vary in different market environments, policy changes, bank strategy adjustments and other factors. Based on this real - world phenomenon, the present invention perturbs and augments the parameters in the network generation algorithm: the parameter is calculated using the formula and is calculated using the formula .
[0033] Among them, the perturbation factor obeys a uniform distribution on (0, 1), and generally takes 0.5, so that the perturbation range of the parameter is between 0.5 and 1.5, at a moderate level.
[0034] Based on the new parameter a new interbank network is generated.
[0035] Through this parameter perturbation method, different network structures can be generated while maintaining the basic topological features of the network.
[0036] S122: Data perturbation augmentation The present invention perturbs the data of interbank lending assets and liabilities. Banks need to make some accounting estimates when preparing annual reports, such as loan loss provisions, asset impairment provisions, etc. These estimates are subjective and there will be deviations in the data size.
[0037] Based on this real - world phenomenon, the present invention adds random Gaussian noise to the interbank asset data of the bank and calculates using the formula where the Gaussian noise obeys .
[0038] According to the perturbed interbank asset data, the interbank network is generated by the minimum density method.
[0039] S2: Construct a generative adversarial hierarchical fusion model The generative adversarial hierarchical fusion model includes: a feature encoding module, a hierarchical fusion network HFN module, a network generation module, and a discriminator module. Its overall architecture is asFigure 2 as shown
[0040] Among them, the feature encoding module, the HFN module, and the network generation module constitute the generator of the generative adversarial network, and the discriminator module is the discriminator of the generative adversarial network.
[0041] Input data is generated by step S1 (including generated data and augmented data), and includes the basic features of the bank (such as equity funds, total liabilities, total assets, inter-bank liabilities, inter-bank assets) and the time-series inter-bank adjacency matrix (representing the inter-bank lending relationship).
[0042] The hierarchical fusion network HFN module includes three modules: Graph Convolutional Network (GCN), Graph Attention Network (GAT), and Transformer.
[0043] First, the HFN module performs feature encoding on the input data. Subsequently, the GCN layer and the GAT layer extract the structural features and important node information of the inter-bank network, and the Transformer encoder captures the temporal dependencies between multiple features.
[0044] Then, the network generation module obtains edges, weights, and bank attributes through an MLP (Multilayer Perceptron), and filters them through SAEDC to generate the prediction results of the inter-bank network (including the inter-bank network structure, network weights, and bank features) at the next moment, denoted as .
[0045] During the training process, the MLP of the discriminator module scores the generated network and the real network respectively, generating Fake Scores (predicted network scores) and Real Scores (real network scores). Calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models.
[0046] The detailed processing flow of the generative adversarial hierarchical fusion model is as Figure 3 shown, and the specific description is as follows: S21. Perform feature encoding on the input data Feature encoding includes two parts: time encoding and node encoding: Time encoding: Use YearEncoder (a prior art encoder for encoding year features) to process the year information in the input data, and use an MLP to further map the time features to the hidden space, that is , where the MLP consists of two stacked modules, each module contains a Linear (linear layer), LayerNorm (layer normalization), and a GELU activation function. After these two modules, LayerNorm (layer normalization) and Dropout (dropout layer) are applied; Node Encoding: Use MLP to encode the input data G, that is, NetEncoding(G) = MLP(G), where the structure of MLP is the same as that of the MLP in the time encoding part; S22: Use the HFN module to process the feature encoding The HFN module captures the temporal features of the interbank network and the attribute features of nodes through hierarchical feature fusion of three modules: GCN, GAT, and Transformer in sequence.
[0047] S221: GCN module The GCN module extracts the local topological features of the interbank network for each time slice based on the input feature encoding, and generates a preliminary node representation for each time slice . By stacking multiple layers of GCN, higher-order neighbor relationships of the interbank network can be captured. The output representations of these layers will be used as the input to the GAT module for further modeling the relationships between nodes in the interbank lending network.
[0048] In this embodiment, three layers of GCN are adopted, and each layer of GCN contains three layers of GCNConv (graph convolutional layer), corresponding to Figure 3 3×GCNConv in
[0049] The GCN input includes the adjacency matrix at time slice t and the node feature matrix .
[0050] The operation of GCN is defined by the following formula: where, is the adjacency matrix after adding self-loops, and self-loops are used to ensure that nodes incorporate their own information; is the identity matrix; Normalize the adjacency matrix to balance the contribution of node features and ensure the stability of training and the rationality of feature propagation. is a learnable projection matrix, represents the dimension of the input node features, the dimension of the hidden layer features, which is a hyperparameter of GCN.
[0051] S222: GAT module The attention mechanism of two-layer GAT is used to weight-model the adjacency relationship of the inter-bank network. Each layer of GAT contains three layers of GATConv (graph attention convolutional layer), corresponding to Figure 3 the 3×GATConv in
[0052] The input of the GAT module includes the output features of GCN at time slice t and the corresponding adjacency matrix , , where N represents the number of banks. The processing formula is .
[0053] Specifically, GAT generates the feature representation of each node through the following steps: 1) For node i and its neighbor j, node feature projection where is a learnable projection matrix that projects the node feature into a new feature space.
[0054] 2) For node i and its neighbor j, use the attention mechanism to calculate the weight of edge (i, j): , where is the feature of node , is the weight between nodes i and j. a is a learnable attention parameter, represents the feature concatenation operation concat.
[0055] 3) Use the Softmax function to normalize the attention coefficients of each node to obtain . According to the normalized attention weights (where N(i) refers to the set of first-order neighbors of node i), aggregate the neighbor features and update the node representation , where is the activation function.
[0056] To enhance the model's expressive power, the GAT layer uses the multi-head attention mechanism and finally takes the average to obtain the final node representation , where K represents the total number of attention heads. In this embodiment, 4 heads are adopted, corresponding to Figure 3 the heads = 4 in
[0057] S223: Transformer module In the time series prediction task of the interbank network, the high-order features of nodes not only depend on the topological relationship of the current time slice, but also need to capture the dynamic changes and time series dependencies between time slices. For this purpose, in this embodiment, a Transformer encoder is introduced after GCN and GAT to model the long-term dependencies of node features between multiple time slices.
[0058] The output of GAT is concatenated with the time-coded features to obtain the input tensor of the Transformer encoder: Z Q, K, and V are generated through linear transformation, which are the query, key, and value vectors respectively, and are used as the input of the Transformer, where is the length of the input time series. Then, by calculating the correlation (attention weight) between any two time slices in the sequence, the dependencies between time slices are captured. The attention formula is: To avoid gradient vanishing, the Transformer adds residual connections after multi-head attention (in this embodiment, 8 heads are used, corresponding to Figure 3 heads = 8 in
[0059] Final output () ). The output features of the Transformer encoder not only contain the global information between time slices, but also retain the local topological features of nodes, providing high-quality feature representations for the subsequent network generation module.
[0060] S23: Network generation module The adjacency matrix for generating the next time slice in the network generation module . The generation principle of the adjacency matrix is to generate possible edges for each node, and these edges form an adjacency matrix. The present invention introduces an adaptive edge density control module SAEDC to control the generation of edges.
[0061] Specifically, the inputs of this module include: the feature matrix ; the historical network , including nodes and weights; time encoding , which is the time encoding information at time t+1 and is used to characterize the time features of the prediction time slice.
[0062] The processing procedure of this module is as follows: 1) To predict the edge at time t+1, it is first necessary to construct candidate edges in the network. Specifically, candidate edges are first constructed for each pair of nodes (i, j). The candidate edges include historical edges and possible new edges, and the dimension is |E| (the number of candidate edges). For each edge, its generated feature representation is , which are the features of nodes respectively, is the time encoding, and
[0063] is the historical weight of the edge (if the edge (i, j) is a new edge, it is initialized to 0). is fed into the edge generator ( ). The edge generator is an MLP, where the MLP contains two stacked modules, and each module consists of a Linear (linear layer), LayerNorm (layer normalization), ReLU activation function, and Dropout (dropout layer); after the processing of these two modules, a linear layer (Linear) is connected. The features of each candidate edge are processed by the MLP, and the output of the edge generator includes three parts: The first part is the connection probability of the edge , indicating whether nodes i and j are connected at time t+1.
[0064] The second part is the weight prediction of the edge , indicating the predicted borrowing amount from bank i to bank j.
[0065] The third part is the prediction of the bank asset features .
[0066] The overall formula is as follows: 3) To generate an adjacency matrix that conforms to the characteristics of the inter-bank network, it is necessary to screen the candidate edges. The present invention proposes the SAEDC method to achieve edge screening, and this method dynamically generates the edge screening threshold in the network according to the average degree of the network and the total number N of network nodes.
[0067] Specifically, first calculate the target number of edges : , where is a learnable parameter.
[0068] Then calculate the threshold , where P is the set of connection probabilities of all candidate edges, |P| is the total number of candidate edges, and Quantile represents the quantile function.
[0069] Next, based on the edge probabilities generated by the MLP, filter these candidate edge sets to obtain the filtered edges, and use the predicted value as the final weight of the edges.
[0070] The finally output adjacency matrix is expressed as: The SAEDC module generates a predicted adjacency matrix that conforms to the topological characteristics of the inter-bank network by dynamically adjusting the number of target edges and a dynamic threshold screening mechanism based on the connection probability distribution. This module can flexibly adapt according to changes in the node scale and the average degree of the network, and at the same time ensure that the number of generated edges is close to the target number of edges through quantile screening. Combining the edge generator's prediction of the connection probability and edge weight effectively captures the dynamic evolution characteristics of the network. In the generative adversarial network framework, this design can improve the authenticity and rationality of the generative network, ensure that the topological structure of the generated graph is closer to the actual network, and at the same time provide high-quality adversarial samples for the discriminator, thereby improving the training stability and prediction performance of the model.
[0071] After the above processing, the inter-bank network is finally predicted, including , which respectively refer to: inter-bank edges and weights, and the features of bank individuals i and j.
[0072] S24: Discriminator module The discriminator is mainly used for model training and is completed by an MLP network (the structure of this MLP is the same as that in the generator). Its input is the predicted network at time slice t + 1 and the real network , and calculates the score through the formula. The score of the predicted network is Fake Scores, and the score of the real network is Real Scores. Calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models.
[0073] Network G (specifically the predicted network or the real network ) The scoring process is as follows: 1) Since network G is the original complex data and cannot be directly used as the input for the subsequent network, after the same feature encoding process as in the generator (see step S21), we obtain as the input of the MLP.
[0074] 2) Then the MLP evaluates each edge and the corresponding node features and outputs a score ; The scores of all edges and nodes are weighted and summed to obtain the total score of network G , where E refers to the set of all existing edges in network G, that is . S3: Divide the time-series data of the inter-bank network after data augmentation into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model.
[0075] In this embodiment, the bank data from 2012 to 2019 in the CSMAR database is used, and the time-series data from 2020 to 2022 is used for the test set. The WGAN framework is used for the joint training of the generator and the discriminator. Through the adversarial optimization of the generator and the discriminator, the distribution difference between the generated network and the real network is minimized. Each time the generator is updated, the discriminator is updated 5 times. WGAN overcomes the problems of instability and mode collapse in GAN training by introducing the Wasserstein distance to replace the Jensen-Shannon (JS) divergence of the traditional GAN.
[0076] The loss function of the generator of this model is set as: .
[0077] The goal of the discriminator is to distinguish the generated network and the real adjacency network , and quantify the gap between the two for the model training of itself and the generator. The loss function is used: , to estimate the Wasserstein distance between the generated distribution and the real distribution, is the score of the real graph, the score of the generated graph.
[0078] Calculating the Wasserstein distance needs to satisfy the continuity condition. For this reason, a gradient penalty term is introduced: , where is the sample generated by interpolation, defined as: , , The final loss function of the discriminator is as follows: .
[0079] The discriminator increases the difference between by minimizing and so that it can more effectively distinguish between real samples and fake samples. The WGAN framework optimizes the generator and discriminator through the Wasserstein distance, and combines the SAEDC module to improve the authenticity and dynamic adaptability of the generation network.
[0080] S4: Analyze the error between the predicted results and the original test set time series data, and conduct model comparison and ablation experiments. To measure the beneficial effects of the present invention, the following metrics are used to measure the error of the model prediction results: S41: Measure the error metric of the node feature results The present invention uses the vertex mean squared error (Vertex Mean Squared Error) to measure the deviation between the predicted value and the true value of the asset attribute: = , where represents the feature values of all nodes at time t + 1, represents the predicted feature values of all nodes at time t + 1, represents the number of nodes.
[0081] S42: Similarity metric for the weight results S421: Measure the similarity between the distribution of the predicted value of the interbank lending amount and the true value distribution The present invention uses the weight similarity Ws (Weight similarity) to represent the similarity between the predicted weight distribution and the true distribution: , where represent the true weight values and predicted weight values of all nodes at time t + 1 respectively; are the weight mean and standard deviation of the graph respectively, are the weight mean and standard deviation of the predicted graph respectively; The α and β parameters represent the degree of emphasis on the weight mean and variance.
[0082] S422: Measure the error between the predicted value and the true value of the interbank lending amount The present invention uses the weight mean squared error W (WeightMean Squared Error, the weighted mean squared error) is used to measure the deviation between the predicted value and the true value of the asset attribute: W = , where represents the true weight value of all nodes at time t + 1, represents the predicted weight value of all nodes at time t + 1, represents the number of nodes.
[0083] S43: Measure the similarity index of node edges The present invention adopts density similarity (Density similarity, density similarity) to measure the similarity between the predicted value and the true value of the interbank lending relationship: . Where , , and respectively represent the true number of edges and the predicted number of edges of all nodes at time t + 1, represents the number of nodes.
[0084] S44: Measure the error of the node connection pattern The present invention adopts Kullback-Leibler Divergence Kullback-Leibler Divergence) to measure the similarity between the predicted result and the true result of the interbank connection pattern: , where represents the probability that the degree of the true network is k, and ϵ is a very small positive number to prevent the probability from being zero, usually taking or a smaller positive number, represents the number of nodes with degree k, represents the total degree.
[0085] S45: Model comparison The present invention illustrates the beneficial effects of the present model by the loss convergence effect of different models during training and the values of different indicators. Specifically as follows: S451: Comparison of model convergence efficiency The comparison of the model convergence effect is as Figure 4As shown in the figure, it can be seen that as the number of iterations increases, the loss function of the traditional GAN (the curve with triangles) oscillates greatly during the training process, and the convergence speed is slow, showing the characteristic of unstable training; while the HFAN-ITF (the curve with circles) converges significantly faster than the traditional GAN, the loss function is more stable, and the training is more stable. This shows that the optimization of HFAN-ITF for the traditional GAN has greatly improved the convergence efficiency and stability; after removing the SAEDC module from HFAN-ITF (the curve with squares), the convergence speed is slower than that of HFAN-ITF, verifying the role of the SAEDC module in further improving the convergence speed. Generally speaking, the WGAN mechanism significantly improves the stability of training, and the SAEDC module further accelerates the convergence speed of the model by optimizing the successive control density control, making the training more efficient.
[0086] S452: Comparison of Model Network Prediction Performance The comparison metrics used are: Ws (weight similarity), Ds (density similarity), KL (KL divergence), W-MSE (weight mean square error), and V-MSE (node feature mean square error).
[0087] The prediction results of the model for the network are as Figure 5 shown. From the experimental results, the HFAN-ITF model shows obvious advantages in various metrics. Especially in the Ws metric (0.3559), it is much higher than other models, indicating that it has the best performance in the accuracy of weight prediction. At the same time, in the Ds metric (0.9997), it reaches the optimal level equivalent to other models, and can well maintain the edge feature consistency between the generated data and the real data. In the KL metric (1.7386), the performance of this model is slightly better than most of the comparison models, indicating that it also has high accuracy in predicting the consistency between the data distribution and the real data network connection mode. Generally speaking, the HFAN-ITF model shows excellent performance in network weight and structure prediction.
[0088] S453: Comparison of Model Node Feature Prediction Performance The prediction results of the model in terms of node features are as Figure 6 shown. We compared with the current mainstream methods, and the results show that the V-MSE value of the HFAN-ITF model is significantly lower than that of LSTM, RNN, and Transformer. Among them, the error of HFAN-ITF is 0.112, which is better than 0.1756 of Transformer, reflecting its superior performance in the node feature prediction task.
[0089] S46: Ablation Experiment In the ablation experiment of the present invention, by gradually removing the key modules in this model, including data augmentation (Augment), graph convolutional network (GCN), graph attention network (GAT), Transformer, and SAEDC, the impact of each module on the model performance is analyzed. The results are as Figure 7 shown. The experimental results show that the removal of the SAEDC module has the greatest impact on performance metrics (such as Ds and KL), while other modules such as GAT, GCN, and Transformer also significantly contribute to the overall performance of the model to varying degrees. This indicates that the modules cooperate with each other in the model to jointly improve the generation effect.
[0090] In summary, on the one hand, the inter-bank network time series prediction method based on the generative adversarial hierarchical fusion model described in the present invention constructs a diverse dataset that conforms to the basic characteristics of the inter-bank network based on the current lack of inter-bank transaction data, making it possible to predict the inter-bank network; on the other hand, it proposes an inter-bank network time series prediction method based on the generative adversarial hierarchical fusion model, and accurately and efficiently predicts the inter-bank network through this method. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0091] The above description is only a description of the preferred embodiments of the present application and does not limit the scope of the present application in any way. Any change or modification made by any person skilled in the art based on the technical content disclosed above shall be regarded as an equivalent effective embodiment and fall within the scope of protection of the technical solution of the present application.
Claims
1. A method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model, characterized in that, It includes the following steps: S1: Generate inter-bank network time series data and perform data augmentation; Using the minimum density method, generate data with inter-bank network characteristics based on inter-bank balance sheet data, and expand the data through data augmentation to ensure data diversity, and then construct a diverse dataset that conforms to the basic characteristics of the inter-bank network; S2: Construct a generative adversarial hierarchical fusion model; S3: Divide the augmented inter-bank network time series data into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model; S4: Conduct error analysis on the predicted results and the original test set time series data, and perform model comparison and ablation experiments.
2. The method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that, In step S1, the specific method of generating inter-bank network time series data by the minimum density method is as follows: The minimum density method generates an inter-bank network according to the characteristics of "sparsity" and "assortativity" of inter-bank lending; The minimum density method is expressed as a constrained optimization problem with P(X) and X as the solution targets, where P(X) represents the inter-bank lending matrix and X represents the inter-bank lending amount; The optimization objective of "sparsity" is :[[-END]] Among them, represents the amount borrowed by bank i from bank j, represents the number of banks, represents the fixed cost of establishing connections among banks, is the total amount of interbank lending for bank , is the total amount of interbank borrowing for bank , where , ; The lending tendency of "assortativity" among banks is quantified by the entropy function to quantify this "assortativity", where ; Finally, combining the two parts of "sparsity" and "assortativity", the total optimization objective function of the minimum density method is designed as , where , represents the weight parameter of the "assortativity" objective; P(X) and X obtained by solving with the minimum density method respectively correspond to those at time t .
3. The method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that, In step S1, the data augmentation strategy includes parameter perturbation augmentation and data perturbation augmentation, which are specifically as follows: S121: Parameter perturbation augmentation Parameter perturbation enhances the key parameters in the minimum density network generation method and perform random perturbations: the parameter is calculated using the formula ; is calculated using the formula ; where the perturbation factor is a uniform distribution on (0, 1); Based on new parameters Generate a new interbank network; S122: Data perturbation augmentation For the bank 's inter-bank asset data Add random Gaussian noise and calculate using the formula , where the Gaussian noise ; According to the perturbed inter-bank asset data, use the minimum density method to generate an inter-bank network.
4. The method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that, In step S2, the generative adversarial hierarchical fusion model includes: a feature encoding module, a hierarchical fusion network HFN module, a network generation module, and a discriminator module; Among them, the feature encoding module, the HFN module, and the network generation module constitute the generator of the generative adversarial network, and the discriminator module is the discriminator of the generative adversarial network.
5. The method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model according to claim 4, characterized in that, The feature encoding module includes two parts: time encoding and node encoding: Time Encoding: The YearEncoder is used to process the year information in the input data, and an MLP is further used to map the time features into the hidden space. ; Node encoding: Use MLP to encode the input data G, NetEncoding(G)=MLP(G); After splicing the features of time encoding and node encoding, input them into the GCN network and the discriminator.
6. The method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model according to claim 4, characterized in that, The hierarchical fusion network HFN module includes three modules: a graph convolutional network GCN, a graph attention network GAT, and a Transformer; the HFN module captures the time series features and node attribute features of the inter-bank network through hierarchical feature fusion of the three modules of GCN, GAT, and Transformer in sequence; specifically as follows: S221: GCN module The GCN module extracts the local topological features of the interbank network for each time slice based on the input feature encoding, and generates preliminary node representations for each time slice. By stacking multiple layers of GCN, the high-order neighbor relationships of the interbank network are captured. The GCN input includes the adjacency matrix at time slice t and the node feature matrix ; The operation of GCN is defined by the following formula: Among them, is the adjacency matrix after adding self-loops, and self-loops are used to ensure that nodes incorporate their own information; is the identity matrix; Normalize the adjacency matrix to balance the contribution of node features and ensure the stability of training and the rationality of feature propagation; is a learnable projection matrix, represents the dimension of the input node features, the dimension of the hidden layer features, which is a hyperparameter of the GCN; S222: GAT module Adopt the attention mechanism of two-layer GAT to perform weighted modeling on the adjacency relationship of the inter-bank network. GAT captures the heterogeneous relationship between nodes by dynamically learning the weights of each node and its neighbors; The input of the GAT module includes the output features of the GCN at time slice t and the corresponding adjacency matrix , , where N represents the number of banks; The processing formula is ; Specifically, GAT generates the feature representation of each node through the following steps: 1) For node i and its neighbor j, node feature projection where is a learnable projection matrix that projects the node feature into a new feature space; 2) For node i and its neighbor j, use the attention mechanism to calculate the weight of edge (i, j): , Among them is a node feature is the weight between nodes i and j; a is a learnable attention parameter represents the feature concatenation operation concat; 3) Use the Softmax function to normalize the attention coefficients of each node to obtain , and aggregate the neighbor features and update the node representation according to the normalized attention weights , wherein is an activation function; To enhance the model's expressive power, the GAT layer uses the multi-head attention mechanism and finally takes the average to obtain the final node representation , where represents the neighbor set of node i, and K represents the total number of attention heads; S223: Transformer module The Transformer encoder is used to model the long-term dependence relationship of node features between multiple time slices; Output of GAT and the time encoding features are concatenated to obtain the input tensor of the Transformer encoder: Z Generate Q, K, and V through linear transformation, which are the query, key, and value vectors respectively, and serve as the input to the Transformer, where is the length of the input time series; then capture the dependencies between time slices by calculating the correlation between any two time slices in the sequence; the attention formula is: To avoid the vanishing gradient, residual connections are added after multi-head attention and the feed-forward network in Transformer, and layer normalization is used to stabilize the training. A two-layer fully connected network is used to process the features of each time slice; Final output .
7. According to the method for predicting the time series of the inter-bank network based on the generative adversarial hierarchical fusion model described in claim 4, it is characterized in that, The network generation module is used to generate the adjacency matrix of the next time slice ; an adaptive edge density control module SAEDC is introduced to control the generation of edges; Specifically, the input of this module Including: feature matrix ; historical network , including nodes and weights; time encoding , which is the time encoding information at time t + 1 and is used to characterize the time features of the prediction time slice; The processing process of this module is as follows: 1) To predict the edges at time t+1, it is first necessary to construct candidate edges in the network; specifically, candidate edges are first constructed for each pair of nodes (i, j). The candidate edges include historical edges and possible new edges, with a dimension of E. For each edge, its generated feature representation is , are the features of nodes respectively, is the time encoding, is the historical weight of the edge. 2) Edge feature matrix is fed into the edge generator; the edge generator is an MLP composed of a three-layer fully connected network, which processes the features of each candidate edge. The output of the edge generator includes three parts: The first part is the connection probability of edges , indicating whether nodes i and j are connected at time t + 1; The second part is the prediction of the edge weights , representing the predicted borrowing amount from bank i to bank j; The third part is the prediction of bank asset characteristics ; The overall formula is as follows: 3) To generate an adjacency matrix that conforms to the characteristics of the inter-bank network, the SAEDC method is used to screen candidate edges. The SAEDC method dynamically generates an edge screening threshold in the network based on the average degree of the network and the total number of network nodes N; specifically as follows: First calculate the number of target edges : , wherein is a learnable parameter; Then calculate the threshold , where P is the set of connection probabilities of all candidate edges, |P| is the total number of candidate edges, and Quantile represents the quantile function; Next, based on the edge probabilities generated by the MLP , these candidate edge sets are screened to obtain the screened edges, and the predicted values are used as the final weights of the edges; The finally output adjacency matrix It is expressed as: Predicted interbank network , including , representing the interbank edge and weight, and the characteristics of bank individuals i and j respectively.
8. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 4, wherein, The discriminator is used for the training of the model. The scoring of the network is completed by an MLP network, and its input is the prediction network at time slice t+1 and the real network , and the score is calculated through the formula ; the predicted network score is FakeScores, and the real network score is Real Scores; calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models; The scoring process of network G is as follows: 1) Since network G is the original complex data and cannot be directly used as the input for the subsequent network, after the same feature encoding processing as in the generator, is obtained and used as the input for the MLP; 2) Next, the MLP evaluates each edge and the corresponding node features and outputs a score. The scores of all edges and nodes are weighted and summed to obtain the total score of network G. where E refers to the set of all existing edges in network G, that is .
9. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, wherein, In step S3, the training of the generative adversarial hierarchical fusion model adopts the WGAN framework for joint training of the generator and the discriminator. Through the adversarial optimization of the generator and the discriminator, the distribution difference between the generative network and the real network is minimized; The loss function of the generator is set as: ; The discriminator aims to distinguish the generative network from the real adjacency network , and quantify the gap between the two for its own and the generator's model training; using the loss function: , to estimate the Wasserstein distance between the generated distribution and the true distribution, is the score of the true graph, the score of the generated graph; The calculation of the Wasserstein distance needs to satisfy the continuity condition, and for this purpose, a gradient penalty term is introduced: , wherein is a sample generated by interpolation and is defined as: , , The final loss function of the discriminator is: 。
Citation Information
Patent Citations
Financial time series prediction system and method based on generative adversarial network
CN114579640A
Long time sequence prediction method based on generative adversarial space-time attention network
CN119537799A
Time series prediction model training method based on data enhancement
CN119831086A
Multivariate time series anomaly detection method for intelligent internet of things system
WO2024207627A1