An Inter-bank Network Time Series Prediction Method Based on a Generative Adversarial Hierarchical Fusion Model

The interbank network data is generated through the minimum density method and the generative adversarial hierarchical fusion model is used, combined with data augmentation and adaptive edge density control, the problems of data loss and low training efficiency in interbank network timing prediction are solved, and more efficient interbank network prediction is achieved.

CN120180049BActive Publication Date: 2025-07-29TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510660707.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-29
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

There are problems in the existing interbank network timing prediction methods such as missing data, insufficient model feature capture capability and inefficient training efficiency. Especially when dealing with the commercial confidentiality and complex interaction relationships of interbank transaction data, traditional methods are difficult to effectively predict.

Method used

The minimum density method is used to generate interbank network feature data, the data set is expanded through data augmentation technology, and a generative adversarial hierarchical fusion model (HFAN-ITF) is constructed, and feature extraction is used using GCN, GAT and Transformer, and the Wasserstein distance and adaptive edge density control module (SAEDC) is optimized for model training.

Benefits of technology

It effectively alleviates the problem of data loss, improves the ability to capture the interaction relationship between interbank network structure and assets, improves prediction accuracy and training efficiency, and ensures the stability and convergence speed of the model in interbank network prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180049B_ABST
    Figure CN120180049B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for inter-bank network time series prediction based on a generative adversarial hierarchical fusion model. The inter-bank network is generated by the minimum density method, and data augmentation is performed by using the methods of parameter perturbation and data perturbation. The generative adversarial hierarchical fusion model includes a WGAN and a hierarchical fusion network HFN, where the HFN includes a graph convolutional network, a graph attention network, and a Transformer. In addition, the training of the model is optimized by the self-adaptive edge density control module SAEDC proposed by the present invention. The effectiveness of the method is verified through model comparison and ablation experiments. The purpose of the present invention is to be used for the time series prediction of the inter-bank network, and at the same time solve the limitations of the time series prediction model in terms of data, training efficiency, and prediction accuracy, so as to provide a reliable basis for the research in the field of bank systemic risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and network time series prediction, and in particular to an inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model. Background Art

[0002] As a core component of the modern financial system, the inter-bank market's dynamic changes directly affect the stability and efficiency of the financial system. It plays a crucial role, especially in aspects such as systemic risk identification, liquidity management, and evaluation of the effectiveness of monetary policies. In recent years, with the rapid development of the global financial market, the complexity of the inter-bank market network has increased significantly. How to effectively predict the inter-bank time series network, explore the bank systemic risks under the evolution of the time series network, and thus enhance the early warning capabilities of banks and regulatory authorities for potential financial crises has become the focus of current research.

[0003] The time series prediction of the inter-bank network plays a crucial role in the study of bank systemic risks. However, there are still the following three deficiencies in the existing inter-bank network time series prediction methods:

[0004] Firstly, for inter-bank network data, inter-bank transaction information is regarded as a commercial secret and is not publicly available. The lack of inter-bank transaction data caused by this commercial competition has directly hindered the research on inter-bank network prediction.

[0005] Secondly, in terms of the feature capture ability of the model, there is a complex interaction relationship between the inter-bank network and bank assets, that is, changes in bank assets will affect the bank network structure and transaction amounts, and changes in the bank network structure and transaction amounts will in turn affect bank assets. However, traditional time series machine learning methods such as RNN, LSTM, etc. cannot effectively capture this complex correlation between the bank network structure and bank asset features. In addition, when the historical data span is large (such as cross-year transaction patterns), the network structure changes significantly. Traditional graph convolutional networks GNN usually average when dealing with node neighbor relationships, and cannot distinguish the importance and contribution degrees of different inter-bank relationships, and have insufficient ability to identify key connections and important nodes in the network.

[0006] Thirdly, in terms of the efficiency of model training, in the initial training process of a model based on traditional GAN, when there is a large deviation between the real data distribution and the generated data distribution, it will lead to unstable training and slow convergence speed. The main manifestations are: in the initial stage, the networks generated by the generator are all judged as fake, resulting in the disappearance of the generator gradient and unstable model training; the initial edge screening threshold is high, resulting in a small number of edges generated by the generator, insufficient positive feedback of the model, and slow convergence speed of model training.

[0007] Therefore, how to provide an effective interbank network time series prediction method to solve the deficiencies of missing interbank network data, insufficient feature capture ability of the model, and low model training efficiency in existing work has become an urgent technical problem for those skilled in the art. Summary of the Invention

[0008] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to propose a solution based on the Hierarchical Fusion Wasserstein Generative Adversarial Network for Interbank Network Temporal Forecasting (HFAN-ITF).

[0009] First, the present invention uses the minimum density method to generate data with interbank network characteristics based on interbank balance sheet data, and expands the data through data augmentation to ensure data diversity, thereby constructing a diversified dataset that conforms to the basic characteristics of the interbank network.

[0010] Secondly, the present invention designs a hierarchical fusion network HFN in the generator part of the WGAN (Wasserstein Generative Adversarial Network) architecture. This network fuses the feature extraction information of GCN, GAT, and Transformer in sequence to achieve the purpose of effective feature extraction. Among them, the GCN network aggregates global information to extract the structural features of the network, and GAT assigns different weights to each node to aggregate neighbor node information, thereby mining the connection relationships of key nodes; the Transformer is used to model the time series network, and by assigning different weights to the feature vectors of different time steps, it fuses global time series features to predict the network of the next time step, so as to better understand the relationships between different time steps and multiple variables.

[0011] In addition, the present invention uses the Wasserstein distance in the discriminator to measure the distribution deviation between the generated network and the real network, making the training of the model smoother and improving the overall training effect of the model; the SAEDC (Self-adaptive Edge Density Control) adopts the self-adaptive threshold edge density control strategy SAEDC, increases the number of edges during initial training, and improves the convergence speed of the model.

[0012] Technical Solution

[0013] An interbank network time series prediction method based on a hierarchical fusion generative adversarial model, comprising the following steps:

[0014] S1: Generate the inter-bank network time series data and perform data augmentation;

[0015] Adopt the minimum density method to generate data with inter-bank network characteristics based on the inter-bank balance sheet data, and expand the data through data augmentation to ensure data diversity, and then construct a diversified data set that conforms to the basic characteristics of the inter-bank network.

[0016] S2: Construct a generative adversarial hierarchical fusion model;

[0017] S3: Divide the data-augmented inter-bank network time series data into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model;

[0018] S4: Conduct error analysis on the predicted results and the original test set time series data, and perform model comparison and ablation experiments.

[0019] Beneficial effects:

[0020] (1) The present invention generates data conforming to the inter-bank network characteristics based on the inter-bank balance sheet data through the minimum density method, and combines data augmentation technology to expand the data set, effectively alleviating the data missing problem caused by the non-disclosure of inter-bank transaction data due to commercial secrets.

[0021] (2) Through the hierarchical fusion network HFN, the present invention significantly improves the model's ability to capture the complex interaction relationship between the inter-bank network structure and bank assets, thereby improving the accuracy of inter-bank network prediction.

[0022] (3) By introducing the WGAN architecture and the SAEDC module, the present invention effectively alleviates the problems of gradient disappearance and training instability caused by excessive distribution deviation in the initial training stage of the traditional GAN model and the problem of sparse edges under the traditional edge density control strategy, thereby accelerating model convergence and improving the overall training efficiency. Description of the drawings

[0023] Figure 1 is the overall flowchart of the present invention;

[0024] Figure 2 is the division diagram of the HFAN-ITF module of the generative adversarial hierarchical fusion model of the present invention;

[0025] Figure 3 is the detailed design diagram of the HFAN-ITF of the generative adversarial hierarchical fusion model of the present invention;

[0026] Figure 4 is the experimental result comparison curve diagram of the model training convergence speed between the embodiment of the present invention and the benchmark model;

[0027] Figure 5 It is a comparison chart of the experimental results of the embodiments of the present invention with other models in terms of graph prediction metrics;

[0028] Figure 6 It is a comparison chart of the experimental results of the embodiments of the present invention with other models in terms of node prediction metrics;

[0029] Figure 7 It is a graph of the ablation experiment results of the embodiments of the present invention. Detailed implementation manners

[0030] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0031] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0032] A method for predicting the time series of the inter-bank network based on a generative adversarial hierarchical fusion model, including the following steps (as Figure 1 ):

[0033] S1: Generate the time series data of the inter-bank network and perform data augmentation

[0034] The present invention adopts the minimum density method to generate data with inter-bank network characteristics based on the inter-bank balance sheet data, and expands the data through data augmentation to ensure the diversity of the data, thereby constructing a diversified data set that conforms to the basic characteristics of the inter-bank network.

[0035] The inter-bank network data includes: the adjacency matrix of the lending network between banks and the node feature vectors in the network (basic features of banks: equity funds, total liabilities, total assets, inter-bank liabilities, inter-bank assets).

[0036] The inter-bank network time series data is represented as , where is the total time series length, and the time unit is years.

[0037] is the inter-bank network at time t, represented as , where is the adjacency matrix of the interbank lending network among N banks. , where represents the amount of borrowing from bank i to bank j . , represents the eigenvector of the nodes in the network (basic characteristics of banks: equity, total liabilities, total assets, interbank liabilities, interbank assets).

[0038] Specifically as follows:

[0039] S11: Generation of the interbank network based on the minimum density method

[0040] The present invention generates interbank network data by the minimum density method, and this network has the characteristics of a real interbank network. And the interbank balance sheet data is used as the characteristic data of banks to conduct time series prediction research on the data-driven interbank network.

[0041] The minimum density method generates an interbank network according to the characteristics of "sparsity" (i.e., the interbank network is a sparse matrix) and "assortativity" (i.e., there is a large difference in the asset sizes between the borrowing and lending banks) of interbank lending.

[0042] The minimum density method can be expressed as a constrained optimization problem with P(X) and X as the solution targets, where P(X) represents the interbank lending matrix and X represents the interbank lending amount.

[0043] The optimization objective of "sparsity" is :

[0044]

[0045] Among them, represents the amount of borrowing from bank i to bank j, represents the number of banks, represents the fixed cost of establishing a connection between banks, is the total interbank lending amount of bank , is the total interbank borrowing amount of bank , where , .

[0046] In addition, it is necessary to satisfy the lending tendency of "assortativity" among banks, and the entropy function is used to quantify this "assortativity", where .

[0047] Finally, combining these two parts of "sparsity" and "assortativity", the total optimization objective function of the minimum density method is designed as , where , a weight parameter representing the goal of "heterogamy".

[0048] The solution method of the minimum density method is a prior art.

[0049] The P(X) and X obtained by solving with the minimum density method respectively correspond to those at time t .

[0050] S12: Data augmentation

[0051] The present invention adopts a dual data augmentation strategy to expand the interbank network data. Specifically, the data augmentation strategy includes parameter perturbation augmentation and data perturbation augmentation, and network samples with diverse features are obtained through these two augmentations.

[0052] S121: Parameter perturbation augmentation

[0053] Parameter perturbation augmentation performs random perturbations on the key parameters and in the minimum density network generation method. In the real interbank market, the transaction costs and transaction propensities vary in different market environments, policy changes, bank strategy adjustments and other factors. Based on this real-world phenomenon, the present invention perturbs and augments the parameters in the network generation algorithm: the parameter is calculated using the formula , is calculated using the formula .

[0054] Among them, the perturbation factor is uniformly distributed on (0, 1), and generally takes 0.5, so that the perturbation range of the parameter is between 0.5 and 1.5, at a moderate level.

[0055] Based on the new parameter a new interbank network is generated.

[0056] Through this parameter perturbation method, different network structures can be generated while maintaining the basic topological features of the network.

[0057] S122: Data perturbation augmentation

[0058] The present invention perturbs the interbank lending assets and liabilities. Banks need to make some accounting estimates when preparing annual reports, such as loan loss provisions, asset impairment provisions, etc. These estimates are subjective and there will be deviations in the data size.

[0059] Based on this real-world phenomenon, the present invention perturbs the interbank asset data of the bank Add random Gaussian noise and use the formula for calculation, where the Gaussian noise .

[0060] Generate an interbank network using the minimum density method based on the perturbed interbank asset data.

[0061] S2: Construct a generative adversarial hierarchical fusion model

[0062] The generative adversarial hierarchical fusion model includes: a feature encoding module, a hierarchical fusion network HFN module, a network generation module, and a discriminator module. Its overall architecture is as shown in Figure 2 .

[0063] Among them, the feature encoding module, the HFN module, and the network generation module constitute the generator of the generative adversarial network, and the discriminator module is the discriminator of the generative adversarial network.

[0064] Input data is generated by step S1 (including generated data and augmented data), and includes the basic features of banks (such as equity funds, total liabilities, total assets, interbank liabilities, interbank assets) and the time-series interbank adjacency matrix (representing interbank lending relationships).

[0065] The hierarchical fusion network HFN module includes three modules: a graph convolutional network (Graph Convolutional Network, GCN), a graph attention network (Graph Attention Network, GAT), and a Transformer.

[0066] First, the HFN module performs feature encoding on the input data. Subsequently, the GCN layer and the GAT layer extract the structural features and important node information of the interbank network, and the Transformer encoder captures the time-series dependencies between multiple features.

[0067] Then, the network generation module obtains edges, weights, and bank attributes through an MLP (Multilayer Perceptron), and filters them through SAEDC to generate the prediction results of the interbank network (including the interbank network structure, network weights, bank features) at the next moment, denoted as .

[0068] During the training process, the MLP of the discriminator module scores the generated network and the real network respectively, generating Fake Scores (predicted network scores) and Real Scores (real network scores). Calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models.

[0069] The detailed processing flow of the generative adversarial hierarchical fusion model is as follows Figure 3 as shown below:

[0070] S21. Perform feature encoding on the input data

[0071] The feature encoding includes two parts: time encoding and node encoding:

[0072] Time encoding: Use YearEncoder (a prior art encoder for encoding year features) to process the year information in the input data, and use MLP to further map the time features to the hidden space, that is , where MLP consists of two stacked modules, each module contains a Linear (linear layer), LayerNorm (layer normalization), and a GELU activation function. After these two modules, apply LayerNorm (layer normalization) and Dropout (dropout layer);

[0073] Node encoding: Use MLP to encode the input data G, that is NetEncoding(G)=MLP(G), where the structure of MLP is the same as that of the MLP in the time encoding part;

[0074] S22: Use the HFN module to process the feature encoding

[0075] The HFN module captures the temporal features of the interbank network and the attribute features of the nodes through hierarchical feature fusion of three modules: GCN, GAT, and Transformer in sequence.

[0076] S221: GCN module

[0077] The GCN module extracts the local topological features of the interbank network for each time slice based on the input feature encoding, and generates a preliminary node representation for each time slice . By stacking multiple layers of GCN, higher-order neighbor relationships of the interbank network can be captured. The output representations of these layers will be used as the input of the GAT module for further modeling the relationships between nodes in the interbank lending network.

[0078] In this embodiment, three layers of GCN are adopted, and each layer of GCN contains three layers of GCNConv (graph convolutional layer), corresponding to Figure 3 3×GCNConv in

[0079] The GCN input includes the adjacency matrix of time slice t and the node feature matrix .

[0080] The operation of GCN is defined by the following formula:

[0081]

[0082] Among them, is the adjacency matrix after adding self-loops, and the self-loops are used to ensure that the nodes incorporate their own information; is the identity matrix; Normalize the adjacency matrix to balance the contribution of node features and ensure the stability of training and the rationality of feature propagation. is a learnable projection matrix, represents the dimension of the input node features, the dimension of the hidden layer features, which are hyperparameters of the GCN.

[0083] S222: GAT module

[0084] Adopt the attention mechanism of two-layer GAT to weight-model the adjacency relationship of the inter-bank network. Each layer of GAT contains three layers of GATConv (graph attention convolutional layer), corresponding to Figure 3 3×GATConv in. Different from GCN, GAT can more flexibly capture the heterogeneous relationships between nodes by dynamically learning the weights of each node and its neighbors.

[0085] The input of the GAT module includes the output features of the GCN at time slice t and the corresponding adjacency matrix , , N represents the number of banks. The processing formula is .

[0086] Specifically, GAT generates the feature representation of each node through the following steps:

[0087] 1) For node i and its neighbor j, project the node features Among them, is a learnable projection matrix that projects the node features into a new feature space.

[0088] 2) For node i and its neighbor j, use the attention mechanism to calculate the weight of edge (i, j):

[0089] ,

[0090] Among them is the feature of node , is the weight between nodes i and j. a is a learnable attention parameter, represents the feature concatenation operation concat.

[0091] 3) Use the Softmax function to normalize the attention coefficients of each node to obtain , and aggregate the neighbor features and update the node representation according to the normalized attention weights (where N(i) refers to the set of first-order neighbors of node i)

[0092] ,

[0093] where is the activation function.

[0094] To enhance the model's expressive power, the GAT layer uses the multi-head attention mechanism, and finally takes the average value to obtain the final node representation , where K represents the total number of attention heads. In this embodiment, 4 heads are adopted, corresponding to Figure 3 heads = 4 in

[0095] S223: Transformer module

[0096] In the time series prediction task of the inter-bank network, the high-order features of nodes not only depend on the topological relationship of the current time slice, but also need to capture the dynamic changes and time series dependencies between time slices. For this reason, this embodiment introduces a Transformer encoder after GCN and GAT to model the long-term dependencies of node features between multiple time slices.

[0097] The output of GAT is concatenated with the time-encoded features to obtain the input tensor of the Transformer encoder:

[0098] Z

[0099] Generate Q, K, and V through linear transformation, which are the query, key, and value vectors respectively, and serve as the input of the Transformer, where is the length of the input time series. Then, capture the dependencies between time slices by calculating the correlation (attention weights) between any two time slices in the sequence. The attention formula is:

[0100]

[0101] To avoid gradient disappearance, Transformer uses multi-head attention (8 heads are adopted in this embodiment, corresponding to Figure 3After adding residual connections to both the multi - head attention (with heads = 8) and the FeedForward network, and using LayerNorm (layer normalization) to stabilize the training, a two - layer fully - connected network is used to process the features of each time slice, further enhancing the feature expression ability. The Transformer encoder can effectively capture the dynamic dependencies in the inter - bank network time series by introducing the multi - head self - attention mechanism and the FeedForward network, thus enhancing the accuracy of time - series prediction.

[0102] Final output ( ). The output features of the Transformer encoder not only contain the global information between time slices but also retain the local topological features of the nodes, providing high - quality feature representations for the subsequent network generation module.

[0103] S23: Network generation module

[0104] The adjacency matrix for generating the next time slice in the network generation module . The principle of generating the adjacency matrix is to generate possible edges for each node, and these edges form an adjacency matrix. The present invention introduces an adaptive edge density control module SAEDC to control the generation of edges.

[0105] Specifically, the inputs of this module include: the feature matrix ; the historical network , including nodes and weights; the time encoding , which is the time - encoding information at time t + 1 and is used to characterize the time features of the predicted time slice.

[0106] The processing process of this module is as follows:

[0107] 1) To predict the edges at time t + 1, it is first necessary to construct candidate edges in the network. Specifically, candidate edges are first constructed for each pair of nodes (i, j). The candidate edges include historical edges and possible new edges, with a dimension of |E| (the number of candidate edges). For each edge, its generated feature representation is , are the features of nodes respectively, is the time encoding, is the historical weight of the edge (if the edge (i, j) is a new edge, it is initialized to 0).

[0108] 2) The edge feature matrix is fed into the edge generator ( ). The edge generator is an MLP, which contains two stacked modules, each module consisting of a Linear (linear layer), LayerNorm (layer normalization), ReLU activation function, and Dropout (dropout layer); after the processing of these two modules, a linear layer (Linear) is connected. The features of each candidate edge are processed by the MLP, and the output of the edge generator includes three parts:

[0109] The first part is the connection probability of the edge , indicating whether nodes i and j are connected at time t+1.

[0110] The second part is the prediction of the edge weight , representing the predicted borrowing amount from bank i to bank j.

[0111] The third part is the prediction of the bank asset characteristics .

[0112] The overall formula is as follows:

[0113]

[0114] 3) To generate an adjacency matrix that conforms to the characteristics of the inter-bank network, it is necessary to screen the candidate edges. The present invention proposes the SAEDC method to achieve edge screening, which dynamically generates the edge screening threshold in the network according to the average degree of the network and the total number of network nodes N.

[0115] Specifically, first calculate the target number of edges :

[0116] ,

[0117] where is a learnable parameter.

[0118] Then calculate the threshold , where P is the set of connection probabilities of all candidate edges, |P| is the total number of candidate edges, and Quantile represents the quantile function.

[0119] Next, according to the edge probability 0generated by the MLP, screen these candidate edge sets to obtain the screened edges, and use the predicted value as the final weight of the edge.

[0120] The finally output adjacency matrix is expressed as:

[0121]

[0122] The SAEDC module generates a predicted adjacency matrix that conforms to the topological characteristics of the interbank network by dynamically adjusting the number of target edges and a dynamic threshold screening mechanism based on the connection probability distribution. This module can flexibly adapt according to changes in the node scale and the average network degree. At the same time, it ensures that the number of generated edges is close to the target number of edges through quantile screening. Combining with the edge generator's prediction of the connection probability and edge weight , it effectively captures the dynamic evolution characteristics of the network. In the framework of the generative adversarial network, this design can improve the authenticity and rationality of the generated network, ensure that the topological structure of the generated graph is closer to the actual network, and at the same time provide high-quality adversarial samples for the discriminator, thus improving the training stability and prediction performance of the model.

[0123] After the above processing, the interbank network is finally predicted , including the interbank edges and weights, and the characteristics of bank individuals i and j.

[0124] S24: Discriminator module

[0125] The discriminator is mainly used for the training of the model. It completes the scoring of the network by an MLP network (the structure of this MLP is the same as that in the generator). Its input is the predicted network at time slice t + 1 and the real network , and calculates the score through the formula . The score of the predicted network is Fake Scores, and the score of the real network is Real Scores. Calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models.

[0126] The scoring process of network G (specifically the predicted network or the real network ) is as follows:

[0127] 1) Since network G is the original complex data and cannot be directly used as the input for the subsequent network, the same feature encoding process as in the generator is adopted (see step S21) to obtain as the input of the MLP.

[0128] 2) Then the MLP evaluates each edge and the corresponding node features and outputs the score ; perform a weighted sum of the scores of all edges and nodes to obtain the total score of network G, where E refers to the set of all existing edges in network G, that is . ​

[0129] S3: Divide the time-series data of the inter-bank network after data augmentation into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model.

[0130] In this embodiment, the bank data of China from 2012 to 2019 in the CSMAR database is used, and the time-series data from 2020 to 2022 is used for the test set. The WGAN framework is used for the joint training of the generator and the discriminator. Through the adversarial optimization of the generator and the discriminator, the distribution difference between the generative network and the real network is minimized. Each time the generator is updated, the discriminator is updated 5 times. WGAN overcomes the problems of instability and mode collapse in GAN training by introducing the Wasserstein distance to replace the Jensen-Shannon (JS) divergence of the traditional GAN.

[0131] The loss function of the generator of this model is set as:

[0132] .

[0133] The goal of the discriminator is to distinguish the generative network and the real adjacency network , and provide the gap between the two for the model training of itself and the generator. The loss function is used:

[0134] ,

[0135] to estimate the Wasserstein distance between the generative distribution and the real distribution, is the score of the real graph, the score of the generated graph.

[0136] Calculating the Wasserstein distance needs to satisfy the continuity condition. For this reason, a gradient penalty term is introduced:

[0137] ,

[0138] where is the sample generated by interpolation, defined as: , ,

[0139] The final loss function of the discriminator is:

[0140] .

[0141] The discriminator increases by minimizing and The differences enable it to more effectively distinguish between real samples and false samples. The WGAN framework optimizes the generator and discriminator through the Wasserstein distance, and combines the SAEDC module to improve the authenticity and dynamic adaptability of the generation network.

[0142] S4: Analyze the error between the predicted results and the original test set time series data, and conduct model comparison and ablation experiments.

[0143] To measure the beneficial effects of the present invention, the following indicators are used to measure the error of the model prediction results:

[0144] S41: Measure the error index of the node feature results

[0145] The present invention uses the vertex mean squared error (Vertex Mean Squared Error) to measure the deviation between the predicted value and the true value of the asset attribute: = , where represents the feature values of all nodes at time t + 1, represents the predicted feature values of all nodes at time t + 1, represents the number of nodes.

[0146] S42: Similarity index for the weight results

[0147] S421: Measure the similarity between the distribution of the predicted value of the interbank borrowing amount and the true value distribution

[0148] The present invention uses the weight similarity Ws (Weight similarity) to represent the similarity between the predicted weight distribution and the true distribution: , where respectively represent the true weight value and the predicted weight value of all nodes at time t + 1; are respectively the weight mean and standard deviation of the graph, are respectively the weight mean and standard deviation of the predicted graph; The α and β parameters represent the degree of emphasis on the weight mean and variance.

[0149] S422: Measure the error between the predicted value and the true value of the interbank borrowing amount

[0150] The present invention uses the weight mean squared error W (WeightMean Squared Error) to measure the deviation between the predicted value and the true value of the asset attribute: W = , where Denote the true weight values of all nodes at time t+1, Denote the predicted weight values of all nodes at time t+1, Denote the number of nodes.

[0151] S43: Measure the similarity index of node edges

[0152] The present invention adopts density similarity (Density similarity) to measure the similarity between the predicted value and the true value of the interbank lending relationship: . Where , , and respectively represent the true edge number and the predicted edge number of all nodes at time t+1, Denote the number of nodes.

[0153] S44: Measure the error of the node connection pattern

[0154] The present invention adopts Kullback-Leibler Divergence (Kullback-Leibler Divergence) to measure the similarity between the predicted result and the true result of the interbank connection pattern: , where represents the probability that the degree in the true network is k, and ϵ is a very small positive number to prevent the probability from being zero, usually taking or a smaller positive number, represents the number of nodes with degree k, represents the total degree.

[0155] S45: Model comparison

[0156] The present invention illustrates the beneficial effects of the present model through the loss convergence effect during training of different models and the values of different indicators. Specifically as follows:

[0157] S451: Comparison of model convergence efficiency

[0158] Compare the model convergence effects as Figure 4As shown in the figure, it can be seen that as the number of iterations increases, the traditional GAN (the curve with triangles) has a large oscillation in the loss function during training, a slow convergence speed, and shows the characteristic of unstable training; while the HFAN-ITF (the curve with circles) has a significantly faster convergence speed compared to the traditional GAN, the loss function is more stable, and the training is more stable. This indicates that the optimization of HFAN-ITF for the traditional GAN has a greater improvement in convergence efficiency and stability; after removing the SAEDC module from HFAN-ITF (the curve with squares), the convergence speed is slower than that of HFAN-ITF, verifying the role of the SAEDC module in further improving the convergence speed. Generally speaking, the WGAN mechanism significantly improves the stability of training, and the SAEDC module further accelerates the convergence speed of the model by optimizing the connection density control, making the training more efficient.

[0159] S452: Comparison of Model Network Prediction Performance

[0160] The comparison metrics used are: Ws (weight similarity), Ds (density similarity), KL (KL divergence), W-MSE (weight mean square error), and V-MSE (node feature mean square error).

[0161] The prediction results of the model for the network are as Figure 5 shown. From the experimental results, the HFAN-ITF model shows obvious advantages in various metrics. Especially in the Ws metric (0.3559), it is much higher than other models, indicating that it has the best performance in weight prediction accuracy. At the same time, in the Ds metric (0.9997), it reaches the optimal level equivalent to other models, and can well maintain the edge feature consistency between the generated data and the real data. In the KL metric (1.7386), the performance of this model is slightly better than most of the comparison models, indicating that it also has high accuracy in predicting the consistency between the data distribution and the real data network connection mode. Generally speaking, the HFAN-ITF model shows excellent performance in network weight and structure prediction.

[0162] S453: Comparison of Model Node Feature Prediction Performance

[0163] The prediction results of the model in terms of node features are as Figure 6 shown. We compared it with the current mainstream methods, and the results show that the V-MSE value of the HFAN-ITF model is significantly lower than that of LSTM, RNN, and Transformer. Among them, the error of HFAN-ITF is 0.112, which is better than 0.1756 of Transformer, reflecting its superior performance in the node feature prediction task.

[0164] S46: Ablation Experiment

[0165] In the ablation experiment of the present invention, by gradually removing the key modules in this model, including data augmentation (Augment), graph convolutional network (GCN), graph attention network (GAT), Transformer, and SAEDC, the impact of each module on the model performance is analyzed. The results are as Figure 7 shown. The experimental results show that the removal of the SAEDC module has the greatest impact on the performance metrics (such as Ds and KL), while other modules such as GAT, GCN, and Transformer also significantly contribute to the overall performance of the model to varying degrees. This indicates that the modules cooperate in the model to jointly improve the generation effect.

[0166] In summary, on the one hand, the inter-bank network time series prediction method based on the generative adversarial hierarchical fusion model described in the present invention constructs a set of diverse data sets that conform to the basic characteristics of the inter-bank network due to the current lack of inter-bank transaction data, making it possible to predict the inter-bank network; on the other hand, it proposes an inter-bank network time series prediction method of the generative adversarial hierarchical fusion model, and accurately and efficiently predicts the inter-bank network through this method. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0167] The above description is only a description of the preferred embodiments of the present application, and is not any limitation on the scope of the present application. Any change or modification made by any person skilled in the art according to the technical content disclosed above shall be regarded as an equivalent effective embodiment, and all belong to the scope protected by the technical solution of the present application.

Claims

1. A time series prediction method for interbank networks based on a generative adversarial hierarchical fusion model, characterized in that It includes the following steps: S1: Generate the inter-bank network time series data and perform data augmentation; Adopt the minimum density method to generate data with inter-bank network characteristics based on the inter-bank balance sheet data, and expand the data through data augmentation to ensure data diversity, and then construct a diverse dataset that conforms to the basic characteristics of the inter-bank network; S2: Construct a generative adversarial hierarchical fusion model; S3: Divide the augmented inter-bank network time series data into a training set and a test set, and use the training set to train the generative adversarial hierarchical fusion model; S4: Conduct error analysis on the predicted results and the original test set time series data, and perform model comparison and ablation experiments; In step S2, the generative adversarial hierarchical fusion model includes: a feature encoding module, a hierarchical fusion network HFN module, a network generation module, and a discriminator module; Among them, the feature encoding module, the HFN module, and the network generation module constitute the generator of the generative adversarial network, and the discriminator module is the discriminator of the generative adversarial network; The hierarchical fusion network HFN module includes three modules: a graph convolutional network GCN, a graph attention network GAT, and a Transformer; the HFN module captures the time series features of the inter-bank network and the attribute features of the nodes through hierarchical feature fusion of the three modules of GCN, GAT, and Transformer in sequence; The network generation module is used to generate the adjacency matrix of the next time slice ; an adaptive edge density control module SAEDC is introduced to control the generation of edges; In step S3, the training of the generative adversarial hierarchical fusion model adopts the WGAN framework for joint training of the generator and the discriminator, and through the adversarial optimization of the generator and the discriminator, the distribution difference between the generative network and the real network is minimized.

2. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that In step S1, the specific method for generating the inter-bank network time series data by the minimum density method is as follows: The minimum density method generates the inter-bank network according to the characteristics of "sparsity" and "assortativity" of inter-bank lending; The minimum density method is expressed as a constrained optimization problem with P(X) and X as the solution targets, where P(X) represents the inter-bank lending matrix and X represents the inter-bank lending amount; The optimization objective of "sparsity" is : Among them, represents the amount borrowed by bank i from bank j, represents the number of banks, represents the fixed cost of establishing connections among banks, is the total amount of interbank lending for bank , is the total amount of interbank borrowing for bank , where , ; The lending preference of "assortativity" among banks is quantified by the entropy function to quantify this "assortativity", where ; Finally, integrating the two parts of "sparsity" and "assortativity", the total optimization objective function of the minimum density method is designed as , where , represents the weight parameter of the "assortativity" objective; P(X) and X obtained by solving with the minimum density method respectively correspond to those at time t .

3. The inter-bank network time series prediction method based on the generative adversarial hierarchical fusion model according to claim 1, characterized in that In step S1, the data augmentation strategy includes parameter perturbation augmentation and data perturbation augmentation, which are specifically as follows: S121: Parameter perturbation augmentation Parameter perturbation enhancement for key parameters in the minimum density network generation method and perform random perturbations: the parameter is calculated using the formula ; is calculated using the formula ; where the perturbation factor is a uniform distribution on (0, 1); Based on new parameters Generate a new interbank network; S122: Data perturbation augmentation For the bank inter-bank asset data Add random Gaussian noise and calculate using the formula where the Gaussian noise ; According to the perturbed inter-bank asset data, the minimum density method is used to generate the inter-bank network.

4. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, wherein The feature encoding module includes two parts: time encoding and node encoding: Time Encoding: The YearEncoder is used to process the year information in the input data, and an MLP is further used to map the time features into the hidden space. ; Node encoding: Use MLP to encode the input data G, NetEncoding(G)=MLP(G); After splicing the features of the time encoding and the node encoding, input them into the GCN network and the discriminator.

5. The inter-bank network time series prediction method based on the generative adversarial hierarchical fusion model according to claim 1, characterized in that, The specific generative adversarial hierarchical fusion model is as follows: S221: GCN module The GCN module extracts the local topological features of the interbank network for each time slice based on the input feature encoding, and generates preliminary node representations for each time slice. By stacking multiple layers of GCN, the high-order neighbor relationships of the interbank network are captured. The GCN input includes the adjacency matrix at time slice t and the node feature matrix ; The operation of GCN is defined by the following formula: Among them, is the adjacency matrix after adding self-loops, and the self-loops are used to ensure that the nodes incorporate their own information; is the identity matrix; Normalize the adjacency matrix to balance the contribution of node features and ensure the stability of training and the rationality of feature propagation; is a learnable projection matrix, represents the dimension of the input node features, the dimension of the hidden layer features, which is a hyperparameter of the GCN; S222: GAT module Adopt the attention mechanism of two-layer GAT to perform weighted modeling on the adjacency relationship of the inter-bank network. GAT captures the heterogeneous relationship between nodes by dynamically learning the weights of each node and its neighbors; The input of the GAT module includes the output features of the GCN at time slice t and the corresponding adjacency matrix , , where N represents the number of banks; The processing formula is ; Specifically, GAT generates the feature representation of each node through the following steps: 1) For node i and its neighbor j, node feature projection where is a learnable projection matrix that projects the node feature into a new feature space; 2) For node i and its neighbor j, use the attention mechanism to calculate the weight of edge (i, j): , Among them is the node feature, is the weight between nodes i and j; a is a learnable attention parameter, represents the feature concatenation operation concat; 3) Use the Softmax function to normalize the attention coefficients of each node to obtain , and aggregate the neighbor features and update the node representation according to the normalized attention weights , Among them is the activation function; To enhance the model's expressive power, the GAT layer uses the multi-head attention mechanism and finally takes the average to obtain the final node representation , where represents the neighbor set of node i, and K represents the total number of attention heads; S223: Transformer Module The Transformer encoder is used to model the long-term dependencies of node features across multiple time slices; Output of GAT and the time encoding features are concatenated to obtain the input tensor of the Transformer encoder: Z Generate Q, K, and V through linear transformation, which are query, key, and value vectors respectively, and serve as the input to the Transformer, where is the length of the input time series; then capture the dependencies between time slices by calculating the correlation between any two time slices in the sequence; the attention formula is: To avoid gradient vanishing, residual connections are added after both the multi-head attention and the feed-forward network in the Transformer, and layer normalization is used to stabilize the training. A two-layer fully connected network is used to process the features of each time slice; Final output 。 6. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that, The inputs of the Adaptive Edge Density Control module SAEDC include: a feature matrix ; a historical network , including nodes and weights; a time encoding , which is the time encoding information at time t+1 and is used to characterize the time features of the prediction time slice; The processing procedure of the adaptive edge density control module is as follows: 1) To predict the edges at time t+1, it is first necessary to construct candidate edges in the network; specifically, candidate edges are first constructed for each pair of nodes (i, j). The candidate edges include historical edges and possible new edges, with a dimension of E; for each edge, its generated feature representation is , are the features of nodes respectively, is the time encoding, is the historical weight of the edge. 2) Edge feature matrix is fed into the edge generator; the edge generator is an MLP composed of a three-layer fully connected network, which processes the features of each candidate edge. The output of the edge generator includes three parts: The first part is the connection probability of edges , indicating whether nodes i and j are connected at time t + 1; The second part is the prediction of the edge weights , representing the predicted borrowing amount from bank i to bank j; The third part is the prediction of the characteristics of bank assets ; The overall formula is as follows: 3) To generate an adjacency matrix that conforms to the characteristics of the interbank network, the SAEDC method is used to screen candidate edges. The SAEDC method dynamically generates an edge screening threshold in the network based on the average degree of the network and the total number of network nodes N; specifically as follows: Calculate the number of target edges first : , wherein is a learnable parameter; Then calculate the threshold , where P is the set of connection probabilities of all candidate edges, |P| is the total number of candidate edges, and Quantile represents the quantile function; Next, based on the edge probabilities generated by the MLP , these candidate edge sets are screened to obtain the screened edges, and the predicted values are used as the final weights of the edges; The finally output adjacency matrix It is expressed as: Predicted interbank network , including , representing the interbank edge and weight, and the characteristics of bank individuals i and j respectively.

7. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that The discriminator is used for the training of the model. The scoring of the network is completed by an MLP network, and its input is the prediction network at time slice t+1 and the real network , and the score is calculated through the formula ; the score of the prediction network is FakeScores, and the score of the real network is Real Scores; calculate the Wasserstein distance between Fake Scores and Real Scores to train the generator and discriminator models; The scoring procedure of network G is as follows: 1) Since network G is the original complex data and cannot be directly used as the input of the subsequent network, after the same feature encoding processing as in the generator, is obtained as the input of the MLP; ​ 2) Next, the MLP evaluates each edge and the corresponding node features and outputs a score. The scores of all edges and nodes are weighted and summed to obtain the total score of network G. where E refers to the set of all existing edges in network G, that is .

8. The inter-bank network time series prediction method based on a generative adversarial hierarchical fusion model according to claim 1, characterized in that, In the described generative adversarial hierarchical fusion model, The generator loss function is set as: ; The discriminator aims to distinguish between the generative network and the real adjacency network and train its own and the generator's models by quantifying the gap between the two; using the loss function: , To estimate the Wasserstein distance between the generated distribution and the true distribution, is the score of the true graph, the score of the generated graph; The calculation of the Wasserstein distance needs to satisfy the continuity condition, and for this purpose, a gradient penalty term is introduced: , wherein is a sample generated by interpolation and is defined as: , , The final loss function of the discriminator is: 。

Citation Information

Patent Citations

  • Long time sequence prediction method based on generative adversarial space-time attention network

    CN119537799A

  • Time series prediction model training method based on data enhancement

    CN119831086A