Risk Prediction Method and System Based on Dynamic Aggregation and Privacy Protection
By real-time acquisition of multi-source heterogeneous data, construction of local risk correlation graphs, and privacy-preserving computation using federated graph neural networks, the problems of data silos, static analysis, and privacy leaks in existing risk data management systems have been solved, enabling real-time risk prediction and improved accuracy.
Patent Information
- Application Number
- CN202510780015.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing risk data management systems suffer from problems such as data silos, static analysis, privacy leaks, and model bias. They struggle to achieve dynamic aggregation of multi-source heterogeneous data and real-time risk prediction, and lack effective feature extraction and fusion mechanisms.
By collecting and standardizing multi-source heterogeneous data in real time, a local risk association map is constructed. Then, a federated graph neural network is used for privacy-preserving computation, multi-level feature extraction and risk level classification are performed to achieve data collaboration and risk knowledge fusion among multiple institutions.
It enables real-time aggregation of multi-source heterogeneous data, improves the timeliness and accuracy of risk response, protects data privacy, and enhances the interpretability of the model and the accuracy of risk prediction.
Smart Images

Figure CN120316649B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a risk prediction method and system based on dynamic aggregation and privacy protection. Background Technology
[0002] With the development of big data technology, risk data management is playing an increasingly important role in high-risk industries such as finance, healthcare, and supply chain. Current risk data management primarily relies on centralized storage and processing architectures, collecting and integrating risk data from different business systems through data warehouses or data lakes to build risk models for risk assessment and prediction. These systems typically employ traditional machine learning algorithms, such as logistic regression, decision trees, or neural networks, training models based on historical data to identify potential risks.
[0003] However, existing risk data management technologies face four key challenges: First, the problem of data silos, where risk data is scattered across different systems and institutions, lacking standardized aggregation methods, resulting in an incomplete risk view; second, the problem of static analysis, where traditional models mainly rely on batch processing of historical data, failing to respond in real time to dynamically changing risk environments; third, the problem of privacy breaches, where data sharing among multiple institutions lacks effective anonymization and encryption mechanisms, posing data privacy risks; and fourth, the problem of model bias, where existing risk assessment models have low interpretability and fail to effectively integrate multi-source heterogeneous data, leading to inaccurate risk assessments.
[0004] Furthermore, as risk scenarios become increasingly complex, existing technologies exhibit limitations in handling multi-dimensional risk characteristics. In particular, when it is necessary to simultaneously consider risk factors at the micro-individual, meso-relational, and macro-system levels, there is a lack of effective feature extraction and fusion mechanisms. Moreover, in multi-institutional collaborative environments, how to achieve joint learning of graph-structured data while ensuring privacy and security, and how to construct predictive systems that can adapt to highly dynamic risk environments while maintaining model interpretability, remain pressing technical challenges that need to be addressed. Summary of the Invention
[0005] This application provides a risk prediction method and system based on dynamic aggregation and privacy protection, which is used to realize the dynamic aggregation of multi-source heterogeneous data, protect data privacy while realizing data collaboration among multiple institutions, and improve the accuracy and interpretability of risk prediction through multi-level feature extraction. It can perform cross-institutional risk knowledge fusion without sharing the original data, and ensure the real-time and accuracy of risk prediction results.
[0006] Firstly, this application provides a risk prediction method based on dynamic aggregation and privacy protection. The method includes: real-time acquisition and standardization of multi-source heterogeneous data to obtain a standardized data stream; constructing a local risk association graph within each participating institution based on the standardized data stream, the local risk association graph containing nodes, edges, and topological attributes; inputting the local risk association graph into a federated graph neural network for privacy-preserving computation to obtain a risk representation integrating risk knowledge from multiple institutions, the privacy-preserving computation using a local differential privacy mechanism; and performing multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal.
[0007] Optionally, the real-time acquisition and standardization processing of multi-source heterogeneous data to obtain a standardized data stream includes:
[0008] The multi-threaded data collector gathers raw data from internal trading systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors to obtain multi-source heterogeneous data.
[0009] Key fields are extracted from the structured data in the multi-source heterogeneous data; nested attributes are extracted from the semi-structured data using JSON / XML parsing algorithms; and semantic information is extracted from the unstructured data using natural language processing to obtain the processed data.
[0010] The processed data is formatted uniformly by converting heterogeneous data into UTF-8 encoding format and applying entity parsing technology to identify and merge different representations of the same entity to obtain data in a unified format.
[0011] The numerical features in the unified format data are standardized using the Z-Score normalization algorithm, the categorical features are converted into numerical representations using one-hot encoding, and the data is then subjected to quality checks and repairs to obtain the standardized data stream.
[0012] Optionally, the construction of a local risk association graph within each participating institution based on the standardized data flow, wherein the local risk association graph includes nodes, edges, and topological attributes, including:
[0013] Node types are defined from the standardized data stream according to business rules, including customer entity nodes, transaction event nodes, asset project nodes and risk signal nodes, and a specific set of attributes is assigned to each type of node to obtain a graph node set;
[0014] Based on the data relationship, the edge types are defined, including transaction relationship edges, association relationship edges, and influence relationship edges. The connection relationship between related nodes is converted into edges in the graph, resulting in a graph edge set.
[0015] Incremental updates are performed on the graph node set and graph edge set. New data is converted into nodes or edges in the graph in real time. Attribute value updates and node state changes are performed on nodes, and weight adjustments and timeliness updates are performed on edges to obtain a dynamically updated graph.
[0016] Calculate time decay weights for nodes and edges in the dynamically updated graph to gradually reduce the influence of historical data over time, thus obtaining a time-sensitive graph.
[0017] The time-sensitive graph is partitioned using parallel computing, and the graph database storage structure is used to optimize graph traversal and query performance, resulting in a structure-optimized graph.
[0018] Calculate the topological attributes of the structure optimization graph, including connectivity, community structure, and centrality measure, and save the calculation results as graph attribute values to obtain the local risk association graph.
[0019] Optionally, the step of inputting the local risk association graph into a federated graph neural network for privacy-preserving computation to obtain a risk representation that integrates risk knowledge from multiple institutions, wherein the privacy-preserving computation uses a local differential privacy mechanism, including:
[0020] In each participating institution, a local graph convolutional structure of a federated graph neural network is constructed, including a risk entity node layer, a risk association layer, and a risk feature layer. The risk association layer performs risk propagation analysis by aggregating risk entities and related transaction information to obtain the local risk network structure.
[0021] The local risk association map is input into the local risk network structure, and an initial risk embedding representation is obtained through layer-by-layer risk information aggregation and nonlinear transformation processing.
[0022] A sensitive information pruning operation is performed on the gradient of the initial risk embedding representation to limit the gradient values related to customer sensitive information within a preset threshold range, thereby obtaining the risk information pruned gradient.
[0023] After cropping the risk information, Laplace noise is added to the gradient to obtain the privacy-preserving risk gradient;
[0024] The privacy protection risk gradient is uploaded to the risk coordination server, and the risk gradients of all participating institutions are weighted and aggregated. Weights are assigned according to the risk data distribution of each institution to obtain the federal risk knowledge update value.
[0025] The updated federal risk knowledge values are distributed to each participating institution to update the parameters of the local risk graph convolutional network. At the same time, anonymous risk transmission relationships are exchanged through a privacy-preserving risk association extension protocol to obtain the risk representation that integrates the risk knowledge of multiple institutions.
[0026] Optionally, the step of inputting the local risk association map into the local risk network structure, and obtaining the initial risk embedding representation through layer-by-layer risk information aggregation and nonlinear transformation processing, includes:
[0027] Node features are extracted from the local risk association graph to obtain an initial node feature matrix;
[0028] Based on the local risk association graph, an adjacency matrix and a degree matrix are generated, and a normalized adjacency matrix is calculated to obtain a standardized graph structure representation.
[0029] The initial node feature matrix is multiplied by the standardized graph structure representation, and then transformed through the first-layer weight matrix to obtain the first-order risk propagation feature;
[0030] A second information propagation and transformation is performed on the first-order risk propagation feature, and the second-order risk propagation feature is obtained by processing through the second-layer weight matrix.
[0031] The second-order risk propagation features are subjected to skip connection processing, and the initial features, first-order features and second-order features are concatenated to obtain a multi-scale risk feature representation;
[0032] The multi-scale risk feature representation is weighted and aggregated by a layer-by-layer attention mechanism, and the importance weights of the features at each layer are calculated and fused to obtain the initial risk embedding representation.
[0033] Optionally, the step of weighted aggregation of the multi-scale risk feature representation through a layer-by-layer attention mechanism, calculating the importance weights of features at each layer and fusing them to obtain the initial risk embedding representation includes:
[0034] For the multi-scale risk feature representation, a query transformation matrix, a key transformation matrix, and a value transformation matrix are constructed respectively. Each matrix contains specific parameter weights. The risk features are projected onto three different feature spaces through matrix multiplication to obtain the risk query vector, the risk key vector, and the risk value vector.
[0035] Calculate the product of the risk query vector and the transpose of the risk key vector, divide the result by the square root of 64 to scale and adjust, construct a quantitative index of the influence between risk nodes, and obtain the risk association strength matrix.
[0036] Perform a softmax operation on each row of the risk association strength matrix to convert each element into a value between 0 and 1, with each row summing to 1, to form the probability distribution of the risk propagation path and obtain the risk attention weight matrix.
[0037] Multiply the risk attention weight matrix with the risk value vector, filter and aggregate the information of each risk node according to its importance, and obtain a single-head risk attention output;
[0038] The original risk feature vector is divided into 8 independent sub-vectors. For each sub-vector, the transformation matrix is repeatedly constructed, the correlation strength is calculated, the transformation is converted into a weight distribution, and weighted aggregation is performed to obtain 8 sets of independent risk feature representations, thus obtaining the multi-head risk attention result.
[0039] The multi-head risk attention results are concatenated along the feature dimension to form complete risk features. The output transformation matrix is then transformed back to the original dimensional space, and the original risk feature vector is added to perform residual connection operations to obtain the initial risk embedding representation.
[0040] Optionally, the step of performing multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal includes:
[0041] Micro-level individual features are extracted from the risk representation that integrates risk knowledge from multiple institutions to obtain a micro-risk feature set, which includes: customer financial status, credit history and behavioral patterns;
[0042] Based on the correlations in the risk representation, meso-level relationship features are extracted to obtain a meso-level risk feature set;
[0043] Based on the overall structure of the risk representation, macro-system hierarchical features are extracted to obtain a macro-risk feature set;
[0044] The risk prediction results are obtained by fusing the micro-risk feature set, meso-risk feature set and macro-risk feature set and calculating the risk probability distribution.
[0045] The risk prediction results are classified into four levels: low risk, medium risk, high risk, and extremely high risk. Corresponding risk management strategies are set for each level to obtain the risk level classification results.
[0046] Based on the risk level classification results, a risk warning signal is generated, including risk score, key risk factors, risk development trend and warning confidence level, and different levels of warning mechanisms are triggered for different risk levels to obtain the comprehensive risk score and warning signal.
[0047] Secondly, this application provides a risk prediction system based on dynamic aggregation and privacy protection, the risk prediction system based on dynamic aggregation and privacy protection includes:
[0048] The acquisition module is used to acquire and standardize multi-source heterogeneous data in real time to obtain a standardized data stream.
[0049] The construction module is used to construct a local risk association graph within each participating institution based on the standardized data flow. The local risk association graph includes nodes, edges, and topological attributes.
[0050] The calculation module is used to input the local risk association graph into the federated graph neural network for privacy-preserving calculation to obtain a risk representation that integrates risk knowledge from multiple institutions. The privacy-preserving calculation uses a local differential privacy mechanism.
[0051] The segmentation module is used to perform multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal.
[0052] Thirdly, a risk prediction device based on dynamic aggregation and privacy protection is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the risk prediction device based on dynamic aggregation and privacy protection to execute the aforementioned risk prediction method based on dynamic aggregation and privacy protection.
[0053] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the aforementioned risk prediction method based on dynamic aggregation and privacy protection.
[0054] The technical solution provided in this application effectively solves key technical problems in existing risk data management and achieves significant technical results through features such as real-time acquisition and standardized processing of multi-source heterogeneous data, construction of local risk association graphs, privacy-preserving computation using federated graph neural networks, and multi-level feature extraction and risk level classification. The real-time acquisition and standardized processing mechanism for multi-source heterogeneous data breaks through the data silo limitations of traditional risk management systems. By integrating risk data scattered across different systems through a unified data standardization process, it achieves millisecond-level data update latency, improving risk response from the traditional hourly level to the second level, significantly enhancing the timeliness of risk monitoring. Secondly, the local risk association graph constructed within each participating institution includes nodes, edges, and topological attributes. This graph-structured data representation not only preserves the complex relationships between entities but also maintains high sensitivity to the latest risk status through a time decay weight mechanism, solving the problem that traditional static analysis models cannot respond to dynamic risks in real time. Third, the local risk correlation graph is input into the federated graph neural network for privacy-preserving computation. Employing a local differential privacy mechanism, it achieves the fusion of risk knowledge from multiple institutions without sharing the original data, thus protecting sensitive data privacy while breaking down data barriers between institutions. Finally, multi-level feature extraction and risk level classification are performed on the risk representation, comprehensively capturing risk signals from three levels: micro-individual, meso-relationship, and macro-system. Visualized risk interpretation reports enhance the model's interpretability.
[0055] Federated graph neural network algorithm, as the core technology, captures high-order relationship patterns between entities through graph convolution operations, identifies key risk transmission paths by combining attention mechanisms, and protects sensitive gradient information through local differential privacy. This algorithm not only ensures data privacy and security, but also gathers risk knowledge from multiple institutions through distributed learning, thus significantly improving risk identification capabilities. Attached Figure Description
[0056] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a schematic diagram of an embodiment of the risk prediction method based on dynamic aggregation and privacy protection in this application.
[0058] Figure 2 This is a schematic diagram of an embodiment of the risk prediction system based on dynamic aggregation and privacy protection in this application.
[0059] Figure 3 This is a schematic block diagram of the risk prediction device based on dynamic aggregation and privacy protection in an embodiment of the present invention. Detailed Implementation
[0060] This application provides a risk prediction method and system based on dynamic aggregation and privacy protection. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0061] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the risk prediction method based on dynamic aggregation and privacy protection in this application includes:
[0062] Step S101: Real-time acquisition and standardization processing of multi-source heterogeneous data to obtain a standardized data stream;
[0063] Step S102: Based on standardized data flow, construct a local risk association graph within each participating institution. The local risk association graph includes nodes, edges, and topological attributes.
[0064] Step S103: Input the local risk association graph into the federated graph neural network for privacy-preserving computation to obtain a risk representation that integrates risk knowledge from multiple institutions. The privacy-preserving computation uses a local differential privacy mechanism.
[0065] Step S104: Perform multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal.
[0066] It is understood that the executing entity of this application can be a risk prediction system based on dynamic aggregation and privacy protection, or it can be a terminal or a server; no specific limitation is made here. This application's embodiments use a server as an example for illustration.
[0067] Specifically, real-time acquisition and standardization of multi-source heterogeneous data are performed. A multi-threaded data collector gathers raw data from internal transaction systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors. The multi-threaded data collector is a parallel data acquisition tool that significantly reduces data acquisition latency by creating multiple data reading threads simultaneously, keeping the latency below 50 milliseconds. The acquired heterogeneous data undergoes format unification processing, converting data with different encodings to UTF-8 format to ensure consistency in subsequent processing. Simultaneously, numerical features are standardized using the Z-Score normalization algorithm. Z-Score normalization transforms the raw data into a standard normal distribution with a mean of 0 and a standard deviation of 1, calculated as (x-μ) / σ, where x is the original value, μ is the mean, and σ is the standard deviation. This processing allows for effective comparison and analysis of features of different magnitudes. Categorical features are converted to numerical representation using one-hot encoding, enabling discrete variables to be effectively processed by machine learning algorithms.
[0068] A local risk correlation graph is constructed based on standardized data flow, defining node types including customer entity nodes, transaction event nodes, asset project nodes, and risk signal nodes. Incremental update technology is applied during graph construction, converting new data into nodes or edges in the graph in real time to maintain the timeliness of risk information. Time decay weight is an influence factor that weakens exponentially over time, calculated as λ^t, where λ is the decay coefficient (typically between 0.9 and 0.99), and t is the time interval, giving higher weight to recent risk events. This construction method solves the static analysis problem in traditional risk assessment, making risk assessment more real-time and dynamically adaptable. In financial risk prediction, this graph can clearly show the flow of funds and the path of risk transmission, such as the correlation between an abnormally large transaction (transaction event node) and multiple high-risk enterprises (customer entity nodes), as well as the evolution trend of these correlations over time.
[0069] Local risk correlation graphs are input into federated graph neural networks for privacy-preserving computation. Federated graph neural networks are an innovative distributed learning framework that combines the advantages of federated learning and graph neural networks, allowing multiple institutions to collaboratively train models without sharing raw data. In the field of financial risk management, multiple banks can use this network to jointly build risk prediction models without disclosing customer information to each other. Privacy-preserving computation employs a local differential privacy mechanism, adding carefully calibrated Laplace noise to the gradients to prevent the leakage of sensitive information. Gradient pruning limits gradient values to a preset threshold range, preventing a single outlier from having an excessive impact on the model, while the addition of Laplace noise ensures that individual privacy is protected during model updates. This approach effectively solves the privacy leakage problem faced by traditional risk data marts, enabling financial institutions to share risk knowledge while protecting data privacy. Multi-level feature extraction and risk level classification are performed on the risk representation that integrates risk knowledge from multiple institutions. Multi-level feature extraction includes three dimensions: micro-level individual, meso-level relationship, and macro-level system, comprehensively capturing risk signals. The micro-level focuses on the risk characteristics of individual entities, the meso-level on the risk transmission relationships between entities, and the macro-level on the overall risk situation. Ensemble learning algorithms fuse risk characteristics from different levels to generate a comprehensive risk score. Ensemble learning combines the results of multiple basic predictors (such as decision trees, neural networks, and logistic regression), leveraging the complementary advantages of different algorithms to improve prediction accuracy. In supply chain risk management, this method can simultaneously consider the supplier's own financial situation (micro-level characteristics), the interdependence between upstream and downstream enterprises (meso-level characteristics), and the overall market environment (macro-level characteristics), thereby generating more comprehensive and accurate risk assessment results. Risk level classification categorizes the prediction results into four levels: low risk, medium risk, high risk, and extremely high risk, triggering corresponding early warning mechanisms to support risk management decisions and effectively addressing the problem of low interpretability in traditional models.
[0070] In one specific embodiment, the process of performing step S101 may specifically include the following steps:
[0071] The multi-threaded data collector gathers raw data from internal trading systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors to obtain multi-source heterogeneous data.
[0072] Key fields are extracted from structured data in multi-source heterogeneous data; nested attributes are extracted from semi-structured data using JSON / XML parsing algorithms; and semantic information is extracted from unstructured data using natural language processing, resulting in processed data.
[0073] The processed data is formatted uniformly by converting heterogeneous data into UTF-8 encoding format and applying entity parsing technology to identify and merge different representations of the same entity to obtain data in a unified format.
[0074] Numerical features in the unified format data are standardized using the Z-Score normalization algorithm, categorical features are converted into numerical representations using one-hot encoding, and the data is then subjected to quality checks and repairs to obtain a standardized data stream.
[0075] Specifically, the multi-threaded data collector gathers raw data from multiple data sources, including internal transaction systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors. Based on thread pool technology, the collector allocates an independent thread to each data source and employs a message queue mechanism to manage data transmission, ensuring data acquisition latency is controlled at the millisecond level. Data collected from the internal transaction system is primarily structured data, including transaction amount, transaction time, and parties involved. Data from customer behavior logs records customer activity patterns, while data from social media platforms mainly consists of unstructured text content. Different processing methods are used for different types of data when processing the collected multi-source heterogeneous data. For structured data, key fields are extracted using a predefined field mapping table, such as extracting transaction amount, transaction time, and parties involved from transaction records. For semi-structured data such as JSON or XML, specialized parsing algorithms are used. JSON parsing algorithms convert JSON strings into in-memory object structures through lexical and syntactic analysis, extracting nested attributes. XML parsing algorithms use DOM or SAX parsers to analyze the elements, attributes, and content of XML documents, forming tree structures or event streams from which the desired information is extracted. For unstructured data, such as social media text, natural language processing techniques are used to extract semantic information, including word segmentation, part-of-speech tagging, and named entity recognition, converting free text into structured feature vectors and capturing risk-related information within the text. When unifying the format of the processed data, data with different encodings are uniformly converted to UTF-8 encoding to ensure correct processing of multilingual text. Then, entity parsing techniques are applied to identify and merge different representations of the same entity. Entity parsing technology determines whether different data records point to the same entity by comparing the attribute similarity of entities. This technology includes four steps: attribute standardization, blocking algorithms, similarity calculation, and decision rules. Attribute standardization converts attributes of different formats into a unified format; blocking algorithms group potentially similar records into the same block, reducing the number of comparisons; similarity calculation uses algorithms such as edit distance and Jaccard coefficient to calculate the degree of similarity between records; decision rules determine whether records belong to the same entity based on a similarity threshold. Through entity resolution, multiple records from different sources belonging to the same customer are identified and merged to form a unified view. When performing feature processing on the unified format data, numerical features are standardized using the Z-Score normalization algorithm, which subtracts the mean from the original value and divides by the standard deviation, making features of different dimensions comparable. Categorical features are converted into numerical representations using one-hot encoding, mapping each category to a binary vector with only one element being 1 and the rest being 0, enabling machine learning algorithms to process categorical data.In addition, data quality inspection and repair are performed, including null value detection, outlier identification, consistency verification and integrity verification. Repair strategies such as filling, correction or filtering are applied to the detected problematic data.
[0076] In one specific embodiment, the process of performing step S102 may specifically include the following steps:
[0077] Node types are defined from the standardized data flow according to business rules, including customer entity nodes, transaction event nodes, asset project nodes and risk signal nodes, and a specific set of attributes is assigned to each type of node to obtain the graph node set;
[0078] Based on the data relationship, the edge types are defined, including transaction relationship edges, association relationship edges, and influence relationship edges. The connection relationship between related nodes is converted into edges in the graph, resulting in a graph edge set.
[0079] Incremental updates are performed on the graph node set and graph edge set. New data is converted into nodes or edges in the graph in real time. Attribute value updates and node state changes are performed on nodes, and weight adjustments and timeliness updates are performed on edges to obtain a dynamically updated graph.
[0080] To dynamically update the nodes and edges in the graph, time decay weights are calculated so that the influence of historical data gradually weakens over time, resulting in a time-sensitive graph.
[0081] Parallel computing is used to partition the time-sensitive graph, and a graph database storage structure is used to optimize graph traversal and query performance, resulting in a structure-optimized graph.
[0082] The topological properties of the computational structure optimization graph, including connectivity, community structure, and centrality measure, are calculated and stored as graph attribute values to obtain a local risk association graph.
[0083] Specifically, node types are defined from the standardized data flow according to business rules. These node types include customer entity nodes, transaction event nodes, asset project nodes, and risk signal nodes. Customer entity nodes represent individuals or institutions participating in transactions, with attributes including basic information, credit rating, and financial status. Transaction event nodes record specific transaction behaviors, with attributes including transaction time, amount, and direction. Asset project nodes represent financial assets or investment projects, with attributes including asset type, value, and risk level. Risk signal nodes are risk warning information extracted from external data sources, with attributes including signal source, strength, and credibility. These nodes are uniquely identified by identifiers (IDs), forming a graph node set. When defining edge types based on data relationships, edge types are mainly divided into transaction relationship edges, association relationship edges, and impact relationship edges. Transaction relationship edges connect customer entity nodes and transaction event nodes, indicating who participated in which transaction. Association relationship edges connect entities with business relationships, such as supplier relationships or holding relationships. Impact relationship edges connect risk signal nodes with potentially affected entities or assets. Each edge type has its own set of attributes. For example, transaction relationship edges include attributes such as transaction direction and amount ratio; association relationship edges include attributes such as relationship type and closeness; and influence relationship edges include attributes such as influence strength and influence path. Edge creation follows predefined relationship mapping rules. For instance, when a financial transaction is detected between two customer entities, a transaction relationship edge is created between them.
[0084] When incrementally updating the graph node set and edge set, an event-driven mechanism is used to process new data in real time. When new transaction data arrives, it is first determined whether the involved entity already exists in the graph. If not, a new node is created; if it already exists, the node attributes are updated. Similarly, based on the relationship information in the new data, new edges are created or the attributes of existing edges are updated. Node attribute value updates include arithmetic updates of numerical attributes (such as increases or decreases in total asset value) and logical updates of status attributes (such as a customer status changing from "normal" to "alert"). Edge weight adjustments are based on the impact of new data on relationship strength; for example, an increase in transaction frequency leads to a stronger edge weight. Timeliness updates consider the time attribute of the data to ensure that the graph reflects the current risk status.
[0085] When calculating time-decay weights for dynamically updated nodes and edges in the graph, an exponential decay function is used. For the weight w of each node or edge, a new weight w' = w is calculated based on the difference between its last update time t and the current time tcurrent. λ is the decay coefficient (0 < λ < 1), typically ranging from 0.9 to 0.99. The longer the time interval, the more pronounced the weight decay, making the graph more sensitive to recent risk information. This mechanism solves the problem of traditional risk models relying on historical data and failing to respond to dynamic risks in real time. When partitioning time-sensitive graphs using parallel computing, the graph's topology is used to divide it into multiple subgraphs. Common partitioning algorithms include edge cutting and vertex cutting; the former minimizes the number of edges across partitions, and the latter minimizes the number of vertices across partitions. The partitioned subgraphs can be processed in parallel, significantly improving the processing efficiency of large-scale graph data. Simultaneously, graph databases (such as Neo4j and JanusGraph) are used to store the graph structure, optimizing graph traversal and query performance. Graph databases use indexing techniques to accelerate node lookups, caching mechanisms to reduce disk access, and optimized query languages (such as Cypher and Gremlin) to support complex graph pattern matching. When calculating the topological properties of the graph for structural optimization, three main attributes are considered: connectivity, community structure, and centrality measures. Connectivity analysis determines whether any two points in the graph are connected by a path, identifying possible paths for risk propagation. Community structure analysis uses algorithms such as maximizing modularity to divide the graph into tightly connected subgroups, discovering risk clustering areas. Centrality measures calculate the importance of nodes in the graph, including degree centrality (number of direct connections), betweenness centrality (number of times the shortest path passes through), and eigenvector centrality (weighted centrality considering the importance of neighbors). These calculation results are saved as attribute values of the graph, forming the final local risk association graph.
[0086] In one specific embodiment, the process of executing step S103 may specifically include the following steps:
[0087] In each participating institution, a local graph convolutional structure of a federated graph neural network is constructed, including a risk entity node layer, a risk association layer, and a risk feature layer. The risk association layer performs risk propagation analysis by aggregating risk entities and related transaction information to obtain the local risk network structure.
[0088] The local risk association map is input into the local risk network structure, and the initial risk embedding representation is obtained through layer-by-layer risk information aggregation and nonlinear transformation processing.
[0089] Specifically, node features are extracted from the local risk association graph to obtain an initial node feature matrix; an adjacency matrix and a degree matrix are generated based on the local risk association graph, and a normalized adjacency matrix is calculated to obtain a standardized graph structure representation; the initial node feature matrix is multiplied by the standardized graph structure representation and transformed through a first-layer weight matrix to obtain first-order risk propagation features; a second information propagation and transformation is performed on the first-order risk propagation features, and processed through a second-layer weight matrix to obtain second-order risk propagation features; skip connection processing is performed on the second-order risk propagation features, and the initial features, first-order features, and second-order features are concatenated to obtain a multi-scale risk feature representation; the multi-scale risk feature representation is weighted and aggregated through a layer-by-layer attention mechanism, the importance weights of features at each layer are calculated and fused to obtain an initial risk embedding representation.
[0090] The process involves layer-by-layer risk information aggregation and nonlinear transformation, including: constructing query transformation matrices, key transformation matrices, and value transformation matrices for multi-scale risk feature representations, each containing specific parameter weights; projecting risk features into three different feature spaces through matrix multiplication to obtain risk query vectors, risk key vectors, and risk value vectors; calculating the product of the transpose of the risk query vector and the risk key vector, scaling the result by dividing by the square root of 64 to construct a quantitative index of the influence between risk nodes, thus obtaining a risk correlation strength matrix; and performing a softmax operation on each row of the risk correlation strength matrix, converting each element to a value between 0 and 1, with each row summing to 1, to form a risk... The probability distribution of risk propagation paths is used to obtain the risk attention weight matrix. The risk attention weight matrix is multiplied by the risk value vector, and the information of each risk node is filtered and aggregated according to importance to obtain the single-head risk attention output. The original risk feature vector is divided into 8 independent sub-vectors. For each sub-vector, the transformation matrix is repeatedly constructed, the association strength is calculated, it is converted into a weight distribution, and weighted aggregation is performed to obtain 8 sets of independent risk feature representations, resulting in multi-head risk attention results. The multi-head risk attention results are spliced in the feature dimension to form a complete risk feature. The output transformation matrix is transformed to the original dimension space, and the original risk feature vector is added to perform residual connection operation to obtain the initial risk embedding representation.
[0091] The gradient of the initial risk embedding representation is subjected to sensitive information pruning operation, which restricts the gradient values related to customer sensitive information to a preset threshold range, and the gradient after risk information pruning is obtained.
[0092] After cropping the risk information, Laplace noise is added to the gradient to obtain the privacy-preserving risk gradient.
[0093] The privacy protection risk gradient is uploaded to the risk coordination server. The risk gradients of all participating institutions are weighted and aggregated. Weights are assigned according to the risk data distribution of each institution to obtain the federal risk knowledge update value.
[0094] The updated federal risk knowledge values are distributed to participating institutions to update the parameters of the local risk graph convolutional network. At the same time, anonymous risk transmission relationships are exchanged through a privacy-preserving risk association extension protocol to obtain a risk representation that integrates risk knowledge from multiple institutions.
[0095] Specifically, a local graph convolutional structure of a federated graph neural network is constructed for each participating institution. This structure includes three key components: a risk entity node layer, a risk association layer, and a risk feature layer. The risk entity node layer is responsible for processing the initial features of various entity nodes; the risk association layer performs risk propagation analysis by aggregating risk entities and related transaction information to capture the risk relationships between entities; and the risk feature layer is responsible for extracting high-level risk feature representations. This three-layer structure design enables the network to effectively extract and transmit risk information from the local risk association graph. When the local risk association graph is input into the local risk network structure, the feature vector of each node is first extracted from the graph to form an initial node feature matrix. For example, for customer entity nodes, features include financial indicators and transaction behavior characteristics; for transaction event nodes, features include transaction amount and frequency. Next, an adjacency matrix is generated based on the structural information of the graph. The adjacency matrix is a two-dimensional matrix describing the connection relationships between nodes in the graph. If there is an edge between node i and node j, the element in the i-th row and j-th column of the matrix is 1; otherwise, it is 0. Simultaneously, the degree matrix is calculated. The degree matrix is a diagonal matrix, and the elements on the diagonal represent the degree of the corresponding node (the number of edges connected to that node). Then, the adjacency matrix is normalized to obtain a standardized graph structure representation. This step is to prevent numerical instability caused by excessive differences in node degree.
[0096] Multiplying the initial node feature matrix with the standardized graph structure representation achieves first-order information propagation, where each node aggregates the feature information of its direct neighbors. This process is equivalent to a risk transmission step in a risk association network, enabling each entity to access the risk information of its directly connected entities. The result is then linearly transformed through the first-layer weight matrix, and an activation function (such as ReLU) is applied to introduce non-linearity, resulting in first-order risk propagation features. Mathematically, this process is equivalent to performing a graph convolution operation on the initial node features. A second information propagation and transformation is then performed based on the first-order risk propagation features. By aggregating neighbor information again and processing it through the second-layer weight matrix, second-order risk propagation features are obtained. Second-order features can capture risk association information within two hops, i.e., the risk transmission relationship between indirectly connected entities. To preserve multi-level risk information, skip connections are applied to the second-order risk propagation features, concatenating the initial features, first-order features, and second-order features to obtain a multi-scale risk feature representation. Skip connections are a technique that can alleviate the gradient vanishing problem in deep networks and preserve original information. By directly connecting features from different layers, the network can simultaneously utilize feature information from different levels of abstraction.
[0097] Subsequently, a weighted aggregation of multi-scale risk feature representations is performed using an attention mechanism. The core idea of the attention mechanism is to enable the model to focus on important information and ignore irrelevant information. In specific implementation, query transformation matrices, key transformation matrices, and value transformation matrices are constructed for the multi-scale risk feature representations, each containing specific parameter weights. Matrix multiplication projects the risk features into three different feature spaces, resulting in risk query vectors, risk key vectors, and risk value vectors. The query vector represents the information requirement of the current node, the key vector represents the information provided by other nodes, and the value vector is the content that actually needs to be aggregated. The product of the transpose of the risk query vector and the risk key vector is calculated to measure the similarity or compatibility between the query and the key. Then, the result is scaled by dividing by the square root of 64 (assuming the feature dimension is 64). This step is to control the numerical range after the dot product operation and prevent gradient explosion. Next, a softmax operation is performed on each row of the obtained risk association strength matrix, converting each element into a value between 0 and 1, with each row summing to 1, forming the probability distribution of the risk propagation path, resulting in the risk attention weight matrix. The softmax function converts the original scores into a probability distribution, giving higher weights to important associations. The risk attention weight matrix is multiplied by the risk value vector, and the value information is aggregated by weighted summation to obtain the single-head risk attention output.
[0098] To capture risk association patterns from different perspectives, the original risk feature vector is divided into eight independent sub-vectors. The attention calculation process described above is performed independently on each sub-vector, resulting in eight independent risk feature representations, forming a multi-head risk attention result. The multi-head mechanism enhances the model's ability to capture complex risk association patterns by computing attention in parallel across different feature subspaces. Finally, the multi-head risk attention result is concatenated along the feature dimension to form a complete risk feature. This is then transformed back to the original dimensional space using the output transformation matrix, and residual connection operations are performed with the original risk feature vector to obtain the initial risk embedding representation. To protect data privacy, the gradient of the initial risk embedding representation undergoes sensitive information pruning, limiting the L2 norm of the gradient to a preset threshold to prevent a single sample from having an excessive impact on the model. This is a fundamental step in achieving differential privacy. Next, Laplace noise is added to the risk-information-pruned gradient. Laplace noise is a probability distribution whose density function decreases exponentially with distance from the mean. This noise ensures that even if an attacker knows all the data except for one entry, they cannot accurately infer the information of that entry, thus achieving data privacy protection.
[0099] After the privacy-preserving risk gradients are uploaded to the risk coordination server, the server performs risk-weighted aggregation of the risk gradients from all participating institutions. Weights are assigned based on the quantity and quality of each institution's risk data, and a global gradient update is calculated to obtain the federated risk knowledge update value. This federated learning approach allows multiple institutions to jointly train the model without sharing raw data, effectively solving the data silo problem in traditional risk assessment. Finally, the federated risk knowledge update value is distributed to each participating institution to update the parameters of the local risk graph convolutional network. Simultaneously, anonymous risk transmission relationships are exchanged through a privacy-preserving risk association extension protocol, resulting in a risk representation that integrates risk knowledge from multiple institutions. This protocol allows institutions to share high-order connection information without disclosing customer identities, further enriching the structural information of the risk graph.
[0100] In one specific embodiment, the process of executing step S104 may specifically include the following steps:
[0101] Micro-level features are extracted from risk representations that integrate risk knowledge from multiple institutions to obtain a set of micro-risk features, which include: customer financial status, credit history, and behavioral patterns.
[0102] Based on the relationships in the risk representation, meso-level relationship features are extracted to obtain the meso-level risk feature set;
[0103] Based on the overall structure of risk representation, macro-system hierarchical features are extracted to obtain a macro-risk feature set;
[0104] The risk prediction results are obtained by fusing the micro-risk feature set, the meso-risk feature set, and the macro-risk feature set and calculating the risk probability distribution.
[0105] The risk prediction results are classified into four levels: low risk, medium risk, high risk and extremely high risk. Corresponding risk management strategies are set for each level to obtain the risk level classification results.
[0106] Risk warning signals are generated based on the risk level classification results, including risk scores, key risk factors, risk development trends and warning confidence levels. Different levels of warning mechanisms are triggered for different risk levels to obtain comprehensive risk scores and warning signals.
[0107] Specifically, micro-level individual features are extracted from the risk representation. These features directly describe the risk status of individual entities. Micro-level individual features include customer financial status (such as debt-to-equity ratio, cash flow, profit margin, etc.), credit history (such as historical defaults, credit rating changes, repayment behavior, etc.), and behavioral patterns (such as transaction frequency, distribution of trading partners, abnormal transaction patterns, etc.). The extraction process employs feature selection algorithms, calculating feature importance scores (such as feature importance based on random forests, information gain, etc.) to select the most predictive features, forming a micro-risk feature set. When extracting meso-level relationship features based on the correlations in the risk representation, attention is paid to the interaction relationships between entities and the risk transmission paths. Meso-level features include the strength of associations in the transaction network (quantified by transaction frequency and amount), transmission paths (measuring the possible paths for risk to spread from one entity to another), and structural similarity (assessing the similarity of entities' positions and connection patterns within the network). Extracting meso-level features requires network analysis calculations, such as degree centrality (the number of edges directly connected to a node), betweenness centrality (the number of shortest paths through a node), and community detection (identifying closely connected clusters of nodes). These calculations reflect the flow and aggregation patterns of risk within the network, forming a meso-level risk feature set. When extracting macro-level system features based on the overall structure of risk representation, the focus is on the overall risk environment and systemic risk. Macro-level features include risk concentration (assessing the degree of risk aggregation in a specific region or industry), contagiousness (measuring the speed and extent of risk spreading from one region to others), and vulnerability (measuring the system's resilience to shocks). These features are obtained through global graph structure analysis, such as calculating global clustering coefficients (assessing overall connectivity), eigenvector centrality (weighted centrality considering the importance of neighbors), and graph density (the ratio of the number of edges to the maximum possible number of edges), forming a macro-level risk feature set.
[0108] When fusing micro, meso, and macro risk feature sets, an ensemble learning approach is employed. Specifically, three basic predictors are first trained: a decision tree forest for micro features, a neural network for meso features, and logistic regression for macro features. Each predictor independently learns risk patterns at its respective level. Then, a meta-learner is trained using a stacking method, taking the outputs of the basic predictors as input to learn the optimal fusion weights. Cross-validation is used in the training of the meta-learner to prevent overfitting and improve generalization ability. The fused model calculates the risk probability distribution for each entity, i.e., the probability values for different risk levels, to obtain the risk prediction results.
[0109] When classifying risk prediction results into risk levels, probability thresholds are set to categorize risk levels into four levels: low risk (probability < 0.25), medium risk (0.25 ≤ probability < 0.5), high risk (0.5 ≤ probability < 0.75), and very high risk (probability ≥ 0.75). Each risk level corresponds to a different handling strategy: low-risk levels are handled with routine monitoring strategies; medium-risk levels increase monitoring frequency and require additional information; high-risk levels trigger audit processes and require risk mitigation measures; and very high-risk levels activate emergency response plans and restrict related transactions. This tiered approach addresses the low interpretability of traditional risk models, making risk management decisions more transparent and understandable.
[0110] When generating risk warning signals based on risk level classification results, a comprehensive risk score is calculated. This score comprehensively considers three factors: risk probability, potential impact, and time urgency. Key risk factors are identified through SHAP (SHapley Additive exPlanations) value analysis, calculating the contribution of each feature to the model's prediction, thereby identifying the core factors dominating the risk. Risk development trends are determined by comparing the rate of change between the current risk score and historical scores to determine whether the risk is rising, stable, or declining. Warning confidence is assessed based on the model's accuracy on historical data and the certainty of the current prediction. Finally, corresponding warning mechanisms are triggered according to different risk levels: low risk sends a regular report, medium risk triggers an email alert, high risk triggers an SMS notification, and extremely high risk triggers multiple channels of warning, including telephone calls and mandatory system blocking, forming a complete comprehensive risk score and warning signal.
[0111] The risk prediction method based on dynamic aggregation and privacy protection in the embodiments of this application has been described above. The risk prediction system based on dynamic aggregation and privacy protection in the embodiments of this application is described below. Please refer to [link / reference]. Figure 2 One embodiment of the risk prediction system based on dynamic aggregation and privacy protection in this application includes:
[0112] The acquisition module 201 is used to acquire and standardize multi-source heterogeneous data in real time to obtain a standardized data stream.
[0113] Module 202 is used to construct a local risk association graph within each participating institution based on a standardized data flow. The local risk association graph includes nodes, edges, and topological attributes.
[0114] The calculation module 203 is used to input the local risk association graph into the federated graph neural network for privacy-preserving calculation to obtain a risk representation that integrates risk knowledge from multiple institutions. The privacy-preserving calculation uses a local differential privacy mechanism.
[0115] The segmentation module 204 is used to perform multi-level feature extraction and risk level segmentation on the risk representation to obtain a comprehensive risk score and early warning signal.
[0116] above Figure 2 The risk prediction system based on dynamic aggregation and privacy protection in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The risk prediction device based on dynamic aggregation and privacy protection in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0117] Figure 3 This is a schematic diagram of a risk prediction device based on dynamic aggregation and privacy protection provided in an embodiment of the present invention. The risk prediction device 300 based on dynamic aggregation and privacy protection can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing applications 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the risk prediction device 300 based on dynamic aggregation and privacy protection. Furthermore, the processor 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media 330 on the risk prediction device 300 based on dynamic aggregation and privacy protection to implement the steps of the aforementioned risk prediction method based on dynamic aggregation and privacy protection.
[0118] The risk prediction device 300 based on dynamic aggregation and privacy protection may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated risk prediction device structure based on dynamic aggregation and privacy protection does not constitute a limitation on the risk prediction device based on dynamic aggregation and privacy protection provided by the present invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0119] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the risk prediction method based on dynamic aggregation and privacy protection.
[0120] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0121] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a risk prediction device based on dynamic aggregation and privacy protection (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0122] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A risk prediction method based on dynamic aggregation and privacy protection, characterized in that, The method includes: Real-time acquisition and standardization of multi-source heterogeneous data are performed. A multi-threaded collector is used to collect raw data from internal transaction systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors to obtain multi-source heterogeneous data. The acquired heterogeneous data is then formatted to unify the format, converting data with different encodings to UTF-8 format. Numerical features are standardized using the Z-Score normalization algorithm to obtain a standardized data stream. Based on the standardized data flow, a local risk association graph is constructed within each participating institution. The local risk association graph includes nodes, edges, and topological attributes. The local risk association graph is input into a federated graph neural network for privacy-preserving computation to obtain a risk representation that integrates risk knowledge from multiple institutions. This privacy-preserving computation uses a local differential privacy mechanism, including: constructing a local graph convolutional structure for each participating institution within the federated graph neural network, comprising a risk entity node layer, a risk association layer, and a risk feature layer. The risk association layer performs risk propagation analysis by aggregating risk entities and related transaction information to obtain a local risk network structure. The local risk association graph is input into the local risk network structure, and an initial risk embedding representation is obtained through layer-by-layer risk information aggregation and nonlinear transformation processing. The gradient of the initial risk embedding representation is then sensitively processed. The process involves a risk information pruning operation, which restricts the gradient values related to sensitive customer information to a preset threshold range to obtain a pruned risk information gradient. Laplace noise is then added to this pruned gradient to obtain a privacy-protected risk gradient. This privacy-protected risk gradient is uploaded to a risk coordination server, where the risk gradients of all participating institutions are weighted and aggregated. Weights are assigned based on the risk data distribution of each institution to obtain a federated risk knowledge update value. This federated risk knowledge update value is then distributed to each participating institution to update the parameters of the local risk graph convolutional network. Simultaneously, anonymous risk transmission relationships are exchanged through a privacy-protected risk association extension protocol to obtain the risk representation that integrates multi-institutional risk knowledge. Multi-level feature extraction and risk level classification are performed on the risk representation to obtain a comprehensive risk score and early warning signal.
2. The risk prediction method based on dynamic aggregation and privacy protection according to claim 1, characterized in that, The process involves unifying the format of the collected heterogeneous data, converting data with different encodings to UTF-8 format; and standardizing numerical features using the Z-Score normalization algorithm to obtain a standardized data stream, including: Key fields are extracted from the structured data in the multi-source heterogeneous data; nested attributes are extracted from the semi-structured data using JSON / XML parsing algorithms; and semantic information is extracted from the unstructured data using natural language processing to obtain the processed data. The processed data is formatted uniformly by converting heterogeneous data into UTF-8 encoding format and applying entity parsing technology to identify and merge different representations of the same entity to obtain data in a unified format. The numerical features in the unified format data are standardized using the Z-Score normalization algorithm, the categorical features are converted into numerical representations using one-hot encoding, and the data is then subjected to quality checks and repairs to obtain the standardized data stream.
3. The risk prediction method based on dynamic aggregation and privacy protection according to claim 1, characterized in that, The local risk association graph is constructed within each participating institution based on the standardized data flow. This local risk association graph includes nodes, edges, and topological attributes, including: Node types are defined from the standardized data stream according to business rules, including customer entity nodes, transaction event nodes, asset project nodes and risk signal nodes, and a specific set of attributes is assigned to each type of node to obtain a graph node set; Based on the data relationship, the edge types are defined, including transaction relationship edges, association relationship edges, and influence relationship edges. The connection relationship between related nodes is converted into edges in the graph, resulting in a graph edge set. Incremental updates are performed on the graph node set and graph edge set. New data is converted into nodes or edges in the graph in real time. Attribute value updates and node state changes are performed on nodes, and weight adjustments and timeliness updates are performed on edges to obtain a dynamically updated graph. Calculate time decay weights for nodes and edges in the dynamically updated graph to gradually reduce the influence of historical data over time, thus obtaining a time-sensitive graph. The time-sensitive graph is partitioned using parallel computing, and the graph database storage structure is used to optimize graph traversal and query performance, resulting in a structure-optimized graph. Calculate the topological attributes of the structure optimization graph, including connectivity, community structure, and centrality measure, and save the calculation results as graph attribute values to obtain the local risk association graph.
4. The risk prediction method based on dynamic aggregation and privacy protection according to claim 1, characterized in that, The step of inputting the local risk association map into the local risk network structure, and obtaining the initial risk embedding representation through layer-by-layer risk information aggregation and nonlinear transformation processing includes: Node features are extracted from the local risk association graph to obtain an initial node feature matrix; an adjacency matrix and a degree matrix are generated based on the local risk association graph, and a normalized adjacency matrix is calculated to obtain a standardized graph structure representation; The initial node feature matrix is multiplied by the standardized graph structure representation, and then transformed through the first-layer weight matrix to obtain the first-order risk propagation feature; A second information propagation and transformation is performed on the first-order risk propagation feature, and the second-order risk propagation feature is obtained by processing through the second-layer weight matrix. The second-order risk propagation features are subjected to skip connection processing, and the initial features, first-order features and second-order features are concatenated to obtain a multi-scale risk feature representation; The multi-scale risk feature representation is weighted and aggregated by a layer-by-layer attention mechanism, and the importance weights of the features at each layer are calculated and fused to obtain the initial risk embedding representation.
5. The risk prediction method based on dynamic aggregation and privacy protection according to claim 4, characterized in that, The initial risk embedding representation is obtained by weighting and aggregating the multi-scale risk feature representation through a layer-by-layer attention mechanism, calculating the importance weights of features at each layer, and fusing them. For the multi-scale risk feature representation, a query transformation matrix, a key transformation matrix, and a value transformation matrix are constructed respectively. Each matrix contains specific parameter weights. The risk features are projected onto three different feature spaces through matrix multiplication to obtain the risk query vector, the risk key vector, and the risk value vector. Calculate the product of the risk query vector and the transpose of the risk key vector, divide the result by the square root of 64 to scale and adjust, construct a quantitative index of the influence between risk nodes, and obtain the risk association strength matrix. Perform a softmax operation on each row of the risk association strength matrix to convert each element into a value between 0 and 1, with each row summing to 1, to form the probability distribution of the risk propagation path and obtain the risk attention weight matrix. The risk attention weight matrix is multiplied by the risk value vector, and the information of each risk node is filtered and aggregated according to its importance to obtain a single-head risk attention output. The original risk feature vector is divided into 8 independent sub-vectors. For each sub-vector, the transformation matrix is repeatedly constructed, the correlation strength is calculated, it is converted into a weight distribution and weighted aggregation is performed to obtain 8 sets of independent risk feature representations, thus obtaining the multi-head risk attention result. The multi-head risk attention results are concatenated along the feature dimension to form complete risk features. The output transformation matrix is then transformed back to the original dimensional space, and the original risk feature vector is added to perform residual connection operations to obtain the initial risk embedding representation.
6. The risk prediction method based on dynamic aggregation and privacy protection according to claim 1, characterized in that, The process of performing multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal includes: Micro-level individual features are extracted from the risk representation that integrates risk knowledge from multiple institutions to obtain a micro-risk feature set, which includes: customer financial status, credit history and behavioral patterns; Based on the correlations in the risk representation, meso-level relationship features are extracted to obtain a meso-level risk feature set; Based on the overall structure of the risk representation, macro-system hierarchical features are extracted to obtain a macro-risk feature set; The risk prediction results are obtained by fusing the micro-risk feature set, meso-risk feature set and macro-risk feature set and calculating the risk probability distribution. The risk prediction results are classified into four levels: low risk, medium risk, high risk, and extremely high risk. Corresponding risk management strategies are set for each level to obtain the risk level classification results. Based on the risk level classification results, a risk warning signal is generated, including risk score, key risk factors, risk development trend and warning confidence level, and different levels of warning mechanisms are triggered for different risk levels to obtain the comprehensive risk score and warning signal.
7. A risk prediction system based on dynamic aggregation and privacy protection, characterized in that, For implementing the risk prediction method based on dynamic aggregation and privacy protection as described in any one of claims 1-6, the risk prediction system based on dynamic aggregation and privacy protection comprises: The data acquisition module is used for real-time acquisition and standardization of multi-source heterogeneous data. It uses a multi-threaded collector to acquire raw data from internal transaction systems, customer behavior logs, external market data sources, social media platforms, and IoT sensors to obtain multi-source heterogeneous data. The module performs format unification processing on the acquired heterogeneous data, converting data with different encodings to UTF-8 format. Numerical features are standardized using the Z-Score normalization algorithm to obtain a standardized data stream. The construction module is used to construct a local risk association graph within each participating institution based on the standardized data flow. The local risk association graph includes nodes, edges, and topological attributes. The computation module is used to input the local risk association graph into a federated graph neural network for privacy-preserving computation to obtain a risk representation that integrates risk knowledge from multiple institutions. The privacy-preserving computation uses a local differential privacy mechanism, including: constructing a local graph convolutional structure for the federated graph neural network at each participating institution, including a risk entity node layer, a risk association layer, and a risk feature layer. The risk association layer performs risk propagation analysis by aggregating risk entities and related transaction information to obtain a local risk network structure; inputting the local risk association graph into the local risk network structure, and obtaining an initial risk embedding representation through layer-by-layer risk information aggregation and nonlinear transformation processing; and performing a stepwise transformation on the initial risk embedding representation. Sensitive information is cropped to limit the gradient values related to sensitive customer information to a preset threshold range, resulting in a cropped risk information gradient. Laplace noise is added to this cropped gradient to obtain a privacy-protected risk gradient. This privacy-protected risk gradient is uploaded to a risk coordination server, where the risk gradients of all participating institutions are weighted and aggregated. Weights are assigned based on the risk data distribution of each institution to obtain a federated risk knowledge update value. This federated risk knowledge update value is then distributed to each participating institution to update the parameters of the local risk graph convolutional network. Simultaneously, anonymous risk transmission relationships are exchanged through a privacy-protected risk association extension protocol to obtain the risk representation that integrates multi-institutional risk knowledge. The segmentation module is used to perform multi-level feature extraction and risk level classification on the risk representation to obtain a comprehensive risk score and early warning signal.
8. A risk prediction device based on dynamic aggregation and privacy protection, characterized in that, It includes a memory and a processor, the memory storing a computer program that can run on the processor, and the processor executing the computer program to implement the risk prediction method based on dynamic aggregation and privacy protection as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, it causes the processor to execute the risk prediction method based on dynamic aggregation and privacy protection as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Risk prediction method and device, equipment and storage medium
CN113822494A
Risk early warning method, device, equipment and medium
CN117495548A