Dynamic graph analysis method, system and equipment for big data
By mirroring the time sequence attribute data flow to the mirror node in dynamic graph analysis and using directed topology networks and cascaded risk conduction map neural networks, the problems of low computing efficiency and insufficient dynamic topology adaptation in the prior art are solved, and efficient real-time risk conduction analysis and cascaded risk warning are achieved.
Patent Information
- Application Number
- CN202510915631.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-03
AI Technical Summary
When handling large-scale timing attribute data flows, the existing technology has low computing efficiency, poor real-time performance, and insufficient dynamic topological adaptation, resulting in lagging risk transmission path mining, and the inability to achieve accurate capture of real-time risk transmission.
By mirroring K entity nodes in the base layer dynamic graph according to the sliding time window, K timing attribute data streams are mirrored to K mirror nodes in the analysis layer dynamic graph, using directed topological network for cascade risk conduction mining, combining with the cascade risk conduction graph neural network for real-time risk conduction path generation and analysis, and returning the results to the source entity node.
It realizes efficient real-time risk transmission analysis, accurate dynamic topology inheritance, active cascading risk warning, optimizes resource consumption, and improves the mining efficiency and accuracy of risk transmission paths.
Smart Images

Figure CN120578764A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field related to graph analysis, and specifically to a dynamic graph analysis method, system and equipment for big data. Background Art
[0002] With the development of big data technology, complex entity relationships have become dynamic, time-series, and networked. Traditional static graph analysis is unable to meet the needs of real-time risk monitoring and transmission path mining. This is especially true in areas such as financial risk control, supply chain management, and network security, where risk transmission between entities often has a cascading effect. Dynamic graph analysis has become a key means of identifying and blocking the spread of risk. However, existing analytical methods often face problems such as low computational efficiency, poor real-time performance, and insufficient dynamic adaptability of topological structures when processing large-scale time-series attribute data streams. This causes the mining of risk transmission paths to lag behind actual business needs, and relies on offline computing or periodic snapshot updates, making it difficult to accurately capture real-time risk transmission. Furthermore, data storage and analysis are often coupled to the same graph layer, resulting in high computing resource consumption and an inability to effectively support the rapid adjustment of dynamic topological structures.
[0003] Therefore, the current related technologies have technical problems such as poor real-time performance, low computing efficiency, insufficient dynamic topology adaptation, and delayed risk transmission path mining. Summary of the Invention
[0004] This application solves the technical problems of poor real-time performance, low computing efficiency, insufficient dynamic topology adaptation, and delayed risk transmission path mining in the existing technology by providing a dynamic graph analysis method, system and equipment for big data, and achieves the technical effects of efficient real-time risk transmission analysis, precise inheritance of dynamic topology, active warning of cascading risks, and optimization of resource consumption.
[0005] The present application provides a dynamic graph analysis method for big data, which includes: K entity nodes in the base layer dynamic graph mirror K time series attribute data streams to K mirror nodes in the analysis layer dynamic graph based on a sliding time window, wherein the K mirror nodes inherit the directed topological network of the base layer dynamic graph; the K mirror nodes load the K time series attribute data streams in the directed topological network to trigger cascade risk conduction mining and output a real-time risk conduction path; after the base layer dynamic graph receives the real-time risk conduction path returned by the analysis layer dynamic graph, the real-time risk conduction path is locally stored, and a node risk disposal instruction is constructed based on the real-time risk conduction path; according to the mirror node composition of the real-time risk conduction path, the node risk disposal instruction is returned to the source entity node.
[0006] In a possible implementation, the big data-oriented dynamic graph analysis method further performs the following processing: heterogeneous data streams are transmitted to the basic layer dynamic graph using a distributed message queue; the basic layer dynamic graph performs attribute decomposition on the heterogeneous data streams to obtain K time series attribute data streams; and according to the data attribute labels of the K time series attribute data streams, the K time series attribute data streams are mapped and loaded into the K entity nodes in the basic layer dynamic graph.
[0007] In a possible implementation, the dynamic graph analysis method for big data further performs the following processing: extracting K time series risk time windows locally from the K entity nodes; combining and enumerating the K entity nodes to obtain After grouping entity nodes, according to the The group entity node mapping combines the K time series risk time windows to obtain Group time series risk time window; by calculating the The risk event overlap coefficient of the group time series risk time window is used to quantify the Group entity node Group risk correlation; based on the preset correlation strength threshold and the The association deviation of the group risk association degree is used to perform directed weighted connection of the K entity nodes to obtain the directed topological network.
[0008] In a possible implementation, the dynamic graph analysis method for big data further performs the following processing: Whether the group risk correlation falls within the preset correlation strength threshold, The group entity nodes are screened to obtain P groups of associated nodes; based on the P groups of associated nodes, the K entity nodes are undirectedly topologically connected to output an initial topological network; the P risk triggering direction frequencies of the P groups of temporal risk time windows are counted, and P risk conduction direction vectors are constructed based on the P risk triggering direction frequencies; the P risk conduction direction vectors are used to perform directed weighting on the initial topological network to output the directed topological network.
[0009] In a possible implementation, the dynamic graph analysis method for big data also performs the following processing: slidingly dividing the heterogeneous data stream into multiple heterogeneous data segments; performing structured analysis on the multiple heterogeneous data segments and outputting multiple standardized attribute vectors; predefining K node attribute vectors of the K entity nodes; calculating the similarity of multiple groups of attribute vectors between the multiple standardized attribute vectors and the K node attribute vectors, and then reorganizing the multiple groups of attribute vectors into K groups of heterogeneous data segments according to dynamic matching rules; and splicing the K groups of heterogeneous data segments in time series to output the K time series attribute data streams.
[0010] In a possible implementation, the big data-oriented dynamic graph analysis method further performs the following processing: using the P risk conduction direction vectors as timing constraints, the P groups of associated nodes as node screening constraints, and locally calling the P groups of timing risk conduction feature data; using the P groups of timing risk conduction feature data as training data, performing cascade risk conduction model training in the directed topology network, and outputting a cascade risk conduction graph neural network, wherein the cascade risk conduction graph neural network includes K node conduction identification models of the K mirror nodes; mapping the K timing attribute data streams and inputting them in parallel into the K node conduction identification models of the cascade risk conduction graph neural network, and collaboratively generating the real-time risk conduction path through topological message passing.
[0011] In a possible implementation, the big data-oriented dynamic graph analysis method further performs the following processing: in the cascade risk conduction graph neural network, after the B mirror node receives the A risk conduction identifier and the A risk conduction probability sent by the A mirror node, the B node conduction identification model is activated, wherein the A mirror node is out-degree connected to the B mirror node; the A risk conduction identifier, the A risk conduction probability and the B time series data stream are loaded into the B node conduction identification model, and risk conduction prediction is performed through the B node conduction identification model, and the B risk conduction identifier and the B risk conduction probability are output; if the D node orientation is parsed from the B risk conduction identifier, the B risk conduction identifier and the B risk conduction probability are sent to the D mirror node according to the D node orientation, and the D node conduction identification model is activated in the cascade risk conduction graph neural network, wherein the B mirror node is out-degree connected to the C mirror node and the D mirror node respectively; and so on, the risk identification and risk probability are recursively conducted in the cascade risk conduction graph neural network until the probability threshold is less than the preset scale, and the real-time risk conduction path is generated by backtracking the node chain.
[0012] In a possible implementation, the big data-oriented dynamic graph analysis method also performs the following processing: parsing the real-time risk transmission path and extracting the source mirror node identifier; locating the source entity node corresponding to the source mirror node identifier in the basic layer dynamic graph based on the cross-layer node mapping relationship; matching the node risk disposal instruction according to the risk identifier of the source entity node, and then returning the node risk disposal instruction to the source entity node through the security instruction channel.
[0013] The present application also provides a dynamic graph analysis system for big data, which includes: a time series attribute data mirroring unit, which is used for K entity nodes in the basic layer dynamic graph to mirror K time series attribute data streams to K mirror nodes in the analysis layer dynamic graph based on a sliding time window, wherein the K mirror nodes inherit the directed topological network of the basic layer dynamic graph; a real-time risk conduction path output unit, which is used for the K mirror nodes to load the K time series attribute data streams in the directed topological network to trigger cascade risk conduction mining and output real-time risk conduction paths; a node risk disposal instruction construction unit, which is used for the basic layer dynamic graph to locally store the real-time risk conduction path returned by the analysis layer dynamic graph after receiving the real-time risk conduction path returned by the analysis layer dynamic graph, and to construct a node risk disposal instruction based on the real-time risk conduction path; a node risk disposal instruction return unit, which is used to return the node risk disposal instruction to the source entity node according to the mirror node composition of the real-time risk conduction path.
[0014] The present application also provides an electronic device, comprising: a memory for storing executable instructions; and a processor for implementing a dynamic graph analysis method for big data when executing the executable instructions stored in the memory.
[0015] The proposed method, system, and device for big data-oriented dynamic graph analysis proposed in this application will mirror K time-series attribute data streams from K entity nodes in the base-layer dynamic graph to K mirror nodes in the analysis-layer dynamic graph based on a sliding time window. Cascade risk conduction mining will be triggered in a directed topology network. The base-layer dynamic graph will receive real-time risk conduction paths and store them locally, constructing node risk disposal instructions based on the real-time risk conduction paths. Based on the mirror node composition of the real-time risk conduction paths, the node risk disposal instructions will be transmitted back to the source entity node. This solves the technical problems of poor real-time performance, low computational efficiency, insufficient dynamic topology adaptation, and delayed risk conduction path mining in existing technologies, achieving the technical effects of efficient real-time risk conduction analysis, precise inheritance of dynamic topology, proactive early warning of cascade risks, and optimized resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0017] Figure 1 A flow chart of the dynamic graph analysis method for big data provided in the embodiment of the present application.
[0018] Figure 2 Schematic diagram of the structure of the dynamic graph analysis system for big data provided in the embodiment of the present application.
[0019] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0020] Explanation of the accompanying symbols: time-series attribute data mirroring unit 10, real-time risk transmission path output unit 20, node risk handling instruction construction unit 30, node risk handling instruction return unit 40, input device 401, processor 402, memory 403, output device 404. DETAILED DESCRIPTION
[0021] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0022] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0023] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0024] The present application provides a dynamic graph analysis method for big data, such as Figure 1 As shown, the method includes:
[0025] In step S100, K entity nodes in the base layer dynamic graph mirror K time series attribute data streams to K mirror nodes in the analysis layer dynamic graph based on a sliding time window, wherein the K mirror nodes inherit the directed topology network of the base layer dynamic graph.
[0026] Step S100 further includes step S101, transmitting the heterogeneous data stream into the basic layer dynamic graph using a distributed message queue; step S102, the basic layer dynamic graph performs attribute decomposition on the heterogeneous data stream to obtain K time series attribute data streams; step S103, according to the data attribute labels of the K time series attribute data streams, mapping the K time series attribute data streams to the K entity nodes in the basic layer dynamic graph.
[0027] Preferably, multi-source heterogeneous data is obtained, which may include database logs, sensor data, transaction records, social network information, etc., and heterogeneous data streams are determined, whose formats, frequencies and structures are different. Distributed message queues such as Kafka and Pulsar are used as data access layers to stream heterogeneous data into the dynamic graph of the basic layer. Among them, distributed message queues can decouple data production and consumption, and cope with sudden peaks of data flows to avoid system overload. At the same time, through message persistence, partitioning and replication mechanisms, data loss is guaranteed, and the reliability of data storage is ensured. Heterogeneous data streams contain information across multiple dimensions. For example, a financial transaction may involve transaction amount, time, accounts of both parties, and risk tags. Using the base-layer dynamic graph, attribute decomposition is performed on these heterogeneous data streams, breaking the composite data into independent time-series attribute data streams. Specifically, the heterogeneous data streams are structured and parsed to extract fields in formats such as JSON and XML. Time-series data streams are then processed, with each attribute sorted by timestamp to form a time-series data stream. The same attributes across different data sources are then standardized to ensure consistent data analysis. Ultimately, K time-series attribute data streams are generated, where K is a positive integer. Each time-series attribute data stream contains a data attribute label. When new entities or attributes are added (such as newly registered users), they are automatically matched based on the labels, eliminating the need for graph reconstruction. Finally, these K time-series attribute data streams are mapped and loaded onto K entity nodes in the base-layer dynamic graph. This involves locating the corresponding entity node in the base-layer dynamic graph based on the data attribute labels, and then updating the state of the node's attribute set by loading the time-series data stream into the node.
[0028] Preferably, the dynamic graph of the base layer is the core data storage layer, which is a dynamic network composed of all entities and their relationships. K entity nodes represent the key entities that need to be monitored and analyzed at present, such as high-risk accounts in the financial system, key nodes in the supply chain, etc.; the dynamic graph of the analysis layer predicts the short-term risk status (risk transmission); the sliding time window is a dynamic data partitioning strategy that only retains data within the most recent period, such as transaction records in the past 5 minutes, and continuously updates the data in the window over time to ensure the timeliness of the analysis. The time series attribute data stream is a sequence of attributes of the entity node that changes over time, such as the change in account balance per minute, the frequency of user logins, etc.; mirroring refers to creating a mirror node in the dynamic graph of the analysis layer that corresponds one-to-one to the entity node of the base layer, but only copies the data within the current time window, rather than the full historical data.
[0029] Preferably, the K entity nodes in the base layer dynamic graph filter the corresponding K time-series attribute data streams according to a sliding time window and copy them to the K corresponding mirror nodes in the analysis layer dynamic graph in real time, retaining only the key data within the window, such as the most recent 10 transactions, to reduce the computational pressure of the analysis layer dynamic graph. The K mirror nodes inherit the directed topology network of the base layer dynamic graph. Specifically, the directed topology network refers to the association relationship and directionality between entity nodes. Inheriting the topology means that the mirror nodes of the analysis layer not only copy the data but also completely retain the network structure between entities in the base layer, such as who is associated with whom and the strength of the relationship. If "user X" and "user Y" have a transfer relationship in the base layer, the mirror nodes "X'" and "Y'" of the analysis layer also retain the directed edge. When the base layer topology changes, such as when a new association relationship is added, the analysis layer topology is synchronously adjusted to ensure the accuracy of the transmission analysis, thereby achieving a balance between the real-time performance and efficiency of the base layer dynamic graph and the analysis layer dynamic graph to ensure optimal resource allocation.
[0030] Furthermore, step S102 also includes step a, slidingly dividing the heterogeneous data stream into multiple heterogeneous data segments; step b, performing structured analysis on the multiple heterogeneous data segments and outputting multiple standardized attribute vectors; step c, predefining K node attribute vectors of the K entity nodes; step d, calculating the similarity of multiple groups of attribute vectors between the multiple standardized attribute vectors and the K node attribute vectors, and then reorganizing the multiple groups of attribute vectors into K groups of heterogeneous data segments according to dynamic matching rules; step e, time-series splicing of the K groups of heterogeneous data segments and outputting the K time-series attribute data streams.
[0031] Preferably, heterogeneous data from different sources are dynamically divided into data streams according to fixed time windows (such as every 5 minutes) to form multiple continuous or overlapping heterogeneous data segments for parallel processing, and then multiple heterogeneous data segments are structured and parsed. Specifically, key fields are extracted from each heterogeneous data segment, and unstructured data is processed, such as extracting entities and events from log texts through NLP, and then the parsed fields are converted into vectors in a unified format, and multiple standardized attribute vectors are output to eliminate the differences in the original data format; then K node attribute vectors of K entity nodes are predefined, that is, the core features of the entity nodes are defined, where each entity node (such as user, device) has a predefined attribute template, such as [account type, region, wind Risk level] as the matching benchmark for data segmentation; for each standardized attribute vector and K node attribute vectors, the similarity of multiple groups of attribute vectors is calculated by cosine similarity or Euclidean distance, and then the multiple groups of attribute vectors are reorganized into K groups of heterogeneous data segments according to the dynamic matching rule, that is, the node attribute vectors with a similarity exceeding a predetermined threshold (such as 0.8) are classified into the corresponding node group. If a segment is similar to multiple nodes, the node with the highest score is selected, and then all heterogeneous data segments are grouped according to the matching results to form K groups of heterogeneous data segments. Finally, the K groups of heterogeneous data segments are spliced in time series, that is, each group of data segments is sorted by timestamp and merged into a continuous time series data stream, and the independent time series attribute data streams of K entity nodes are output to reflect their dynamic behavior changes.
[0032] Furthermore, step S100 further includes step S110, extracting K time series risk time windows locally from the K entity nodes; step S120, combining and enumerating the K entity nodes to obtain After grouping entity nodes, according to the The group entity node mapping combines the K time series risk time windows to obtain Group temporal risk time window; Step S130, by calculating the The risk event overlap coefficient of the group time series risk time window is used to quantify the Group entity node Step S140, according to the preset correlation strength threshold and the The association deviation of the group risk association degree is used to perform directed weighted connection of the K entity nodes to obtain the directed topological network.
[0033] Preferably, each entity node, such as a bank account or supply chain node, dynamically calculates a risk time window based on its historical behavior data, that is, a risk indicator sequence within a sliding time interval, which represents the risk status change of the entity within a specific time period. Through time series analysis, such as fluctuation detection and anomaly marking, key risk events are extracted; from K entity nodes, two pairs are paired (such as node A and node B, node A and node C...), and a total of For example, three nodes (A, B, C) are combined into (A, B), (A, C), and (B, C). Each node pair is associated with its own temporal risk time window to form Group risk time window.
[0034] Preferably, the overlap ratio of risk events in the two time windows is calculated, or the temporal similarity is quantified by correlation coefficient (such as Pearson coefficient). The risk event overlap coefficient of the group time series risk time window is used to measure the synchronization of risk events in the two groups of risk time windows. The final output is Group entity node corresponding to Group risk correlation indicates the intensity of risk transmission between two groups of nodes.
[0035] Preferably, according to the preset correlation strength threshold The association deviation of the group risk association degree is used to perform directed weighted connections of K entity nodes, where the association strength threshold is a preset threshold value, and only node pairs with an association degree higher than this value are retained. If the association degree of group (A, B) is 0.8 (exceeding the threshold), an edge A→B is added to the graph with a weight of 0.8. The directionality is determined by the chronological order of the risk events. For example, if the risk event of A is always earlier than that of B, the edge direction is A→B, and a directed topological network is finally obtained.
[0036] Furthermore, step S140 further includes step S141, according to the Whether the group risk correlation falls within the preset correlation strength threshold, The group entity nodes are screened to obtain P group associated nodes; in step S142, the K entity nodes are topologically connected in an undirected manner according to the P group associated nodes, and an initial topological network is output; in step S143, the P risk trigger direction frequencies of the P group temporal risk time windows are counted, and P risk conduction direction vectors are constructed according to the P risk trigger direction frequencies; in step S144, the P risk conduction direction vectors are used to perform directed weighting of the initial topological network, and the directed topological network is output.
[0037] Preferably, the preset correlation strength threshold is set to 0.7, and the comparison Whether the group risk correlation falls within the preset correlation strength threshold, Among the entity node pairs of the group, the valid association pairs whose risk correlation reaches the preset association strength threshold are screened out and recorded as P group association nodes. Then, the P group association nodes are temporarily stored with K entity nodes as undirected edges (i.e., bidirectional connections) to form an initial topological network. For example, if the P group is (A, B), (A, C), (B, D), then the initial network contains edges AB, AC, BD (no direction). Then, the frequency of P risk trigger directions in the P group temporal risk time window is counted, that is, for each group of association nodes, the number of times its risk events occur in sequence is counted. For example, if A→B (A's risk event occurs earlier than B) occurs 5 times and B→A occurs 2 times, then the frequency of the A→B direction is 5 and the frequency of the B→A direction is 2. Then, according to the P risk trigger direction frequencies, P risk transmission direction vectors are constructed; finally, the P risk transmission direction vectors are used to perform directed weighting of the initial topological network. When the frequency of the A→B direction is significantly higher than that of B→A, the edge A→B is retained, otherwise it is regarded as bidirectional or non-directional, and the proportion of the direction frequency is used as the corresponding weight. Finally, P directed edges are constructed and combined to generate a directed topological network.
[0038] In step S200, the K mirror nodes load the K time series attribute data streams in the directed topology network to trigger cascade risk conduction mining and output a real-time risk conduction path.
[0039] Step S200 further includes step S210, taking the P risk conduction direction vectors as timing constraints, taking the P group of associated nodes as node screening constraints, and locally calling the P group of timing risk conduction feature data; step S220, taking the P group of timing risk conduction feature data as training data, performing cascade risk conduction model training in the directed topology network, and outputting a cascade risk conduction graph neural network, wherein the cascade risk conduction graph neural network includes K node conduction identification models of the K mirror nodes; step S230, mapping the K timing attribute data streams and inputting them in parallel into the K node conduction identification models of the cascade risk conduction graph neural network, and collaboratively generating the real-time risk conduction path through topology message passing.
[0040] Preferably, P risk conduction direction vectors are used as time series constraints, and P groups of associated nodes are used as node screening constraints. That is, according to the conduction direction vector preferences of P group node pairs, P groups of significantly associated node pairs are screened out, and then the time series risk feature data of these node pairs within the sliding time window are extracted from the storage, such as the transaction amount, risk score, event trigger time, etc. of node A and node B in the past hour, focusing on the dynamic behavior data of highly associated node pairs, reducing noise interference, and ensuring that the conduction direction of the training data is consistent with the topological network; a graph neural network (GNN) architecture is constructed based on a directed topological network to support message passing, i.e., information interaction between nodes. Specifically, the P group of time series Risk conduction feature data is used as training data to perform a cascade risk conduction model, learn to predict the probability of the risk conduction path, obtain K-node conduction identification models of K mirror nodes, output the cascade risk conduction graph neural network, and retain the directionality and weight of the network topology; then the K time series attribute data streams are mapped and input in parallel into the K-node conduction identification model of the cascade risk conduction graph neural network. Each node model judges the current risk status based on the input data, and then transmits the risk signal through the directed edge. Finally, the output of all nodes is integrated to generate a complete conduction path, realizing real-time dynamic risk tracking and responding to changes in seconds. Parallel computing and message passing mechanisms support low-latency analysis.
[0041] Furthermore, step S230 also includes step S231, in the cascade risk transmission graph neural network, after the B mirror node receives the A risk transmission identifier and the A risk transmission probability sent by the A mirror node, the B node transmission identification model is activated, wherein the A mirror node is out-degree connected to the B mirror node; step S232, the A risk transmission identifier, the A risk transmission probability and the B time series data stream are loaded into the B node transmission identification model, risk transmission prediction is performed through the B node transmission identification model, and the B risk transmission identifier and the B risk transmission probability are output; step S233, if the D node orientation is parsed from the B risk transmission identifier, the B risk transmission identifier and the B risk transmission probability are sent to the D mirror node according to the D node orientation, and the D node transmission identification model is activated in the cascade risk transmission graph neural network, wherein the B mirror node is out-degree connected to the C mirror node and the D mirror node respectively; step S234, and so on, recursively transmit the risk identification and risk probability in the cascade risk transmission graph neural network until the probability threshold is less than the preset scale, and then the real-time risk transmission path is generated by backtracking the node chain.
[0042] Preferably, dynamic propagation and path generation of risk signals in the cascade risk conduction graph neural network are performed through recursive interaction between nodes, wherein there is a preset out-degree connection between mirror node A and mirror node B, that is, A and B have an actual business relationship in the basic layer dynamic graph, such as a transfer link. Specifically, when mirror node A detects a risk, such as an abnormal transaction, its local model generates a risk conduction identifier, marks the risk type, such as "fund borrowing" and the conduction probability (such as 0.9), and sends it to mirror node B through the out-degree connection. After receiving the signal from mirror node A, mirror node B activates its own conduction identification model, loads A risk conduction identifier, A risk conduction probability and B time series data stream into the conduction identification model of node B for risk conduction prediction. If it is determined that the risk may continue to be transmitted, the B risk conduction identifier and the updated probability, i.e., B risk conduction probability, are output.
[0043] Preferably, if the B mirror node is connected to the C mirror node and the D mirror node at the same time, but only the direction of the D mirror node meets the conduction condition (such as the probability threshold > 0.5), the signal is only transmitted to the D mirror node, and the mirror node is not activated. According to the D node guidance, the B risk conduction identifier and the B risk conduction probability are sent to the D mirror node, and the D node conduction identification model is activated in the cascade risk conduction graph neural network to perform risk conduction prediction. Similarly, the risk identification and risk probability are recursively conducted in the cascade risk conduction graph neural network. When the conduction probability decays layer by layer to a preset scale (such as <0.3), the recursion is stopped to avoid invalid calculations. Then, by tracing back the node chain in reverse, that is, tracing back the signal source from the termination node to splice the node chain, a real-time risk conduction path is generated and the probability value of each segment is marked to ensure the accuracy and efficiency of path generation.
[0044] Preferably, the effectiveness of risk transmission is judged by a probability threshold, and the source of the risk is traced back to generate a complete path. Specifically, when the risk transmission probability received by node B is lower than a preset threshold, it is determined that the risk has been significantly attenuated after being transmitted to B, and no longer continues to transmit signals to subsequent nodes. At this time, node B serves as the termination point of the transmission path; then starting from the termination node B, trace back layer by layer along the input edge of the directed topological network (i.e., the direction of the risk source), including locating the direct upstream node A of B, and determining the risk transmission path from A to B. If node A itself also receives signals from other nodes (such as X), continue to trace back to X until the risk source node is found, that is, the node without upstream input; finally, the traced node chain is spliced in sequence to generate a complete risk transmission path, associate the risk transmission relationship between node A and node B, and mark the probability value of each segment to form an explainable transmission evidence chain.
[0045] In step S300, after the base layer dynamic graph receives the real-time risk transmission path transmitted back by the analysis layer dynamic graph, the base layer dynamic graph locally stores the real-time risk transmission path and constructs a node risk handling instruction based on the real-time risk transmission path.
[0046] Preferably, the base layer dynamic graph stores the real-time risk transmission path transmitted back by the analysis layer dynamic graph locally and persistently, forming two types of data assets: a historical path library, which is archived in time series for long-term pattern analysis, such as high-frequency transmission path mining; and topological enhancement data, which dynamically updates the path information to the edge weights of the base layer dynamic graph; and then constructs node risk disposal instructions based on the real-time risk transmission path, for example, generating blocking instructions for source nodes, monitoring and strengthening instructions for key transit nodes, and tracing and verifying instructions for transmission edges, and setting instruction priorities according to path probability values, and then pushing the instructions to business systems, such as bank core systems and risk control platforms, through APIs or message queues to trigger automated disposal. At the same time, through the metadata of the path nodes, such as account type and risk level, the preset disposal rule library is matched, and the effect of the node risk disposal instructions is fed back for analysis, optimizing subsequent transmission predictions, and realizing active early warning of cascading risks and optimization of resource consumption.
[0047] Step S400: Based on the mirror node structure of the real-time risk transmission path, the node risk handling instruction is transmitted back to the source entity node.
[0048] Step S400 further includes step S410, parsing the real-time risk transmission path and extracting the source mirror node identifier; step S420, locating the source entity node corresponding to the source mirror node identifier in the basic layer dynamic graph based on the cross-layer node mapping relationship; step S430, matching the node risk disposal instruction according to the risk identifier of the source entity node, and returning the node risk disposal instruction to the source entity node through the security instruction channel.
[0049] Preferably, according to the mirror node composition of the real-time risk transmission path, the node risk disposal instruction is returned to the source entity node, that is, the risk transmission path generated by the analysis layer dynamic map is accurately mapped back to the entity node of the basic layer dynamic map, and the disposal measures for the risk source are triggered. Specifically, the first node identifier, that is, the source mirror node identifier, is extracted from the real-time risk transmission path transmitted back by the analysis layer dynamic map, and the corresponding source entity node in the basic layer dynamic map is located through the pre-established cross-layer node mapping relationship, that is, the cross-layer node mapping table, and then the risk identifier of the source entity node is read, and the node risk disposal instruction is matched according to the risk identifier of the source entity node. Specifically, the predefined disposal rules are matched according to the risk type. The database generates targeted node risk disposal instructions, including high-risk identifiers for immediate blocking; medium-risk identifiers for enhanced verification; and low-risk identifiers for monitoring tags. Finally, the node risk disposal instructions are transmitted back to the source entity node corresponding to the dynamic graph of the basic layer through secure instruction channels such as encrypted APIs or dedicated message queues, while ensuring that the instructions are tamper-proof and leak-proof during transmission and execution. Furthermore, the closed-loop control of risk discovery-source positioning-instruction execution is achieved through cross-layer node association. Through the two-way mapping of mirror nodes and entity nodes, it is ensured that the virtual analysis results can accurately act on real business objects, realize differentiated risk response strategies, and improve the accuracy and timeliness of cascade risk warnings and reduce resource consumption.
[0050] In the above, refer to Figure 1 The dynamic graph analysis method for big data according to the embodiment of the present invention is described in detail. Figure 2 A dynamic graph analysis system for big data according to an embodiment of the present invention is described.
[0051] The dynamic graph analysis system for big data according to the embodiment of the present invention is used to solve the technical problems of poor real-time performance, low computing efficiency, insufficient dynamic topology adaptation, and delayed risk transmission path mining in the existing technology, and achieves the technical effects of efficient real-time risk transmission analysis, precise inheritance of dynamic topology, active warning of cascading risks, and optimized resource consumption. Figure 2 As shown, the dynamic graph analysis system for big data includes: a time series attribute data mirroring unit 10, a real-time risk transmission path output unit 20, a node risk disposal instruction construction unit 30, and a node risk disposal instruction return unit 40.
[0052] A time series attribute data mirroring unit 10 is used for mirroring K time series attribute data streams from K entity nodes in the base layer dynamic graph to K mirror nodes in the analysis layer dynamic graph based on a sliding time window, wherein the K mirror nodes inherit the directed topological network of the base layer dynamic graph; a real-time risk conduction path output unit 20 is used for the K mirror nodes to load the K time series attribute data streams in the directed topological network to trigger cascade risk conduction mining and output a real-time risk conduction path; a node risk disposal instruction construction unit 30 is used for locally storing the real-time risk conduction path after the base layer dynamic graph receives the real-time risk conduction path returned by the analysis layer dynamic graph, and constructing a node risk disposal instruction based on the real-time risk conduction path; a node risk disposal instruction return unit 40 is used for returning the node risk disposal instruction to the source entity node according to the mirror node composition of the real-time risk conduction path.
[0053] The specific configuration of the time series attribute data mirroring unit 10 will be described in detail below. The time series attribute data mirroring unit 10 further includes: transmitting heterogeneous data streams to the base layer dynamic graph using a distributed message queue; performing attribute decomposition on the heterogeneous data streams by the base layer dynamic graph to obtain K time series attribute data streams; and mapping and loading the K time series attribute data streams to the K entity nodes in the base layer dynamic graph based on the data attribute labels of the K time series attribute data streams.
[0054] The specific configuration of the time series attribute data mirroring unit 10 will be described in detail below. The time series attribute data mirroring unit 10 further includes: extracting K time series risk time windows locally from the K entity nodes; combining and enumerating the K entity nodes to obtain After grouping entity nodes, according to the The group entity node mapping combines the K time series risk time windows to obtain Group time series risk time window; by calculating the The risk event overlap coefficient of the group time series risk time window is used to quantify the Group entity node Group risk correlation; based on the preset correlation strength threshold and the The association deviation of the group risk association degree is used to perform directed weighted connection of the K entity nodes to obtain the directed topological network.
[0055] The specific configuration of the time series attribute data mirror unit 10 will be described in detail below. The time series attribute data mirror unit 10 further includes: Whether the group risk correlation falls within the preset correlation strength threshold, The group entity nodes are screened to obtain P groups of associated nodes; based on the P groups of associated nodes, the K entity nodes are undirectedly topologically connected to output an initial topological network; the P risk triggering direction frequencies of the P groups of temporal risk time windows are counted, and P risk conduction direction vectors are constructed based on the P risk triggering direction frequencies; the P risk conduction direction vectors are used to perform directed weighting on the initial topological network to output the directed topological network.
[0056] The specific configuration of the time series attribute data mirroring unit 10 will be described in detail below. The time series attribute data mirroring unit 10 further includes: slidingly splitting the heterogeneous data stream into multiple heterogeneous data segments; performing structured parsing on the multiple heterogeneous data segments to output multiple standardized attribute vectors; predefining K node attribute vectors for the K entity nodes; calculating the similarity of multiple groups of attribute vectors between the multiple standardized attribute vectors and the K node attribute vectors, and then reorganizing the multiple groups of attribute vectors into K groups of heterogeneous data segments according to dynamic matching rules; and splicing the K groups of heterogeneous data segments in a time series manner to output the K time series attribute data streams.
[0057] The specific configuration of the real-time risk conduction path output unit 20 will be described in detail below. The real-time risk conduction path output unit 20 further includes: using the P risk conduction direction vectors as timing constraints and the P groups of associated nodes as node screening constraints, locally calling the P groups of temporal risk conduction feature data; using the P groups of temporal risk conduction feature data as training data, training a cascade risk conduction model in the directed topology network, and outputting a cascade risk conduction graph neural network, wherein the cascade risk conduction graph neural network includes K node conduction identification models of the K mirror nodes; mapping the K temporal attribute data streams and inputting them in parallel into the K node conduction identification models of the cascade risk conduction graph neural network, and collaboratively generating the real-time risk conduction path through topology message passing.
[0058] The specific configuration of the real-time risk transmission path output unit 20 will be described in detail below. The real-time risk transmission path output unit 20 further includes: in the cascaded risk transmission graph neural network, after receiving the A risk transmission identifier and A risk transmission probability from the A mirror node, the B mirror node activates the B node transmission identification model, wherein the A mirror node is out-degree connected to the B mirror node; the A risk transmission identifier, the A risk transmission probability, and the B time series data stream are loaded into the B node transmission identification model, and the B node transmission identification model performs risk transmission prediction and outputs the B risk transmission identifier and the B risk transmission probability; if the D node orientation is parsed from the B risk transmission identifier, the B risk transmission identifier and the B risk transmission probability are sent to the D mirror node based on the D node orientation, and the D node transmission identification model is activated in the cascaded risk transmission graph neural network, wherein the B mirror node is out-degree connected to the C mirror node and the D mirror node respectively; and so on, recursively transmitting risk identification and risk probability in the cascaded risk transmission graph neural network until the probability threshold is less than a preset scale, at which time the real-time risk transmission path is generated by backtracking the node chain.
[0059] The specific configuration of the node risk handling instruction return unit 40 will be described in detail below. The node risk handling instruction return unit 40 further includes: parsing the real-time risk transmission path to extract the source mirror node identifier; locating the source entity node corresponding to the source mirror node identifier in the base layer dynamic graph based on the cross-layer node mapping relationship; matching the node risk handling instruction according to the risk identifier of the source entity node, and returning the node risk handling instruction to the source entity node via a secure instruction channel.
[0060] The dynamic graph analysis system for big data provided by the embodiment of the present invention can execute the dynamic graph analysis method for big data provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0061] Figure 3 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing an embodiment of the present invention. Figure 3 The electronic device shown is merely an example and should not limit the functionality and scope of use of the embodiments of the present invention. The electronic device is in the form of a general-purpose computing device, and its components may include, but are not limited to, an input device 401, a processor 402, a memory 403, and an output device 404. The processor 402 may be one or more; the memory 403 may include a computer-readable medium and at least one program product, which has a set (at least one) of program modules configured to perform the functions of the various embodiments of the present application.
[0062] The memory 403 shown in the embodiment of the present invention may adopt any combination of one or more computer-readable media; the computer-readable storage medium may be, but is not limited to, an infrared, semiconductor system, device or component, or any combination of the above, for storing software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the dynamic graph analysis method for big data in the embodiment of the present invention. The processor 402 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 403, thereby realizing the above-mentioned dynamic graph analysis method for big data.
[0063] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.
[0064] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A dynamic graph analysis method for big data, characterized by: The method comprises: The K entity nodes in the base layer dynamic graph mirror the K time series attribute data streams to the K mirror nodes in the analysis layer dynamic graph based on the sliding time window, wherein the K mirror nodes inherit the directed topology network of the base layer dynamic graph; The K mirror nodes load the K time series attribute data streams to trigger cascade risk conduction mining in the directed topology network, and output a real-time risk conduction path; After receiving the real-time risk transmission path transmitted back by the analysis layer dynamic graph, the base layer dynamic graph locally stores the real-time risk transmission path and constructs a node risk handling instruction based on the real-time risk transmission path; According to the mirror node composition of the real-time risk transmission path, the node risk handling instruction is returned to the source entity node.
2. The big data-oriented dynamic graph analysis method according to claim 1, characterized in that: The method further comprises: The heterogeneous data streams are transmitted to the base layer dynamic graph using a distributed message queue; The base layer dynamic graph performs attribute decomposition on the heterogeneous data stream to obtain K time series attribute data streams; According to the data attribute labels of the K time series attribute data streams, the K time series attribute data streams are mapped and loaded into the K entity nodes in the base layer dynamic graph.
3. The big data-oriented dynamic graph analysis method according to claim 1, characterized in that: The method further comprises: Extracting K timing risk time windows locally at the K entity nodes; Combine and enumerate the K entity nodes to obtain After grouping entity nodes, according to the The group entity node mapping combines the K time series risk time windows to obtain Group temporal risk time window; By calculating the The risk event overlap coefficient of the group time series risk time window is used to quantify the Group entity node Group risk association; According to the preset correlation strength threshold and the The association deviation of the group risk association degree is used to perform directed weighted connection of the K entity nodes to obtain the directed topological network.
4. The big data-oriented dynamic graph analysis method according to claim 3, characterized in that: According to the preset correlation strength threshold and the The method further comprises: performing directed weighted connections on the K entity nodes based on the association deviation of the group risk association degree to obtain the directed topological network. According to the Whether the group risk correlation falls within the preset correlation strength threshold, The group entity nodes are filtered to obtain P group associated nodes; Performing undirected topological connections on the K entity nodes according to the P groups of associated nodes, and outputting an initial topological network; Count the P risk trigger direction frequencies of P groups of time series risk time windows, and construct P risk transmission direction vectors based on the P risk trigger direction frequencies; The P risk conduction direction vectors are used to perform directed weighting on the initial topology network, and the directed topology network is output.
5. The big data-oriented dynamic graph analysis method according to claim 2, characterized in that: The base layer dynamic graph performs attribute decomposition on the heterogeneous data stream to obtain K time-series attribute data streams, and the method includes: Slidingly dividing the heterogeneous data stream into a plurality of heterogeneous data segments; Performing structural analysis on the multiple heterogeneous data segments and outputting multiple standardized attribute vectors; Predefine K node attribute vectors of the K entity nodes; After calculating the similarities between the plurality of normalized attribute vectors and the plurality of groups of attribute vectors of the K node attribute vectors, the plurality of groups of attribute vectors are reorganized into K groups of heterogeneous data segments according to a dynamic matching rule; The K groups of heterogeneous data segments are spliced in time series to output the K time series attribute data streams.
6. The method for dynamic graph analysis of big data according to claim 4, wherein: The K mirror nodes load the K time-series attribute data streams to trigger cascade risk conduction mining in the directed topology network, and output a real-time risk conduction path. The method includes: Using the P risk conduction direction vectors as time series constraints and the P groups of associated nodes as node screening constraints, locally calling the P groups of time series risk conduction feature data; Using the P groups of temporal risk conduction feature data as training data, performing cascade risk conduction model training on the directed topological network, and outputting a cascade risk conduction graph neural network, wherein the cascade risk conduction graph neural network includes K node conduction recognition models of the K mirror nodes; The K time series attribute data streams are mapped and input in parallel into the K node conduction identification models of the cascade risk conduction graph neural network, and the real-time risk conduction path is collaboratively generated through topological message passing.
7. The big data-oriented dynamic graph analysis method according to claim 6, characterized in that: Mapping the K time-series attribute data streams and inputting them in parallel into the K node conduction recognition models of the cascaded risk conduction graph neural network, and collaboratively generating the real-time risk conduction path through topological message passing, the method comprising: In the cascade risk conduction graph neural network, after the B mirror node receives the A risk conduction identifier and the A risk conduction probability sent by the A mirror node, the B node conduction identification model is activated, wherein the A mirror node is out-degree connected to the B mirror node; Loading the A risk transmission identifier, A risk transmission probability, and B time series data stream into the B node transmission identification model, performing risk transmission prediction via the B node transmission identification model, and outputting the B risk transmission identifier and B risk transmission probability; If the D node orientation is parsed from the B risk transmission identifier, the B risk transmission identifier and the B risk transmission probability are sent to the D mirror node according to the D node orientation, and the D node transmission identification model is activated in the cascade risk transmission graph neural network, wherein the B mirror node is out-degree connected to the C mirror node and the D mirror node respectively; By analogy, the risk identification and risk probability are recursively conducted in the cascade risk conduction graph neural network until the probability threshold is less than a preset scale, and the real-time risk conduction path is generated by backtracking the node chain.
8. The big data-oriented dynamic graph analysis method according to claim 1, characterized in that: According to the mirror node structure of the real-time risk transmission path, the node risk handling instruction is transmitted back to the source entity node, and the method includes: Parsing the real-time risk transmission path and extracting the source mirror node identifier; Based on the cross-layer node mapping relationship, locate the source entity node corresponding to the source mirror node identifier in the base layer dynamic graph; After matching the node risk handling instruction according to the risk identifier of the source entity node, the node risk handling instruction is returned to the source entity node through a secure instruction channel.
9. A dynamic graph analysis system for big data, characterized by: The system is used to implement the big data-oriented dynamic graph analysis method according to any one of claims 1 to 8, and the system includes: A time series attribute data mirroring unit is used to mirror K time series attribute data streams from K entity nodes in the base layer dynamic graph to K mirror nodes in the analysis layer dynamic graph based on a sliding time window, wherein the K mirror nodes inherit the directed topology network of the base layer dynamic graph; A real-time risk conduction path output unit, configured to load the K time series attribute data streams on the K mirror nodes to trigger cascade risk conduction mining in the directed topology network and output a real-time risk conduction path; A node risk handling instruction construction unit is configured to locally store the real-time risk transmission path transmitted back by the analysis layer dynamic graph after the base layer dynamic graph receives the real-time risk transmission path, and to construct a node risk handling instruction based on the real-time risk transmission path; The node risk handling instruction return unit is used to return the node risk handling instruction to the source entity node according to the mirror node structure of the real-time risk transmission path.
10. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; A processor is configured to implement the big data-oriented dynamic graph analysis method according to any one of claims 1 to 8 when executing the executable instructions stored in the memory.
Citation Information
Patent Citations
Risk conduction probability knowledge graph generation method and device thereof, equipment and storage medium
CN114048330A
Risk conduction prediction method, device, equipment and medium
CN115034596A
Network risk processing method and system based on atlas
CN119766505A
Risk conduction association map optimization method and apparatus, computer device and storage medium
WO2020232879A1