A subgraph partitioning equity penetration method and device based on equity graph data
Patent Information
- Application Number
- CN202410693582.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-05-31
AI Technical Summary
但当股权信息数据量较大,指定层级较多,通过对点迭代计算时需要查询的边数据量随之增大,将耗费较多的时间和计算资源
[0029]目前,常用的金融股权关系图计算方法包括企业股东逐层穿透和深度优先遍历。企业股东逐层穿透方法基于公开披露的金融数据,逐层追溯企业的股东信息,并输出指定层级下的股东详情。然而,在股东数量众多和需要指定多个层级的情况下,计算的难度将增加。深度优先遍历方法可以获取一个机构对另一个机构持股的所有路径,但其缺点是很难全面反映某个机构的股东信息和结构层次。本发明采用子图分割多跳股权穿透计算方法,能较为全面地反映金融实体的股权架构信息、减少计算时间、减少数据传递开销并节约计算资源,实现简化股权穿透流程,提高穿透计算效率。
Smart Images

Figure CN118467648B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equity calculation and analysis technology, and proposes a method and apparatus for equity penetration based on subgraph segmentation of equity graph data. Background Technology
[0002] Financial equity relationships, as an important component of the financial sector, are crucial for understanding the equity structure, shareholder relationships, and market risks among companies. However, due to the complexity of financial graphs and the massive amount of financial data, traditional equity penetration methods and tools often fall short of the need for comprehensive and accurate analysis of equity relationships.
[0003] Currently, commonly used methods for calculating financial equity relationship graphs include layer-by-layer shareholder penetration, depth-first search, and multi-level equity graph calculation methods. Layer-by-layer shareholder penetration, based on publicly disclosed financial data, traces the company's shareholder information layer by layer and outputs shareholder details at specified levels. However, the computational complexity increases when there are numerous shareholders and multiple levels need to be specified. Depth-first search can obtain all paths of one institution's shareholding in another, but its drawback is that it struggles to comprehensively reflect an institution's shareholder information and structural hierarchy. Multi-level equity graph calculation methods, based on corporate registration data, abstract corporate entities as nodes and equity relationships as edges. By iterating through the edges to each node, equity information is transferred, achieving higher efficiency than traditional layer-by-layer shareholder penetration methods. However, when the amount of equity information data is large and there are many specified levels, the amount of edge data that needs to be queried during point-to-point iteration increases, consuming significant time and computational resources. Therefore, there is an urgent need for a method that can reduce computation time, data transfer overhead, and computational resources to better perform penetration calculation and analysis of multi-hop equity relationships within a company. Summary of the Invention
[0004] In order to simplify the process and improve computational efficiency when transmitting equity information, this invention provides a subgraph segmentation equity penetration method based on equity graph data.
[0005] The present invention adopts the following technical solution:
[0006] A subgraph segmentation and equity penetration method based on equity graph data includes the following steps:
[0007] Step 1: Read the enterprise's business registration data, store the enterprise's business registration data into a flexible distributed data structure, obtain the initial vertex data and initial edge data, and construct the initial equity graph data;
[0008] Step 2: Extract connected components from the initial equity graph data in Step 1, and group, sort, split, and combine the connected components to divide multiple connected components into multiple subgraphs;
[0009] Step 3: Traverse all ternary views under each subgraph in Step 2. Based on the edge recorded in the ternary view and its corresponding enterprise node and shareholder node information, aggregate equity information to generate enterprise node data containing direct shareholder equity information. Nodes without direct shareholders have their equity information initialized to empty.
[0010] Step 4: Construct graph data containing one-hop relationships using enterprise node data containing direct shareholder node equity information and initial edge data, where each node information carries one-hop equity information;
[0011] Step 5: For each ternary view in the graph data of the one-hop relationship constructed in Step 4, obtain the shareholder node based on the current aggregated hop count information, and determine whether the shareholder node has equity information. If it does, update the equity information of the shareholder node to the enterprise node; if the shareholder node does not have equity information, write the equity information recorded in the ternary view directly to the enterprise node.
[0012] Step 6: Deduplicate the equity information set and combine the processed equity node data with the equity relationship edge data to form a new equity graph data;
[0013] Step 7: Repeat steps 4 to 6, updating the equity information carried by the node data in each repetition.
[0014] Furthermore, in step 1, the enterprise business registration data is cleaned and preprocessed and then stored in the form of a node dataset and an edge dataset. The vertex dataset contains vertex ID, entity name, entity type and other information, and the edge dataset contains source node ID, target node ID and equity information. The graph data is composed of vertex data and edge data connected together.
[0015] Furthermore, in step 2, the number of edges contained in each connected component in the connected component set is counted, the connected components in the connected component set are sorted in ascending order according to the number of edges contained, and the sorting results are then divided and combined to form multiple subgraphs.
[0016] Furthermore, in step 3, the equity information includes the shareholder node ID, shareholder node name, shareholder node type, shareholder shareholding ratio, shareholder shareholding hop count, shareholder shareholding path, and other information.
[0017] Furthermore, the enterprise node data format generated in step 3, which includes direct shareholder equity information, is (entity id, (entity name, entity type, entity information, entity shareholder array)). The data format of the entity shareholder array is [(shareholder id, shareholder name, shareholder type, shareholder information, shareholder shareholding ratio, hop distance, shareholder shareholding path)]. The intermediate nodes for equity relationship transmission are stored in the shareholder shareholding path, with the format [(path entity id, path shareholding ratio)].
[0018] Furthermore, when steps 4 to 6 are not repeated, the nodes of the obtained graph data have one-hop equity information. When steps 4 to 6 are repeated once, the obtained entity nodes have two-hop equity information, that is, the second-level shareholder information of the enterprise. When steps 4 to 6 are repeated n times, the obtained entity nodes have (n+1) hop equity information, that is, the (n+1) level shareholder information of the enterprise.
[0019] Furthermore, in step 7, steps 4 to 6 are repeated, i.e., the equity penetration method is used to perform multi-hop equity information penetration, and the attributes of the entity node are updated in the calculation of each hop.
[0020] On the other hand, the present invention also provides a subgraph segmentation equity penetration device based on equity graph data, including a data reading module, a graph data construction module, a subgraph segmentation module, an equity information aggregation module, an equity information deduplication module, a loop calculation module, and an equity structure extraction module;
[0021] The data reading module is used to read in the cleaned and preprocessed enterprise business registration data, and generate node data and edge data for constructing graph data respectively; the node data must include enterprise nodes and shareholder nodes, enterprise nodes include sole proprietorships, limited liability companies or joint-stock companies, and shareholder nodes include natural persons and others;
[0022] The graph data construction module is used to construct a graph representing equity relationships using node data and edge data. When the input is initial node data and initial edge set data, the graph constructed by this module represents a one-hop equity structure. When the input is node data with equity information and equity relationship edge data, the graph constructed by this module contains equity graph data with a certain number of hops in the equity relationship.
[0023] The subgraph segmentation module is used to statistically analyze and segment connected components based on equity graph data, dividing multiple connected components into subgraphs, and ensuring that the computational load of each subgraph is as similar as possible. This transforms the equity penetration operation from the original complete equity graph data into equity penetration of multiple subgraphs separately, allowing the device to complete equity penetration calculations with more hops in the large financial equity graph with a small amount of resources. Furthermore, when the subgraphs are properly segmented, it can achieve higher efficiency than the full graph equity penetration method.
[0024] The equity information aggregation module is used to traverse all ternary view relationships in the subgraph data, analyze node types, and send the equity relationships between nodes to the nodes in the form of messages to update the equity information in the node data.
[0025] The equity information deduplication module is used to deduplicat equity information in node data and output new vertex data with equity information for construction.
[0026] The loop calculation module is used to make the graph data construction module, the equity information aggregation module, and the equity information deduplication module perform repeated operations. The number of repetitions is the number of hops for equity penetration, and finally an equity data graph with a specified number of hops is obtained, whose node data carries equity information under the specified number of hops.
[0027] The equity structure extraction module is used to extract the equity structure information related to an entity from a single point data point carrying equity information.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] Currently, commonly used methods for calculating financial equity relationship graphs include layer-by-layer shareholder penetration and depth-first traversal. Layer-by-layer shareholder penetration, based on publicly disclosed financial data, traces the company's shareholder information layer by layer and outputs shareholder details at a specified level. However, the computational complexity increases when there are numerous shareholders and multiple levels need to be specified. Depth-first traversal can obtain all paths of one institution's shareholding in another, but its drawback is that it struggles to comprehensively reflect the shareholder information and structural hierarchy of an institution. This invention employs a subgraph segmentation multi-hop equity penetration calculation method, which can more comprehensively reflect the equity structure information of financial entities, reduce computation time, reduce data transmission overhead, and save computational resources, thereby simplifying the equity penetration process and improving penetration calculation efficiency. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort. Attached image description:
[0032] Figure 1 This is a schematic diagram of the subgraph segmentation equity penetration method based on equity graph data proposed in this invention.
[0033] Figure 2This is a schematic diagram of the initial graph data structure constructed by the present invention based on example point data and edge data.
[0034] Figure 3 This is a schematic diagram of the data format of the equity diagram data proposed in this invention.
[0035] Figure 4 This is a flowchart illustrating the implementation process of constructing equity diagram data as proposed in this invention.
[0036] Figure 5 This is a schematic diagram illustrating the process of extracting connected components and partitioning subgraphs from graph data according to the present invention.
[0037] Figure 6 This is a schematic diagram of equity subgraph data segmentation proposed in this invention.
[0038] Figure 7 This is a schematic diagram of the subgraph segmentation equity penetration device based on equity graph data proposed in this invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0040] Example 1
[0041] Use the following example point data, formatted as (entity id, entity name, entity type, other information).
[0042] 1A type_A i nfo_A
[0043] 2B type_B i nfo_B
[0044] 3C type_C info_C
[0045]
[0046] The following example edge data is used, in the format (shareholder ID, company ID, shareholding percentage).
[0047]
[0048] Please see Figure 1 The steps of the subgraph segmentation equity penetration method based on equity graph data proposed in this invention include:
[0049] Step 1: Read the enterprise's business registration data, store the enterprise's business registration data in an elastic distributed data structure, obtain the initial vertex data and initial edge data, and construct the initial equity graph data based on the initial vertex data and initial edge data;
[0050] Step 2: Extract connected components from the initial equity graph data in Step 1, group and sort the connected components, and then divide and combine the sorting results to divide multiple connected components into multiple subgraphs.
[0051] Step 3: Traverse all ternary views under each subgraph in Step 2. Based on the edge recorded in the ternary view and its corresponding enterprise node and shareholder node information, aggregate equity information to generate enterprise node data containing direct shareholder equity information. Nodes without direct shareholders have their equity information initialized to empty.
[0052] Step 4: Construct graph data containing one-hop relationships using enterprise node data containing direct shareholder node equity information and initial edge data; compared with the initial graph data, in the graph data containing one-hop relationships, each node information carries one-hop equity information.
[0053] Step 5: For each ternary view in the graph data constructed in Step 4, obtain the shareholder node based on the current aggregation hop count information, determine whether the shareholder node has equity information, if so, it means that an equity chain has been generated, and update the equity information of the shareholder node to the enterprise node; if the shareholder node does not have equity information, then write the equity information recorded in the ternary view directly to the enterprise node.
[0054] Step 6: Deduplicate the equity information set and combine the processed equity node data with the equity relationship edge data to form a new equity graph data;
[0055] Step 7: Repeat steps 4 to 6, updating the equity information carried by the node data in each repetition.
[0056] Please see Figure 2 Step 1 of the method involves reading the data and storing it in a resilient distributed dataset. The data format for point data is (entity id, (entity name, entity type, other information)), and the data format for edge data is (shareholder id, company id, shareholding percentage). The initial graph data constructed using the point and edge data is shown below. Figure 2-4 As shown.
[0057] Please see Figure 5Step 2 of the method extracts connected components from the initial graph data constructed in Step 1, obtaining connected component 1, connected component 2, connected component 3, and connected component 4. The number of edges in each connected component is counted, revealing that connected component 1 has 6 edges, connected component 2 has 2 edges, and connected components 3 and 4 each have 1 edge. These are then sorted in ascending order of edge count, resulting in connected component 3, connected component 4, connected component 2, and connected component 1. Connected component 1 is then combined individually to form a sub-component. Figure 1 Connected component 2, connected component 3, and connected component 4 are combined to form a sub-component. Figure 2 For the child Figure 1 To obtain the final shareholders of all nodes, multi-hop equity penetration calculation is required, and for sub-nodes... Figure 2 By performing a single-hop equity penetration calculation, the final shareholders of all nodes can be obtained. This step divides the complete graph data, forming multiple subgraphs for subsequent processing.
[0058] Please see Figure 6 Step 2, another way to extract connected components from the initial graph data constructed in Step 1, is as follows: Obtain connected components A, B, C, and D. Count the number of edges in each connected component; connected component A has 2 edges, connected component B has 3 edges, connected component C has 1 edge, and connected component D has 2 edges. Sort them in ascending order of edge count, resulting in connected components B, A, D, and C. Combine connected components A and D to form a sub-component. Figure 1 Combine connected component B and connected component C to form a sub-component. Figure 2 For the child Figure 1 A single-hop equity penetration calculation can determine the ultimate shareholders of all nodes, while for sub-nodes... Figure 2 This requires multi-hop equity penetration calculation to obtain the final shareholders of all nodes. This step segments the complete graph data, forming multiple subgraphs for subsequent processing.
[0059] Step 3 of the method involves preliminary traversal and aggregation result processing. The input data for step 3 is used as an example. Figure 5 sub Figure 1 For example, the method pairs Figure 1The process iterates through all ternary views, aggregating equity information based on the edges recorded in the ternary views and their corresponding enterprise and shareholder nodes. This equity information is then written into the enterprise nodes, generating enterprise node data containing direct shareholder equity information. Nodes without direct shareholders have their equity information initialized to empty. After this process, the aggregated node data format is (entity id, (entity name, entity type, entity information, entity shareholder array)). The entity shareholder array format is [(shareholder id, shareholder name, shareholder type, shareholder information, shareholder shareholding ratio, hop distance, shareholder shareholding path)]. If an entity has no shareholders, it is filled with [(0, null, null, null, 0.0, 0, null)] as the default shareholder array, ensuring consistency across all node data formats. The shareholder shareholding path stores intermediate nodes for equity relationship transmission, formatted as [(path entity id, path shareholding ratio)]. Figure 1 The data set containing equity points generated after equity penetration aggregation is as follows:
[0060] (1,(A,type_A,i nfo_A,[(0,nu ll,nu ll,nu ll,0.0,0,nu ll)]))
[0061] (2,(B,type_B,i nfo_B,[(1,A,type_A,i nfo_A,0.5,1,[(2,0.5)])]))
[0062] (3,(C,type_C,i nfo_C,[(1,A,type_A,i nfo_A,0.5,1,[(3,0.5)])]))
[0063] (4,(D,type_D,i nfo_D,[(2,B,type_B,i nfo_B,0.5,1,[(4,0.5)])]))
[0064] (5,(E,type_E,i nfo_E,[(2,B,type_B,i nfo_B,0.5,1,[(5,0.5)])]))
[0065] (6,(F,type_F,i nfo_F,[(5,E,type_E,i nfo_E,0.5,1,[(6,0.5)])]))
[0066] (7,(G,type_G,i nfo_G,[(5,E,type_E,i nfo_E,0.5,1,[(7,0.5)])]))
[0067] The above data includes information on direct shareholders. Since entity A has no shareholders, the shareholder array is a default value.
[0068] Step 4 involves constructing graph data containing equity relationships. Node data containing shareholder node equity information and initial edge data are used to construct graph data containing equity relationships, providing a data foundation for the next step of aggregating equity node data.
[0069] Step 5 involves aggregating equity node data. For the graph data containing equity relationships created in Step 4, each ternary view is traversed. Based on the current equity penetration hop count, it is determined whether the shareholder node has equity information. If it does, it indicates that an equity chain has been generated. Therefore, the equity information of the shareholder node is updated to the enterprise node to save the equity information passed by the shareholder. If the shareholder node does not have equity information, the equity information recorded in the ternary view is directly written to the enterprise node.
[0070] Step 6: Process equity node data. For equity node data, traverse all equity nodes and deduplicate them in the equity information set to ensure that the shareholding path of the same shareholder in the enterprise is not repeated. The reason for duplicate content in the equity information set is that each time Step 5 traverses to a specific ternary view, it will generate equity information that may already exist in the node data, which needs to be deduplicated to ensure the accuracy of the shareholder information array in the node data.
[0071] Step 7: Repeat steps 4 through 6, for the specified number of hops in the equity penetration calculation. In each repetition, the equity information carried by the node will be updated. After completion, the equity penetration calculation result for the specified number of hops will be saved in the node data.
[0072] Next, based on the equity-bearing point data generated in step 3, the first round of steps 4, 5, and 6 is performed sequentially, as described in step 7.
[0073] After the first loop, a point dataset containing 2-hop equity information is obtained, as follows:
[0074] (1,(A,type_A,i nfo_A,[(0,nu ll,nu ll,nu ll,0.0,0,nu ll)]))
[0075] (2,(B,type_B,i nfo_B,[(1,A,type_A,i nfo_A,0.5,1,[(2,0.5)])]))
[0076] (3,(C,type_C,i nfo_C,[(1,A,type_A,i nfo_A,0.5,1,[(3,0.5)])]))
[0077] (4,(D,type_D,i nfo_D,[(2,B,type_B,i nfo_B,0.5,1,[(4,0.5)]),(1,A,type_A,i nfo_A,0.25,2,[(2,0.5),(4,0.5)])]))
[0078] (5,(E,type_E,i nfo_E,[(2,B,type_B,i nfo_B,0.5,1,[(4,0.5)]),(1,A,type_A,i nfo_A,0.25,2,[(2,0.5),(5,0.5)])]))
[0079] (6,(F,type_F,i nfo_F,[(5,E,type_E,i nfo_E,0.5,1,[(6,0.5)]),(2,B,type_B,i nfo_B,0.25,2,[(5,0.5),(6,0.5)])]))
[0080] (7,(G,type_G,i nfo_G,[(5,E,type_E,i nfo_E,0.5,1,[(7,0.5)]),(2,B,type_B,i nfo_B,0.25,2,[(5,0.5),(7,0.5)])]))
[0081] The explanation for the above data is as follows: For points A, B, and C, since there are no shareholders holding shares outside the two-hop distance, no second-hop shareholder information was written after the first cycle, and the data for these points was not updated. However, for points D, E, F, and G, there are shareholders holding shares outside the two-hop distance; therefore, after the first cycle, second-hop shareholder information was added to the data for these points.
[0082] Next, based on the equity-bearing point data generated in the first cycle, the second round of steps 4, 5, and 6 is repeated as described in step 7.
[0083] After the second loop, a point dataset containing 3-hop equity information is obtained, with the following details:
[0084] (1,(A,type_A,i nfo_A,[(0,nu ll,nu ll,nu ll,0.0,0,nu ll)]))
[0085] (2,(B,type_B,i nfo_B,[(1,A,type_A,i nfo_A,0.5,1,[(2,0.5)])]))
[0086] (3,(C,type_C,i nfo_C,[(1,A,type_A,i nfo_A,0.5,1,[(3,0.5)])]))
[0087] (4,(D,type_D,i nfo_D,[(2,B,type_B,i nfo_B,0.5,1,[(4,0.5)]),(1,A,type_A,i nfo_A,0.25,2,[(2,0.5),(4,0.5)])]))
[0088] (5,(E,type_E,i nfo_E,[(2,B,type_B,i nfo_B,0.5,1,[(4,0.5)]),(1,A,type_A,i nfo_A,0.25,2,[(2,0.5),(5,0.5)])]))
[0089] (6,(F,type_F,i nfo_F,[(5,E,type_E,i nfo_E,0.5,1,[(6,0.5)]),(2,B,type_B,i nfo_B,0.25,2,[(5,0.5),(6,0.5)]),(1,A,type_A,i nfo_A,0.125,3,[(2,0.5),(5,0.5),(6,0.5)])]))
[0090] (7,(G,type_G,i nfo_G,[(5,E,type_E,i nfo_E,0.5,1,[(7,0.5)]),(2,B,type_B,i nfo_B,0.25,2,[(5,0.5),(7,0.5)]),(1,A,type_A,i nfo_A,0.125,3,[(2,0.5),(5,0.5),(7,0.5)])]))
[0091] The explanation for the above data is as follows: For points A, B, C, D, and E, since there are no shareholders holding shares beyond the three-hop distance, no three shareholder information entries were written after the second cycle, and the data for these points was not updated. However, for points F and G, there are shareholders holding shares beyond the three-hop distance; therefore, after the second cycle, three-hop shareholder information was added to the data for these points.
[0092] Example 2
[0093] Please see Figure 7 The equity penetration device based on equity graph data proposed in this invention includes a data reading module M1, a graph data construction module M2, a subgraph segmentation module M3, an equity information aggregation module M4, an equity information deduplication module M5, a loop calculation module M6, and an equity structure extraction module M7.
[0094] The data reading module M1 can read in the cleaned and preprocessed enterprise business data and generate node data and edge data for constructing graph data respectively; the node data must include enterprise nodes and shareholder nodes. Enterprise nodes include sole proprietorships, limited liability companies or joint-stock companies, and shareholder nodes include natural persons and others.
[0095] The graph data construction module M2 uses node data and edge data to construct a graph representing equity relationships. When the input is initial node data and initial edge set data, the graph constructed by this module represents a one-hop equity structure. When the input is node data and initial edge set data with equity information, this module constructs equity graph data containing certain equity penetration calculation results.
[0096] The subgraph segmentation module M3 extracts and analyzes the connected components of the complete equity graph data, sorts and combines the connected components according to the number of equity edges they contain to form multiple equity subgraph data, thus completing the subgraph segmentation.
[0097] The equity information aggregation module M4 updates the equity information in the node data based on each ternary view of the input graph data.
[0098] The equity information deduplication module M5 deduplicates the equity information in the node data, outputs node data containing equity information, and ensures the correctness of the equity information.
[0099] The loop calculation module M6 is used to make the graph data construction module M2, the equity information aggregation module M4, and the equity information deduplication module M5 repeat their operations. The number of repetitions is the number of hops for the equity penetration calculation, and finally the node data containing the equity information of the specified number of hops is obtained.
[0100] The equity structure extraction module M7 can extract the equity structure information related to an entity by using point data of a single aggregated equity information.
[0101] Please see Figure 5The equity graph data structure proposed in this invention is divided into node data, edge data, and a ternary view. In the edge-incremental graph, nodes represent entities, i.e., enterprise or shareholder information; node data contains entity data and equity information. Edges in the edge-incremental graph represent equity relationships between entities; edge data contains shareholder information on their holdings in enterprises. The ternary view of the edge-incremental graph represents entity information and equity relationships between entities.
[0102] Please see Figure 6 The equity graph data construction process proposed in this invention is as follows: For existing entity point data and equity edge data in the text, the graph data construction module reads them in to form corresponding point data variables and edge data variables. When reading entity point data, each line of the point data is checked. If the point data format is correct, a point data variable is formed; if the point data format is incorrect, usable data is filtered to form point data variables. When reading equity edge data, each line of the edge data is checked. If the edge data format is correct, an edge data variable is formed. If the edge data format is incorrect, default data that does not affect subsequent calculation results is filled in to form an edge data variable. Equity graph data variables are then constructed based on the point data variables and edge data variables.
[0103] Please see Figure 7 The equity subgraph data segmentation method proposed in this invention can extract connected components from complete equity graph data. For a set of connected components, it is grouped according to the number of edge data contained in the set, and then sorted according to the number of edges, outputting the sorting result. Based on the sorting result and the number of edges, it is determined how to merge the groups to form multiple equity subgraph data.
[0104] This invention employs a subgraph segmentation equity penetration method and apparatus based on equity graph data. The method reduces computation time, data transmission overhead, and computational resource consumption during multi-hop equity penetration calculations, thereby simplifying the equity penetration process and improving penetration calculation efficiency. The apparatus can be deployed in a single-machine environment for simple, small-scale equity penetration calculations, or in a distributed environment for multi-hop equity mining of massive datasets. It exhibits good performance in both single-machine and cluster environments.
[0105] It should be understood that any parts not described in detail in this specification belong to the prior art.
[0106] It should be understood that the above description of the preferred embodiments is quite detailed, but this should not be construed as limiting the scope of protection of this invention. It is neither necessary nor possible to exhaustively describe all possible implementations. Those skilled in the art, guided by this invention, can make substitutions or modifications without departing from the scope of the claims, all of which fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A method for subgraph segmentation and equity penetration based on equity graph data, characterized in that, Includes the following steps: Step 1: Read the enterprise's business registration data, store the enterprise's business registration data into a flexible distributed data structure, obtain the initial vertex data and initial edge data, and construct the initial equity graph data; Step 2: Extract connected components from the initial equity graph data in Step 1, and group, sort, split, and combine the connected components to divide multiple connected components into multiple subgraphs; Step 3: Traverse all ternary views under each subgraph in Step 2. Based on the edges recorded in the ternary views and their corresponding enterprise nodes and shareholder nodes, aggregate equity information to generate enterprise node data containing direct shareholder equity information. Nodes without direct shareholders have their equity information initialized to empty. Equity information includes shareholder node ID, shareholder node name, shareholder node type, shareholder shareholding ratio, shareholder shareholding hop count, shareholder shareholding path, and other information. The generated enterprise node data containing direct shareholder equity information is formatted as (entity id, (entity name, entity type, entity information, entity shareholder array)). The entity shareholder array is formatted as [(shareholder id, shareholder name, shareholder type, shareholder information, shareholder shareholding ratio, hop count distance, shareholder shareholding path)]. The intermediate nodes for equity relationship transmission are stored in the shareholder shareholding path, formatted as [(path entity id, path shareholding ratio)]. Step 4: Construct graph data containing one-hop relationships using enterprise node data containing direct shareholder node equity information and initial edge data, where each node information carries one-hop equity information; Step 5: For each ternary view in the graph data of the one-hop relationship constructed in Step 4, obtain the shareholder node based on the current aggregated hop count information, determine whether the shareholder node has equity information, and if so, update the equity information of the shareholder node to the enterprise node. If the shareholder node does not have equity information, the equity information recorded in the ternary view will be directly written into the enterprise node. Step 6: Deduplicate the equity information set and combine the processed equity node data with the equity relationship edge data to form a new equity graph data; Step 7: Repeat steps 4 to 6, updating the equity information carried by the node data in each repetition.
2. The subgraph segmentation and equity penetration method based on equity graph data according to claim 1, characterized in that, In step 1, the enterprise business registration data is cleaned and preprocessed and then stored in the form of a node dataset and an edge dataset. The vertex dataset contains vertex ID, entity name, entity type and other information, and the edge dataset contains source node ID, target node ID and equity information. The graph data is composed of vertex data and edge data connected together.
3. The subgraph segmentation and equity penetration method based on equity graph data according to claim 1, characterized in that, In step 2, the number of edges contained in each connected component in the connected component set is counted, the connected components in the connected component set are sorted in ascending order according to the number of edges contained, and the sorting results are then divided and combined to form multiple subgraphs.
4. The subgraph segmentation and equity penetration method based on equity graph data according to claim 1, characterized in that, When steps 4 through 6 are not repeated, the nodes in the obtained graph data have one-hop equity information. When steps 4 through 6 are repeated once, the obtained entity nodes have two-hop equity information, i.e., the company's second-tier shareholder information. n In steps 4 through 6, the obtained entity nodes possess ( n +1) Jump to equity information, that is, the company's ( n +1) Level shareholder information.
5. The subgraph segmentation and equity penetration method based on equity graph data according to claim 1, characterized in that, In step 7, steps 4 to 6 are repeated, i.e., the equity penetration method is used to perform multi-hop equity information penetration. In the calculation of each hop, the attributes of the entity node are updated.
6. A subgraph segmentation and equity penetration device based on equity graph data, characterized in that, It includes a data reading module, a graph data construction module, a subgraph segmentation module, an equity information aggregation module, an equity information deduplication module, a loop calculation module, and an equity structure extraction module; The data reading module is used to read in the cleaned and preprocessed enterprise business registration data, and generate node data and edge data for constructing graph data respectively; the node data must include enterprise nodes and shareholder nodes, enterprise nodes include sole proprietorships, limited liability companies or joint-stock companies, and shareholder nodes include natural persons and others; The graph data construction module is used to construct a graph representing equity relationships using node data and edge data. When the input is initial node data and initial edge set data, the graph constructed by this module represents a one-hop equity structure. When the input is node data with equity information and equity relationship edge data, the graph constructed by this module contains equity graph data with a certain number of hops in the equity relationship. The subgraph segmentation module is used to statistically analyze and segment connected components based on equity graph data, dividing multiple connected components into subgraphs, and ensuring that the computational load of each subgraph is as similar as possible. This transforms the equity penetration operation from the original complete equity graph data into equity penetration of multiple subgraphs separately, allowing the device to complete equity penetration calculations with more hops in the large financial equity graph with a small amount of resources. Furthermore, when the subgraphs are properly segmented, it can achieve higher efficiency than the full graph equity penetration method. The equity information aggregation module is used to traverse all ternary view relationships in the subgraph data, analyze node types, and send the equity relationships between nodes to the nodes in the form of messages to update the equity information in the node data. The equity information deduplication module is used to deduplicat equity information in node data and output new vertex data with equity information for construction. The loop calculation module is used to make the graph data construction module, the equity information aggregation module, and the equity information deduplication module perform repeated operations. The number of repetitions is the number of hops for equity penetration, and finally an equity data graph with a specified number of hops is obtained, whose node data carries equity information under the specified number of hops. The equity structure extraction module is used to extract the equity structure information related to the entity from a single point data carrying equity information; The subgraph segmentation equity penetration device based on equity graph data is used to perform the steps in the subgraph segmentation equity penetration method based on equity graph data according to any one of claims 1-5.
Citation Information
Patent Citations
Abnormal account identification method, data scheduling platform and graph computing platform
CN117216736A
Distributed system generating rule compiler engine apparatuses, methods, systems and media
US20200293916A1