Graph data storage and processing optimization method for fund transaction analysis
By introducing the E-finance Chunk block diagram representation model into the graph data processing system, the problem of inefficiency of the existing system is solved and efficient graph data processing and analysis is realized, especially in the review and analysis of the fund trading network, the efficiency is improved by 4 times.
Patent Information
- Application Number
- CN202510042452.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-27
AI Technical Summary
The existing graph data processing system relies on complex data cleaning and graph traversal operation calculation models, resulting in inefficiency of related data processing applications.
A representation model based on E-finance Chunk block diagram is adopted, which consists of classified and hierarchical vertex storage components and block layout optimization components. Vertex storage components optimize vertex storage through metadata storage and block storage, and block layout optimization components optimize block layout through sorting and two-level BFS traversal.
It significantly improves the I/O efficiency of large-scale graph data, and the review and processing efficiency of massive transaction data is increased by 4 times, allowing judicial personnel to flexibly allocate review data and mine the criminal activities behind fund transactions.
Smart Images

Figure CN120045130A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph data processing, and specifically relates to a method for optimizing graph data storage and processing for fund transaction analysis. Background Art
[0002] In financial crimes, the scale of fund transaction data is huge and the types are complex, which leads to problems such as low efficiency in reviewing account natures and the roles of entities during case handling. In response to these problems, judicial organs and relevant scientific research institutions use graph models of the fund transaction behaviors of involved entities to represent the relationship characteristics between transaction data and conduct large-scale analysis of the rules of transaction data. However, existing graph data processing systems rely on complex data cleaning and graph traversal operation calculation models, which also result in the inefficiency of relevant data processing application programs. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] The technical problem to be solved by the present invention is how to provide a method for optimizing graph data storage and processing for fund transaction analysis to solve the inefficiency problem of relevant data processing application programs caused by the dependence of existing graph data processing systems on complex data cleaning and graph traversal operation calculation models.
[0005] (2) Technical Solutions
[0006] To solve the above technical problems, the present invention proposes a method for optimizing graph data storage and processing for fund transaction analysis. This method adopts a representation model based on the E-finance Chunk graph, which consists of two key components: one is a classified and hierarchical vertex storage component, and the other is a block layout optimization component;
[0007] Component 1: Classified and Hierarchical Vertex Storage Component
[0008] S11. The metadata of all vertices is stored in memory. The metadata of vertices refers to the relevant attribute information of each vertex itself, and the metadata includes: the unique identifier of the vertex, the degree of the vertex, the classification label of the vertex, and other business information related to the vertex; the degree of the vertex represents the number of connections between this vertex and other vertices; the classification label of the vertex refers to classifying vertices according to the degree, including: Min node, Med node, Mas node; Let d i represent the degree of vertex v i , the vertex with degree d i ≤k 1 is a Min node, the vertex with degree k 1 <d i ≤k 2 is a Med node, and the vertex with degree d i >k2 The vertices of 2 are Mas nodes;
[0009] S12. For Min nodes, directly store the vertex index in an 8-byte space, which is used to store the metadata of the vertex, including information about its connections to other nodes;
[0010] S13. For Med nodes, organize them into aligned blocks of appropriate size, divide the storage space into multiple segments, each segment stores a series of blocks of the same size, and each block consists of the adjacency lists of multiple vertices; if the remaining space in the current block is too small to accommodate all the adjacent points of the next vertex, keep this blank space and store the entire adjacency list of the next vertex in the next block;
[0011] S14. For Mas vertices, use the direct Huge Page technology for management, where an 8-byte vertex index is used to store the pointer to the adjacency list of these Mas vertices;
[0012] S15. After the classification processing, perform hierarchical storage, placing vertices of different degrees in blocks of different sizes;
[0013] S16. When performing a query, if querying a Med vertex υ, first obtain the block level according to its edge degree, then obtain the corresponding block according to the block ID cid of this level, and then obtain the adjacency list of this υ from the block offset coff
[0014] υ = chunk_start_addr + cid × chunk_size + coff;
[0015] S17. The allocation strategy for the block buffer area is to divide the buffer space into four parts in proportion according to the size of each layer of block files as the corresponding block buffers;
[0016] Component Two. Block Layout Optimization Component
[0017] S21. First, sort the vertices, sort all vertices in descending order of their in-degree to generate a list P;
[0018] S22. Perform a two-level BFS traversal, select a vertex from the head of the list P as the root, perform a two-level breadth-first search BFS, and then add it to the list Q, where the list P is generated after sorting by the in-degree of the vertices;
[0019] S23. The system checks the number of vertices in the list Q. If the number of vertices in Q is less than the preset threshold R, continue to select a new root vertex from the list P and repeat the BFS traversal until the number of vertices in Q reaches or exceeds R;
[0020] S24. Place consecutive vertices in the new order into the same chunk.
[0021] (III) Advantageous Effects
[0022] The present invention proposes a method for optimizing the storage and processing of graph data for financial transaction analysis. The present invention proposes a block graph processing model called E-finance Chunk, which is a processing system for improving the I / O efficiency of large-scale graph data on SSD hard disks. E-finance Chunk introduces a new block-based graph representation model with classified and hierarchical vertex storage and efficient block layout optimization features. This model is used to optimize the analysis of fund transaction network graphs for money laundering and the crime of illegally absorbing public deposits. The processing efficiency of this model for reviewing and processing massive transaction data is 4 times higher than that of external graph systems and in-memory graph systems. The application of this model enables judicial personnel to flexibly configure the reviewed data and uncover criminal activities behind transactions. Description of the Drawings
[0023] Figure 1 It is the schematic diagram of the method for optimizing the storage and processing of graph data for financial transaction analysis of the present invention;
[0024] Figure 2 It is the schematic diagram of the E-finance Chunk block graph processing model system of the present invention;
[0025] Figure 3 It is the effect diagram of fund loop analysis. Detailed Embodiments
[0026] To make the objectives, content, and advantages of the present invention clearer, the following further describes the detailed embodiments of the present invention in conjunction with the drawings and embodiments.
[0027] To solve the problem of low efficiency of related data processing application programs caused by the existing graph data processing system relying on complex data cleaning and graph traversal operation calculation models. The present invention introduces a block graph processing model called E-finance Chunk, which is a processing system designed to improve the I / O efficiency of large-scale graph data on SSD hard disks. E-finance Chunk introduces a new block-based graph representation model with classified and hierarchical vertex storage and efficient block layout optimization features. After laboratory tests, the E-finance Chunk block graph is superior to existing external graph systems and in-memory graph systems that rely on general cache systems in large-scale fund transaction network analysis.
[0028] To solve the above technical problems, the present invention proposes a method for optimizing the storage and processing of graph data for capital transaction analysis. This method adopts a representation model based on the E-finance Chunk graph, which focuses on the effectiveness of system I / O of the chunk graph and user-friendliness, and consists of two key components: one is a classified and hierarchical vertex storage component, and the other is a chunk layout optimization component. The present invention is implemented through the following technical solutions:
[0029] Component 1, Classified and Hierarchical Vertex Storage Component
[0030] S11. The metadata of all vertices is stored in memory. The metadata of vertices refers to the relevant attribute information of each vertex (such as a capital account, lending relationship). These metadata include: the unique identifier of the vertex, the degree of the vertex, the classification label of the vertex, and other business information related to the vertex.
[0031] 1. The unique identifier of the vertex: Usually it is an account number or a transaction number.
[0032] 2. The degree of the vertex: It represents the number of connections between this vertex and other vertices.
[0033] This degree can be further divided into: in-degree: the number of edges pointing to this vertex from other vertices, representing the number of transactions in which the account receives funds. Out-degree: the number of edges pointing from this vertex to other vertices, representing the number of times the account initiates capital transactions.
[0034] 3. The classification label of the vertex: Classify the vertices according to the degree, such as Min nodes, Med nodes, Mas nodes (see the explanation later).
[0035] 4. Other business information related to the vertex: Such as transaction type, account status, etc.
[0036] After these vertex metadata are stored in memory, the vertices are classified. The classification rules are: (1) Vertices with a degree of 1 or 2 are classified as Min nodes; (2) Vertices with a degree range of [3, d θ are classified as Med nodes; (3) Vertices with a degree greater than d θ are classified as Mas vertices. Where d θ is used to set and adjust the size of Med vertices.
[0037] S12. Different data structures are used for different vertex storages. For Min nodes, the present invention directly stores the vertex index in an 8-byte space, and these 8 bytes are used to store the metadata of the vertex, including the information of its connection with other nodes. This design optimizes the space efficiency by reducing the pointer consumption in the traditional adjacency list. The experimental results show that this block graph storage method is superior to the traditional graph storage structure in terms of I / O performance, especially when dealing with large-scale fund transaction networks, showing better efficiency.
[0038] The 8-byte metadata field is directly used to store the indices of the vertices they are connected to. In this model, the size of the vertex index is 4 bytes. The first 4 bytes in the 8-byte space of each Min node are used to store the index of the vertex itself, and the last 4 bytes are used to store the indices of its adjacent vertices. This step optimizes the traditional way of pointing to the adjacency list with pointers.
[0039] S13. For nodes classified as Med with the vertex degree range of [3, d θ , the present invention organizes them into aligned blocks of appropriate sizes, divides the storage space into multiple segments, each segment stores a series of blocks of the same size, and each block consists of the adjacency lists of multiple vertices. (The stored adjacency lists are the vertices connected to the current vertex, that is, the accounts that have transacted with this vertex. In the graph data structure, the adjacency list is used to represent all the vertices connected to a vertex, that is, the transaction objects of this account in the fund transaction network). If the remaining space in the current block is too small to accommodate all the adjacent points of the next vertex, the present invention retains this blank space and stores the entire adjacency list of the next vertex in the next block.
[0040] Each entry in the adjacency list corresponds to the vertex number (index) of an adjacent vertex. This is because vertices are referenced by their unique identifiers (i.e., indices) when stored. That is to say, the adjacency list directly reflects which other vertices the current vertex is connected to and identifies the targets of these connections through the indices of the vertices.
[0041] S14. For Mas vertices with vertex degree greater than d θ , the present invention uses the direct Huge Page (DHP) technology to manage them to reduce the TLB miss rate. The present invention uses an 8-byte vertex index to store the pointer to the adjacency list of these Mas vertices for querying them;
[0042] S15. After the classification process is completed, hierarchical storage is performed. To accommodate different vertices of different sizes, the present invention places vertices of different levels in blocks of different sizes. The block sizes of each layer are 4KB, 32KB, 256KB, and 2MB respectively. If the size of the adjacency list of a vertex is less than the block size, the vertex is stored in a lower-level block. And the present invention incorporates an adaptive hash hierarchical storage algorithm, which dynamically maps data to different storage layers. Through a hash function and a load adjustment mechanism, data is distributed among different storage layers. When the data volume of a certain layer exceeds the load threshold, the data will adaptively migrate to a more suitable storage layer. This method can achieve data balance and dynamic adjustment in distributed storage systems and large-scale graph data systems.
[0043] S16. When performing a query, if querying a Med vertex υ, first obtain its block level according to its edge degree (i.e., the number of connections between this vertex and other vertices), and then obtain the corresponding block according to the block ID cid of this level (cid, i.e., chunk ID, is the block ID, indicating the block number where the adjacency list of the current vertex is stored, and each block has a unique block ID). Then we can obtain the adjacency list of this υ from its block offset coff (i.e., the block offset, indicating the specific position where the adjacency list of this vertex is stored in a certain block).
[0044] υ = chunk_start_addr + cid × chunk_size + coff;
[0045] S17. The allocation strategy for the block buffer area is to divide the buffer space into four parts proportionally according to the size of each layer's block file, and use them as the corresponding block buffers. For example, there are M bytes of memory for all block buffers, and the block file sizes of each layer are S0, S1, S2, and S3 respectively. The memory allocated to the block buffer of the S i layer's block buffer is
[0046] Component 2. Block Layout Optimization Component
[0047] S21. First, sort the vertices. Sort all vertices in descending order of their in-degree to generate a list P. Vertices with higher in-degree usually have a higher probability of accessing their out-neighbors, so they are ranked at the front of the list. The purpose of sorting is to preferentially store frequently accessed vertices in adjacent storage locations (in the same block), which can reduce the number of I / O operations during querying and improve query efficiency.
[0048] S22. Perform a two-level BFS traversal. Select a vertex from the head of list P as the root, perform a two-level breadth-first search (BFS), and then add it to list Q, where list P is generated by sorting the vertices according to their in-degrees. By sorting in descending order of in-degree, the system preferentially processes those vertices with greater influence. Vertices with high in-degrees (such as key accounts in money laundering activities) are often the focus of the analysis.
[0049] This search method can obtain the vertices at the same level as the root vertex. Add the vertices traversed by BFS to another list Q, and at the same time remove these vertices from list P. The process is as follows:
[0050] S221. Select the root vertex: Select a vertex from the head of list P as the root vertex (the vertex with the highest in-degree).
[0051] S222. First-level BFS traversal: Starting from this root vertex, use BFS to traverse all its directly connected vertices, that is, the first-level neighbors (i.e., the vertices that have direct transaction relationships with the root vertex).
[0052] S223. Second-level BFS traversal: Continue to traverse the direct neighbors of each first-level neighbor, that is, the second-level neighbors (these vertices are the transaction counterparts of the first-level neighbors but have no direct transaction relationships with the root vertex).
[0053] S224. Add to list Q: All vertices found in the first-level and second-level BFS traversals will be added to a new list Q, and at the same time these vertices will be removed from P.
[0054] The purpose of this operation is that the two-level BFS aims to gather the vertices with stronger associations with the root vertex as much as possible, forming a vertex set that includes the root vertex, first-level neighbors, and second-level neighbors. Since these vertices are associated with each other in the transaction network, they are also likely to be accessed together during actual queries. Through BFS, the system can gather them together and reduce random access during subsequent queries.
[0055] S23. The system checks the number of vertices in list Q. If the number of vertices in Q is less than the preset threshold R (R is 95% of the total number of vertices), then continue to select a new root vertex from list P and repeat the BFS traversal until the number of vertices in Q reaches or exceeds R. Setting the threshold R is to ensure that after each traversal, the generated vertex set (list Q) will not be too large or too small. A list Q of an appropriate size helps to reasonably allocate vertices to the same block, while avoiding waste of space or a decrease in storage efficiency due to too much or too little data in the block.
[0056] S24. Place consecutive vertices in the new order into the same chunk. If these vertices are stored in different chunks, when querying these vertices, the system needs to frequently switch between multiple chunks, increasing the number of random I / O operations and reducing the system performance. By storing them in the same chunk, the overhead of chunk switching can be reduced. By default, the threshold R is set to 95% × |V|, where |V| is the total number of vertices. This means that most vertices will be reordered to ensure that the chunk partitioning of the graph can cover most vertices.
[0057] Through the vertex sorting of S21 - S24 and the two - level BFS traversal, the system can preferentially store vertices with strong correlation and frequent access in the same chunk. This layout optimization process reduces the random access during query, enabling most queries to be completed in a few chunks, significantly improving the locality of data access and the overall I / O efficiency of the system. These steps together optimize the chunk layout to ensure more efficient query operations.
[0058] Example 1:
[0059] In recent years, financial crimes have shown a continuous high - incidence trend, and their complexity has made it difficult for judicial organs to recover illegal funds. Due to the large number of people involved and the huge amount of funds involved in related cases, the electronic data of funds involved in the cases often has a large scale, is multi - source heterogeneous, has complex fund flows, and hidden fund transaction behaviors. Facing the massive electronic data of funds, how to conduct fast and flexible analysis and establish a review and analysis model is a common difficulty in the work of judicial organs to recover stolen money and losses:
[0060] To solve the above - mentioned technical problems, the present invention proposes a chunk graph processing model called E - finance Chunk to improve the I / O efficiency of large - scale graph data. The following will specifically describe each step in detail with reference to specific embodiments.
[0061] S1. Data pre - processing. Perform data cleaning and data standardization before data entry for the electronic evidence data of funds provided by financial institutions, and process different types of transaction amounts and time formats. Suppose the transaction records T = {t 1 ,t 2 ,…,t n} of a money - laundering case, where each t i represents a transaction, including transaction amount A i , transaction time τ i , and information such as both parties to the transaction (such as account A and account B). Its processing steps include:
[0062] Abnormal data cleaning: Remove or correct outliers in the data, such as negative values or non - numerical - formatted amounts.
[0063] Time format standardization: Convert all transaction times τ i to a unified format (such as UNIX timestamp).
[0064] Data merging: Merge multiple transactions involving the same account into one vertex, and calculate the out-degree (i.e., the number of transactions) and in-degree (i.e., the number of receptions) of the vertex.
[0065] S2. Classification of graph data vertices. Classify according to the degree of each vertex (i.e., account). Let d i represent the degree of vertex v i (i.e., the number of connected edges), and the classification rules are as follows:
[0066] Min node (minimum node): Vertices with degree d i ≤k 1 , usually low-frequency trading accounts.
[0067] Med node (middle node): Vertices with degree k 1 <d i ≤k 2 , representing accounts with moderate trading frequency.
[0068] Mas node (maximum node): Vertices with degree d i >k 2 , and these vertices usually represent high-frequency trading accounts, which may involve a large amount of capital flow.
[0069] S3. Optimization of vertex storage structure. Select a suitable storage structure according to the characteristics of different types of vertices. For example, assume that a vertex v1 represents an account A, and its adjacency list stores the accounts that have transactions with account A (such as account B and account C). If the vertex indices of account B and account C are 100 and 105 respectively, then the adjacency list of v1 will contain these two indices, so that the specific information of these accounts can be quickly found during data processing. In the present invention, we select 3 types of nodes for storage structure selection:
[0070] Min nodes: Since the degrees of these nodes are relatively low, we directly store their neighbor nodes in the 8-byte vertex metadata field. For example, if v 1 is a Min node and its degree is 2, then its neighbor information (nb 0 , nb 1 ) can be directly stored in the metadata field.
[0071] Med nodes: For Med nodes with higher degrees, use the block ID (cid) and block offset (coff) in the chunk buffer to locate their neighbor lists. The neighbor information can be obtained through the formula:
[0072] The neighbor address is calculated and accessed as neighbor address = chunk_start_addr + cid × chunk_size + coff.
[0073] Mas nodes: For super vertex Mas nodes, the v_foff field (super vertex offset) is used to store and access a large amount of neighbor information of these nodes. The storage structure of such nodes needs to be specially designed to optimize access.
[0074] S4. Optimized layout of tiles. Optimize the storage of vertices by sorting and tiling. Assume that we sort the vertices by their in-degree d in Sorting, arranging all vertices in descending order to obtain a list P = {v 1 , v 2 , …, v n}. Then, perform the following operations on the vertices in the list:
[0075] Select a vertex as the root, perform a two-layer breadth-first search (BFS) on the electronic data of funds in financial cases, add the traversed vertices to the list Q, and remove these vertices from P at the same time.
[0076] If the number of vertices in Q is less than the preset threshold R, continue to select a new root node for BFS until the number of vertices in Q reaches or exceeds R.
[0077] Allocate the sorted vertices to each chunk in order to improve the locality and efficiency of data access.
[0078] S5. Storage allocation strategy. Allocate memory resources according to the sizes S i of chunk files at different levels. Let the total memory be M, then the allocated memory for each layer of chunk buffer L i is:
[0079]
[0080] Among them, S 0 , S 1 , S 2 , S 3 are the sizes of chunk files at different levels respectively, ensuring reasonable memory allocation between different levels to optimize storage performance. As the data scale and access frequency change, monitor the load conditions of each storage layer. When the load of a certain layer exceeds the set threshold, the data adaptively migrates to a higher-level storage layer. When the load decreases, the data can also be dynamically downgraded to a lower level to maintain the balance of storage and access.
[0081] S6: Performance Optimization and Evaluation. Improve the overall system performance by reducing the number of I / O operations and optimizing CPU cache performance. For small nodes, read neighbor information by directly accessing metadata fields to avoid additional I / O operations. Use the following cache hit rate formula to evaluate the optimization effect:
[0082]
[0083] Observe the change of cache hit rate through experiments and further adjust system parameters to achieve the best performance.
[0084] S7: Result Analysis and Feedback. After completing all processing, use a mathematical model to analyze the clustering and liquidity of transaction data. Evaluate the analysis effect of the system by calculating indicators such as graph connectivity and degree distribution, and make adjustments according to the results. For example, calculate the standard deviation σ of the vertex degrees d to evaluate the uniformity of the transaction network:
[0085] where σ d is the average degree of all vertices, that is, the average number of transaction connections of all accounts; n is the number of vertices in the graph, that is, the total number of accounts in the network; d i is the degree of the i-th vertex, representing the number of transaction connections (i.e., the number of edges) of account i; refers to the average total degree of all vertices in the graph, and this average degree reflects the average level of the connection situation of vertices in the whole graph.
[0086] The purpose of this formula is to evaluate the uniformity of the entire transaction network by calculating the dispersion degree of vertex degrees in the graph. The standard deviation σ d reflects whether the transaction behaviors of each account in the network are evenly distributed. If the standard deviation is large, it means that the transaction frequencies of some accounts (vertices) are significantly higher than those of other accounts, and there may be a few high-frequency accounts dominating most transactions; if the standard deviation is small, it indicates that the transaction frequencies of all accounts are relatively balanced.
[0087] The present invention evaluates the uniformity in the following ways:
[0088] (1) If the standard deviation of the transaction network is small (σ d close to 0), it means that the transaction behaviors of most accounts are relatively uniform, and the transaction frequencies of all accounts are close. Such a network is relatively uniform.
[0089] (2) If the standard deviation is large, it indicates that there are significant differences in the network, and some accounts may participate in a large number of transactions while the transaction frequencies of other accounts are low. Such a network often has super nodes (high-frequency trading accounts), and these accounts play an important role in the network and may be key accounts for illegal activities such as money laundering.
[0090] By calculating the standard deviation of vertex degrees, we can analyze whether there are abnormally high-frequency accounts in the transaction network, as well as the distribution of transaction behaviors in the entire network. This is of great significance for the analysis of financial crimes.
[0091] Finally, according to the analysis results of S1-S7, adjust the classification threshold k 1 ,k 2 And parameters such as chunk size to further optimize system performance.
[0092] Beneficial Effects
[0093] The present invention proposes a block graph processing model called E-finance Chunk, which is a processing system for improving the I / O efficiency of large-scale graph data on SSD hard disks. E-finance Chunk introduces a new block-based graph representation model with classified and hierarchical vertex storage and efficient block layout optimization features. This model is used to analyze and optimize the financial transaction network graph for money laundering and illegal absorption of public deposits. The model is 4 times more efficient in reviewing and processing massive transaction data than external graph systems and memory graph systems. The application of this model allows judicial personnel to flexibly configure review data and dig out criminal activities behind transactions.
[0094] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A graph data storage and processing optimization method for fund transaction analysis, characterized in that: This method adopts a representation model based on the E-finance Chunk block graph, which consists of two key components: one is the classified and hierarchical vertex storage component, and the other is the block layout optimization component; Component 1: Classified and hierarchical vertex storage components S11. The metadata of all vertices is stored in the memory. The metadata of the vertex refers to the relevant attribute information of each vertex itself. The metadata includes: the unique identifier of the vertex, the edge degree of the vertex, the classification label of the vertex and other business information related to the vertex; the edge degree of the vertex indicates the number of connections between the vertex and other vertices; the classification label of the vertex refers to the classification of the vertex according to the degree, including: Min node, Med node, Mas node; let d i Represents vertex v i The degree of i Vertices ≤ k1 are Min nodes, with degree k1 <d i Vertices ≤ k2 are Med nodes with degree d i >The vertex of k2 is a Mas node; S12. For the Min node, the vertex index is directly stored in an 8-byte space. The 8 bytes are used to store the metadata of the vertex, including the information about its connection with other nodes. S13. For the nodes of Med, organize them into aligned blocks of appropriate size, divide the storage space into multiple segments, each segment stores a series of blocks of the same size, and each block consists of the adjacency lists of multiple vertices; if the remaining space of the current block is too small to accommodate all the adjacent points of the next vertex, then reserve this empty space and store the entire adjacency list of the next vertex in the next block; S14. Use direct Huge Page technology to manage the Mas vertices, where an 8-byte vertex index is used to store pointers to the adjacency lists of these Mas vertices; S15, after the classification process is completed, hierarchical storage is performed, and vertices of different degrees are placed in blocks of different sizes; S16. When querying, if you query a Med vertex υ, first get the block level according to its edge degree, then get the corresponding block according to the block ID cid of the level, and then get the adjacency list of this υ from the block offset coff υ=chunk_start_addr+cid×chunk_size+coff; S17, the allocation strategy for the block cache area is to divide the buffer space into four parts in proportion according to the size of each layer of block files as corresponding block buffers; Component 2: Block Layout Optimization Component S21, first sort the vertices, sort all the vertices in descending order of their in-degree, and generate a list P; S22, perform a two-level BFS traversal, select a vertex from the head of list P as the root, perform a two-level breadth-first search BFS, and then add it to list Q, where list P is generated by sorting the in-degree of the vertices; S23, the system checks the number of vertices in the list Q; if the number of vertices in Q is less than a preset threshold R, it continues to select a new root vertex from the list P and repeats the BFS traversal until the number of vertices in Q reaches or exceeds R; S24. Place consecutive vertices in the new order into the same chunk.
2. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In S11, the in-degree indicates the number of edges from other vertices to the vertex, representing the number of transactions in which the account receives funds; the out-degree indicates the number of edges from the vertex to other vertices, representing the number of fund transactions initiated by the account; vertices with an edge degree of 1 or 2 are classified as Min nodes, and the edge degree range of the vertex is [3, d θ ] is classified as a Med node, and its edge degree is greater than d θ The edge degree of the vertex is classified as Mas vertex, where d θ Used to set the size of Med vertices.
3. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In the S12, the 8-byte metadata field is directly used to store the indexes of the vertices connected to them. In this model, the size of the vertex index is 4 bytes. The first 4 bytes of the 8-byte space of each Min node are used to store the vertex's own index, and the last 4 bytes are used to store the indexes of its adjacent vertices.
4. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In S15, if the size of the adjacency list of a vertex is smaller than the block size, the vertex is stored in a low-level block; and an adaptive hash tiered storage algorithm is integrated, which dynamically maps data to different storage layers, and distributes data between different storage layers through hash functions and load adjustment mechanisms; when the amount of data in a certain layer exceeds the load threshold, the data will be adaptively migrated to a more suitable storage layer.
5. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In S15, the block size of each layer is 4KB, 32KB, 256KB and 2MB respectively.
6. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In S17, there are M bytes of memory for all block buffers, and the block file sizes of each layer are S0, S1, S2 and S3 respectively, which are allocated to the Sth layer. i The memory of the layer's block buffer is 7. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: The S22 includes: S221, select the root vertex: select a vertex from the head of the list P as the root vertex, that is, the vertex with the highest in-degree; S222, first-level BFS traversal: Starting from this root vertex, use BFS to traverse all its directly connected vertices, that is, first-level neighbors, that is, vertices that have a direct transaction relationship with the root vertex; S223, Secondary BFS traversal: Continue to traverse the direct neighbors of each first-level neighbor, that is, the second-level neighbors. These vertices are the transaction objects of the first-level neighbors, but have no direct transaction relationship with the root vertex; S224. Add to list Q: All vertices found in the first-level and second-level BFS traversals are added to a new list Q, and these vertices are removed from P.
8. The graph data storage and processing optimization method for fund transaction analysis according to claim 1, characterized in that: In S23, R is 95% of the total number of vertices.
9. The method for optimizing graph data storage and processing for fund transaction analysis according to any one of claims 1 to 8, characterized in that: The method also includes: performance optimization and evaluation; Improve overall system performance by reducing the number of I / O operations and optimizing CPU cache performance; For small nodes, neighbor information is read by directly accessing the metadata field to avoid additional I / O operations. The following cache hit rate formula is used to evaluate the optimization effect: Through experiments, the changes in cache hit rate are observed and system parameters are further adjusted to achieve optimal performance.
10. The graph data storage and processing optimization method for fund transaction analysis according to claim 9, characterized in that: The method also includes: result analysis and feedback; After all processing is completed, mathematical models are used to analyze the aggregation and liquidity of transaction data; the system's analysis effect is evaluated by calculating the connectivity and degree distribution indicators of the graph, and adjustments are made based on the results; Calculate the standard deviation σ of vertex degree d To evaluate the uniformity of the transaction network: Among them, σ d The average degree of all vertices, that is, the average number of transaction connections among all accounts; n is the number of vertices in the graph, that is, the total number of accounts in the network; d i is the degree of the i-th vertex, indicating the number of transaction connections of account i; It refers to the average total degree of all vertices in the graph. This average degree reflects the average level of connectivity of the vertices in the entire graph. Standard deviation d It reflects whether the transaction behaviors of each account in the network are evenly distributed. If the standard deviation is large, it means that the transaction frequency of some accounts is significantly higher than that of other accounts, and a few high-frequency accounts may dominate most transactions; if the standard deviation is small, it means that the transaction frequency of all accounts is relatively balanced; Adjust the classification thresholds k1, k2 and chunk size parameters to further optimize system performance.