Distributed financial data processing system

Through the distributed financial data processing system, the problems of low efficiency and poor accuracy in data collection, integration, verification and resource scheduling of traditional systems are solved, and efficient, safe and compliant financial transaction processing is achieved.

CN120494991AInactive Publication Date: 2025-08-15FANGSHENG SCIENCE & TECHNOLOGY IND PARK MANAGEMENT (NANJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510586073.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional financial data processing systems have problems such as low efficiency, poor accuracy and insufficient security in data collection, integration, computing verification, business flow modeling and resource scheduling, especially in a multi-source heterogeneous data environment, which is difficult to meet the real-time and compliance needs of financial transactions.

Method used

The distributed financial data processing system is adopted, and the multi-source data acquisition module uses distributed edge computing nodes and dynamic load balancing algorithms. The heterogeneous data integration module performs cross-platform format analysis. The distributed computing verification module builds a trusted verification model based on a hybrid consensus mechanism. The business flow modeling module uses spatiotemporal causal convolution network to model the timing dependency of capital flow. The resource scheduling module dynamically adjusts resource allocation through a multi-objective optimization algorithm.

Benefits of technology

It improves data collection efficiency and accuracy, ensures data consistency and security, enhances risk prediction capabilities and resource utilization efficiency, and ensures the legality and compliance of financial transactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494991A_ABST
    Figure CN120494991A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed data processing, and discloses a distributed financial data processing system. The system comprises a multi-source data acquisition module, a heterogeneous data integration module, a distributed calculation verification module, a service flow modeling module, a resource scheduling module and the like. The multi-source data acquisition module captures multi-dimensional financial transaction data in real time by using a distributed edge computing node, and distributes an acquisition path through a dynamic load balancing algorithm; the heterogeneous data integration module analyzes the cross-platform data format and generates a standardized data stream; the distributed calculation verification module constructs a credible verification model based on a hybrid consensus mechanism; the business flow modeling module predicts a risk conduction path; and the resource scheduling module optimizes resource allocation. The system effectively solves the problems of a traditional financial data processing system in the aspects of data acquisition, integration, calculation verification and the like, improves the data processing efficiency, the risk prediction capability and the resource utilization efficiency, and guarantees the safety and compliance of financial transactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed data processing, and in particular to a distributed financial data processing system. Background Art

[0002] With the booming financial industry, the volume of financial data is exploding, with increasingly diverse sources and complex formats. Traditional centralized financial data processing systems are no longer able to meet actual business needs and are exhibiting numerous drawbacks.

[0003] From a data collection perspective, traditional systems rely on a small number of fixed nodes to collect data. This makes data loss and delayed collection highly susceptible to massive amounts of multi-source financial transaction data. In the securities market, for example, where numerous transactions occur every second, centralized collection nodes are unable to quickly capture all transaction information, leading to the omission of critical data and compromising the accuracy of subsequent analysis. Furthermore, this collection approach lacks consideration for the load on various data sources, often causing some nodes to process data slowly due to excessive volumes, thus impacting overall data collection efficiency.

[0004] When it comes to data integration, data formats vary significantly across financial platforms. Data formats in banking systems may be based on specific financial standards, while data formats on internet financial platforms may prioritize user convenience. These differences are significant. Traditional systems lack effective cross-platform format parsing capabilities, making it difficult to convert this heterogeneous data into a unified format. This makes data integration and analysis difficult, hindering the full realization of its value. For example, when conducting comprehensive financial risk assessments, the inability to integrate data from different platforms results in only partial analysis, making it impossible to accurately assess overall risk.

[0005] In the computational verification phase, traditional systems' verification mechanisms lack reliability in distributed environments. As financial services expand globally, transactions involve numerous distributed nodes, making it difficult for traditional verification methods to ensure data consistency and transaction legitimacy across all nodes. For example, in cross-border payments, due to varying regulations and systems across regions, traditional verification mechanisms are prone to vulnerabilities, leading to increased risks of transaction disputes or fraud.

[0006] Traditional systems also have serious shortcomings in business flow modeling and resource scheduling. Traditional modeling methods cannot accurately capture the temporal dependencies of capital flows, resulting in significant errors when predicting the transmission paths of business risks. Furthermore, resource allocation often uses static strategies that cannot be adjusted in real time to adapt to business changes, resulting in wasted or insufficient resources. For example, during e-commerce promotions, demand for financial transactions surges, and traditional resource scheduling cannot allocate sufficient resources to related businesses in a timely manner, affecting transaction processing speed and user experience. Summary of the Invention

[0007] The object of the present invention is to provide a distributed financial data processing system to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a distributed financial data processing system, comprising:

[0009] Multi-source data acquisition module: used to capture multi-dimensional financial transaction data in real time through distributed edge computing nodes and allocate data acquisition paths based on a dynamic load balancing algorithm;

[0010] Heterogeneous data integration module: This module parses the financial transaction data in a cross-platform format to generate standardized data streams, including transaction subject association maps, capital flow topology features, risk label distribution matrices, and compliance rule constraints;

[0011] Distributed computing verification module: Builds a data trust verification model based on a hybrid consensus mechanism, inputs the standardized data stream into the parallel computing unit, and generates transaction legitimacy determination results and distributed ledger update instructions;

[0012] Business flow modeling module: This module uses a spatiotemporal causal convolutional network to model the temporal dependencies of capital flows and combines it with a hierarchical attention mechanism to generate a business risk transmission path prediction matrix.

[0013] Resource scheduling module: Based on the prediction matrix, dynamically adjust the computing node resource allocation strategy through a multi-objective optimization algorithm.

[0014] Preferably, in the multi-source data acquisition module, the dynamic load balancing algorithm includes: dynamically allocating data acquisition tasks using a greedy strategy based on communication delay constraints and data shard integrity requirements.

[0015] Preferably, the heterogeneous data integration module includes:

[0016] A graph embedding network is used to extract high-order relationship features from the transaction subject association graph, and a random walk strategy is used to align cross-platform entity identifiers;

[0017] The capital flow topology features are modeled using a bidirectional gated recurrent unit to model the capital reflux cycle pattern, and the abnormal transaction detection threshold is generated in combination with the gradient boosting tree algorithm.

[0018] Preferably, in the distributed computing verification module, the hybrid consensus mechanism includes:

[0019] A zero-knowledge proof algorithm is used to verify the integrity of transaction privacy data, generate verifiable statements and broadcast them to consensus nodes; a multi-stage voting mechanism is built based on an asynchronous Byzantine fault-tolerant protocol, and the node voting weight is dynamically adjusted through a weight decay function.

[0020] Preferably, the business flow modeling module further includes:

[0021] A timestamp encoding algorithm is used to construct a causal dependency graph for the timing dependency of the capital flow, and a multi-head self-attention mechanism is used to capture the correlation of the cross-account capital chain;

[0022] Based on the risk transmission path prediction matrix, a generative adversarial network is used to simulate the cascading impact of black swan events on the capital network.

[0023] Preferably, the resource scheduling module also includes: constructing a computing node resource status monitoring matrix and using a density peak clustering algorithm to identify resource bottleneck areas; based on the prediction matrix and the resource status matrix, scheduling cross-node computing tasks through a dynamic priority queue and generating resource reservation instructions.

[0024] Preferably, the system further comprises: performing conflict detection on the distributed ledger update instructions, and using a version vector clock algorithm to mark the timing of concurrent operations; when a ledger state conflict is detected, triggering a rollback compensation mechanism, and generating an atomic undo operation sequence based on a transaction dependency graph.

[0025] Preferably, the conflict detection further comprises: constructing a transaction operation impact domain diffusion model, and using a streaming topological sorting algorithm to identify dependency chains of unfinished transactions.

[0026] Preferably, the system also includes: constructing an adaptive matching model for regulatory rules, extracting the constraints of compliance policy clauses based on a semantic parsing network; generating a rule violation risk score through a logical reasoning engine, and embedding it into the prediction matrix of the business flow modeling module.

[0027] Preferably, the regulatory rule adaptive matching model further includes:

[0028] Use knowledge graph embedding technology to align policy terms with transaction behavior characteristics and construct a multi-dimensional compliance judgment space;

[0029] Dynamically optimize rule matching thresholds through reinforcement learning strategies and generate regulatory report generation templates.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The distributed financial data processing system of the present invention has many significant beneficial effects. In the data collection stage, the multi-source data collection module greatly improves the efficiency and accuracy of data collection with the help of distributed edge computing nodes and dynamic load balancing algorithms. Distributed edge computing nodes are widely distributed and can capture multi-dimensional financial transaction data from all corners in real time to ensure that no data is missed. The dynamic load balancing algorithm is based on communication delay constraints and data sharding integrity requirements, and adopts a greedy strategy to dynamically allocate data collection tasks, avoiding data collection delays or failures caused by uneven node loads. This is like building a tight financial data collection network that does not miss any transaction information, providing a comprehensive and accurate data foundation for subsequent analysis and decision-making.

[0032] The heterogeneous data integration module effectively solves the problem of inconsistent data formats. By parsing financial transaction data across different platforms, it generates standardized data streams, including key information such as the transaction entity association graph and the topological characteristics of capital flows. A graph embedding network is used to extract high-order relationship features from the transaction entity association graph, and a random walk strategy is used to align cross-platform entity identifiers. This clearly demonstrates the complex relationships between different transaction entities and explores potential business collaborations or risks. A bidirectional gated recurrent unit is used to model the capital return cycle pattern for the capital flow topological characteristics. Combined with the gradient boosting tree algorithm to generate anomaly transaction detection thresholds, this method can promptly detect abnormal transactions and ensure the security and stability of financial transactions.

[0033] The distributed computing verification module builds a trusted data verification model based on a hybrid consensus mechanism, ensuring data credibility and consistency in a distributed environment. It uses a zero-knowledge proof algorithm to verify the integrity of private transaction data, generating verifiable statements that are broadcast to consensus nodes. This protects transaction privacy while enabling other nodes to verify data integrity. A multi-stage voting mechanism, based on an asynchronous Byzantine fault-tolerant protocol, dynamically adjusts node voting weights through a weight decay function, effectively preventing individual nodes from excessively influencing voting results. This ensures fairer and more accurate determination of transaction legitimacy and more reliable distributed ledger updates.

[0034] The business flow modeling module uses a spatiotemporal causal convolutional network to model the temporal dependencies of capital flows. Combined with a hierarchical attention mechanism, it generates a business risk transmission path prediction matrix, significantly improving risk prediction capabilities. A timestamp encoding algorithm is used to construct a causal dependency graph based on the temporal dependencies of capital flows. A multi-head self-attention mechanism captures cross-account capital chain correlations, enabling precise analysis of potential risk points and risk transmission paths in the capital flow process. Based on the risk transmission path prediction matrix, a generative adversarial network is employed to simulate the cascading impact of black swan events on the capital network, helping financial institutions prepare for extreme situations.

[0035] Based on the prediction matrix generated by the business flow modeling module, the resource scheduling module dynamically adjusts the compute node resource allocation strategy through a multi-objective optimization algorithm, significantly improving system resource utilization efficiency. A compute node resource status monitoring matrix is constructed, and a density peak clustering algorithm is used to identify resource bottlenecks, enabling timely detection of resource shortfalls in the system. Based on the prediction matrix and the resource status matrix, cross-node computing tasks are scheduled through dynamic priority queues and resource reservation instructions are generated, achieving rational resource allocation and efficient utilization, avoiding resource waste and task backlogs.

[0036] The system also features mechanisms for ledger conflict detection and resolution, as well as adaptive regulatory rule matching capabilities. Conflict detection is performed on distributed ledger update instructions, using a version vector clock algorithm to mark the timing of concurrent operations. When a ledger state conflict is detected, a rollback compensation mechanism is triggered, generating an atomic undo operation sequence based on the transaction dependency graph, ensuring the consistency and integrity of ledger data. An adaptive regulatory rule matching model is constructed, extracting the constraints of compliance policy clauses based on a semantic parsing network. A logical reasoning engine generates rule violation risk scores, which are then embedded in the prediction matrix of the business flow modeling module. This allows the system to strictly adhere to regulatory requirements while processing business, mitigating compliance risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a working principle diagram of the distributed financial data processing system of the present invention;

[0038] Figure 2 This is the working principle diagram of the heterogeneous data integration module;

[0039] Figure 3 This is the working principle diagram of the resource scheduling module;

[0040] Figure 4 A diagram showing the working principle of conflict detection and handling for distributed ledger update instructions. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] See also Figures 1-4 The present invention provides a distributed financial data processing system, the overall implementation of which is as follows:

[0043] The system primarily comprises a multi-source data acquisition module, a heterogeneous data integration module, a distributed computing verification module, a business flow modeling module, and a resource scheduling module. In actual operation, the multi-source data acquisition module captures multi-dimensional financial transaction data in real time through distributed edge computing nodes. These distributed edge computing nodes are located in different geographic locations or network environments and can collect data from a wide range of financial transaction scenarios, such as online payment platforms, bank transfer systems, and securities markets. To ensure efficient and stable data collection, the module allocates data collection paths based on a dynamic load balancing algorithm. This approach allows collection tasks to be rationally assigned to the most appropriate nodes based on their load and network status, preventing overloaded nodes from impacting data collection efficiency and ensuring a smooth data collection process.

[0044] Collected financial transaction data comes in a variety of formats. The heterogeneous data integration module is responsible for parsing this data across various platforms. This parsing process generates a standardized data stream, including a graph of transaction entity relationships, topological features of capital flows, a risk label distribution matrix, and compliance rule constraints. This module converts data from different platforms and formats into a standardized format that the system can process, laying the foundation for subsequent analysis and processing.

[0045] The distributed computing verification module builds a data trust verification model based on a hybrid consensus mechanism. This model feeds the standardized data stream processed by the heterogeneous data integration module into a parallel computing unit, generating transaction legitimacy determination results and distributed ledger update instructions. The hybrid consensus mechanism ensures data credibility and consistency in a distributed environment, ensuring accurate determination of transaction legitimacy while simultaneously updating the distributed ledger and recording transaction information.

[0046] The business flow modeling module uses a spatiotemporal causal convolutional network to model the temporal dependencies of capital flows. This network effectively captures the patterns of capital flow across different temporal and spatial dimensions, and, combined with a layered attention mechanism, generates a business risk transmission path prediction matrix. This approach allows for the prediction of potential risk paths in capital flows and proactive risk prevention measures.

[0047] The resource scheduling module uses a multi-objective optimization algorithm to dynamically adjust the resource allocation strategy for computing nodes based on the prediction matrix generated by the business flow modeling module. Based on the system's current computing task requirements and the resource status of each node, computing resources are rationally allocated to improve the overall operational efficiency of the system.

[0048] The specific implementation of the present invention is further described in detail below through five examples.

[0049] Example 1:

[0050] This embodiment elaborates on the specific implementation of some technologies in the multi-source data acquisition module and the heterogeneous data integration module. In the multi-source data acquisition module, the dynamic load balancing algorithm uses a greedy strategy to dynamically allocate data acquisition tasks based on the communication delay constraint and the data sharding integrity requirements. In an actual financial data acquisition scenario, assume that there are multiple distributed edge computing nodes N1, N2, N3..., each node has different communication delays and processing capabilities. The communication delay constraint D represents the maximum delay time allowed for data transmission between nodes, and the data sharding integrity requirement ensures that the collected data shards can be processed completely without data loss or incompleteness. When allocating tasks, the greedy strategy will give priority to nodes with the smallest communication delay and that can meet the data sharding integrity requirements. For example, when a batch of new financial transaction data needs to be collected, the system will first evaluate the communication delay of each node. choose And the node N that can ensure the complete processing of data shards j This can minimize data transmission time and improve collection efficiency while ensuring data collection integrity.

[0051] In the heterogeneous data integration module, a graph embedding network is used to extract high-order relationship features from the transaction entity association graph and a random walk strategy is used to align cross-platform entity identifiers. The transaction entity association graph is a complex network structure consisting of numerous transaction entities and their relationships. The graph embedding network can map this complex graph structure into a low-dimensional vector space and extract high-order relationship features. For example, in a graph containing multiple transaction entities such as enterprises, banks, and individuals, the graph embedding network can mine indirect relationships between different entities, such as when enterprise A has financial transactions with enterprise C through bank B, thereby uncovering deeper business partnerships or potential risks. The random walk strategy is used to align cross-platform entity identifiers because the same entity may have different identifiers on different platforms. During the random walk process, starting from a node, the next adjacent node is selected based on a certain probability. During this walk, by comparing the attributes and relationships of nodes on different platforms, node identifiers with similar characteristics are aligned, thus achieving unified entity identifiers across platforms and facilitating subsequent in-depth analysis and processing of the transaction entity association graph.

[0052] Example 2:

[0053] This embodiment focuses on another part of the heterogeneous data integration module and the related technical implementation of the distributed computing verification module. In the heterogeneous data integration module, the bidirectional gated recurrent unit (Bi-GRU) is used to model the capital flow cycle pattern for the capital flow topology feature, and the gradient boosting tree algorithm is combined to generate the abnormal transaction detection threshold. The capital flow topology feature reflects the flow path and direction of funds between different transaction entities. Bi-GRU is a special recurrent neural network structure that can learn the time series data of capital flow from both the forward and reverse directions at the same time, so as to capture the capital flow cycle pattern more comprehensively. Assume that the capital flow data is a time series S = [s1, s2,…, s t ], where s t represents the capital flow state at time t. By processing this time series, Bi-GRU can learn the inflow and outflow patterns of funds in different time periods and determine the cycle of capital return.

[0054] The gradient boosting tree algorithm generates an abnormal transaction detection threshold based on the capital flow cycle pattern learned by the Bi-GRU. The gradient boosting tree is an ensemble learning algorithm that gradually improves the model's predictive capabilities by continuously fitting residuals. In this process, capital flow data from normal trading conditions serves as a training set and is input into the gradient boosting tree model. The model learns the characteristics and patterns of normal trading data and calculates a threshold value, T. When the capital flow in an actual transaction deviates from the normal cycle pattern to a certain extent, that is, exceeds this threshold value, T, the transaction is classified as abnormal. This allows for the timely detection of potentially risky transactions and ensures the security of financial transactions.

[0055] In the distributed computing verification module, the hybrid consensus mechanism uses a zero-knowledge proof algorithm to verify the integrity of private transaction data, generate verifiable statements, and broadcast them to consensus nodes. A multi-stage voting mechanism is constructed based on an asynchronous Byzantine fault-tolerant protocol, dynamically adjusting node voting weights through a weight decay function. The zero-knowledge proof algorithm allows the prover to prove the correctness of a proposition without revealing any useful information to the verifier. In financial transactions, private transaction data contains a large amount of sensitive information, such as the transaction amount and the identities of both parties. Using the zero-knowledge proof algorithm, the verifier can verify the integrity of this data without having access to this specific private data. For example, the prover can generate a proof P using a specific algorithm. Based on this proof and related verification rules, the verifier can determine the integrity of the private transaction data without having to know the specific transaction content.

[0056] The multi-stage voting mechanism built on the asynchronous Byzantine fault-tolerant protocol is designed to ensure that all nodes reach a consensus on the legitimacy of transactions in a distributed environment. In this mechanism, the voting process is divided into multiple stages, and each stage each node votes based on its own judgment. The weight decay function W(n) is used to dynamically adjust the node voting weight, where n represents the number of times a node participates in voting. As the number of times a node participates in voting increases, its voting weight will gradually decay, that is, Here, W0 is the node's initial voting weight, and α is the decay coefficient. This prevents certain nodes from excessively influencing voting results due to frequent voting, ensuring the fairness and objectivity of voting, and thus improving the accuracy of transaction legitimacy determination.

[0057] Example 3:

[0058] This embodiment mainly describes the specific implementation of some technologies in the business flow modeling module and the relevant content of the resource scheduling module. In the business flow modeling module, a timestamp encoding algorithm is used to construct a causal dependency graph for the temporal dependency of capital flow, and a multi-head self-attention mechanism is used to capture the correlation of cross-account capital chains. In financial transactions, each capital flow is accompanied by a timestamp. The timestamp encoding algorithm will encode these timestamp information and convert them into coded information that can reflect the order of capital flow. Assume that there is a series of capital flow records R = [r1, r2, ..., r m ], each record r i All contain the timestamp t of the funds flow i Through the timestamp coding algorithm, these timestamps are converted into coding vectors E = [e1, e2, ..., e m ], and construct a causal dependency graph based on these encoding vectors. In the causal dependency graph, nodes represent fund flow records, and edges represent the causal relationship between fund flows, that is, the flow that occurs earlier in time may affect the flow that occurs later in time.

[0059] The multi-head self-attention mechanism is used to capture cross-account fund chain correlations. Using multiple attention heads, the multi-head self-attention mechanism focuses on and analyzes fund flow data from different perspectives. For example, in a complex financial network, funds flow between multiple accounts. Different attention heads can focus on information such as fund flow paths and amount fluctuation trends between different accounts, thereby more comprehensively capturing cross-account fund chain correlations. This approach can uncover hidden relationships within complex fund flows, providing a more accurate basis for risk prediction.

[0060] In the resource scheduling module, a computing node resource status monitoring matrix is constructed, and a density peak clustering algorithm is used to identify resource bottleneck areas. The computing node resource status monitoring matrix M records the resource usage of each computing node, such as CPU utilization, memory utilization, network bandwidth, and other information. The matrix M can be expressed as:

[0061]

[0062] where m ij Indicates the usage status of the jth resource of the i-th computing node, n is the number of computing nodes, and k is the number of resource types.

[0063] The density peak clustering algorithm is a clustering algorithm based on data point density. Based on the computing node resource status monitoring matrix, the algorithm calculates the local density ρ of each data point (i.e., the resource status of the computing node) i and the relative distance δ i To identify resource bottleneck areas. Local density ρ i It represents the number of data points in a certain neighborhood centered on node i, and the calculation formula is:

[0064]

[0065] where d ij is the distance between node i and node j, d c Is a cutoff distance used to control the size of the neighborhood. i Represents the distance between node i and the nearest node with a greater density than it, that is If the density of node i is the largest, then δ i is the maximum distance between it and all other nodes. By analyzing ρ i and δ i , identifies nodes with high local density and large relative distances as resource bottlenecks. Based on these identified resource bottlenecks and the prediction matrix generated by the traffic flow modeling module, a dynamic priority queue is used to schedule cross-node computing tasks and generate resource reservation instructions to rationally allocate computing resources and improve overall system performance.

[0066] Example 4:

[0067] This embodiment focuses on conflict detection and related compensation mechanisms for distributed ledger update instructions in the system. The system will perform conflict detection on distributed ledger update instructions and use the version vector clock algorithm to mark the timing of concurrent operations. In a distributed ledger environment, multiple nodes may update the ledger at the same time, which may lead to conflicts. The version vector clock algorithm assigns a version vector to each ledger operation. The version vector is a vector composed of multiple elements, each element corresponding to a node. Assume that there are n nodes, and the version vector V = [v1, v2, ..., v n ], where v i Represents the version number of the operation on node i. When a node initiates a ledger update, it increments its version number by 1 and broadcasts the update operation and version vector to other nodes. Upon receiving the update operation, other nodes compare their own version vectors with the received version vector. If the corresponding node's version number in the received version vector is smaller than their own, the operation is outdated and requires appropriate processing. If the corresponding node's version number in the received version vector is larger than their own, their own ledger needs to be updated. This method can mark the timing of concurrent operations and determine the order of operations.

[0068] When a ledger state conflict is detected, the rollback compensation mechanism is triggered, generating a sequence of atomic undo operations based on the transaction dependency graph. The transaction dependency graph records the dependencies between transactions, meaning that the execution of one transaction may depend on the completion of other transactions. When a conflict occurs, the transaction dependency graph is used to find all related transactions, starting from the point of conflict and working backwards. For example, suppose transaction T1 depends on transaction T2, and T2's update conflicts with another transaction T3's update. During a rollback, T3's operations are rolled back first. Then, based on the transaction dependencies, other transactions affected by T3 are rolled back in sequence, generating an atomic undo sequence. This ensures that the ledger state can be restored to its correct state before the conflict, guaranteeing the consistency and integrity of the ledger data.

[0069] Conflict detection also involves constructing a transaction impact diffusion model and using a streaming topological sorting algorithm to identify the dependency chains of unfinished transactions. The transaction impact diffusion model describes the scope and extent of a transaction's impact on other transactions and the ledger state. By analyzing the transaction's operational content and the ledger's structure, the impact domain of each transaction is determined. For example, the impact domain of a transaction involving a funds transfer may include the account balances and transaction records of both parties. The streaming topological sorting algorithm is used to process the continuous flow of transactions in a distributed environment. In the transaction flow, unfinished transactions are sorted based on their dependencies. If all of a transaction's predecessors have completed, it can be queued for execution; otherwise, it must wait for its predecessors to complete. This approach allows for the timely identification of unfinished transaction dependency chains, avoiding conflicts and errors caused by unclear dependency relationships, and further improving the system's ability to manage and control distributed ledger updates.

[0070] Example 5:

[0071] This embodiment focuses on describing the specific implementation of the adaptive matching model for regulatory rules in the system and its related functions. The system constructs an adaptive matching model for regulatory rules and extracts the constraints of compliance policy clauses based on the semantic parsing network. In the financial field, there are a large number of compliance policy clauses, which are expressed in the form of natural language. The semantic parsing network extracts the key constraints by analyzing and understanding these natural language clauses. For example, for a policy clause that "prohibits large loans to high-risk industry enterprises", the semantic parsing network can identify key information such as "high-risk industry enterprises" and "large loans", and convert them into constraints that the system can handle, such as setting a list of high-risk industry enterprises and a threshold for the amount of large loans.

[0072] A rule violation risk score is generated through a logical reasoning engine and embedded in the prediction matrix of the business flow modeling module. The logical reasoning engine performs inference and judgment based on extracted constraints and actual financial transaction data. In a specific loan business scenario, assuming a loan application, the logical reasoning engine first determines whether the loan applicant is from a high-risk industry list. If so, it then determines whether the loan amount exceeds the set large loan threshold. If these two conditions are met, a rule violation risk score is calculated based on a specific algorithm. This risk score is embedded in the prediction matrix of the business flow modeling module, allowing compliance risk factors to be fully considered when predicting business risk transmission paths, improving the accuracy and comprehensiveness of risk predictions.

[0073] The adaptive regulatory rule matching model also uses knowledge graph embedding technology to align policy terms with transaction characteristics, constructing a multidimensional compliance judgment space. A knowledge graph is a semantic network that contains a rich set of entities and their relationships. Using this technology, entities in policy terms and transaction behaviors are mapped into the same low-dimensional vector space. Within this vector space, policy terms and transaction characteristics are aligned by calculating similarities between these vectors. For example, the vector representation of "high-risk industry enterprises" in the knowledge graph is compared with the vector representation of the enterprise in actual transactions to determine whether the enterprise meets the characteristics of a high-risk industry enterprise. Based on these aligned features, a multidimensional compliance judgment space is constructed. Within this space, compliance assessments of transactions are conducted across multiple dimensions, such as the nature of the enterprise, transaction amount, and transaction time, enabling a more comprehensive and accurate assessment of transaction compliance.

[0074] A reinforcement learning strategy dynamically optimizes the rule matching threshold and generates a regulatory report generation template. Reinforcement learning is a machine learning method that learns optimal strategies by interacting with the environment and earning rewards. In the adaptive regulatory rule matching model, the rule matching threshold is used as an optimizable parameter. By continuously trying different thresholds and assigning rewards or penalties based on the actual compliance judgment results, the reinforcement learning algorithm can find the optimal rule matching threshold. For example, if a threshold is set too high, some transactions that should be classified as violations may be missed; if it is set too low, some normal transactions may be mistakenly classified as violations. Through reinforcement learning, the threshold is continuously adjusted to achieve more accurate rule matching. At the same time, a regulatory report generation template is generated. Based on different compliance judgment results and relevant data, standardized regulatory reports are generated according to the template, facilitating regulatory authorities' supervision and review of financial transactions.

[0075] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0076] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A distributed financial data processing system, characterized in that: include: Multi-source data acquisition module: used to capture multi-dimensional financial transaction data in real time through distributed edge computing nodes and allocate data acquisition paths based on a dynamic load balancing algorithm; Heterogeneous data integration module: This module parses the financial transaction data in a cross-platform format to generate standardized data streams, including transaction subject association maps, capital flow topology features, risk label distribution matrices, and compliance rule constraints; Distributed computing verification module: Builds a data trust verification model based on a hybrid consensus mechanism, inputs the standardized data stream into the parallel computing unit, and generates transaction legitimacy determination results and distributed ledger update instructions; Business flow modeling module: This module uses a spatiotemporal causal convolutional network to model the temporal dependencies of capital flows and combines it with a hierarchical attention mechanism to generate a business risk transmission path prediction matrix. Resource scheduling module: Based on the prediction matrix, dynamically adjust the computing node resource allocation strategy through a multi-objective optimization algorithm.

2. The distributed financial data processing system according to claim 1, characterized in that: In the multi-source data acquisition module, the dynamic load balancing algorithm includes: based on communication delay constraints and data sharding integrity requirements, a greedy strategy is used to dynamically allocate data acquisition tasks.

3. The distributed financial data processing system according to claim 1, characterized in that: The heterogeneous data integration module includes: A graph embedding network is used to extract high-order relationship features from the transaction subject association graph, and a random walk strategy is used to align cross-platform entity identifiers; The capital flow topology features are modeled using a bidirectional gated recurrent unit to model the capital reflux cycle pattern, and the abnormal transaction detection threshold is generated in combination with the gradient boosting tree algorithm.

4. The distributed financial data processing system according to claim 1, characterized in that: In the distributed computing verification module, the hybrid consensus mechanism includes: A zero-knowledge proof algorithm is used to verify the integrity of transaction privacy data, generate verifiable statements and broadcast them to consensus nodes; a multi-stage voting mechanism is built based on an asynchronous Byzantine fault-tolerant protocol, and the node voting weight is dynamically adjusted through a weight decay function.

5. The distributed financial data processing system according to claim 1, characterized in that: The business flow modeling module also includes: A timestamp encoding algorithm is used to construct a causal dependency graph for the timing dependency of the capital flow, and a multi-head self-attention mechanism is used to capture the correlation of the cross-account capital chain; Based on the risk transmission path prediction matrix, a generative adversarial network is used to simulate the cascading impact of black swan events on the capital network.

6. The distributed financial data processing system according to claim 1, characterized in that: The resource scheduling module also includes: constructing a computing node resource status monitoring matrix and using a density peak clustering algorithm to identify resource bottleneck areas; scheduling cross-node computing tasks through a dynamic priority queue based on the prediction matrix and the resource status matrix, and generating resource reservation instructions.

7. The distributed financial data processing system according to claim 1, characterized in that: The system further includes: performing conflict detection on the distributed ledger update instructions and using a version vector clock algorithm to mark the timing of concurrent operations; when a ledger state conflict is detected, triggering a rollback compensation mechanism and generating an atomic undo operation sequence based on a transaction dependency graph.

8. The distributed financial data processing system according to claim 7, characterized in that: The conflict detection further includes: constructing a transaction operation impact domain diffusion model, and using a streaming topological sorting algorithm to identify dependency chains of unfinished transactions.

9. The distributed financial data processing system according to claim 1, characterized in that: The system also includes: building a regulatory rule adaptive matching model, extracting the constraints of compliance policy clauses based on a semantic parsing network; generating a rule violation risk score through a logical reasoning engine, and embedding it into the prediction matrix of the business flow modeling module.

10. The distributed financial data processing system according to claim 9, characterized in that: The regulatory rule adaptive matching model also includes: Use knowledge graph embedding technology to align policy terms with transaction behavior characteristics and construct a multi-dimensional compliance judgment space; Dynamically optimize rule matching thresholds through reinforcement learning strategies and generate regulatory report generation templates.

Citation Information

Cited By

  • Abnormal root cause positioning method and device based on order business system, and electronic equipment

    CN121095591A