Bank batch processing system and working method thereof
By designing a bank batch processing system and utilizing modules for task scheduling, data partitioning, fault tolerance, monitoring, and logging, the system solves the problems of low efficiency, poor fault tolerance, poor scalability, and high maintenance costs of traditional bank batch processing systems, achieving efficient, reliable, and secure bank batch processing.
Patent Information
- Application Number
- CN202510885758.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional bank batch processing systems suffer from low processing efficiency, poor fault tolerance, poor scalability, and high maintenance costs, making it difficult to meet the demands of multi-core CPU performance utilization and the rapid growth of financial business.
A bank batch processing system was designed, including a task scheduling module, a task sharding module, a task execution and fault tolerance module, a monitoring and logging module, and a performance optimization module. It adopts the Quartz or XXL-job scheduling framework, supports centralized and distributed job collaboration, and combines rack awareness and dynamic node performance evaluation to implement dynamic sharding and fine-grained processing strategies. It integrates exception handling, provides encrypted storage, data monitoring, log analysis, and automatic alarm notification functions, achieves data encryption and transmission, and supports multi-environment deployment and dynamic resource scheduling.
It significantly improves processing efficiency, enhances fault tolerance, reduces operation and maintenance costs, meets financial-grade security and compliance requirements, adapts to complex banking business scenarios, optimizes heterogeneous storage environments, and enables efficient parallel processing and fine-grained anomaly management.
Smart Images

Figure CN120973510A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bank batch processing technology, specifically to a bank batch processing system and its working method. Background Technology
[0002] As banking operations continue to expand, the number of batch tasks that need to be processed in banking systems is increasing. Traditional batch processing methods have the following problems: Low processing efficiency: Traditional batch processing systems typically use single-threaded processing, which cannot fully utilize the performance of multi-core CPUs, resulting in low processing efficiency.
[0003] Poor fault tolerance: In batch processing, if an exception occurs, traditional batch processing systems often need to start from scratch and reprocess, resulting in wasted resources.
[0004] Poor scalability: Traditional batch processing systems struggle to cope with rapid growth in business volume and have poor scalability.
[0005] High maintenance costs: Traditional batch processing systems typically rely on complex scripts and manual operations, resulting in high maintenance costs.
[0006] Spring Batch is a lightweight batch processing framework that offers rich functionality and flexible configuration, effectively addressing the aforementioned issues. However, existing bank batch processing systems still have room for improvement in areas such as task scheduling, fault tolerance, and performance optimization. Summary of the Invention
[0007] The purpose of this invention is to address the shortcomings of existing technologies by providing a bank batch processing system and its operating method.
[0008] To achieve the above objectives, in a first aspect, the present invention provides a bank batch processing system, comprising: The task scheduling module is configured to automatically trigger batch processing tasks according to a preset scheduling strategy, and integrates a unified configuration platform for dynamic registration and distribution of batch tasks. The task sharding module is configured to divide batch processing tasks into multiple subtasks for parallel processing according to the sharding strategy, and dynamically adjust the number of shards based on system resources. The task execution and fault tolerance module is configured to capture task execution exceptions, retry or skip tasks with execution exceptions based on preset exception execution strategies, and integrate exception log analysis and automatic alarm notification functions. The monitoring and logging module is configured to monitor the task execution status in real time and record log information, providing a visual interface and connecting to the bank's operation and maintenance platform for real-time observability, alarms and backtracking. The performance optimization module is configured to optimize task execution performance, including thread pool parameter tuning, database index strategy optimization, batch submission control strategy, and scheduling strategy for heterogeneous storage environments. The scheduling strategy combines rack awareness, node dynamic performance evaluation, and access popularity scheduling.
[0009] Furthermore, the task scheduling module is based on the Quartz or xxl-job scheduling framework, supports centralized scheduling and distributed job collaboration, is compatible with multi-environment deployments in the banking industry (Dev / Test / Prod), and uses Spring Cloud Config for unified management of the configuration center.
[0010] Furthermore, the scheduling strategy is timed scheduling and / or event-triggered scheduling.
[0011] Furthermore, the sharding strategy includes configurable sharding strategies based on data volume, time window, and customer dimension.
[0012] Furthermore, the task execution and fault tolerance module integrates the Prometheus+Alertmanager monitoring and alarm platform to provide real-time notifications for abnormal states and record the operator, job status, and data summary.
[0013] Furthermore, the preset abnormal execution strategy includes a retry limit and an abnormal skipping rule.
[0014] Furthermore, the monitoring and logging module supports quick retrieval based on task, batch number, and user ID. The logging system contains batch-level operation records for data backtracking and compliance auditing.
[0015] Furthermore, the dimensions of the log information include task, batch number, and user ID.
[0016] Furthermore, it also includes a security mechanism module, configured to encrypt and store task data, transmit and process link data using TLS secure encryption, configure job operation permissions based on user roles, and support whitelist control and approval mechanisms.
[0017] In a second aspect, the present invention provides a method for operating the aforementioned bank batch processing system, comprising: The scheduling module automatically triggers batch processing tasks based on timed or event-triggered strategies, supporting dynamic job configuration hot loading and multi-environment deployment adaptation. The task is divided into multiple subtasks by the sharding module, and parallel processing is carried out based on configurable strategies of data volume, time window and customer dimension, and the number of shards is dynamically adjusted. Subtasks are executed in parallel through the task execution and fault tolerance module, and execution exceptions are captured in real time and recorded in detail. Exception tasks are retried or skipped. The monitoring and logging modules monitor task status in real time, record traceable audit logs, and connect to the bank's operation and maintenance platform to achieve visualized monitoring and alarms. The performance optimization module is used for thread pool parameter tuning, database index optimization, batch commit control, rack awareness, node dynamic performance evaluation, and access heat scheduling strategies in heterogeneous storage environments.
[0018] Beneficial effects: 1. Significantly improved processing efficiency: Parallel processing and dynamic sharding: The task sharding module supports configurable sharding strategies based on multiple dimensions such as data volume, time window, and customer dimensions. It can also dynamically adjust the number of shards based on system resources, fully leveraging the advantages of multi-core CPUs and distributed computing resources. This enables the system to efficiently handle massive data processing demands and significantly shorten task processing time in various business scenarios.
[0019] Storage and Database Optimization: For heterogeneous storage systems such as HDFS and Ceph, a "rack awareness + dynamic node performance evaluation + access heat scheduling" strategy is adopted to solve the shortcomings of traditional static allocation and effectively improve data access efficiency. Simultaneously, the database reduces the number of I / O interactions and significantly shortens database operation time through composite index optimization and batch commit control.
[0020] 2. Enhanced fault tolerance across the board: Fine-grained anomaly handling: The fault tolerance module features retry limit and anomaly skipping strategies, combined with a real-time alarm mechanism, to avoid restarting all tasks. When encountering anomalies, the system can handle them precisely, addressing only the problematic parts, saving significant amounts of redundant computational resources and greatly improving system reliability.
[0021] Closed-loop alarm and audit traceability: The system records batch-level audit logs such as operator, job status, and data summary, and supports quick retrieval by task, batch number, user ID, etc., which can quickly locate the cause of anomalies and meet the strict requirements of financial supervision for data traceability.
[0022] 3. Enhanced scalability and flexibility: Dynamic resource scheduling: The task scheduling module supports the Quartz / xxl-job framework and is adaptable to both centralized and distributed collaborative scheduling modes. Spring Cloud Config provides unified management of configurations for different environments (development, testing, and production), enabling dynamic registration and hot reloading of batch tasks, adapting to environment changes without requiring a system restart. Furthermore, the number of shards can be automatically adjusted based on system resource metrics, easily handling fluctuations in business volume and expanding processing capacity without code modification.
[0023] Multi-scenario adaptation: The system supports various business scenarios such as batch deduction, payroll disbursement, reconciliation, and statement generation. Through differentiated sharding strategies, it can meet the data organization and processing logic needs of different businesses.
[0024] 4. Significantly reduced operation and maintenance costs: Automated scheduling and monitoring: The task scheduling module enables automated operations triggered by specific times or events, greatly reducing manual operations and avoiding human error. The monitoring and logging module provides a visual interface that displays task progress, resource usage, and anomaly trends in real time, enabling rapid troubleshooting, significantly shortening troubleshooting time, and reducing maintenance manpower investment.
[0025] Standardized configuration and management: A unified configuration platform and dynamic hot-reload mechanism simplify the complexity of multi-environment deployment and reduce configuration management costs. The logging system supports the automatic generation of compliance audit reports to meet regulatory inspection requirements.
[0026] 5. Financial-grade security and compliance guarantees: Data security mechanism: Sensitive data is stored and transmitted in encrypted form, and key fields such as bank account numbers and transaction amounts are anonymized to meet relevant laws and regulations as well as industry security standards.
[0027] Access and Audit Controls: Role-based access control is employed, with dual approval and whitelist controls implemented for critical operations to prevent unauthorized actions. Batch-level audit logs fully record the operation process, ensuring data immutability and traceability.
[0028] 6. Technological innovation and industry adaptation: Component integration architecture optimization: It breaks through the limitations of SpringBatch native components, integrates Prometheus alerting, SpringCloudConfig configuration center, etc., and builds a collaborative architecture of "scheduling-sharding-fault tolerance-monitoring-optimization", which effectively solves the problem of component adaptation in complex business scenarios of banks.
[0029] Deep optimization for heterogeneous environments: For the hybrid cloud architecture of banks, a storage strategy combining rack awareness and node dynamic performance was designed to improve the data storage and access efficiency in heterogeneous environments and provide strong technical support for distributed batch processing. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of a bank batch processing system according to an embodiment of the present invention. Detailed Implementation
[0031] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solutions of the present invention, and it should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0032] like Figure 1 As shown, this embodiment of the invention provides a bank batch processing system, including a task scheduling module 1, a task sharding module 2, a task execution and fault tolerance module 3, a monitoring and logging module 4, and a performance optimization module.
[0033] Task scheduling module 1 is configured to automatically trigger batch processing tasks according to a preset scheduling strategy, and integrates a unified configuration platform for dynamic registration and distribution of batch tasks. The aforementioned scheduling strategy includes timed scheduling and / or event-triggered scheduling. Task scheduling module 1 is based on the Quartz or XXL-Job scheduling framework, supports centralized scheduling and distributed job collaboration, is adaptable to multi-environment deployments in the banking sector (Dev / Test / Prod), and uses Spring Cloud Config for unified configuration center management.
[0034] Task sharding module 2 is configured to divide batch processing tasks into multiple subtasks for parallel processing according to a sharding strategy, and dynamically adjust the number of shards based on system resources. The aforementioned sharding strategies include configurable sharding strategies based on data volume, time window, and client dimensions.
[0035] The task execution and fault tolerance module 3 is configured to capture task execution exceptions and, based on preset exception execution strategies, retry or skip tasks with execution exceptions. It also integrates exception log analysis and automatic alarm notification functions. The preset exception execution strategies include retry limit and exception skipping rules. The task execution and fault tolerance module 3 integrates the Prometheus+Alertmanager monitoring and alarm platform, providing real-time notifications of abnormal states and recording the operator, job status, and data summary.
[0036] The monitoring and logging module 4 is configured to monitor task execution status in real time and record log information, providing a visual interface and connecting to the bank's operations and maintenance platform for real-time observability, alerts, and backtracking. Log information dimensions include task, batch number, and user ID. The monitoring and logging module 4 supports quick retrieval based on task, batch number, and user ID. The log system contains batch-level operation records for data backtracking and compliance auditing.
[0037] The performance optimization module is configured to optimize task execution performance, including thread pool parameter tuning, database index strategy optimization, batch submission control strategy, and scheduling strategy for heterogeneous storage environments. The scheduling strategy combines rack awareness, node dynamic performance evaluation, and access popularity scheduling.
[0038] It also includes a security mechanism module, configured to encrypt and store task data, transmit and process link data using TLS secure encryption, configure job operation permissions based on user roles, and support whitelist control and approval mechanisms.
[0039] Based on the above embodiments, those skilled in the art can easily understand that the present invention also provides a method for operating the above-mentioned bank batch processing system, including: The scheduling module automatically triggers batch processing tasks based on timed or event-triggered strategies, supporting dynamic job configuration, hot reloading, and multi-environment deployment adaptation. Task scheduling steps support task dependency configuration, allowing subsequent tasks to be executed only after a preceding task succeeds, and supports failure retry and rescheduling strategies.
[0040] The task is divided into multiple subtasks by the sharding module, and parallel processing is carried out based on configurable strategies of data volume, time window and customer dimension, and the number of shards is dynamically adjusted.
[0041] The task execution and fault tolerance module executes subtasks in parallel, captures execution exceptions in real time and records detailed logs, and retryes or skips abnormal tasks.
[0042] The monitoring and logging modules provide real-time monitoring of task status, record traceable audit logs, and connect to the bank's operations and maintenance platform to achieve visualized monitoring and alerts.
[0043] The performance optimization module performs thread pool parameter tuning, database index optimization, batch commit control, and implements rack awareness, dynamic node performance evaluation, and access heat scheduling strategies in heterogeneous storage environments. For HDFS, Ceph, or GlusterFS storage systems, a rack awareness combined with dynamic node performance evaluation and access heat scheduling strategies is used to optimize data storage and access efficiency.
[0044] Example 1: Bank's bulk payroll disbursement service: (a) Application scenarios: On the 10th of each month, the bank conducts a "bulk payroll disbursement" service, which requires processing the payroll of 150,000 employees from 200 branches. This involves interbank transfers, account verification, tax withholding, and other operations, and must be completed within 4 hours while meeting the requirements of data security compliance, automatic anomaly handling, and dynamic resource scheduling.
[0045] (II) System Module Implementation and Methodological Steps: Task scheduling module 1: Implementation details: The Quartz scheduling framework is used to trigger tasks on a scheduled basis at 08:00 on the 10th of each month. Manual triggering via the management backend is also supported. Spring Cloud Config is used to uniformly manage configurations across Dev / Test / Prod environments. For example, the database connection parameters for the production environment are jdbc:mysql: / / prod-db:3306 / batch_db, while the test environment automatically switches to jdbc:mysql: / / test-db:3306 / batch_db. Task dependencies are configured so that the task is triggered only after successful "payroll data reconciliation." If a preceding task fails, the current task is automatically added to the retry queue, with a maximum of 5 retries.
[0046] Execution steps: At 08:00, the scheduling module triggers the task on time and loads real-time job parameters from the configuration center without requiring a system restart.
[0047] It automates scheduling, reduces human error, and solves the problem of "high maintenance costs"; it supports centralized and distributed collaborative scheduling, adapts to the complex IT environment of banks, and improves system scalability.
[0048] Task Sharding Module 2: Implementation details: Data is sharded by customer dimension (branch offices + departments), dividing the data of 150,000 employees into 2,000 sub-tasks. Then, every 100 transactions are merged into one shard, generating a total of 1,500 parallel sub-tasks. System resources are monitored in real time; when CPU utilization exceeds 80%, the number of shards is automatically adjusted to 2,000, distributed across 20 servers, with each server configured with 5 processing threads.
[0049] Execution steps: The sharding module parses the payroll data file and generates subtasks based on "branch code + department ID". Each subtask contains 100 payroll records.
[0050] Parallel processing leverages the performance of multi-core CPUs, reducing single-task processing time from 30 seconds / 100 transactions to 5 seconds / 100 transactions, improving efficiency by 6 times and solving the problem of "low processing efficiency"; dynamic sharding strategy adapts to fluctuations in business volume, avoiding resource waste and processing timeouts.
[0051] Task Execution and Fault Tolerance Module 3: Implementation details: The account verification process captures exceptions such as "Account Frozen" (error code: ACCT_LOCKED) and "Interbank Transfer Failed" (error code: TRANSFER_ERR). For "Interbank Transfer Failed," a default retry of 3 times is performed, with a 5-minute interval between each attempt. "Account Frozen" is marked as "Exception Skipped," and an SMS alert is sent to the operations team via Prometheus + Alertmanager, recording the operator of the exception, the job status, and a data summary. During subtask execution, 100 outgoing transfer requests are submitted in batches using Spring Batch's Chunk-orientedProcessing, with transaction control enabled in the database.
[0052] Execution steps: The subtask executes batch submission requests. If the third retry of the transfer still fails, the record is marked as "pending manual processing" and skipped to continue with subsequent tasks.
[0053] Fine-grained anomaly handling avoids full task restarts, saves 97% of redundant computing resources, and solves the problem of "poor fault tolerance"; automatic alarms and audit logs realize closed-loop management of anomalies, reduce operation and maintenance inspection costs, and comply with financial compliance requirements.
[0054] Monitoring and logging module 4: Implementation details: Provides a visual interface to display task progress, resource usage, and anomaly trends in real time. Supports querying logs by combination of "batch number," "user ID," and "task status," returning results within 10 seconds. Task logs are synchronized to the bank's unified operations and maintenance system, supporting the review of 3 months of historical data and generating compliance audit reports.
[0055] Execution steps: For every 1,000 disbursements completed, push data updates to the monitoring module's dashboard; maintenance personnel can query the reasons for disbursement failures by entering "batch number + user ID" into the bank's internal system.
[0056] Real-time monitoring shortens the time for anomaly location from hours to minutes, reducing "maintenance costs"; compliance audit logs meet regulatory requirements and facilitate the retrieval of records by tax authorities.
[0057] Performance optimization module: Implementation details: Create a composite index (branch_id, batch_no, status) for the "Date Submission Log" and set the batch submission parameter fetchSize=500. For HDFS and Ceph storage systems, adopt a "rack awareness + node performance evaluation" strategy. HDFS stores hot data on high-performance nodes, and Ceph dynamically adjusts the data shard storage location based on node IOPS. Set the thread pool core thread count to 20, the maximum thread count to 50, and the queue capacity to 100.
[0058] Execution steps: When reading data, select the three nodes with the highest IOPS from the Ceph cluster to load the salary file in parallel. When writing to the database, write 1000 transaction records at a time.
[0059] Storage scheduling improved data access efficiency by 30%, database operation time was reduced from 2 hours to 1.2 hours, and overall task processing time was controlled within 3.5 hours; thread pool and batch submission optimizations improved system throughput to cope with business growth.
[0060] Security mechanism module: Implementation details: Payroll data files are encrypted with AES-256 before transmission, and the "Bank Account Number" and "Payment Amount" fields are anonymized during storage. The task processing chain uses TLS 1.3 encryption and verifies digital certificates. Based on role-based permission assignment, "Batch Task Administrators" can trigger payroll tasks, while "Operations Engineers" can only view monitoring data; operations require dual approval.
[0061] Execution steps: Payroll data transmission automatically triggers a TLS encrypted channel; administrator login requires dual authentication via SMS verification code and fingerprint.
[0062] Encryption and data anonymization prevent data leakage and meet the requirements of the Personal Information Protection Law; access control avoids unauthorized operations and reduces operational risks.
[0063] In Example 1: Improved processing efficiency: Task sharding and performance optimization reduced the processing time for 150,000 dispatch tasks from 6 hours to 3.5 hours, improving efficiency by 41.7%, making full use of multi-core CPUs and distributed resources.
[0064] Enhanced fault tolerance: Fine-grained anomaly handling and real-time alarms reduce the cost of handling anomaly tasks by more than 90%, and improve system reliability from 95% to 99.5%.
[0065] Enhanced scalability: Dynamic sharding and multi-environment adaptation allow the system to easily handle doubling data volumes, expanding processing capabilities without modifying the code.
[0066] Reduced maintenance costs: Automated scheduling reduces manual operations by 90%, and monitoring visualization and log retrieval shorten troubleshooting time from 2 hours to 15 minutes, reducing maintenance manpower input by 60%.
[0067] Security and compliance met: Data encryption, transmission encryption, access control and audit logs meet the security and compliance requirements of financial systems, and are risk-free through security certifications such as PCI-DSS.
[0068] It solves the efficiency, fault tolerance, scalability, maintenance, and security problems of traditional batch processing systems.
[0069] Example 2: Bank Batch Reconciliation Transaction (a) Application scenarios: Jiangsu Sushang Bank needs to reconcile massive amounts of transaction data daily, involving data interaction with multiple external institutions and different internal business systems. The data volume reaches 2 million transactions per day, and the reconciliation must be completed within 3 hours to ensure the accuracy of transaction data, promptly identify and process discrepancies, and meet regulatory requirements for data accuracy and timeliness.
[0070] (II) System Module Implementation and Methodological Steps: Task scheduling module 1: Implementation Details: The xxl-job scheduling framework is used, triggering batch reconciliation tasks daily at 2:00 AM. Manual triggering by the administrator is also supported in special circumstances (such as data consistency checks after system upgrades). Configuration is managed through Spring Cloud Config, allowing different data sources and target addresses to be configured for different environments (Dev / Test / Prod). For example, the production environment obtains data from prod-external-data-source, while the test environment obtains data from test-external-data-source. Task dependencies are configured to ensure that the reconciliation task only starts after the preceding data collection task has successfully completed. If the preceding task fails, the reconciliation task is automatically added to the retry queue, with a maximum of 3 retries.
[0071] Execution steps: At 2:00 AM, the scheduling module triggers the task and loads the real-time configuration from the configuration center, including information such as data collection address and reconciliation rules.
[0072] Automated scheduling reduces the tedious manual task triggering, avoids human error, and lowers maintenance costs; it supports centralized scheduling and distributed job collaboration, adapts to the scheduling needs of different business scenarios in the bank, and improves system scalability.
[0073] Task Sharding Module 2: Implementation details: Data is sharded according to the source institution and business type. The 2 million transaction data are divided into 50 institutions × 10 business types = 500 sub-tasks. Then, based on the data volume, each shard consists of 4000 transactions, ultimately generating 500 parallel sub-tasks. The number of shards is automatically adjusted by monitoring system resources (e.g., memory usage > 70%). If memory is scarce, the number of shards is increased to 600, distributed across 30 servers for parallel processing, with each server configured with 4 processing threads.
[0074] Execution steps: The sharding module parses the transaction data and generates subtasks based on the "data source institution code + business type" tag. Each subtask contains 4,000 transaction data.
[0075] Parallel processing leverages the performance of multi-core CPUs to reduce the processing time of a single task from 60 seconds / 4000 transactions in the traditional single-threaded model to 10 seconds / 4000 transactions, improving overall efficiency by 6 times and solving the problem of low processing efficiency. The dynamic sharding strategy can be adjusted according to the system resource status and changes in data volume to avoid resource waste or processing timeouts and enhance the system's ability to cope with business fluctuations.
[0076] Task Execution and Fault Tolerance Module 3: Implementation Details: During the reconciliation process, exceptions such as "data format inconsistency" (error code: DATA_FORMAT_ERR) and "transaction amount discrepancy" (error code: AMOUNT_MISMATCH) are captured. For "data format inconsistency," if it is a repairable format issue (such as an incorrect date format), two retryes are preset, with a 3-minute interval between each attempt to repair the data before reconciliation. For "transaction amount discrepancy," it is marked as "exception skipped," a detailed exception report is generated, and an email alert is sent to relevant business departments and the operations team via Prometheus + Alertmanager, recording detailed information such as the transaction serial number and amount difference involved in the exception. Simultaneously, the operator of each exception is recorded (automatically marked as "BATCH_JOB_002" by the system), the job status (retrying / skipped), and a data summary (e.g., "Transaction serial number: 202506150001, amount due: 5000 yuan, actual amount received: 4500 yuan"). When the subtask is executed, 4000 transactions are submitted in batches for reconciliation using Spring Batch's Chunk-orientedProcessing. Transaction control is enabled at the database layer to ensure that only the data of the current shard is rolled back in case of failure.
[0077] Execution steps: The subtask executes batch reconciliation requests. If data format issues still exist after the second retry, the record is marked as "pending manual processing" and skipped to continue with subsequent tasks.
[0078] Fine-grained anomaly handling avoids a full task restart. For example, if 5 out of 4000 transactions in a certain segment have discrepancies in amount, only these 5 transactions are skipped and marked as pending processing. Compared with the traditional method of retrying from the beginning, this saves a lot of computing resources and solves the problem of poor fault tolerance. Automatic alarms and detailed audit logs realize closed-loop management of anomalies, which facilitates quick location and resolution of problems and meets the requirements of financial systems for data accuracy and traceability.
[0079] Monitoring and logging module 4: Implementation details: A visual interface is provided to display the real-time progress of reconciliation tasks (e.g., "1.5 million transactions reconciled, discrepancy rate 0.1%"), resource usage (CPU: 65%, memory: 60%), and abnormal trends (100 discrepancies today, mainly due to amount discrepancies). Log queries are supported by combining "batch number (20250615001)," "data source institution (Institution A)," and "task status (FAILED)," with results returned within 10 seconds. Task logs are synchronized to the bank's unified operations and maintenance system, supporting the review of historical reconciliation task data within the past 6 months and generating compliant audit reports to meet regulatory requirements for data auditing.
[0080] Execution steps: After every 10,000 reconciliations are completed, real-time data is pushed to the monitoring module and the dashboard is updated; business personnel and maintenance personnel can use the bank's internal system to enter "batch number + data source institution" to query the specific details of the reconciliation discrepancies of a certain institution (such as "transaction serial number: 202506150005, the amount recorded by the other institution differs from the amount recorded by our bank by 500 yuan").
[0081] Real-time monitoring allows for timely tracking of reconciliation task progress, reducing anomaly location time from hours to minutes and lowering maintenance costs; compliance audit logs meet regulatory requirements, providing complete reconciliation records and anomaly handling information quickly during regulatory inspections.
[0082] Performance optimization module: Implementation details: A composite index (data_source, transaction_id, reconciliation_status) was created for the "Reconciliation Record Table," improving query efficiency by 35%; the batch submission parameter was set to fetchSize=3000 to reduce the number of database interactions. For the backend storage system, a "rack awareness + dynamic node performance evaluation + access heat scheduling" strategy was adopted, storing frequently accessed recent reconciliation data on high-performance nodes. The thread pool was configured with a core thread count of 15, a maximum thread count of 40, and a queue capacity of 80 to adapt to the load characteristics of batch reconciliation tasks.
[0083] Execution steps: During the data reading phase, appropriate nodes are selected from the storage system based on access frequency and node performance to load transaction data in parallel; during database writing, 3,000 reconciliation result records are written at once using a batch commit strategy to reduce network overhead.
[0084] Storage scheduling strategies improve data access efficiency by 25%, reduce database operation time from 1.5 hours to 1 hour, and control the overall reconciliation task processing time within 2.5 hours, meeting timeliness requirements; thread pool and batch submission optimizations avoid resource contention and support system throughput increases from 3,000 transactions / minute to 8,000 transactions / minute, coping with data volume growth.
[0085] Security mechanism module: Implementation details: Transaction data is encrypted using AES-256 before transmission. Sensitive fields such as "transaction amount" and "customer account" are anonymized when stored in the database (e.g., displayed as "5000 yuan" or "622812**"). The task processing chain uses TLS 1.3 encryption, and digital certificates are verified during client-server communication to prevent data tampering. Permissions are assigned based on roles: "Reconciliation Administrators" can initiate reconciliation tasks and view detailed reconciliation results, while "Regular Business Personnel" can only view summary reconciliation information, and any operation requires dual approval (whitelist + work order system).
[0086] Execution steps: When transaction data is transmitted from external institutions or internal systems to the batch processing system, a TLS encrypted channel is automatically triggered; when administrators log in, they need to use SMS verification code + fingerprint dual authentication to ensure that operation permissions are compliant.
[0087] Encryption and data masking prevent the leakage of sensitive data and meet data security requirements; access control prevents unauthorized operations, such as preventing ordinary business personnel from arbitrarily starting reconciliation tasks or viewing sensitive information, thus reducing operational risks.
[0088] Compared to Embodiment 1, this embodiment is specifically configured for the characteristics of batch reconciliation transactions, with targeted improvements in modules such as task scheduling, sharding, and fault tolerance. For example, the XXL-job scheduling framework is used to accommodate the timed and manual triggering requirements of reconciliation tasks, and sharding by data source institution and business type better suits the data organization format of reconciliation transactions. Both embodiments fully utilize the functions of each module in the system, achieving efficient processing, fault tolerance, monitoring, performance optimization, and security. In terms of processing efficiency, through task sharding and performance optimization, the processing time for reconciliation tasks is controlled within 2.5 hours, significantly improving processing speed, similar to payroll processing. Regarding fault tolerance, both employ fine-grained anomaly handling and automatic alarm mechanisms to ensure task reliability. In terms of security and compliance, both meet the stringent requirements of financial systems through encryption, data masking, and access control, further demonstrating the effectiveness and adaptability of this application in different business scenarios.
[0089] Example 3: Bank Batch Statement Generation Service: (a) Application scenarios: Banks are required to generate monthly statements for 1 million customers, which include all transaction details for the month. The statements must be generated and sent to the customer's designated email address within 24 hours, while also meeting data security and formatting requirements. This ensures that customers can clearly view their transaction records and complies with regulatory requirements for information disclosure by financial institutions.
[0090] (II) System Module Implementation and Methodological Steps: Task scheduling module 1: Implementation Details: Based on the Quartz scheduling framework, a batch statement generation task is triggered at 23:00 on the last day of each month. Manual triggering by the administrator is also supported in special circumstances (such as system data update delays requiring manual adjustment of the generation time). Spring Cloud Config is used for unified management of multi-environment configurations, with different email server addresses and ports configured for different environments. For example, the production environment uses prod-mail-server.com with port 25, while the test environment uses test-mail-server.com with port 2525. Task dependencies are configured to ensure that the statement generation task only begins after the preceding transaction data aggregation task succeeds. If the preceding task fails, the statement generation task is automatically added to the retry queue, with a maximum of 4 retries.
[0091] Execution steps: At 23:00 at the end of each month, the scheduling module triggers a task to load real-time configuration information from the configuration center, including email sending server configuration, statement template path, etc.
[0092] It automates scheduling, reduces human error, and lowers maintenance costs; its centralized scheduling and distributed collaboration capabilities can adapt to task scheduling needs in different environments and improve system scalability.
[0093] Task Sharding Module 2: Implementation details: Customers are sharded according to their ID range, with 1 million customers divided into 1000 shards, each containing 1000 customers. System resources are monitored in real time (e.g., disk I / O utilization > 85%). If resources are strained, the number of shards is automatically adjusted to 1200, distributed across 25 servers for parallel processing, with each server configured with 3 processing threads. Simultaneously, sharding is dynamically adjusted based on customer activity (e.g., number of transactions in the past month), prioritizing shards with a high concentration of active customers on higher-performance servers.
[0094] Execution steps: The sharding module generates subtasks based on the customer ID range, and each subtask is responsible for generating account statements for 1000 customers.
[0095] Parallel processing fully utilizes the performance of multi-core CPUs, reducing the processing time for a single task from 10 minutes / 1000 customers in the traditional single-threaded model to 2 minutes / 1000 customers, improving overall efficiency by 5 times and solving the problem of low processing efficiency; the dynamic sharding strategy, combined with resource monitoring and customer activity, effectively avoids resource bottlenecks and adapts to fluctuations in business volume.
[0096] Task Execution and Fault Tolerance Module 3: Implementation Details: During the statement generation process, exceptions such as "Data Missing" (error code: DATA_MISSING) and "Format Conversion Failure" (error code: FORMAT_ERR) are captured. For "Data Missing," if the missing data can be obtained from the backup system, a default of 3 retries are performed, each with a 10-minute interval, regenerating the statement after retrieving the data. For "Format Conversion Failure," it is marked as "Exception Skipped," an exception report is generated, and an instant messaging alert is sent to the operations team via Prometheus + Alertmanager, recording detailed information such as the range of customer IDs involved in the exception and the error type. The system records the operator for each exception (automatically marked as "BATCH_JOB_003"), job status (retrying / skipped), and data summary (e.g., "Customer ID range: 10000-10999, 10 missing transaction records"). During subtask execution, Spring Batch's Chunk-oriented Processing batch processes data from 1000 customers, with transaction control enabled at the database layer to ensure that only the current shard data is rolled back in case of failure.
[0097] Execution steps: The subtask executes batch generation of reconciliation statement requests. If data loss issues still exist after the third retry, the shard is marked as "pending manual processing" and skipped to continue subsequent tasks.
[0098] Fine-grained anomaly handling avoids restarting all tasks. For example, if 10 out of 1000 customers in a certain shard fail to generate data due to missing data, only the tasks of these 10 customers are retried. Compared with the traditional method of retrying from the beginning, this saves a lot of computing resources and solves the problem of poor fault tolerance. Automatic alarms and detailed audit logs realize closed-loop management of anomalies, which facilitates quick location and resolution of problems and meets the requirements of financial systems for data accuracy and traceability.
[0099] Monitoring and logging module 4: Implementation details: A visual interface is provided to display the real-time progress of statement generation tasks (e.g., "800,000 statements generated, success rate 98%), resource usage (CPU: 70%, memory: 55%, disk I / O: 75%), and abnormal trends (50 format conversion failures and 20 data loss incidents today). Log queries are supported by combining "batch number (20250731001)," "customer ID (6228****5678)," and "task status (FAILED)," with results returned within 10 seconds. Task logs are synchronized to the bank's unified operations and maintenance system, supporting the review of historical statement generation task data up to 12 months, generating compliant audit reports, and meeting regulatory requirements for data auditing.
[0100] Execution steps: After every 10,000 statements are generated, real-time data is pushed to the monitoring module and the dashboard is updated; the maintenance personnel can enter "batch number + customer ID" through the bank's internal system to query the specific reason for the failure of a customer's statement generation (e.g., "Customer ID: 6228****5678, data missing, still failed after 3 retries").
[0101] Real-time monitoring allows for timely tracking of task progress, reducing anomaly location time from hours to minutes and lowering maintenance costs; compliance audit logs meet regulatory requirements, providing complete task records and anomaly handling information quickly during regulatory inspections.
[0102] Performance optimization module: Implementation details: A composite index (customer_id, transaction_date, transaction_type) was created for the "Transaction Record Table," improving query efficiency by 45%; the batch submission parameter was set to fetchSize=2000 to reduce the number of database interactions. For the storage system, a "rack awareness + dynamic node performance evaluation + access heat scheduling" strategy was adopted to store frequently used statement templates and frequently accessed customer transaction data on high-performance nodes. The thread pool was configured with a core thread count of 12, a maximum thread count of 30, and a queue capacity of 60 to adapt to the load characteristics of batch statement generation tasks.
[0103] Execution steps: During the data reading phase, appropriate nodes are selected from the storage system based on access frequency and node performance to load transaction data and account statement templates in parallel; during database writing, 2000 account statement generation result records are written at once using a batch submission strategy to reduce network overhead.
[0104] Storage scheduling strategies improve data access efficiency by 30%, reduce database operation time from 12 hours to 8 hours, and control the overall statement generation task processing time within 20 hours to meet timeliness requirements; thread pool and batch submission optimizations avoid resource contention and support system throughput increases from 2,000 documents / minute to 5,000 documents / minute to cope with data volume growth.
[0105] Security mechanism module: Implementation details: Statement data is encrypted using AES-256 before transmission. When stored in the database, sensitive fields such as "transaction amount" and "customer ID number" are anonymized (e.g., displayed as "8000 yuan" or "32011234567890**"). The task processing chain uses TLS 1.3 encryption, and digital certificates are verified during client-server communication to prevent data tampering. Permissions are assigned based on roles: "Statement Generation Administrator" can initiate generation tasks and view detailed task progress, while "Regular Employees" can only view summary information, and operations require dual approval (whitelist + work order system).
[0106] Execution steps: When the statement data is read from the database to the generation system, the TLS encrypted channel is automatically triggered; when the administrator logs in, they need to use SMS verification code + fingerprint dual authentication to ensure that the operation permissions are compliant.
[0107] Encryption and data masking prevent the leakage of sensitive data and meet data security requirements; access control prevents unauthorized operations, such as preventing ordinary employees from arbitrarily starting statement generation tasks or viewing sensitive information, thus reducing operational risks.
[0108] Compared to Embodiment 1 (batch payroll processing) and Embodiment 2 (batch reconciliation processing), Embodiment 3 is adapted to the characteristics of batch reconciliation statement generation. In task scheduling, it is configured based on the time characteristics at the end of each month and potential special circumstances; task sharding is based on customer ID ranges and combined with customer activity levels, a sharding method different from the previous two embodiments, which better meets the needs of reconciliation statement generation. However, it maintains consistency in overall architecture and core functions, achieving automation through task scheduling, improving processing efficiency through task sharding, ensuring task reliability through fault tolerance module 3, performing real-time monitoring and auditing through monitoring and logging module 4, improving system performance through performance optimization module, and ensuring data security through security mechanism module. These three embodiments fully demonstrate the effectiveness and adaptability of this application in different banking business scenarios, effectively solving the problems of efficiency, fault tolerance, scalability, maintenance, and security of traditional batch processing systems.
[0109] The above description is merely a preferred embodiment of the present invention. It should be noted that for those skilled in the art, other parts not specifically described are existing technology or common knowledge. Several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A bank batch processing system, characterized in that, include: The task scheduling module is configured to automatically trigger batch processing tasks according to a preset scheduling strategy, and integrates a unified configuration platform for dynamic registration and distribution of batch tasks. The task sharding module is configured to divide batch processing tasks into multiple subtasks for parallel processing according to the sharding strategy, and dynamically adjust the number of shards based on system resources. The task execution and fault tolerance module is configured to capture task execution exceptions, retry or skip tasks with execution exceptions based on preset exception execution strategies, and integrate exception log analysis and automatic alarm notification functions. The monitoring and logging module is configured to monitor the task execution status in real time and record log information, providing a visual interface and connecting to the bank's operation and maintenance platform for real-time observability, alarms and backtracking. The performance optimization module is configured to optimize task execution performance, including thread pool parameter tuning, database index strategy optimization, batch submission control strategy, and scheduling strategy for heterogeneous storage environments. The scheduling strategy combines rack awareness, node dynamic performance evaluation, and access popularity scheduling.
2. The bank batch processing system according to claim 1, characterized in that, The task scheduling module is based on the Quartz or xxl-job scheduling framework, supports centralized scheduling and distributed job collaboration, is compatible with multi-environment deployments in banks (Dev / Test / Prod), and uses Spring Cloud Config for unified management of the configuration center.
3. A bank batch processing system according to claim 1, characterized in that, The scheduling strategy is timed scheduling and / or event-triggered scheduling.
4. A bank batch processing system according to claim 1, characterized in that, The sharding strategy includes configurable sharding strategies based on data volume, time window, and customer dimensions.
5. A bank batch processing system according to claim 1, characterized in that, The task execution and fault tolerance module integrates the Prometheus+Alertmanager monitoring and alarm platform, which provides real-time notifications for abnormal states and records the operator, job status, and data summary.
6. A bank batch processing system according to claim 1, characterized in that, The preset exception execution strategy includes retry limit and exception skipping rules.
7. A bank batch processing system according to claim 1, characterized in that, The monitoring and logging module supports quick retrieval based on task, batch number, and user ID. The logging system contains batch-level operation records for data backtracking and compliance auditing.
8. A bank batch processing system according to claim 1, characterized in that, The dimensions of the log information include task, batch number, and user ID.
9. A bank batch processing system according to claim 1, characterized in that, It also includes a security mechanism module, configured to encrypt and store task data, transmit and process link data using TLS secure encryption, configure job operation permissions based on user roles, and support whitelist control and approval mechanisms.
10. A method for operating a bank batch processing system as described in claim 1, characterized in that, include: The scheduling module automatically triggers batch processing tasks based on timed or event-triggered strategies, supporting dynamic job configuration hot loading and multi-environment deployment adaptation. The task is divided into multiple subtasks by the sharding module, and parallel processing is carried out based on configurable strategies of data volume, time window and customer dimension, and the number of shards is dynamically adjusted. Subtasks are executed in parallel through the task execution and fault tolerance module, and execution exceptions are captured in real time and recorded in detail. Exception tasks are retried or skipped. The monitoring and logging modules monitor task status in real time, record traceable audit logs, and connect to the bank's operation and maintenance platform to achieve visualized monitoring and alarms. The performance optimization module is used for thread pool parameter tuning, database index optimization, batch commit control, rack awareness, node dynamic performance evaluation, and access heat scheduling strategies in heterogeneous storage environments.
Citation Information
Cited By
Parallel computing performance detection method and device, storage medium and computing equipment
CN121785716A