File co-processing method and system based on distributed multithreading and dynamic sub-library
By employing a distributed multithreaded and dynamically sharded file collaborative processing method, the performance bottleneck and resource waste issues in massive file processing are resolved, resulting in an efficient and reliable file processing system.
Patent Information
- Application Number
- CN202510967936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies suffer from performance bottlenecks, resource scheduling imbalances, lack of dynamic scheduling mechanisms, weak transaction control, and difficulty in elastic scaling when processing massive amounts of files, resulting in low processing efficiency and resource waste.
A file collaborative processing method based on distributed multithreading and dynamic database sharding is adopted. Through a multithreaded hierarchical parallel model, database sharding and table partitioning mechanism, dynamic routing and two-stage state machine, efficient processing of file data and optimized resource allocation are achieved.
It achieves an exponential improvement in processing efficiency, supports high concurrency and high availability, reduces operation and maintenance costs, and ensures data integrity and system stability.
Smart Images

Figure CN121092508A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of file data processing, in particular to a file collaborative processing method and system based on distributed multi-threading and dynamic sub-library. BACKGROUND
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute the prior art.
[0003] Under the background of rapid development of digital finance, massive file processing has become an important support link of core business systems. Typical application scenarios include cross-institutional transaction reconciliation files (TB-level daily processing volume), real-time clearing instruction files (millisecond-level delay requirement), regulatory reporting data files (strong format compliance verification), and customer asset proof files (sensitive data encryption processing), etc.
[0004] The traditional method mainly uses a single-thread-based file processing mechanism, but the single-thread file processing mechanism presents multi-dimensional technical limitations in the new business environment, as follows: (1) Performance bottleneck aggravation: The single-thread sequential processing mode presents non-linear growth in processing time when facing composite files (such as nested compressed packages containing CSV, XML multi-format files). Actual measurement data shows that a single thread needs more than 50 minutes to parse a 10GB transaction file, resulting in more than 30% processing backlog risk in the gold trading period, which seriously affects the execution of time-sensitive businesses.
[0005] (2) Resource scheduling imbalance: The resource scheduling mechanism of the existing file processing system has significant defects, mainly manifested in the periodic idling of computing and I / O resources. The typical processing flow contains four stages: file download (network I / O intensive), format parsing (CPU intensive), data verification (memory intensive), and persistent storage (disk I / O intensive). When single-threaded, the resource demand of each stage presents obvious fluctuations: the memory occupancy peaks at 32GB in the data verification stage, while the CPU utilization is less than 15%; in the persistent stage, the disk throughput reaches 800MB / s, but the CPU utilization is only 8%. This resource mismatch limits the overall throughput, and the comprehensive utilization rate of resources is less than 35% for a long time. More seriously, the traditional scheme lacks intelligent scheduling capability and cannot dynamically allocate resources according to task characteristics, and fails to use GPU acceleration for computing-intensive encrypted files, still relying on CPU software implementation.
[0006] (3) Dynamic scheduling mechanism is missing: the multi-source heterogeneous characteristics of financial business put forward higher requirements for task scheduling. The typical system needs to connect 20+ data channels such as SWIFT, UnionPay, and Netlink. Each channel has differentiated characteristics: SWIFT files need to be processed first (regulatory compliance requirements), UnionPay files need to be processed separately, and third-party payment files need to be dynamically routed to different levels of computing nodes. The traditional static hash distribution strategy leads to serious load imbalance: files are concentrated in 2 processing nodes due to hash collision, while the other 8 nodes are in idle state, and the overall resource waste rate is 64%. In addition, the lack of health state awareness of service instances may cause batch file processing to stop and fail if the fault node is not isolated in time.
[0007] (4) Weak transaction control: financial file processing involves multi-stage state coordination and needs to meet the ACID transaction requirements. The typical process includes file download, data parsing, business verification, formal storage, and state rewriting. The traditional solution lacks distributed transaction control, which may cause a large number of transaction records to be retained in temporary tables for several hours due to node failure, causing a large difference between the core system and the accounting system. In addition, network interruption may occur during the commit phase of updating the business table, causing inconsistent record states, which requires manual intervention for repair. The transaction rollback mechanism of the existing solution is also not perfect, and only database operations can be rolled back when an exception occurs, without synchronously cleaning up temporary files on the OSS.
[0008] (5) Difficult to expand: the vertical expansion mode cannot adapt to the explosive growth of business volume. Single-node architecture cannot break through the 1000TPS throughput limit, and expansion in a multi-cloud environment is not met. In addition, the existing solution lacks fine-grained scaling capability, and needs to maintain a hardware size of 3 times the daily demand throughout the year, resulting in huge operation and maintenance costs. SUMMARY
[0009] To solve the above problems, the present disclosure proposes a file collaborative processing method and system based on distributed multi-threading and dynamic database splitting, which introduces dynamic routing and transaction control through a multi-thread hierarchical parallel model and a database and table splitting mechanism, uses a two-stage state machine and a TCC compensation transaction model to ensure atomicity of cross-database operations, and builds a high-performance, high-availability, and easy-to-maintain file processing system to achieve exponential improvement in processing efficiency.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions: A file collaborative processing method based on distributed multi-threading and dynamic database splitting, comprising: Scheduling parameter initialization, parameter parsing and verification, obtaining files to be processed, and generating a file list according to batch information; Adopting the producer-consumer model, the producer main thread scans the file list through the storage engine API to generate a download task queue, and the consumer thread pool concurrently executes the download task, and the maximum concurrency is controlled through a semaphore; Based on the sharding routing algorithm, a hash value is generated to uniformly distribute file data to multiple physical storage nodes, and a dynamic routing table is used to manage real-time updates of the mapping relationship between shards and instances, resources are preferentially allocated to high-priority file tasks based on a business routing strategy, and tasks are dynamically allocated to specified database shards and processing nodes. Through a three-level parallel model, multi-file parallel processing of file parsing, data cleaning and warehousing on the processing node is realized, dynamic Feign calling is performed on dynamic routing and weight load of service instances, and multi-state table state synchronization updating is performed based on a two-stage state machine, and processing logs are saved.
[0011] According to some embodiments, the present disclosure adopts the technical scheme as follows: A file collaborative processing system based on distributed multi-threading and dynamic database splitting, comprising: An initialization module for scheduling parameter initialization, parameter parsing and verification, obtaining files to be processed, and generating a file list according to batch information; A file scanning and downloading module for adopting the producer-consumer model, the producer main thread scanning the file list through the storage engine API to generate a download task queue, and the consumer thread pool concurrently executing the download task, and the maximum concurrency being controlled through a semaphore; A data distribution module for generating a hash value based on the sharding routing algorithm to uniformly distribute file data to multiple physical storage nodes, and using a dynamic routing table to manage real-time updates of the mapping relationship between shards and instances, resources being preferentially allocated to high-priority file tasks based on a business routing strategy, and tasks being dynamically allocated to specified database shards and processing nodes; A multi-thread processing and calling module for realizing multi-file parallel processing of file parsing, data cleaning and warehousing on the processing node through a three-level parallel model, and performing dynamic Feign calling on dynamic routing and weight load of service instances; A state management module for performing multi-state table state synchronization updating based on a two-stage state machine, and saving processing logs.
[0012] According to some embodiments, the present disclosure adopts the technical scheme as follows: A computer program product comprising a computer program, which, when executed by a processor, implements the file collaborative processing method based on distributed multi-threading and dynamic database splitting.
[0013] According to some embodiments, the present disclosure adopts the technical scheme as follows: A non-transitory computer readable storage medium for storing computer instructions, which, when executed by a processor, implement the file collaborative processing method based on distributed multi-threading and dynamic library splitting.
[0014] According to some embodiments, the present disclosure adopts the technical solutions as follows: An electronic device, comprising a processor, a memory and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the file collaborative processing method based on distributed multi-threading and dynamic library splitting.
[0015] Compared with the prior art, the present disclosure has the beneficial effects that: The file collaborative processing method based on distributed multi-threading and dynamic library splitting of the present disclosure realizes exponential improvement of processing efficiency through a multi-thread hierarchical parallel model (file level, shard level, record level) and a library and table splitting mechanism. In a stress test, the time consumed for processing a 10GB XML file by a single node is reduced from 58 minutes in the traditional scheme to 9 minutes, and the throughput reaches 680MB / s (5.6 times higher). The library and table splitting design supports horizontal expansion to 128 physical shards, can bear write requests of 200 million transaction records (about 40TB data) per day, and the processing capacity grows linearly.
[0016] The file collaborative processing method based on distributed multi-threading and dynamic library splitting of the present disclosure realizes millisecond-level dynamic selection of service instances based on shard routing algorithm and dynamic routing table management. In the gray release scenario, 5% of the traffic can be accurately controlled to the new version instance, reducing the business risk. Multi-tenant isolation is realized through K8s namespace and database Schema level permission control. The adaptive routing strategy can identify the business type (AM / BM) and automatically switch the processing link: AM type high priority task is directly connected to the SSD storage cluster, and BM type batch task is routed to the elastic computing pool.
[0017] The file collaborative processing method based on distributed multi-threading and dynamic library splitting of the present disclosure adopts a two-stage state machine (preprocessing + submission) and a TCC compensation transaction model to ensure the atomicity of cross-library operations. In the production environment, the problem of 2400 million account differences caused by network jitter is successfully solved, and the transaction rollback time is shortened from 2 hours of manual intervention to 8 seconds of automatic processing. The global state table realizes high-concurrency update through optimistic locking, and the actual test supports 3000+TPS concurrent write. The automatic cleaning module scans the temporary table through a timing task to release the storage space occupied by invalid data.
[0018] The file collaborative processing method based on distributed multithreading and dynamic database partitioning of the present disclosure prevents dirty data from being stored in the database by checking the data integrity of the ok file. The method supports automatic anomaly detection and alarm, and the failed files are automatically sent to the isolation area. The robot pushes the alarm in real time, and the response time of operation and maintenance is shortened from 30 minutes to 1 minute. The intelligent retry strategy combined with the exponential backoff algorithm improves the self-healing rate of temporary network errors to 92%.
[0019] The file collaborative processing method based on distributed multithreading and dynamic database partitioning of the present disclosure realizes dynamic adjustment of 23 core parameters through the Apollo configuration center, including thread pool size (default 16, expandable to 1024), number of shards (online expansion to 256), and JDBC batch size (flexible setting of 50-1000). Using this function, different strategies are implemented for different file types: large files use small batches and high concurrency, and the processing efficiency is improved by 38%; small files use large batches and low consumption, and the CPU utilization rate is reduced by 27%. The strategy hot-activated capability (without restarting the service) enables operation and maintenance personnel to complete business optimization within 5 minutes, and the fault recovery efficiency is improved by 6 times.
[0020] The file collaborative processing method based on distributed multithreading and dynamic database partitioning of the present disclosure reconstructs the traditional file processing mode, introduces core technologies such as multithreading, database partitioning, dynamic routing, and transaction control, and builds a high-performance, high-availability, and easy-to-maintain file processing system. Not only does it solve the performance bottleneck of the current system, but it also provides a good architectural foundation for future business expansion. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings, which form a part of the present disclosure, are used to provide further understanding of the present disclosure, and the schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure, and do not constitute improper limitations on the present disclosure.
[0022] Figure 1 The flowchart of the file collaborative processing method based on distributed multithreading and dynamic database partitioning of the present disclosure embodiment; Figure 2 The schematic diagram of the producer and consumer executing the task process of the present disclosure embodiment; Figure 3 The process diagram of the task parallel processing process of the present disclosure embodiment; Figure 4 The process diagram of the present disclosure embodiment based on the dynamic allocation of data to the specified database shard and processing node according to the business rules, realizing load balancing and resource isolation. DETAILED DESCRIPTION
[0023] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is illustrative only, and is intended to provide further description of the disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0026] Embodiment 1 In one embodiment of the present disclosure, a distributed multi-thread and dynamic database-based file collaborative processing method is provided, comprising the following steps: Step 1: Initialize the scheduling parameters, perform parameter analysis and verification, obtain the files to be processed, and generate a file list according to the batch information; Step 2: Use the producer-consumer model, the producer main thread scans the file list through the storage engine API to generate a download task queue, the consumer thread pool concurrently executes the download task, and the maximum number of concurrent tasks is controlled by the semaphore; Step 3: Generate a hash value based on the sharding routing algorithm, uniformly distribute the file data to multiple physical storage nodes, and use the dynamic routing table to manage the real-time update of the mapping relationship between shards and instances, allocate resources to high-priority file tasks based on the business routing strategy, and dynamically allocate tasks to specified database shards and processing nodes; Step 4: Implement multi-file parallel processing of file parsing, data cleaning and warehousing on the processing node through a three-level parallel model, dynamically route and weight the service instances through Feign calling, and update the state synchronization of the multi-state table based on a two-stage state machine, and save the processing log.
[0027] As an embodiment, the distributed multi-thread and dynamic database-based file collaborative processing method of the present disclosure introduces core technologies such as multi-threading, database and table splitting, dynamic routing, and transaction control, and builds a high-performance, high-availability, and easy-maintenance file processing method. The specific implementation process is as follows: Step 1: Initialize the scheduling parameters, perform parameter analysis and verification, obtain the files to be processed, and generate a file list according to the batch information; Specifically, the key parameters such as account set, date, and path are extracted from the scheduling context and subjected to legality verification. The file list to be processed is obtained through a remote file server, and the batch information is written into a control table.
[0028] Step 2: Using the producer-consumer model, the producer main thread scans the file list through the storage engine API to generate a download task queue, and the consumer thread pool concurrently executes the download tasks, with the maximum number of concurrent threads controlled by a semaphore. The purpose is to efficiently and concurrently download target files and corresponding ok check files from multi-cloud object storage, ensuring data integrity and transmission efficiency. This includes: building a multi-level task queue, using the producer-consumer model, and the main thread scanning the file list (supporting wildcard matching) through the storage engine API to generate a download task queue; the consumer thread pool concurrently executes the download tasks, with the maximum number of concurrent threads controlled by a semaphore to avoid resource exhaustion. In the multi-level task queue, the core of controlling the maximum number of concurrent threads through the semaphore (Semaphore) is to coordinate the producer and consumer threads, ensuring that the number of threads executing tasks simultaneously does not exceed the preset threshold. The specific implementation steps are as follows: Step 21: Initialize the semaphore, including: create a Semaphore object and specify the maximum number of concurrent threads.
[0029] For example, Semaphore semaphore = new Semaphore(MAX_CONCURRENT_THREADS); This value represents the upper limit of the number of threads that can execute download tasks simultaneously, preventing resource overload.
[0030] Step 22: The producer thread generates tasks, including: the main thread scans the file list (supports wildcard matching) through the storage engine API, encapsulates the metadata of the files that meet the conditions into download tasks, and puts them into the task queue. At this time, the task queue is only responsible for storage and does not involve concurrent control.
[0031] Step 23: The consumer thread executes the task, including: (1) Acquire semaphore permission: after the consumer thread obtains the task from the queue, it immediately calls . If the current number of concurrent threads has not reached the upper limit, the semaphore counter decreases by 1, and the thread continues to execute; if it has reached the upper limit, the thread is blocked and waits.
[0032] (2) Execute the download task: after the thread acquires permission, it performs file download, processing, and other operations. This phase needs to ensure that the task logic does not block other threads.
[0033] (3) Release semaphore permission: after the task is completed, whether it is abnormal or not, it must call in the finally block, the counter is incremented by 1, and the waiting thread is awakened.
[0034] Step 24: Dynamic management of concurrency, including: semaphore automatically adjusts concurrency state through counter. When task execution time varies, the timing difference of releasing permissions naturally forms dynamic scheduling. For example, if 10 threads are executed simultaneously, subsequent threads will be blocked until a thread releases permission, ensuring that system resources are always within a controllable range.
[0035] Step 25: Exception handling and resource recovery, including: if an exception is thrown during task execution, the release mechanism in the finally block can prevent the semaphore from being permanently occupied. At the same time, the thread pool needs to be configured with a reasonable rejection policy to avoid task queue overflow.
[0036] Step 26: Performance optimization and monitoring, including: extensible semaphore logic records current concurrency, blocked thread count, and other statistical information, which can be adjusted in real time through monitoring tools to adapt to dynamic load changes.
[0037] As an embodiment, it can support resume and block download, and for large files (>1GB) request block download (each block 256MB), download progress is persisted to Redis, and after an exception interruption, it can be recovered from the breakpoint, and the actual test resume and block download success rate is improved to 99.7%.
[0038] Further, integrity verification is performed, and immediately after downloading, the file SHA-256 hash value is calculated and compared with the verification code in the ok file. If it fails, it triggers automatic retry (up to 3 times, configurable), and if it still fails, it is marked as "verification exception" and notifies the operation and maintenance platform.
[0039] As an embodiment, it supports multi-cloud compatible design, supports multi-cloud vendor adaptation through abstract storage interface, and the actual test single node download throughput reaches 2.1Gbps (16 threads), which is 8 times higher than single thread.
[0040] Step 3: Generate hash value based on sharding routing algorithm, evenly distribute file data to multiple physical storage nodes, and use dynamic routing table to manage real-time update of sharding and instance mapping relationship, based on business routing strategy, preferentially allocate resources to high priority file tasks, dynamically allocate tasks to specified database shards and processing nodes, the purpose is to dynamically allocate data to specified database shards and processing nodes based on business rules, achieve load balancing and resource isolation. The specific implementation process is as follows: Step 31: The core goal of the sharding routing algorithm is to evenly distribute data to multiple physical storage nodes, while supporting smooth data migration during online expansion. By applying a hash algorithm to the loan number to generate a 32-bit hash value, it is mapped to a physical library by taking the modulus of the number of shards, and supports smooth migration of data during online expansion through a consistent hash ring (read-write compatible during migration). The specific implementation steps are as follows: First, hash each loan number. Select a high-dispersity hash algorithm (MD5 or improved hash function) to generate a 32-bit hash value. This hash value is used as the key basis for subsequent routing, and needs to ensure that the hash results of different loan numbers are as evenly distributed as possible to avoid hot spot problems.
[0041] Next, according to the total number of shards of the current system (configured as 128 shards when initially deployed), the hash value is subjected to a modulo operation (hash value % total number of shards) to obtain the target shard ID. This shard ID is directly mapped to the corresponding physical database instance, completing the routing of data writing or querying. When online expansion, the consistent hashing ring mechanism is used to achieve smooth migration. When a new node is added, it is assigned a virtual node (100 virtual nodes for each physical node), and the virtual nodes are evenly distributed on the hash ring through hash calculation. The original data shards find the nearest virtual node clockwise according to the hash value to determine the home instance. When an instance is expanded or goes offline, only the shards in the virtual node interval managed by it are affected, and other shards are not affected. When node A goes offline, the shards it manages will automatically migrate to the next surviving node in the clockwise direction. During the migration process, the double-write mechanism (writing to the original node and the target node at the same time) is used to ensure data consistency, and after the migration is completed, the routing is switched to achieve read-write compatibility.
[0042] In addition, to support read-write compatibility during migration, a shard state machine needs to be maintained: when the shard is marked as "in migration", read requests can query both the original node and the target node, and take the latest version of the data; write requests are only written to the target node. After the migration is completed, update the routing table and notify the relevant components to finally complete the transfer of shard ownership. The entire process is executed through asynchronous background tasks to avoid affecting online business traffic.
[0043] Step 32: Maintain shard-instance mapping relationship through ZooKeeper, automatically update routing table when instance goes online / offline, support traffic distribution in proportion in gray release scenario. Dynamic routing table management realizes real-time update of shard-instance mapping relationship through ZooKeeper, ensuring that routing information is automatically synchronized when system topology changes. The specific implementation process is as follows: Create a hierarchical data model in ZooKeeper, for example Each node stores a list of instances (IP+port) under the corresponding shard ID. When an instance is online, it creates a temporary node through the ZooKeeper client and registers a listener. When it is offline, the node is automatically deleted, triggering an update to the global routing table. In a gray release scenario, each instance is configured with a weight (new version instances are set to 30%, and old version instances are set to 70%). ZooKeeper generates a random number according to the weight ratio to dynamically allocate traffic. When a request arrives, the routing component calculates the weight interval of the request allocation (0-30% for new instances and 31%-100% for old instances), achieving gradual gray switching. In addition, a version control mechanism is introduced: a globally unique version number is generated each time the routing table is changed, and each client periodically polls ZooKeeper to obtain the latest version. When a version update is detected, the local cache is immediately refreshed to ensure that all requests are routed based on the latest topology information. For instance failures, the ZooKeeper session timeout mechanism is used to automatically detect node abnormalities, trigger the routing table to exclude invalid instances, and redistribute the shards responsible for the instances to other surviving nodes, avoiding the impact of single-point failures on system availability.
[0044] Step 33: The business routing strategy is based on file header metadata analysis, achieving differentiated processing of different business types and ensuring that resources are allocated to high-priority tasks first. The specific implementation process is as follows: When uploading a file, the business type field (such as AM / BM) in the header metadata is parsed. For AM files that require low-latency processing, they are routed to a high-priority Kafka queue. Here, the Kafka partition strategy is configured to write AM messages to a dedicated partition and set higher consumption priority and lower latency thresholds. Using Kafka's priority partition allocation algorithm ensures that AM messages are prioritized for consumption at the Broker end, while the consumer end uses an exclusive thread pool for processing to avoid resource competition with other types of messages. For BM files that are suitable for batch processing, they are distributed to an elastic computing cluster. The cluster management module dynamically adjusts resource allocation based on current load conditions: a small number of computing nodes are started to save costs during light loads, and instances are automatically scaled during peak times. The task scheduling uses a "first come, first served + load balancing" strategy to distribute files to idle computing nodes and uses a distributed lock mechanism to prevent duplicate processing. After batch processing is complete, the result data can be further routed to persistent storage or downstream analysis systems.
[0045] In addition, a dynamic threshold adjustment mechanism is introduced: real-time monitoring of Kafka queue latency and computing cluster CPU utilization, automatic resource quota increase for related instances when AM queue latency exceeds a predetermined threshold, and automatic scaling process triggered when BM cluster load is too high. Through a feedback loop, routing decisions are optimized to ensure stable overall system performance.
[0046] As an embodiment, its performance guarantee is that the routing decision time consumption is controlled within 3ms (JVM internal cache routing table), the actual measured single node routing throughput reaches 12,000 TPS, and the error rate of sharding load balancing is less than 5%.
[0047] Step 4: Multi-file parallel processing of file parsing, data cleaning and warehousing on processing nodes is realized through a three-level parallel model; Specifically, the three-level parallel model realizes efficient processing of file parsing, data cleaning and warehousing, which adopts hierarchical design, including file-level parallelism, sharding-level parallelism and record-level batch submission. File-level parallelism processes multiple files by independent threads to avoid serial bottlenecks. Sharding-level parallelism splits a single large file into multiple Segments according to a fixed block size, and each Segment is parsed by a different thread. Record-level batch submission groups parsed data according to sharding rules, and each group uses JDBC Batch to insert in batches, combined with connection pool to reuse database connections.
[0048] As an embodiment, the implementation process of file-level parallel processing is as follows: (1) Task queue construction: the main thread obtains the list of files to be processed (such as / data / *.csv) through the file scanning module, encapsulates each file path as an independent task, and puts it into the shared task queue.
[0049] (2) Thread pool scheduling: create a fixed-size thread pool (ExecutorService pool = Executors.newFixedThreadPool(8)), and each thread gets a file task from the queue to execute the parsing process independently.
[0050] (3) Parallel execution: thread A processes file1.csv and thread B processes file2.csv, and each thread is independent of each other. If there are 10 files and 4 threads are processing 4 files simultaneously, the remaining tasks will wait for idle threads in the queue.
[0051] (4) Progress monitoring: real-time track processing progress through the getCompletedTaskCount() and getActiveCount() methods of the thread pool, and dynamically adjust the priority of subsequent tasks.
[0052] As an embodiment, the implementation process of sharding-level parallel processing is as follows: (1) File segmentation: For a single large file (GB level data), it is divided into multiple Segments according to a fixed block size (10MB). A 100MB file is divided into 10 Segments, and each Segment is parsed independently.
[0053] (2) Sub-task assignment: create parsing sub-tasks for each Segment and put them into the thread pool's dedicated queue. Thread T1 is responsible for Segment 1 and Segment 4 of file A, and thread T2 is responsible for Segment 2 and Segment 5.
[0054] (3) Parallel parsing: each thread uses RandomAccessFile or NIO's FileChannel to locate the start position of the Segment and asynchronously reads the data. Segment 1 thread parses data from 0-10MB, and Segment 2 parses data from 10-20MB.
[0055] (4) Data buffering and synchronization: the parsing results are temporarily stored in the thread's local buffer, and after all Segments are completed, the results are synchronized and merged through CountDownLatch or CyclicBarrier to ensure data integrity.
[0056] As an embodiment, the specific implementation process of record-level batch submission is as follows: (1) Data grouping: parsed records are grouped according to the sharding rule. For example, data with the same shard ID is placed in the same List <record>, reach threshold (2000 rows) trigger batch commit.
[0057] (2) Database connection reuse: Get database connection from connection pool, accumulate SQL statements using the addBatch() method of PreparedStatement.
[0058] (3) Asynchronous submission and exception handling: Submit the task as an asynchronous Future, allowing the thread to continue parsing other records. If the batch operation fails, capture the exception and split the failed records for single retry, without affecting other batches.
[0059] (4) Performance optimization: Dynamically adjust batch size, combined with the idle connection recycling mechanism of the connection pool, to avoid resource idling.
[0060] As an embodiment, the specific implementation and application process of the three-level parallel model is further explained, assuming that the system needs to process 100 files and the thread pool size is 16. The main thread distributes the file list to the thread pool, thread T1 processes file_001.csv, T2 processes file_002.csv,..., and T16 processes the respective files simultaneously. Each thread internally starts the shard-level parallel processing to achieve non-blocking between files.
[0061] file_001.csv is a 5GB large file, divided into 5MB segments. Thread pool 8 threads process Segment1-8 respectively, each thread calls the parseSegment(startByte, endByte) method, and after parsing, the result is written to the shared memory segment cache area, and finally assembled into full file data by the main thread.
[0062] Segment1 parses 10000 records, divided into 8 pieces according to the segmentation rule. Segment 0 contains 1200 records, segment 1 contains 1000 records,..., the thread inserts 1200 records of segment 0 in batch, and 1000 records of segment 1 are submitted in another batch, and finally 8 segments are inserted into the database in parallel.
[0063] Step 5: Dynamic Feign call for service instance dynamic routing and weight load, and multi-state table state synchronization update based on two-stage state machine, and save processing log.
[0064] Step 51: The function of dynamic Feign call is to dynamically select microservice instances according to business requirements, realize intelligent load balancing and fault tolerance of service call. It integrates Nacos registration center to obtain service instance metadata in real time, and realizes weighted round robin combined with Ribbon.
[0065] Route rules are issued through Apollo configuration center to support route condition configuration based on file feature transaction flow / metering flow. The specific content is as follows: (1) By introducing Nacos as a service registration and discovery component, each microservice instance automatically registers its own information (IP, port, etc.) with Nacos when starting, and the client can obtain the list of available service instances from Nacos. (2) By inheriting the AbstractLoadBalancerRule class, a weighted round robin strategy WeightedRoundRobinRule is defined to implement a weighted round robin algorithm, supporting scenarios with different instance processing capabilities. In runtime, the weight of an instance is dynamically adjusted based on monitoring data or health status.
[0066] @Resource(name = "weightWeightedRoundRobinRule") private WeightedRoundRobinRule weightedRule; / / Set the weight of an instance weightedRule.setServerWeight("xxx.xxx.xx.xx:8080", 5); (3) Real-time update of route rules by listening to Apollo configuration change events.
[0067] (4) Dynamically modify the service name serviceId of Feign call in runtime; private void updateDataStatus(List <string>list, String bathid, intindex) { String serviceId = getFeignClientServiceId(index, bathid); AmRestService amRestService = this.getAmRestService(serviceId); RPCResult <void>result = amRestService.saveCashflow(request); RpcResultUtil.checkRestScuccess(result); } private AmRestService getAmRestService(String serviceId) { / / Build FeignClientBuilder FeignClientBuilder feignClientBuilder = new FeignClientBuilder(applicationContext); / / Set interface class and service name FeignClientBuilder.Builder <amrestservice>builder = feignClientBuilder.forType(AmRestService.class, serviceId); / / Build service interface AmRestService amRestService = builder.build(); return amRestService; } (5) Dynamically route according to file processing logic, in the FileDealServiceImpl class, the getFeignClientServiceId method is responsible for determining the specific service ID according to the batch number and index.
[0068] private String getFeignClientServiceId(int index, String bathid) { if (bathid.contains("loan")) { return getBmServiceNames().get(index); / / BM service } else { return getAmServiceNames().get(index); / / AM service } }。
[0069] Step 52: The two-stage state machine includes preprocessing, submission confirmation, and consistency guarantee process, the specific content is as follows: (1) Preprocessing: write data to temporary table and mark as in progress state, write global intermediate table record operation serial number.
[0070] (2) Submission confirmation: after passing business verification, trigger asynchronous update of balance table and batch table through transaction synchronizer.
[0071] (3) Final consistency guarantee: for cross-database operations, use TCC mode (Try-Confirm-Cancel), and when an exception occurs, use a timing task to scan the temporary table to trigger reverse compensation. The TCC mode includes Try phase, Confirm phase and Cancel phase, and the specific implementation process of TCC mode is as follows: A. Try phase (resource reservation and inspection) After initiating the download file request, the transaction manager generates a global transaction ID. Each participant performs a lightweight resource reservation: the file storage service generates a temporary download credential and locks the file permission, the permission verification service records temporary authorization information, and the log service writes a try log. After all services successfully return, the transaction manager aggregates the results and triggers a global rollback if any service fails. This phase ensures resource availability, avoids subsequent operation conflicts, and reserves the necessary state for the subsequent phase.
[0072] B. Confirm phase (resource formal submission) The transaction manager asynchronously sends a Confirm instruction to the participants that successfully tried. The file storage service converts the temporary credential to a permanent one, releases the lock, and updates the download status; the permission service clears the temporary authorization and formally records the log; and the log service marks the transaction as complete. After all Confirm successes, the transaction manager commits the global transaction to ensure that the resources are formally effective. The asynchronous mechanism avoids blocking the main process and ensures consistency.
[0073] C. Cancel phase (resource rollback) When the Try phase fails, the Confirm times out, or the participant fails to execute, the transaction manager asynchronously triggers Cancel. Each service rolls back the reserved resources: the file storage service deletes the temporary credential and unlocks the file, the permission service restores the original permission state, and the log service records the rollback information. After completion, the global transaction is marked as failed. Through the compensation mechanism, the resource state is restored to the initial state in any abnormal situation, avoiding data inconsistency.
[0074] As an embodiment, the design ensures idempotency (global transaction ID prevents duplication), uses a reliable message queue to asynchronously notify Confirm / Cancel, and sets a resource timeout for automatic release. The monitoring system tracks the transaction state in real time, alarms or manually intervenes for hanging transactions, and implements exception handling and protection.
[0075] A fault-tolerant mechanism is designed, and the global state table adds a version number optimistic lock to avoid concurrent update conflicts and automatically clean up associated temporary data during transaction rollback.
[0076] Step 53: Set up an exception handling process to achieve full-link exception capture, automatic recovery, and precise tracing. This process is a hierarchical processing strategy that handles retryable exceptions (such as network jitter), data exceptions (such as field format errors), system-level failures (such as database downtime), full-link tracking, and data preservation.
[0077] (1) Retryable exceptions (such as network jitter): automatically retry 3 times (with exponential backoff intervals), and mark the task as "pending manual intervention" after retrying fails.
[0078] (2) Data anomaly (such as field format error): record original error data to isolation table, and notify data governance platform synchronously.
[0079] (3) System-level failure (such as database downtime): trigger global transaction rollback, clean up all temporary table data, and push alarm through Dingding robot.
[0080] (4) Full-link tracking: integrate SkyWalking to generate unique TraceID, connect all logs from file download to storage, and combine Grafana dashboard to monitor error rate in real time (threshold alarm > 1%).
[0081] (5) Data security: persist abnormal logs to Elasticsearch, retain for 180 days for audit traceability, and shorten average abnormal processing time from 2 hours manually to 8 minutes automatically.
[0082] Embodiment 2 In an embodiment of the present disclosure, a file collaborative processing system based on distributed multi-threading and dynamic database partitioning is provided, comprising: An initialization module is configured to schedule parameter initialization, perform parameter parsing and verification, obtain a file to be processed, and generate a file list according to batch information; A file scanning and downloading module is configured to use a producer-consumer model, in which a producer main thread scans the file list through a storage engine API to generate a download task queue, and a consumer thread pool concurrently executes the download tasks, and the maximum number of concurrent executions is controlled through a semaphore; A data distribution module is configured to generate a hash value based on a sharding routing algorithm, distribute file data uniformly to multiple physical storage nodes, and use a dynamic routing table to manage real-time updates of sharding and instance mapping relationships, preferentially allocate resources to high-priority file tasks based on a business routing strategy, and dynamically allocate tasks to specified database shards and processing nodes; A multi-thread processing and calling module is configured to implement multi-file parallel processing of file parsing, data cleaning, and storage on a processing node through a three-level parallel model, and dynamically route and weight load service instances through dynamic Feign calling; A state management module is configured to perform multi-state table state synchronization updates based on a two-stage state machine, and save processing logs.
[0083] Embodiment 3 In an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the file collaborative processing method based on distributed multi-threading and dynamic database partitioning.
[0084] Embodiment 4 In an embodiment of the present disclosure, a non-transitory computer readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the file collaborative processing method based on distributed multi-thread and dynamic sub-library.
[0085] Embodiment 5 In an embodiment of the present disclosure, an electronic device is provided, comprising a processor, a memory, and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the file collaborative processing method based on distributed multi-thread and dynamic sub-library.
[0086] The present disclosure is described with reference to the flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices generate a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus for performing the functions specified in one or more flows and / or blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device to cause a series of operational steps to be performed on the computer or other programmable data processing device to produce a computer-implemented process, so that the instructions executed by the computer or other programmable data processing device provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 a step for performing the functions specified in one or more flows and / or blocks.
[0088] The above description of the specific embodiments of the present disclosure in conjunction with the accompanying drawings is not a limitation on the scope of protection of the present disclosure, and those skilled in the art should understand that various modifications or changes made to the technical solutions of the present disclosure without inventive labor are still within the scope of protection of the present disclosure.< / amrestservice> < / void> < / string> < / record>
Claims
1. A file collaborative processing method based on distributed multithreading and dynamic database sharding, characterized in that, include: The scheduling parameters are initialized, the parameters are parsed and verified, the files to be processed are obtained, and a file list is generated based on the batch information. The producer-consumer model is adopted. The producer main thread scans the file list through the storage engine API to generate a download task queue. The consumer thread pool executes the download tasks concurrently, and the maximum concurrency is controlled by a semaphore. Hash values are generated based on the sharding routing algorithm, and file data is evenly distributed to multiple physical storage nodes. A dynamic routing table is used to manage the real-time update of the mapping relationship between shards and instances. Based on the business routing strategy, resources are preferentially allocated to high-priority file tasks, and tasks are dynamically allocated to specified database shards and processing nodes. A three-level parallel model is used to achieve parallel processing of multiple files, including file parsing, data cleaning, and data storage on the processing nodes. Dynamic Feign calls are made for dynamic routing and weighted load balancing of service instances. Multi-state table state synchronization is performed based on a two-stage state machine, and processing logs are saved.
2. The file collaborative processing method based on distributed multithreading and dynamic database sharding as described in claim 1, characterized in that, The core of controlling the maximum concurrency through semaphores lies in coordinating the producer and consumer threads to ensure that the number of threads executing tasks does not exceed a preset threshold. The producer thread, as the main thread, scans the file list through the storage engine API, encapsulates the metadata of files that meet the conditions into download tasks, and puts them into the task queue. At this time, the task queue is only responsible for storage and does not involve concurrency control. The consumer thread obtains tasks from the queue, and after obtaining permission, the thread executes file download and processing operations.
3. The file collaborative processing method based on distributed multithreading and dynamic database sharding as described in claim 1, characterized in that, Semaphores automatically adjust the concurrency state through counters. When task execution times vary, the timing of releasing permits creates dynamic scheduling. If an exception is thrown during task execution, the release mechanism of the finally block can prevent the semaphore from being permanently occupied. At the same time, the thread pool is configured with a rejection policy to avoid task queue overflow.
4. The file collaborative processing method based on distributed multithreading and dynamic database sharding as described in claim 1, characterized in that, The sharding routing algorithm first performs a hash calculation on each IOU number, selects a highly discrete hash algorithm to generate a 32-bit hash value, and performs a modulo operation on the hash value based on the total number of shards in the current system to obtain the target shard ID. This target shard ID is directly mapped to the corresponding physical database instance to complete the routing of data writing or querying. When expanding online, a consistent hash ring mechanism is used to achieve smooth migration. When a new node is added, a virtual node is assigned to it, and the virtual nodes are evenly distributed on the hash ring through hash calculation.
5. The file collaborative processing method based on distributed multithreading and dynamic database sharding as described in claim 1, characterized in that, Dynamic routing table management uses ZooKeeper to achieve real-time updates of the mapping relationship between shards and instances. A hierarchical data model is created in ZooKeeper, with a list of corresponding instances stored under each shard ID. When an instance comes online, a temporary node is created through the ZooKeeper client, and a listener is registered. When an instance goes offline, the node is automatically deleted, triggering a global routing table update. In canary deployment scenarios, a weight is configured for each instance, and ZooKeeper generates random numbers based on the weight ratio to dynamically allocate traffic. When a request arrives, the routing component calculates the weight range assigned to the request, achieving a gradual canary deployment switch.
6. The file collaborative processing method based on distributed multithreading and dynamic database sharding as described in claim 1, characterized in that, Based on a two-phase state machine and TCC compensation transactions, cross-database operations are guaranteed. For cross-database operations, the TCC mode is adopted. In case of an exception, a scheduled task scans the temporary table to trigger reverse compensation. The TCC mode includes a Try phase, a Confirm phase, and a Cancel phase. When the Try phase fails, the Confirm phase times out, or the participant fails to execute, the transaction manager asynchronously triggers Cancel. The file storage service deletes the temporary credentials and unlocks the file. The permission service restores the original permission state. The log service records the rollback information. After completion, the global transaction is marked as failed. The compensation mechanism ensures that the resource state is restored to the initial state under any abnormal circumstances.
7. A file collaborative processing system based on distributed multithreading and dynamic database sharding, characterized in that, include: The initialization module is used to initialize scheduling parameters, parse and verify parameters, obtain files to be processed, and generate a file list based on batch information. The file scanning and download module uses a producer-consumer model. The producer main thread scans the file list through the storage engine API to generate a download task queue. The consumer thread pool executes the download tasks concurrently, and the maximum concurrency is controlled by a semaphore. The data distribution module is used to generate hash values based on the sharding routing algorithm, distribute file data evenly to multiple physical storage nodes, manage the real-time update of the mapping relationship between shards and instances using a dynamic routing table, prioritize the allocation of resources to high-priority file tasks based on business routing strategies, and dynamically allocate tasks to specified database shards and processing nodes. The multi-threaded processing and invocation module is used to implement parallel processing of multiple files, including file parsing, data cleaning, and database insertion on the processing nodes through a three-level parallel model, and to dynamically invoke Feign for dynamic routing and weighted load balancing of service instances. The state management module is used to synchronize and update the states of multiple state tables based on a two-phase state machine and save the processing logs.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the file collaborative processing method based on distributed multithreading and dynamic library sharding as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the file collaborative processing method based on distributed multithreading and dynamic database sharding as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform the file collaborative processing method based on distributed multithreading and dynamic library partitioning as described in any one of claims 1-6.
Citation Information
Cited By
Model file processing method and electronic device
CN122470375A