File processing method, device and system and storage medium
By converting files to be processed into file tasks and leveraging the mutual exclusion relationship of Redis locks, the problem of duplicate transaction processing in distributed file processing systems is solved, achieving higher processing accuracy.
Patent Information
- Application Number
- CN202511066362.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
In existing distributed file processing systems, the risk of transaction duplication is high and the accuracy of transaction processing is low.
By converting files into file tasks and leveraging the mutual exclusion of Redis locks, we ensure that each file task is processed by only one server and mark the processing status to avoid duplicate processing.
It effectively reduces the risk of duplicate transaction processing and improves the accuracy of transaction processing in distributed file processing systems.
Smart Images

Figure CN120950475A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology and can be applied to the financial and medical fields. In particular, it relates to a document processing method, apparatus, system, and storage medium. Background Technology
[0002] With the continuous development of computer technology, distributed file processing systems are increasingly being applied in various business areas (such as finance or healthcare) to batch process multiple distributed transactions within these areas. In the process of distributing transactions using such systems, the transaction files for each transaction are typically distributed to multiple servers in batches. Each server processes the transaction files it detects before moving on to the next. However, this batch processing method has a high risk of duplicate transaction processing because it triggers transaction processing based solely on the detection of a file, resulting in low accuracy in transaction processing within the distributed file processing system. Summary of the Invention
[0003] This invention provides a file processing method, apparatus, system, and storage medium to at least solve the problems of high risk of duplicate transaction processing and low transaction processing accuracy in distributed file processing systems in related technologies. The technical solution of this invention is as follows:
[0004] According to a first aspect of the present invention, a file processing method is provided, employing a file processing system. The system includes multiple servers, and the Redis locks of each Redis instance on each of the multiple servers are mutually exclusive. The method includes: constructing multiple file tasks according to file identifiers of multiple files to be processed, wherein each file task corresponds one-to-one with a file to be processed and includes a corresponding file identifier and an initial state; distributing each of the multiple file tasks to a target server among the multiple servers; and triggering each target server to perform the following preset operations: identifying and receiving a target file task whose task state is in the initial state among the monitored multiple file tasks, locking the target file task, marking the target file task as being processed, and determining that the target file task has been processed and marking the processed target file task as completed before releasing the target file task.
[0005] The target server mentioned above can be any one of multiple servers.
[0006] In some implementations, a file task includes a pending file containing payroll invoices, and the file task is used to instruct payment to be made on the included payroll invoices.
[0007] In one implementation, identifying and receiving a target file task whose task status is in the initial state among multiple monitored file tasks includes: sequentially monitoring the task status of each unlocked file task among the multiple file tasks until an unlocked file task whose task status is in the initial state is detected, and identifying the file task whose task status is in the initial state and is not locked as the target task file of the target server; wherein, if the task status is determined to be in the processing state or the completed state, the task status of other unlocked file tasks continues to be monitored.
[0008] In this implementation, the method for determining the target file task for each target server is specifically defined. Each target server monitors the task processing status and locked status of each file task in a loop, and selects the file tasks that meet both status conditions as target file tasks. This ensures that the target file tasks are in the initial state and are not locked, thereby avoiding duplicate processing of file tasks that are locked, in the process of processing, or have been completed.
[0009] In another implementation, the method further includes: after the target server releases the target file task, automatically triggering the target server to continue executing the preset operation until it receives a completion instruction or a stop instruction for multiple file tasks, triggering the target server to stop executing the preset operation.
[0010] The completion command indicates that all task files have been completed, and all target servers should stop executing file task operations.
[0011] The stop execution command instructs a single target server to stop performing file task operations.
[0012] In this implementation, if the target server completes a file task but does not receive a completion instruction or a stop instruction, it will automatically trigger the monitoring of the task processing status and the locked status to ensure that each server can continue to perform preset operations on unprocessed file tasks, thus ensuring that each server can work automatically and continuously to improve file processing efficiency.
[0013] In another implementation, the method further includes: triggering each target server to perform the following preset operations: determining that the target file task is completed within a preset time range and marking the target file task as completed; determining that the target file task is not completed within the preset time range, marking the target file task as failed and releasing the target file task; and sending a fault warning instruction indicating that the target server has failed, the fault warning instruction including the target identification information of the target server and the file identification of the target file task.
[0014] In some implementations, the target identification information of the faulty target server is stored in a preset fault table. Furthermore, this target identification information is deleted after the faulty target server has been successfully repaired.
[0015] In this implementation, the target server can mark different target file tasks with different processing results, so that other servers can effectively identify the processing status of the target file tasks during the monitoring process, thereby facilitating other servers to execute the unsuccessfully processed target file tasks.
[0016] In another implementation, the method further includes triggering each target server to perform the following preset operation: detecting a file task that is not locked and whose task status is in a processing failure state, and identifying the file task whose task status is in a processing failure state and is not locked as the target task file of the target server.
[0017] In this implementation, the target server also monitors the status of file tasks that have failed to process and re-executes file tasks that have not been successfully processed by other servers, so that multiple file tasks can be completed.
[0018] In another implementation, the method further includes: controlling the target server in the fault warning instruction to stop monitoring each file task in the initial state and the processing failure state; and, according to the received fault warning instruction, selecting each target server in the non-fault state from multiple servers.
[0019] In this implementation, to prevent a failed target server from continuing to operate, the failed target server is controlled to stop monitoring, thereby stopping it from executing preset operations and reducing erroneous operations on file tasks. Simultaneously, to ensure the execution efficiency of file tasks on each target server, servers in a failed state are excluded from the pool of servers, ensuring that all target servers are in a non-faulty state when automatically restarted.
[0020] In another implementation, after determining that the target file task has been completed, the method further includes: writing the file to be processed corresponding to the target file task into the Redis database, and accumulating the number of times the file to be processed has been written to the Redis database; determining the total number of tasks for multiple file tasks; and sending a completion command when the number of writes reaches the total number of tasks.
[0021] In this implementation, once the number of writes reaches the total number of tasks, it indicates that all task files have been completed, thus causing all target servers to stop executing file task operations.
[0022] In another implementation, the method further includes: identifying the file formats of the plurality of files to be processed; determining the files to be converted whose file formats differ from the preset file format; and converting the files to be converted into the preset file format.
[0023] In this implementation, multiple files to be processed are converted into the same preset file format so that the content structure of each file task is consistent. This makes it easier to use the same monitoring and processing methods when monitoring and processing each file task in the future, thereby improving monitoring efficiency and task processing efficiency.
[0024] According to a second aspect of the present invention, a file processing apparatus is provided, which applies a file processing system. The system includes multiple servers, and the Redis locks of each Redis instance on each of the multiple servers are mutually exclusive. The apparatus includes: a construction unit for constructing multiple file tasks according to file identifiers of multiple files to be processed, wherein each file task corresponds one-to-one with a file to be processed and includes a corresponding file identifier and an initial state; a distribution unit for distributing the multiple file tasks to each target server among the multiple servers; and a triggering unit for triggering each target server to perform the following preset operations: determining and receiving a target file task whose task state is in the initial state among the monitored multiple file tasks, locking the target file task, marking the target file task as being processed, and determining that the target file task has been processed and marking the processed target file task as completed before releasing the target file task.
[0025] According to a third aspect of the present invention, a file processing system is provided, which is configured to perform a file processing method as described in the first aspect and any possible implementation thereof.
[0026] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, such that when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform a file processing method as described in the first aspect and any possible implementation thereof.
[0027] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including computer instructions that, when executed on an electronic device, cause the electronic device to perform the file processing method described in the first aspect and any possible implementation thereof.
[0028] The technical solution provided by the embodiments of the present invention brings at least the following beneficial effects: It converts files to be processed into file tasks, enabling each server to process these tasks using Redis. Simultaneously, newly generated file tasks, i.e., unprocessed file tasks, are marked as initial states to distinguish them from subsequently processed file tasks. Given the mutual exclusion relationship between Redis locks, a file task will only be monitored by one server and will not be monitored by two or more servers simultaneously. This prevents each server from monitoring file tasks monitored by other servers, thus avoiding the simultaneous duplicate allocation of the same file task. Furthermore, the status of monitored file tasks—whether in processing or completed—is marked to prevent other servers from processing in progress or completed file tasks again. This effectively avoids duplicate processing of file tasks, significantly reduces the risk of duplicate transaction processing, and improves the transaction processing accuracy of the distributed file processing system.
[0029] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0030] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0031] Figure 1 This is a schematic diagram illustrating a file processing system according to an exemplary embodiment;
[0032] Figure 2 This is a flowchart illustrating a file processing method according to an exemplary embodiment;
[0033] Figure 3 This is a block diagram illustrating a file processing apparatus according to an exemplary embodiment;
[0034] Figure 4 This is a schematic diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0036] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0037] Before providing a detailed description of the document processing method provided in the embodiments of this application, let's briefly introduce the application scenarios and implementation environment involved in the embodiments of this application.
[0038] First, a brief introduction to the application scenarios involved in this application will be given.
[0039] With the continuous development of computer technology, distributed file processing systems are increasingly being applied in various business areas (such as finance or healthcare) to batch process multiple distributed transactions within these areas. In the process of distributing transactions using such systems, the transaction files for each transaction are typically distributed to multiple servers in batches. Each server processes the transaction files it detects before moving on to the next. However, this batch processing method has a high risk of duplicate transaction processing because it triggers transaction processing based solely on the detection of a file, resulting in low accuracy in transaction processing within the distributed file processing system.
[0040] In some embodiments of the current application system, payment companies frequently calculate and distribute commissions on behalf of agents. This can lead to situations similar to batch production line overruns.
[0041] Specifically, some sales personnel in Guangdong province reported in a merchant's WeChat group that some agents found two transaction records showing the same transaction date and amount for a merchant's commission rebate activity when checking the corresponding financial processing app. Normally, each agent should only have one commission rebate record displayed per month. An investigation was launched, and it was confirmed that the issue did indeed involve duplicate commission payments for some agents in April, and was not a problem with the data display on the corresponding financial processing app. This merchant's business relies on life insurance agents for offline merchant expansion. Agents receive a merchant onboarding bonus after successfully registering a merchant, and can also receive a merchant activity bonus and a maintenance bonus (commission based on the merchant's transaction volume) based on the merchant's transaction performance. In short, there's a document specifying how much to pay each person each month, and at the beginning of each month, a batch process is needed to import the data into the database for payment. If this process is repeated even once, the payment is doubled.
[0042] The overall commission payment process is as follows: A merchant business relies on life insurance agents to expand offline merchant networks. After an agent successfully completes a merchant's application, they can receive a merchant application bonus. Based on the merchant's transaction performance, they can also receive a merchant activity bonus and a maintenance bonus (commission is calculated based on the merchant's transaction volume). The bonus amount is statistically summarized on a monthly basis and provided to the life insurance side in the form of a document on the 20th of the following month. The completed bonus amount is then distributed to the agent's corresponding financial processing app account.
[0043] Research has revealed that concurrency and mutual exclusion issues are ubiquitous in the aforementioned distributed systems. From large-scale concurrent processing of business logic and concurrent service requests to small-scale concurrent read / write operations on database tables and multi-threaded access to Java objects, concurrency is everywhere in distributed systems. Databases are the most critical storage resource in a system, subject to a large number of concurrent operations. Improper operations on database rows, especially database tables, are most likely to lead to financial losses.
[0044] To address database locking issues and prevent common database concurrency problems such as dirty reads and writes, and non-repeatable reads, locking mechanisms are commonly used. However, improper use of locks can lead to risks. For example, when updating transaction records, if the WHERE clause does not include the original transaction status and does not check the number of records updated, non-repeatable reads can easily occur.
[0045] At the system code level, based on the MSAP system's file processing framework, when importing files, the batch status of files is changed from pending to processing. However, when locking and updating the batch status, the original status is not updated with new data, causing files that have already been successfully processed to be updated back to processing, resulting in duplicate processing.
[0046] To address the aforementioned issues, this application proposes a file processing method that converts files into file tasks, allowing each server to process these tasks using Redis. Newly generated file tasks, i.e., unprocessed tasks, are marked as initial to distinguish them from subsequently processed tasks. Since the Redis locks are mutually exclusive, a file task will only be monitored by one server and not by two or more servers simultaneously. This prevents servers from monitoring file tasks monitored by other servers, thus avoiding duplicate allocation of the same file task. Furthermore, the status of monitored file tasks—whether in progress or completed—is marked to prevent other servers from processing these tasks again, effectively avoiding duplicate processing, significantly reducing the risk of transaction duplication, and improving the accuracy of transaction processing in the distributed file processing system.
[0047] Secondly, the implementation architecture involved in this application will be briefly introduced below.
[0048] Figure 1 This is a schematic diagram of a file processing system 10 provided in this disclosure. For example... Figure 1 As shown, the file processing system includes multiple servers 101 and multiple client terminals 102. The file processing system 10 is a distributed system. Servers 101 and client terminals 102 can establish connections via wired or wireless networks. Furthermore, the Redis locks of each Redis instance in the multiple servers 101 are mutually exclusive.
[0049] In some embodiments, the user terminal 102 can be a terminal device.
[0050] Each user terminal 102 sends a file processing request to the file processing system. The file processing system receives each file to be processed from the file processing requests and converts each file into a file task. The file processing system distributes the file tasks, and each server 101 monitors and marks the four task statuses of each distributed file task: initialization, processing, processing completion, and processing failure. When each server 101 detects a target file task that meets the preset task status, it processes the target file task and sends the processed task result to the Redis database corresponding to that server 101.
[0051] In some embodiments, the file task includes a billing task. Server 101 contains a Redis database or is connected to a Redis database. Billing results from different business areas are categorized and stored in this Redis database; for example, financial bills are stored in a first database, and medical bills are stored in a second database. Each client 102 can access the billing results in the Redis database of server 101.
[0052] In other embodiments, server 101 may be a single server, or it may be a server cluster consisting of multiple servers. In some embodiments, the server cluster may also be a distributed cluster. This application does not limit the specific implementation of server 101.
[0053] The terminal device can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc., that can install and use content community applications (such as Kuaishou). This disclosure does not impose any special restrictions on the specific form of the terminal device. It can interact with users through one or more methods such as keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device.
[0054] The server 101 described above can be connected to at least one terminal device. Furthermore, this application does not limit the number or type of terminal devices.
[0055] The document processing method provided in this application embodiment can be applied to the aforementioned... Figure 1 The document processing system in the implementation architecture shown is illustrated below. For ease of understanding, the document processing method provided in this application will be described in detail below with reference to the accompanying drawings.
[0056] Figure 2 This is a flowchart illustrating a file processing method according to an exemplary embodiment, such as... Figure 2 As shown, the file processing method includes the following steps.
[0057] S21, construct multiple file tasks according to the file identifiers of multiple files to be processed.
[0058] Each file task corresponds to a file to be processed, and includes the corresponding file identifier and initial state.
[0059] File identifiers enable clearer identification and differentiation of individual file tasks.
[0060] S22, distributes multiple file tasks to their respective target servers in multiple servers.
[0061] S23, trigger each target server to perform the following preset operations: identify and receive a target file task whose task status is in the initial state among multiple monitored file tasks, lock the target file task, mark the target file task as being processed, and determine that the target file task has been processed and mark the processed target file task as completed before releasing the target file task.
[0062] The target server mentioned above can be any one of multiple servers.
[0063] In some implementations, a file task includes a pending file containing payroll invoices, and the file task is used to instruct payment to be made on the included payroll invoices.
[0064] Through the above implementation method, the file to be processed is converted into a file task, which can then be processed by each server using Redis. Newly generated file tasks, i.e., unprocessed file tasks, are marked as initial state to distinguish them from subsequently processed file tasks. Since the Redis locks are mutually exclusive, a file task will only be monitored by one server and will not be monitored by two or more servers simultaneously. This prevents each server from monitoring file tasks monitored by other servers, thus avoiding the duplicate allocation of the same file task. Furthermore, the status of the monitored file task—whether it is in progress or completed—is marked to prevent other servers from processing in progress or completed file tasks again. This effectively avoids duplicate processing of file tasks, significantly reduces the risk of duplicate transaction processing, and improves the transaction processing accuracy of the distributed file processing system.
[0065] As one implementation method, the specific process of determining the target file task to be processed by each target server after each target server is automatically triggered in step S23 above is as follows.
[0066] The task status of each unlocked file task in a multi-file task is monitored sequentially to determine whether the task status of an unlocked file task is in the initial state. If the task status is determined to be in a processing or completed state, the monitoring of the task status of other unlocked file tasks continues until an unlocked file task is detected to be in the initial state. The file task with the initial state and no lock is then identified as the target task file on the target server. If the task status is determined to be in the initial state, the unlocked file task in the initial state is identified as the target task file on the target server.
[0067] In this implementation, the method for determining the target file task for each target server is specifically defined. Each target server monitors the task processing status and locked status of each file task in a loop, and selects the file tasks that meet both status conditions as target file tasks, so that the target file tasks are in the initial state and are not locked, thereby avoiding the repeated processing of file tasks that are locked, in the process of processing, or have been completed.
[0068] In one implementation, the automatic triggering of preset operations is performed in a loop, and the automatic triggering of preset operations is stopped by a loop stop condition. After the target server releases the target file task, the target server is automatically triggered to continue executing the preset operations until a completion instruction or a stop instruction for multiple file tasks is received, at which point the target server is triggered to stop executing the preset operations.
[0069] The loop stopping condition for each target server is receiving a completion command or a stop execution command.
[0070] The completion command indicates that all task files have been completed, and all target servers should stop executing file task operations.
[0071] The stop execution command instructs a single target server to stop performing file task operations.
[0072] In this implementation, if the target server completes a file task but does not receive a completion instruction or a stop instruction, it will automatically trigger the monitoring of the task processing status and the locked status to ensure that each server can continue to perform preset operations on unprocessed file tasks, thereby ensuring that each server can work automatically and continuously to improve file processing efficiency.
[0073] In one implementation, the aforementioned triggering of each target server also performs the following preset operations: determining that the target file task is completed within a preset time range and marking the target file task as completed; determining that the target file task is not completed within the preset time range, marking the target file task as failed and releasing the target file task, and sending a fault warning instruction indicating that the target server has failed, the fault warning instruction including the target identification information of the target server and the file identification of the target file task.
[0074] In some implementations, the target identification information of the malfunctioning target server is stored in a preset fault table. Furthermore, the target identification information is deleted after the malfunctioning target server has been successfully repaired.
[0075] In this implementation, the target server can mark different target file tasks with different processing results, so that other servers can effectively identify the processing status of the target file tasks during the monitoring process, thereby facilitating other servers to execute the unsuccessfully processed target file tasks.
[0076] As one implementation method, the target servers are also triggered to perform the following preset operation: detect a file task that is not locked and whose task status is a processing failure state, and determine the file task with the task status of processing failure and not locked as the target task file of the target server.
[0077] In this implementation, the target server also monitors the status of file tasks that have failed to process and re-executes file tasks that have not been successfully processed by other servers, so that multiple file tasks can be completed.
[0078] As one implementation method, the target server in the control fault warning instruction stops monitoring each file task in the initial state and the processing failure state; and, based on the received fault warning instruction, selects each target server in a non-fault state from multiple servers.
[0079] In this implementation, to prevent a failed target server from continuing to operate, the failed target server is controlled to stop monitoring, thereby stopping the failed target server from executing preset operations and reducing erroneous operations on file tasks. Simultaneously, to ensure the execution efficiency of file tasks on each target server, servers in a failed state are excluded from the pool of servers, ensuring that each target server is in a non-faulty state when automatically triggered to start.
[0080] As one implementation method, after determining that the target file task has been completed, the following operations are also performed: The file to be processed corresponding to the target file task is written to the Redis database, and the number of times the file to be processed is written to the Redis database is accumulated; the total number of tasks for multiple file tasks is determined; when the number of writes reaches the total number of tasks, a completion command is sent.
[0081] In this implementation, once the total number of writes is reached, it indicates that all task files have been completed, thus causing all target servers to stop executing file task operations.
[0082] As one implementation method, to ensure that multiple files to be processed are in a uniform file format, the following operations are performed before executing S21: The file formats of the multiple files to be processed are identified; files whose file formats differ from the preset file format are identified and converted; the files to be converted are converted to the preset file format.
[0083] In this embodiment, multiple files to be processed are converted into the same preset file format so that the content structure of each file task is consistent. This makes it easier to use the same monitoring method and the same task processing method when monitoring and processing each file task in the future, thereby improving monitoring efficiency and task processing efficiency.
[0084] As a specific implementation method, taking the processing of a single file in a file task as an example, the following detailed explanation of the distributed file processing process is provided according to the following control logic.
[0085] 1. Monitor whether the file format of the file conforms to the preset file format.
[0086] 2. Implement Redis locking through software code, that is, lock the file.
[0087] 3. Modify the file status of the file in the file task to put the file into the processing state, so that other servers cannot operate on the file.
[0088] 4. Read the contents of the files in the file task in a loop and write the bills in the file to the database.
[0089] 5. Modify the file status of the file in the file task to "processed successfully" so that other servers can query the file status of this file task to indicate that it has been processed.
[0090] 6. Release the Redis lock so that other servers can access the file for this task.
[0091] The file processing method described above for preventing duplicate file processing can be based on the standard operation of updating the database status: lock, check, and modify, to process the file.
[0092] First, locks: Locking mechanisms ensure that single-point data is not processed concurrently.
[0093] The locking mechanism is the Redis lock in software development. It is single-process and single-threaded. Once acquired, it cannot be accessed by any other process until it is released.
[0094] Secondly, determine the number of updates if the original state is updated.
[0095] Update to the original state. Simply put, this batch of data will be aggregated into a single task, with multiple data entries attached to the task. The status is applied to the task, not the individual data entries.
[0096] Thirdly, the data records are updated based on the lock judgment mechanism.
[0097] The lock used is a Redis lock. Redis locks are single-process and single-threaded. Simply put, once a lock is acquired, others cannot use it until the task is completed and the lock is actively released, at which point others can then process the data.
[0098] When a database update is committed, it is necessary to determine whether the status of the updated data meets the requirements (database data status changes from four states: initialization, processing in progress, processing completed, and processing failed). Only tasks that have failed or been initialized can be acquired and executed. The execution steps are the six steps of a file task processing procedure described in the above specific implementation, in order to avoid non-repeatable reads.
[0099] In some embodiments, the status of the existing transaction must be added to the WHERE clause of the SQL statement, and the number of records returned by the UPDATE method must be checked to ensure it meets expectations. After updating the status of the completed task, the number of updated records will be returned. Only if the number of updated records matches the number of tasks retrieved will the expected result be met. If they do not match, it indicates a failure, meaning some problem has occurred and the task needs to be retried and processed.
[0100] When updating a state machine, the preceding state value in the WHERE clause should be based on the state transition, such as `WHERE status = xx` or `WHERE status in(xx, yy)`. The update result affects the number of rows; the handling logic for 0, 1, and N rows all needs to be implemented. (The reason for the preceding state value in the WHERE clause is to prevent concurrency. For example, if two people receive a task at the same time, and one person has already completed and successfully updated it, and the second person finishes and tries to update it but finds that the task's state has changed, then the second person's processing must be invalidated. Of course, this is a double guarantee; in fact, Redis locks already ensure that only one person can process the task at a time.)
[0101] Based on the above implementation, batch file processing ensures that single-point data is processed without concurrency. The locking mechanism is the same as the Redis lock in software development, which is single-process and single-threaded. Once acquired, no other process can access it until it is released. Simultaneously, it updates with the original state, checks the number of updated records, avoids secondary updates for completed events, and ensures that locking precedes reading to prevent dirty reads. Events arriving concurrently are read from the database only after the locking mechanism is in place to prevent dirty reads.
[0102] To achieve the above functions, the file processing apparatus includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art will readily recognize that, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0103] This disclosure also provides an embodiment such as Figure 3 The file processing apparatus shown is used in a file processing system that includes multiple servers, and the Redis locks of each Redis instance on each server are mutually exclusive. The apparatus includes: a construction unit 301, a distribution unit 302, and a triggering unit 303.
[0104] The construction unit 301 is used to construct multiple file tasks according to the file identifiers of multiple files to be processed. Each file task corresponds to a file to be processed and includes the corresponding file identifier and initial state.
[0105] The distribution unit 302 is used to distribute multiple file tasks to each target server in multiple servers.
[0106] The triggering unit 303 is used to trigger each target server to perform the following preset operations: identify and receive a target file task whose task status is in the initial state among multiple monitored file tasks, lock the target file task, mark the target file task as being processed, and determine that the target file task has been processed and mark the processed target file task as being completed before releasing the target file task.
[0107] In one implementation, the triggering unit 303 is used to: sequentially monitor the task status of each unlocked file task among multiple file tasks until the task status of an unlocked file task is detected to be in the initial state, and determine the file task with the initial state and not locked as the target task file of the target server; wherein, if the task status is determined to be in the processing state or the completed state, the task status of other unlocked file tasks continues to be monitored.
[0108] In another implementation, the triggering unit 303 is also used to: automatically trigger the target server to continue executing the preset operation after the target server releases the target file task, until it receives a completion instruction or a stop instruction for the multiple file tasks, and then triggers the target server to stop executing the preset operation.
[0109] In another implementation, the triggering unit 303 is further configured to: determine that the target file task has been processed within a preset time range and mark the target file task as completed; determine that the target file task has not been processed within the preset time range and mark the target file task as failed and release the target file task; and send a fault warning instruction indicating that the target server has failed, the fault warning instruction including the target identification information of the target server and the file identification of the target file task.
[0110] In another implementation, the triggering unit 303 is also used to: detect that a file task is not locked and the task status is a processing failure status, and determine the file task with the task status of processing failure and not locked as the target task file of the target server.
[0111] In another implementation, the triggering unit 303 is also used to: control the target server of the target identification information in the fault warning instruction to stop monitoring each file task in the initial state and the processing failure state; and, according to the received fault warning instruction, filter out each target server in the non-fault state from multiple servers.
[0112] In another implementation, after determining that the target file task has been completed, the device is further used to: write the file to be processed corresponding to the target file task into the Redis database, and accumulate the number of times the file to be processed is written to the Redis database; determine the total number of tasks for multiple file tasks; and send a completion command when the number of writes reaches the total number of tasks.
[0113] In another embodiment, the building unit 301 is further configured to: identify the file formats of the plurality of files to be processed; determine the files to be converted whose file formats are different from the preset file format; and convert the files to be converted into the preset file format.
[0114] Regarding the apparatus in the above embodiments, the specific manner in which each unit module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0115] Figure 4 This is a schematic diagram of an electronic device provided in this application. (For example...) Figure 4 The electronic device 50 may include at least one processor 501 and a memory 503 for storing processor-executable instructions. The processor 501 is configured to execute the instructions in the memory 503 to implement the file processing method described in the following embodiments.
[0116] In addition, electronic device 50 may also include communication bus 502, at least one communication interface 504, input device 506 and output device 505.
[0117] The processor 501 may be a processor (central processing unit, CPU), a microprocessor unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application.
[0118] The communication bus 502 may include a path for transmitting information between the aforementioned components.
[0119] Communication interface 504 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0120] Input device 506 is used to receive input signals and output device 505 is used to output signals.
[0121] Memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory may exist independently and be connected to the processing unit via a bus. Memory may also be integrated with the processing unit.
[0122] The memory 503 stores instructions for executing the scheme of this application, and the processor 501 controls the execution. The processor 501 executes the instructions stored in the memory 503 to implement the functions of the method of this application.
[0123] In a specific implementation, as one example, the processor 501 may include one or more CPUs, for example... Figure 4 CPU0 and CPU1 in the CPU.
[0124] In a specific implementation, as one example, the electronic device 50 may include multiple processors, such as... Figure 4 Processors 501 and 507 are shown in the diagram. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0125] The electronic device is as follows Figure 4 The diagram includes a processor 501 and a memory 503 for storing executable instructions of the processor 501; wherein the processor 501 is configured to execute the executable instructions to implement the file processing method as described in any of the possible embodiments above. And it can achieve the same technical effect, so to avoid repetition, it will not be described again here.
[0126] This application also provides a computer-readable storage medium, which, when executed by a processor of a file processing device or electronic device, enables the file processing device or electronic device to perform a file processing method as described in any of the possible embodiments above. The same technical effects can be achieved, and to avoid repetition, further details are omitted here.
[0127] This application also provides a computer program product, including a computer program or instructions, which are executed by a processor as a file processing method according to any of the possible implementations described above. The same technical effects can be achieved, and to avoid repetition, further details are omitted here.
[0128] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0129] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A file processing method, characterized in that, An application file processing system, comprising multiple servers, wherein each Redis lock on each server is mutually exclusive; the method includes: Multiple file tasks are constructed based on the file identifiers of multiple files to be processed. Each file task corresponds one-to-one with a file to be processed and includes the corresponding file identifier and initial state. Each of the multiple file tasks is distributed to a target server among the multiple servers; Each of the target servers is triggered to perform the following preset operations: identify and receive a target file task whose task status is in the initial state among the monitored multiple file tasks, lock the target file task, mark the target file task as being processed, determine that the target file task has been processed and mark the processed target file task as being completed, and then release the target file task.
2. The file processing method according to claim 1, characterized in that, The process of identifying and receiving a target file task whose task status is in the initial state among the monitored plurality of file tasks includes: The task status of each unlocked file task in the plurality of file tasks is monitored sequentially until an unlocked file task is detected to be in the initial state. The file task in the initial state and which is not locked is identified as the target task file of the target server. If the task status is determined to be in a processing state or a completed state, the task status of other unlocked file tasks will continue to be monitored.
3. The file processing method according to claim 1, characterized in that, The method further includes: After the target server releases the target file task, it is automatically triggered to continue executing the preset operation until it receives a completion instruction or a stop instruction for the multiple file tasks, which triggers the target server to stop executing the preset operation.
4. The document processing method according to any one of claims 1 to 3, characterized in that, The method further includes: The target servers are also triggered to perform the following preset operations: Once the target file task is determined to be completed within a preset time range, the target file task is marked as completed. If the target file task is not completed within a preset time range, the target file task is marked as a processing failure and released. A fault warning instruction indicating that the target server has failed is also sent. The fault warning instruction includes the target identification information of the target server and the file identification of the target file task.
5. The document processing method according to claim 4, characterized in that, The method further includes: The target servers are also triggered to perform the following preset operations: A file task is detected that is not locked and its status is "processing failed". This file task with the status "processing failed" and not locked is identified as the target task file of the target server.
6. The file processing method according to claim 4, characterized in that, The method further includes: The target server, which controls the target identifier information in the fault warning instruction, stops monitoring each file task in the initial state and the processing failure state; Furthermore, based on the received fault warning instruction, select each target server in a non-faulty state from the plurality of servers.
7. The document processing method according to claim 4, characterized in that, After determining that the target file task processing is complete, the method further includes: Write the files to be processed corresponding to the target file task into the Redis database, and record the cumulative number of times the files to be processed are written to the Redis database; Determine the total number of tasks for the multiple file tasks; When the number of writes reaches the total number of tasks, the completion command is sent.
8. A document processing device, characterized in that, An application file processing system, comprising multiple servers, wherein each Redis lock on each server is mutually exclusive; the apparatus includes: The construction unit is used to construct multiple file tasks according to the file identifiers of multiple files to be processed. Each file task corresponds one-to-one with the file to be processed and includes the corresponding file identifier and initial state. The distribution unit is used to distribute each of the plurality of file tasks to the respective target servers among the plurality of servers. The triggering unit is used to trigger each of the target servers to perform the following preset operations: determine and receive a target file task whose task status is in the initial state among the monitored multiple file tasks, lock the target file task, mark the target file task as being processed, determine that the target file task has been processed and mark the processed target file task as being completed, and then release the target file task.
9. A document processing system, characterized in that, The file processing system is used for distributed file processing tasks. The system includes multiple servers, and the Redis locks of each Redis instance on each of the multiple servers are mutually exclusive. The system is configured to execute the file processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the file processing method as described in any one of claims 1-7.