Distributed processing method and system
By splitting the main task into sub-tasks and distributing and processing through message queues, the problem that traditional stand-alone processing mode is difficult to meet the efficiency of large-scale data processing is solved, and efficient distributed data processing and system robustness are achieved.
Patent Information
- Application Number
- CN202411993681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The traditional stand-alone processing model is difficult to meet the efficiency requirements of modern enterprises for large-scale data processing, especially in tasks with large computing and long processing time.
The distributed processing method is adopted to split the main task into multiple subtasks and distributed to the task processing server for asynchronous processing through a message queue. The timing task server performs a timing scan to ensure that all subtasks are completed and notify the administrator to manually process when the number of execution failures reaches the preset number.
It realizes efficient processing of large-scale data, improves the system's concurrent processing capabilities and robustness, and ensures that all tasks can be finally processed.
Smart Images

Figure CN119938735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet applications, and in particular to a distributed processing method and system. Background Art
[0002] With the rapid development of information technology, the demand for large-scale data processing is growing, and the traditional single-machine processing model can no longer meet the efficiency requirements of modern enterprises for data processing. In practical applications, it is often necessary to calculate, count and analyze a large amount of data. Such tasks are characterized by large amount of calculation and long processing time. Summary of the invention
[0003] To achieve the above objectives and other related objectives, the present invention discloses a distributed processing method, comprising: The user initiates a request to the business server, which creates a main task based on the business request and inserts it into the database for storage, recording the execution status of the main task as pending; The business server splits the main task into multiple subtasks according to the preset rules and stores them in the database, and records the execution status of the subtasks as unexecuted; The business server distributes the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as failed and the reason and number of failures are recorded. The scheduled task server periodically scans whether all subtasks under each main task have been completed. If all are completed, the main task status is updated to completed; If the scheduled task server scans and finds that the status of a subtask under the main task is not executed or failed, it will resend the task to the message queue to wait for the task processing server to receive and process it; when the number of task execution failures reaches the preset number, it will send a message to notify the administrator for manual processing; The scheduled task server queries the completed main task and makes statistics based on the subtask processing results, and outputs the statistical result data that users need to query in Excel.
[0004] Furthermore, the main task is split into multiple subtasks according to preset rules, including: Determine the number of subtasks based on the complexity of the main task; Among them, the main task complexity is obtained as follows: C=α1×W+α2×D+α3×M+α4×T; Where: C: task complexity; W: data processing weight; D: data complexity; M: memory consumption estimation; T: historical average processing time; α1~α4: weight coefficients of each dimension, and Σαi=1; The way to determine the number of subtasks based on the complexity of the main task is: N = roundup ((C × S) / P); Where: N: number of subtasks; C: task complexity; S: current system load factor (0~1); P: processing capacity threshold of a single processing node; Roundup(): rounding up function.
[0005] Further, the resending of the task to the message queue to wait for the task processing server to receive and process the task includes: The time interval for reprocessing the task processing server is determined based on the current number of retries, including: T(n)=min(Tmax,Tbase×(1+r)^(n-1)); Where: T(n): waiting time for the nth retry; Tbase: basic waiting time; Tmax: maximum waiting time; r: increasing coefficient; n: current number of retries.
[0006] Furthermore, when the number of task execution failures reaches a preset number, a message is sent to notify the administrator for manual processing, and the setting of the preset number includes: Adjust the preset times according to the task level and real-time data: Preset number = roundup (basic retry number × (1 + ∑Fi)); Among them, Fi includes: F1: historical success rate influencing factor; F2: system load influencing factor; F3: task complexity influencing factor; F4: time urgency factor; Among them, the basic retry times are set according to the task level.
[0007] On the other hand, the present invention also provides a distributed processing system, comprising: The main task creation module is used for the user to initiate a request to the business server, and the business server creates the main task according to the business request and inserts it into the database for storage, recording the execution status of the main task as pending; The task splitting module is used by the business server to split the main task into multiple subtasks according to preset rules and store them in the database, and record the execution status of the subtask as unexecuted; The task execution module is used by the business server to distribute the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as an execution failure state, and the failure reason and number are recorded; The scheduled scanning module is used by the scheduled task server to perform scheduled scanning to see whether all subtasks under each main task have been completed. If all subtasks have been completed, the main task status is updated to completed. The message notification module is used for the scheduled task server to resend the task to the message queue to wait for the task processing server to receive and process it if it scans that the subtask under the main task is in the state of not being executed or failed; when the number of task execution failures reaches the preset number, a message is sent to notify the administrator for manual processing; The data output module is used by the scheduled task server to query the completed main tasks and perform statistics based on the subtask processing results, and output the statistical result data that users need to query in Excel or other formats.
[0008] By adopting the above technical solution, the business system can be helped to effectively perform distributed batch processing operations, thereby meeting the needs of users for batch data operations, management of operation records, viewing results and other batch operations in certain scenarios. The subtasks after the main task is split are distributed through message queues for asynchronous processing. The application system can appropriately expand the number of task processing servers as needed to improve the concurrent processing capabilities of tasks; at the same time, the results are counted through scheduled task pattern scanning, and it is detected whether there are subtasks with abnormal processing, to ensure that all tasks under the application system can eventually be processed, thereby improving the robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0010] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0011] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0012] Reference Figure 1 , an embodiment of the present invention provides a distributed processing method, including: S1: The user initiates a request to the business server, which creates a main task based on the business request and inserts it into the database for storage, recording the execution status of the main task as pending; S2: The business server splits the main task into multiple subtasks according to the preset rules and stores them in the database, and records the execution status of the subtasks as not executed; Among them, the main task is divided into multiple subtasks according to preset rules, including: Determine the number of subtasks based on the complexity of the main task; Among them, the main task complexity is obtained as follows: C=α1×W+α2×D+α3×M+α4×T; Where: C: task complexity; W: data processing weight; D: data complexity; M: memory consumption estimation; T: historical average processing time; α1~α4: weight coefficients of each dimension, and Σαi=1; The way to determine the number of subtasks based on the complexity of the main task is: N = roundup ((C × S) / P); Where: N: number of subtasks; C: task complexity; S: current system load factor (0~1); P: processing capacity threshold of a single processing node; Roundup(): rounding up function.
[0013] S3: The business server distributes the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as an execution failure state, and the failure reason and number are recorded. S4: The scheduled task server periodically scans whether all subtasks under each main task have been completed. If all subtasks have been completed, the main task status is updated to completed. S5: If the scheduled task server scans and finds that a subtask under the main task is in the state of not being executed or failed, it will resend the task to the message queue to wait for the task processing server to receive and process it; when the number of task execution failures reaches the preset number, it will send a message to notify the administrator for manual processing; Specifically, resending the task to the message queue to wait for the task processing server to receive and process it includes: The time interval for reprocessing the task processing server is determined based on the current number of retries, including: T(n)=min(Tmax,Tbase×(1+r)^(n-1)); Where: T(n): waiting time for the nth retry; Tbase: basic waiting time; Tmax: maximum waiting time; r: increasing coefficient; n: current number of retries.
[0014] When the number of task execution failures reaches the preset number, a message is sent to notify the administrator for manual processing. The preset number of settings includes: Adjust the preset times according to the task level and real-time data: Preset number = roundup (basic retry number × (1 + ∑Fi)); Among them, Fi includes: F1: historical success rate influencing factor; F2: system load influencing factor; F3: task complexity influencing factor; F4: time urgency factor; Among them, the basic retry times are set according to the task level.
[0015] The task level is set by technical personnel in this field. In this embodiment, the task level is divided into four types, namely critical tasks, important tasks, ordinary tasks and low-priority tasks. The basic retry number of critical tasks is 8 times, the basic retry number of important tasks is 6 times, the basic retry number of ordinary tasks is 4 times, and the retry number of low-priority tasks is 2 times.
[0016] S6: The scheduled task server queries the completed main task and performs statistics based on the subtask processing results, and outputs the statistical result data to Excel or other formats that the user needs to query.
[0017] Through the implementation of the above technical solutions, the business system can effectively perform distributed batch processing operations, thereby meeting the needs of users for batch data operations, management of operation records, viewing results and other batch operations in certain scenarios. The subtasks after the main task is split are distributed through message queues for asynchronous processing. The application system can appropriately expand the number of task processing servers as needed to improve the concurrent processing capabilities of tasks; at the same time, the results are counted through scheduled task pattern scanning, and it is detected whether there are subtasks with abnormal processing, to ensure that all tasks under the application system can eventually be processed, thereby improving the robustness of the system.
[0018] An embodiment of the present invention further provides a system, comprising: The main task creation module is used for the user to initiate a request to the business server, and the business server creates the main task according to the business request and inserts it into the database for storage, recording the execution status of the main task as pending; The task splitting module is used by the business server to split the main task into multiple subtasks according to preset rules and store them in the database, and record the execution status of the subtask as unexecuted; The task execution module is used by the business server to distribute the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as an execution failure state, and the failure reason and number are recorded; The scheduled scanning module is used by the scheduled task server to perform scheduled scanning to see whether all subtasks under each main task have been completed. If all subtasks have been completed, the main task status is updated to completed. The message notification module is used for the scheduled task server to resend the task to the message queue to wait for the task processing server to receive and process it if it scans that the subtask under the main task is in the state of not being executed or failed; when the number of task execution failures reaches the preset number, a message is sent to notify the administrator for manual processing; The data output module is used by the scheduled task server to query the completed main tasks and perform statistics based on the subtask processing results, and output the statistical result data that users need to query in Excel or other formats.
[0019] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined.
[0020] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0021] It can be known from the description of the above implementation modes that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in the various implementation modes of the present application or certain parts of the implementation modes.
[0022] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distributed processing method, characterized in that: include: The user initiates a request to the business server, which creates a main task based on the business request and inserts it into the database for storage, recording the execution status of the main task as pending; The business server splits the main task into multiple subtasks according to the preset rules and stores them in the database, and records the execution status of the subtasks as unexecuted; The business server distributes the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as failed and the reason and number of failures are recorded. The scheduled task server periodically scans whether all subtasks under each main task have been completed. If all are completed, the main task status is updated to completed; If the scheduled task server scans and finds that the subtask under the main task is in the state of not being executed or failed, it will resend the task to the message queue to wait for the task processing server to receive and process it; When the number of task execution failures reaches the preset number, a message is sent to notify the administrator for manual processing; The scheduled task server queries the completed main task and makes statistics based on the subtask processing results, and outputs the statistical result data that users need to query in Excel.
2. The method according to claim 1, characterized in that The main task is split into multiple subtasks according to preset rules, including: Determine the number of subtasks based on the complexity of the main task; Among them, the main task complexity is obtained as follows: C=α1×W+α2×D+α3×M+α4×T; Where: C: task complexity; W: data processing weight; D: data complexity; M: memory consumption estimation; T: historical average processing time; α1~α4: weight coefficients of each dimension, and Σαi=1; The way to determine the number of subtasks based on the complexity of the main task is: N = roundup ((C × S) / P); Where: N: number of subtasks; C: task complexity; S: current system load factor (0~1); P: processing capacity threshold of a single processing node; Roundup(): rounding up function.
3. The method according to claim 1, characterized in that The step of resending the task to the message queue and waiting for the task processing server to receive and process the task includes: The time interval for reprocessing the task processing server is determined based on the current number of retries, including: T(n)=min(Tmax,Tbase×(1+r)^(n-1)); Where: T(n): waiting time for the nth retry; Tbase: basic waiting time; Tmax: maximum waiting time; r: increasing coefficient; n: current number of retries.
4. The method according to claim 1, characterized in that: When the number of task execution failures reaches a preset number, a message is sent to notify the administrator for manual processing. The setting of the preset number includes: Adjust the preset times according to the task level and real-time data: Preset number = roundup (basic retry number × (1 + ∑Fi)); Among them, Fi includes: F1: historical success rate influencing factor; F2: system load influencing factor; F3: task complexity influencing factor; F4: time urgency factor; Among them, the basic retry times are set according to the task level.
5. A distributed processing system, characterized in that: include: The main task creation module is used for the user to initiate a request to the business server, and the business server creates the main task according to the business request and inserts it into the database for storage, recording the execution status of the main task as pending; The task splitting module is used by the business server to split the main task into multiple subtasks according to preset rules and store them in the database, and record the execution status of the subtask as unexecuted; The task execution module is used by the business server to distribute the subtask to the task processing server through the message queue. After receiving the subtask, the task processing server marks the subtask as being executed. After the subtask is completed, it marks the subtask as completed again. If the execution fails, it is marked as an execution failure state, and the failure reason and number are recorded; The scheduled scanning module is used by the scheduled task server to perform scheduled scanning to see whether all subtasks under each main task have been completed. If all subtasks have been completed, the main task status is updated to completed. The message notification module is used for the scheduled task server to resend the task to the message queue to wait for the task processing server to receive and process it if it scans that the subtask under the main task is in the state of not being executed or failing to execute; When the number of task execution failures reaches the preset number, a message is sent to notify the administrator for manual processing; The data output module is used by the scheduled task server to query the completed main tasks and perform statistics based on the subtask processing results, and output the statistical result data that users need to query in Excel or other formats.