A distributed task scheduling method and system
By using database optimistic locking and Zookeeper auxiliary checks in distributed task scheduling, the problems of Redis lock failure and complex Zookeeper deployment are solved, achieving high reliability and simplified deployment of distributed systems.
Patent Information
- Application Number
- CN201910163677.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-05
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2039-03-05
AI Technical Summary
In existing distributed task scheduling solutions, improper key expiration time settings in the Redis solution can cause lock failure, while the Zookeeper solution is complex to deploy and lacks reliability, making it difficult to ensure the reliability of the distributed system.
We use database optimistic locking instead of the Redis solution, and combine it with Zookeeper as an auxiliary check. By using the optimistic locking mechanism in the database and automatically releasing the lock when Docker is restarted, and using Zookeeper to detect the Docker heartbeat to assist in releasing the lock, we avoid lock failure and deployment complexity.
It improves the reliability of distributed systems, reduces the risk of lock failure, simplifies the deployment process, and enhances the stability and reliability of the system.
Smart Images

Figure CN111666134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for distributed task scheduling. Background Art
[0002] There are many existing distributed task scheduling solutions. Commonly used approaches include acquiring distributed locks through Redis or Zookeeper to address multi-host multi-tasking scheduling. Distributed locks are exclusive; only one process can acquire the lock and execute a task at a time; other processes cannot. Upon completion of a task, the process releases the lock, but computers are not 100% reliable, so lock release failures can occur.
[0003] Acquiring distributed locks through Redis is currently the most widely used technology. The basic principle is that multiple threads acquire the lock through Redis's atomic operations, with only one thread ultimately securing the lock. With this solution, the typical solution to lock release failures is to set an expiration time for the Redis key. This ensures that even if the lock fails, the key can be released at the expiration date. However, setting a specific expiration time for the key is difficult. Acquiring distributed locks through Zookeeper uses its temporary ordered nodes, but this approach is relatively complex to deploy and is rarely used in production environments.
[0004] In the process of implementing the present invention, the inventors discovered that the prior art has at least the following problems:
[0005] (1) Improper setting of the key expiration time in the Redis solution will cause the lock to fail, but it is difficult to define the appropriate time.
[0006] (2) The Zookeeper solution is highly dependent on Zookeeper, so the reliability of the distributed system cannot be guaranteed and the deployment of the Zookeeper solution is relatively troublesome.
[0007] Therefore, the problem of lock failure in distributed systems is not well solved under the current solutions, which makes it difficult to ensure the reliability of distributed systems. Summary of the Invention
[0008] In light of this, embodiments of the present invention provide a method and system for distributed task scheduling that replaces traditional Redis solutions with optimistic database locking. This eliminates the need to set expiration times, avoiding the problem of lock failure caused by improper key expiration times. Furthermore, by using Zookeeper as a secondary check solution based on optimistic locking, this method avoids the deployment difficulties associated with directly using Zookeeper solutions and eliminates a strong dependency on Zookeeper. Furthermore, while avoiding the problems of existing solutions, it effectively resolves the issue of lock failure in distributed task scheduling.
[0009] To achieve the above objective, according to one aspect of an embodiment of the present invention, a method for distributed task scheduling is provided.
[0010] A distributed task scheduling method according to an embodiment of the present invention includes:
[0011] receiving one or more query requests for a task lock object of a task from one or more servers;
[0012] Query the task lock object in the database;
[0013] In response to querying the task lock object, reading a record related to the task lock object from the database and sending the record to the one or more servers, the record including a task lock state and a specific version number;
[0014] receiving first update data for the task from a first server among the one or more servers, the first update data including a first version number;
[0015] In response to determining that the first version number is the same as the specific version number, performing a first update on the database using the first update data; and
[0016] When an error event occurs on a second server among the one or more servers, other servers among the one or more servers are triggered to trigger a monitoring event, wherein the first server and the second server are the same or different.
[0017] Optionally, before receiving one or more query requests for the task lock object of the task from one or more servers, the method further includes:
[0018] Creating a parent node corresponding to the check server; and
[0019] One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
[0020] Optionally, the one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
[0021] Optionally, the parent node and the one or more child nodes are EPHEMERAL type nodes.
[0022] Optionally, causing other servers among the one or more servers to trigger a monitoring event further includes:
[0023] Deleting the second child node corresponding to the second server from the list to obtain a new list;
[0024] Sending the new list to the child nodes in the new list;
[0025] receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock;
[0026] querying the database according to the lock query request; and
[0027] In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
[0028] Optionally, when the second child node holds an unreleased lock, the task lock state of the task is 1.
[0029] Optionally, releasing the lock further includes: setting the task lock state of the task to 0.
[0030] Optionally, before receiving one or more query requests for the task lock object of the task from one or more servers, the method further includes:
[0031] receiving IP addresses of the one or more servers;
[0032] According to the IP address, query the database for tasks with a task lock status of 1; and
[0033] In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
[0034] Optionally, after querying the task lock object in the database, the method further includes:
[0035] In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
[0036] Optionally, the record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
[0037] According to another aspect of an embodiment of the present invention, a distributed task scheduling system is provided.
[0038] A distributed task scheduling system according to an embodiment of the present invention includes:
[0039] A query request receiving module, configured to receive one or more query requests for a task lock object of a task from one or more servers;
[0040] A lock object query module, used for querying the task lock object in the database;
[0041] a lock object processing module, configured to, in response to querying the task lock object, read a record related to the task lock object from the database and send the record to the one or more servers, the record including a task lock state and a specific version number;
[0042] an update receiving module, configured to receive first update data for the task from a first server among the one or more servers, wherein the first update data includes a first version number;
[0043] an update execution module, configured to, in response to determining that the first version number is the same as the specific version number, execute a first update on the database using the first update data; and
[0044] An auxiliary checking module is used to enable other servers among the one or more servers to trigger a monitoring event when an error event occurs in the second server among the one or more servers, wherein the first server and the second server are the same or different.
[0045] Optionally, the auxiliary inspection module is further configured to:
[0046] Creating a parent node corresponding to the check server; and
[0047] One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
[0048] Optionally, the one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
[0049] Optionally, the parent node and the one or more child nodes are EPHEMERAL type nodes.
[0050] Optionally, the auxiliary inspection module is further configured to:
[0051] Deleting the second child node corresponding to the second server from the list to obtain a new list;
[0052] Sending the new list to the child nodes in the new list;
[0053] receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock;
[0054] querying the database according to the lock query request; and
[0055] In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
[0056] Optionally, when the second child node holds an unreleased lock, the task lock state of the task is 1.
[0057] Optionally, the auxiliary checking module is further configured to: set the task lock state of the task to 0.
[0058] Optionally, the system further comprises:
[0059] A server startup module, configured to receive the IP addresses of the one or more servers;
[0060] According to the IP address, query the database for tasks with a task lock status of 1; and
[0061] In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
[0062] Optionally, the locked object processing module is further configured to:
[0063] In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
[0064] Optionally, the record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
[0065] According to another aspect of an embodiment of the present invention, a distributed task scheduling electronic device is provided.
[0066] An electronic device for distributed task scheduling according to an embodiment of the present invention includes:
[0067] A distributed task scheduling electronic device, characterized by comprising:
[0068] one or more processors;
[0069] a storage system for storing one or more programs,
[0070] When the one or more programs are executed by the one or more processors, the one or more processors implement the distributed task scheduling method provided by the first aspect of the embodiment of the present invention.
[0071] According to yet another aspect of an embodiment of the present invention, a computer-readable medium is provided.
[0072] According to the computer-readable medium of the embodiment of the present invention, a computer program is stored thereon, and when the program is executed by a processor, the distributed task scheduling method provided by the first aspect of the embodiment of the present invention is implemented.
[0073] One embodiment of the above invention has the following advantages or beneficial effects: because database optimistic locking is used to schedule tasks in a distributed system while Zookeeper is used for auxiliary inspection, the technical problems of lock failure caused by improper key expiration time setting and strong dependence on Zookeeper are overcome, thereby achieving the technical effect of improving the reliability of the distributed system and reducing the complexity of solution deployment.
[0074] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.
[0076] Figure 1 is a schematic diagram of the main process of the distributed task scheduling method according to an embodiment of the present invention;
[0077] Figure 2 This is a schematic diagram of an exemplary process of an exemplary Docker startup phase according to an embodiment of the present invention;
[0078] Figure 3 is a schematic diagram of an exemplary flow chart of an exemplary task execution phase according to an embodiment of the present invention;
[0079] Figure 4 is a schematic diagram of an example flow chart of an exemplary program checking stage according to an embodiment of the present invention;
[0080] Figure 5 is a schematic diagram of an exemplary flow chart of an exemplary Zookeeper checking phase according to an embodiment of the present invention;
[0081] Figure 6 is a schematic diagram of the main process of another distributed task scheduling method according to an embodiment of the present invention;
[0082] Figure 7 is a schematic diagram of main modules of a distributed task scheduling system according to an embodiment of the present invention;
[0083] Figure 8 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;
[0084] Figure 9 It is a schematic diagram of the structure of a computer system of a terminal device or server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0085] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0086] Figure 1 FIG. 1 is a schematic diagram of the main process of the distributed task scheduling method according to an embodiment of the present invention. Figure 1 As shown, the method for distributed task scheduling according to an embodiment of the present invention includes steps S101, S102, S103, S104, S105 and S106.
[0087] Step S101: receiving one or more query requests for a task lock object of a task from one or more servers.
[0088] This solution uses optimistic locking in the database to release the lock through the Docker execution program. Steps S101 to S105 are the normal execution flow of the optimistic locking solution. In this article, the term "Docker" refers to a server, such as a Linux server, on which the code program for executing the method can be deployed. Throughout this article, the term "server" and "Docker" can be used interchangeably without affecting the implementation of the solution and the technical effect.
[0089] This technical solution can be divided into two phases: (1) normal execution phase; (2) auxiliary inspection phase. The normal execution phase includes the Docker startup phase and the task execution phase. The auxiliary inspection phase includes program inspection and Zookeeper inspection. Among them, the Docker startup phase can be completed when each server is restarted in the task execution phase. The main purpose is to solve the problem that the lock cannot be released after the restart when Docker still holds an unreleased lock, further improving the reliability of the distributed system. Zookeeper inspection mainly solves the problem that the lock cannot be released due to Docker crashes in the task execution phase. The program inspection will run regularly, and can be checked every 10 minutes, starting synchronously with the normal execution phase. It mainly solves the problem that the optimistic lock needs to be released but fails to be released when the task is completed in the task execution phase.
[0090] During normal execution, when a task is running, the optimistic lock on the database is not released until it completes. However, when we frequently have to restart the Docker container to meet new requirements, the optimistic lock on the database is no longer released. Therefore, we added Docker startup processing logic to address this issue. When the Docker container starts, it automatically executes a listener to release the optimistic lock on the database held by the container.
[0091] Preferably, before receiving one or more query requests for the task-locked object of the task from one or more servers (step S101), the method further includes:
[0092] receiving IP addresses of the one or more servers;
[0093] According to the IP address, query the database for tasks with a task lock status of 1; and
[0094] In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
[0095] The following is combined with Figure 2 Describe the main steps of an exemplary Docker startup phase.
[0096] Figure 2 FIG. 1 is a schematic diagram of an exemplary process of an exemplary Docker startup phase according to an embodiment of the present invention. Figure 2 As shown, an exemplary Docker startup phase 200 according to an embodiment of the present invention includes steps S201, S202, S203, S204 and S205.
[0097] Step S201: Docker starts.
[0098] In one embodiment, Docker startup includes starting the Docker application, initializing the Web container, and executing a task lock release listener.
[0099] Step S202: Execute the task lock release listener.
[0100] In one embodiment, the task lock release listener obtains the local IP address and queries the database through the IP address to see whether there is a locked task, that is, a task with a lock status of 1 in the database.
[0101] Step S203: Check whether there is any task locked by the local machine.
[0102] Step S204: Release the lock.
[0103] If there is a locked task ("yes" at step S203), the lock of the task is released and the lock state of the task in the database is set to 0. If there is no locked task ("no" at step S203), no operation is performed and the process goes to step S205.
[0104] Step S205: Docker startup is completed.
[0105] The Docker startup phase solves the problem that optimism cannot be released when restarting Docker, thereby enhancing the reliability of the distributed system. In some embodiments, the Docker startup phase may not be included before step S101.
[0106] Step S102: Query the task locking object in the database.
[0107] If a Docker wants to execute a task, the first step is to check whether the task is locked by other Dockers. Therefore, after receiving one or more query requests for the task lock object of the task from one or more servers in step S101, the task lock object will be queried in the database to determine whether the task is currently locked by other servers.
[0108] In one embodiment, the Docker application queries the task lock object by task type.
[0109] Step S103: In response to finding the task lock object, read a record related to the task lock object from the database and send the record to the one or more servers, where the record includes a task lock state and a specific version number.
[0110] Preferably, after querying the task lock object in the database, the method further includes:
[0111] In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
[0112] Optionally, the record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
[0113] In one embodiment, if the task lock object is null, it means that there is no task of this type stored in the database. Therefore, a record is inserted into the database with the following fields: task type (int), task description (varchar), task lock status (int, value 0, indicating that the task is not locked), and version number (long, value 1). Then, the local IP address is obtained.
[0114] In one embodiment, if the task lock object is not null, the task lock status is obtained from the task lock object. If the task lock status value is 1 (indicating that the task is locked), the process ends. If the task lock status value is 0 (indicating that the task is not locked), the local IP address is obtained.
[0115] Step S104: receiving first update data for the task from a first server among the one or more servers, where the first update data includes a first version number.
[0116] The task will only be executed if Docker is sure that the task is not locked. At the same time, there may be multiple Dockers that have confirmed that the task is not locked and want to execute it. The optimistic locking mechanism is actually implemented by introducing a version number (version) field in the database table. When we want to read data from the database, we also read the version field. If we want to update the read data and write it back to the database, we need to increase the version by 1 and update the new data and the new version to the data table. At the same time, we must check whether the version value in the current database is the previous version. If it is, we will update normally. If not, the update fails, indicating that another process has updated the data during this process.
[0117] Therefore, each Docker application begins updating the task data in the database. Update criteria include: task type, task lock status (value 0), and version number (1 if coming from step 2; obtained from the lock object if coming from step 3). Updated fields include: task lock status (updated to 1), version number (incremented by 1 from the original version number), local IP address, and update time (the current time). Due to optimistic locking in the database, only one Docker application can successfully modify the task.
[0118] Step S105: In response to determining that the first version number is the same as the specific version number, performing a first update on the database using the first update data.
[0119] In this case, step S104 receives the data from the server that completes the update data earliest among the multiple servers that want to execute the task, so that it successfully obtains the optimistic lock and completes the update.
[0120] As mentioned above, the normal execution phase includes the Docker startup phase and the task execution phase. Generally, each of the multiple Docker containers will go through the Docker startup phase when it is powered on. After each Docker container is powered on, it enters the task execution phase for that Docker container. In one embodiment, during the Docker task execution phase, when a task reaches an executable point, multiple Docker containers will simultaneously execute the task. At this point, all Docker containers will attempt to acquire the optimistic lock on the database. Only the Docker container that has acquired the optimistic lock can execute the task.
[0121] The following is combined with Figure 3 Describe the main steps of an exemplary task execution phase.
[0122] Figure 3 FIG. 1 is a flow chart of an exemplary task execution phase according to an embodiment of the present invention. Figure 3 As shown, an exemplary task execution phase 300 according to an embodiment of the present invention includes steps S301 , S302 , S303 , S304 , S305 , S306 and S307 .
[0123] Step S301: One or more Dockers prepare to execute tasks.
[0124] In one embodiment, step S301 may further include sub-steps S301_1 to S301_N corresponding to the one or more Dockers, respectively representing steps in which Docker_1 to Docker_N prepare to execute tasks.
[0125] In one embodiment, the Docker application queries the task lock object by task type.
[0126] If the task lock object is null, it means that there is no task of this type stored in the database. Therefore, a record is inserted into the database with the following fields: task type (int), task description (varchar), task lock status (int, value 0, indicating that the task is not locked), and version number (long, value 1). Then, the IP address of the local machine is obtained.
[0127] If the task lock object is not null, the task lock status is obtained from the task lock object. If the task lock status value is 1 (indicating that the task is locked), the process ends. If the task lock status value is 0 (indicating that the task is not locked), the local IP address is obtained.
[0128] Step S302: The one or more Dockers obtain a database optimistic lock through the task number and version number.
[0129] Each Docker application begins updating the task data in the database. Update conditions include: task type, task lock status (value 0), and version number (if the task lock object is null in step S301, the value is 1; if the task lock object is not null in step S301, the value can be obtained from the lock object). Updated fields include: task lock status (updated to 1), version number (incremented by 1 from the original version number), local IP address, and update time (current time).
[0130] Step S303: One of the one or more Dockers successfully acquires an optimistic lock.
[0131] Based on database optimistic locking, only one Docker application can be modified successfully.
[0132] Step S304: The Docker that obtains the optimistic lock starts executing the task.
[0133] The Docker application that obtains the optimistic lock starts to execute the task. When the task is completed, the lock is released. The operation of releasing the lock is to update the task lock status to 0.
[0134] Step S305: The Docker that has obtained the optimistic lock completes the task and releases the optimistic lock.
[0135] Step S306: The optimistic lock is released successfully.
[0136] Step S307: Optimistic lock release fails, and an alarm email is sent.
[0137] In step S305, the database is not 100% reliable, so there is a risk of failure to release the optimistic lock. Furthermore, a Docker server crash can also cause a failure to release the optimistic lock. This can also happen if a task is currently executing when the project goes live. If the optimistic lock release fails, additional human intervention is required. When an optimistic lock release fails, an alert email and text message are sent to notify the developer to manually release the optimistic lock.
[0138] As mentioned above, the auxiliary check phase includes program checks and Zookeeper checks. During normal task execution, it's impossible to guarantee 100% service reliability, so the auxiliary check phase is added. The auxiliary check phase includes two types of checks: program checks and Zookeeper checks. Program checks primarily address issues that occur during task execution. They are somewhat similar to step S307 in the task execution phase described above, but the processing occurs at a different time point. In step S307 of the task execution phase, when the task has completed but there's a problem releasing the optimistic lock, manual intervention can be performed to release the optimistic lock. However, program checks primarily address issues that occur during task execution, leading to task interruptions.
[0139] The following is combined with Figure 4 Describe the major steps of an exemplary program review phase.
[0140] Figure 4 FIG. 1 is a flow chart of an exemplary program checking phase according to an embodiment of the present invention. Figure 4 As shown, an exemplary program checking stage 400 according to an embodiment of the present invention includes steps S401 , S402 , S403 and S404 .
[0141] Step S401: The inspection program starts.
[0142] In one embodiment, after the check program is started, it obtains the executing task object from the task table.
[0143] Step S402: Obtain the holding time of the optimistic lock by the task.
[0144] In one embodiment, the time when the task acquires the optimistic lock is subtracted from the current time to obtain the time when the current task holds the optimistic lock.
[0145] Step S403: Determine whether the holding time exceeds 30 minutes.
[0146] In other embodiments, the value of 30 minutes can be configured according to the execution length of the task. It is generally configured as the time that the task with the longest execution time among all tasks holds the optimistic lock. Therefore, this condition will not be triggered under normal circumstances unless there is a problem with the program.
[0147] Step S404: If it is determined that the holding time exceeds 30 minutes, an alarm email and text message are sent.
[0148] If a task holds an optimistic lock for more than 30 minutes, an alert email and SMS message will be sent. Upon receiving the alert, the developer will check whether the task has been interrupted. If so, they will manually release the optimistic lock. If the task is still executing, they will ignore it.
[0149] Step S106: When an error event occurs in the second server among the one or more servers, other servers among the one or more servers are enabled to trigger a monitoring event, wherein the first server and the second server are the same or different.
[0150] Step S106 is a Zookeeper check that is synchronized with the normal execution process of the optimistic locking solution described in steps S101 to S105.
[0151] When using database optimistic locking to schedule a distributed system, if a periodic task (for example, executed every hour) is being executed, then the task has been locked in the database, and other Docker servers executing tasks cannot execute the task. If the Docker executing the task suddenly crashes when the task is halfway through, the lock for the task in the database cannot be released. If no measures are taken to release the lock, the task will never be executed again (it was originally executed every hour). The Zookeeper check provided in this article is to solve the problem of the lock being unable to be released caused by Docker crashes. In one embodiment, the "release" action here is performed by one of the Dockers.
[0152] Zookeeper checks primarily address Docker downtime issues. For example, if Docker_1 is executing Task A and then crashes, but Task A has not yet completed, the optimistic lock held by Task A will not be released. When Docker_2 attempts to execute Task A, it discovers that the lock has not been released and abandons Task A, causing it to never execute.
[0153] Preferably, before receiving one or more query requests for the task lock object of the task from one or more servers, the method further includes:
[0154] Creating a parent node corresponding to the check server; and
[0155] One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
[0156] Optionally, the one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
[0157] Optionally, the parent node and the one or more child nodes are EPHEMERAL type nodes.
[0158] Preferably, causing other servers among the one or more servers to trigger a monitoring event further comprises:
[0159] Deleting the second child node corresponding to the second server from the list to obtain a new list;
[0160] Sending the new list to the child nodes in the new list;
[0161] receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock;
[0162] querying the database according to the lock query request; and
[0163] In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
[0164] Optionally, when the second child node holds an unreleased lock, the task lock state of the task is 1.
[0165] Optionally, releasing the lock further includes: setting the task lock state of the task to 0.
[0166] The following is combined with Figure 5 Describe the main steps of an exemplary Zookeeper check phase.
[0167] Figure 5 FIG. 1 is a flow chart of an exemplary Zookeeper inspection phase according to an embodiment of the present invention. Figure 5 As shown, an exemplary Zookeeper checking phase 500 according to an embodiment of the present invention includes steps S501, S502, S503, S504, S505, S506 and S507.
[0168] Step S501: Zookeeper detects the heartbeat of one or more Dockers.
[0169] All Docker containers executing tasks register with Zookeeper. The registration logic is as follows: First, a node / SERVERS is created on the Zookeeper server. Then, each Docker container creates an EPHEMERAL node under this node upon startup. For example, Docker_1 creates / SERVERS / ${Docker_1's IP}, and Docker_2 creates / SERVERS / ${Docker_2's IP}. Docker_1, Docker_2, ..., and Docker_n all watch the parent node / SERVERS. An important characteristic of EPHEMERAL nodes is that if the connection between the client and the Zookeeper server is lost, the node disappears. Therefore, if a Docker container crashes, its corresponding node disappears, and all clients in the cluster watching / SERVERS are notified.
[0170] Step S502: It is found that one of the one or more Dockers is down.
[0171] Step S503: The surviving Docker executes the watcher monitor.
[0172] If Docker_1 goes down, other Dockers (except Docker_1) will trigger a monitoring event and perform the following operations: First, get all child nodes under the / SERVERS node (a list of all surviving Docker IP addresses). Because Docker_1 is down, all child nodes under the / SERVERS node do not contain Docker_1's IP address. Therefore, the IP address of Docker_1 can be obtained by comparing it with the previous IP address list.
[0173] Step S504: The surviving Docker checks whether the downed Docker holds a lock.
[0174] Step S505: In response to determining that the downed Docker holds a lock for a specific task, the lock is released.
[0175] After obtaining Docker_1's IP address, other Docker machines query the database for tasks that Docker_1 has not yet completed (that is, Docker_1 still holds the optimistic lock for the task and has not released it, but Docker_1 has crashed and will never release the optimistic lock), and release the optimistic lock for the task held by Docker_1.
[0176] The terms in this application have their general meanings. The optimistic locking mechanism adopts a more relaxed locking mechanism. Most of them are implemented based on the data version recording mechanism. The term "Zookeeper" is a distributed, open source distributed application coordination service. It is an open source implementation of Google's Chubby and an important component of Hadoop and Hbase. It is a software that provides consistency services for distributed applications. The functions provided include: configuration maintenance, domain name services, distributed synchronization, group services, etc. The term "Docker" is an open source application container engine that allows developers to package their applications and dependency packages into a portable container, and then publish them to any popular Linux machine. It can also achieve virtualization. Containers use a complete sandbox mechanism and there will be no interface between them.
[0177] The present invention adopts database optimistic locking technology and Zookeeper to complete distributed task scheduling, in the task acquisition lock stage, it is ensured that each task will only be assigned to a thread at the same time by database optimistic locking, and the lock is released again after the task is executed. Because of the reliability and ease of use of the database, the scheme is simple and reliable to implement, but there is a problem. When Docker breaks down, the lock obtained by the task cannot be released. At this time, we can detect the Docker heartbeat by Zookeeper. When Docker goes down, Zookeeper can detect the Docker of the downtime and execute watcher monitor to release the lock. In this scheme, when Docker breaks down, the release of the lock is completed by Zookeeper. Zookeeper is just a kind of auxiliary equipment in the whole process, and does not strongly rely on Zookeeper, which solves the problem of Redis lock failure and the reliability problem of Zookeeper. In some embodiments, the embodiment of the present application can also be applied to pessimistic lock.
[0178] One embodiment of the above invention has the following advantages or beneficial effects: because database optimistic locking is used to schedule tasks in a distributed system while Zookeeper is used for auxiliary inspection, the technical problems of lock failure caused by improper key expiration time setting and strong dependence on Zookeeper are overcome, thereby achieving the technical effect of improving the reliability of the distributed system and reducing the complexity of solution deployment.
[0179] Figure 6 FIG. 1 is a schematic diagram of the main process of another distributed task scheduling method according to an embodiment of the present invention. Figure 6As shown, another distributed task scheduling method according to an embodiment of the present invention includes steps S601, S602, S603, S604, S605, S606, S607, S608, S609 and S610.
[0180] Step S601: Create a parent node corresponding to the inspection server.
[0181] Step S602: creating one or more child nodes corresponding to the one or more servers under the parent node, wherein the parent node maintains all currently surviving child nodes under it in a list.
[0182] Step S603: Receive one or more query requests for the task lock object of the task from one or more servers.
[0183] Step S604: Query the task lock object in the database.
[0184] Step S605: In response to not finding the task lock object, writing a record related to the task lock object into the database.
[0185] Step S606: In response to querying the task lock object, read a record related to the task lock object from the database and send the record to the one or more servers, where the record includes a task lock state and a specific version number.
[0186] Step S607: Receive first update data for the task from a first server among the one or more servers, where the first update data includes a first version number.
[0187] Step S608: In response to determining that the first version number is different from the specific version number, a notification is sent to the first server.
[0188] Step S609: In response to determining that the first version number is the same as the specific version number, performing a first update on the database using the first update data.
[0189] Step S610: When an error event occurs in a second server among the one or more servers, other servers among the one or more servers are enabled to trigger a monitoring event, wherein the first server and the second server are the same or different.
[0190] Preferably, causing other servers among the one or more servers to trigger a monitoring event further comprises:
[0191] Deleting the second child node corresponding to the second server from the list to obtain a new list;
[0192] Sending the new list to the child nodes in the new list;
[0193] receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock;
[0194] querying the database according to the lock query request; and
[0195] In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
[0196] One embodiment of the above invention has the following advantages or beneficial effects: because database optimistic locking is used to schedule tasks in a distributed system while Zookeeper is used for auxiliary inspection, the technical problems of lock failure caused by improper key expiration time setting and strong dependence on Zookeeper are overcome, thereby achieving the technical effect of improving the reliability of the distributed system and reducing the complexity of solution deployment.
[0197] Figure 7 FIG. 1 is a schematic diagram of the main modules of the distributed task scheduling system according to an embodiment of the present invention. Figure 7 As shown, the distributed task scheduling system 700 according to an embodiment of the present invention includes:
[0198] The query request receiving module 701 is configured to receive one or more query requests for a task locking object of a task from one or more servers.
[0199] The lock object query module 702 is used to query the task lock object in the database.
[0200] The lock object processing module 703 is configured to, in response to querying the task lock object, read records related to the task lock object from the database and send the records to the one or more servers, wherein the records include a task lock state and a specific version number.
[0201] The update receiving module 704 is configured to receive first update data for the task from a first server among the one or more servers, where the first update data includes a first version number.
[0202] The update execution module 705 is configured to, in response to determining that the first version number is the same as the specific version number, execute a first update on the database using the first update data.
[0203] The auxiliary checking module 706 is configured to enable other servers among the one or more servers to trigger a monitoring event when an error event occurs in the second server among the one or more servers, wherein the first server and the second server are the same or different.
[0204] Optionally, the auxiliary inspection module 706 is further configured to:
[0205] Creating a parent node corresponding to the check server; and
[0206] One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
[0207] Optionally, the one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
[0208] Optionally, the parent node and the one or more child nodes are EPHEMERAL type nodes.
[0209] Optionally, the auxiliary inspection module 706 is further configured to:
[0210] Deleting the second child node corresponding to the second server from the list to obtain a new list;
[0211] Sending the new list to the child nodes in the new list;
[0212] receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock;
[0213] querying the database according to the lock query request; and
[0214] In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
[0215] Optionally, when the second child node holds an unreleased lock, the task lock state of the task is 1.
[0216] Optionally, the auxiliary checking module is further configured to: set the task lock state of the task to 0.
[0217] Optionally, the distributed task scheduling system 700 further includes:
[0218] A server startup module 707 is configured to receive the IP addresses of the one or more servers;
[0219] According to the IP address, query the database for tasks with a task lock status of 1; and
[0220] In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
[0221] Optionally, the locked object processing module 703 is further configured to:
[0222] In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
[0223] Optionally, the record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
[0224] One embodiment of the above invention has the following advantages or beneficial effects: because database optimistic locking is used to schedule tasks in a distributed system while Zookeeper is used for auxiliary inspection, the technical problems of lock failure caused by improper key expiration time setting and strong dependence on Zookeeper are overcome, thereby achieving the technical effect of improving the reliability of the distributed system and reducing the complexity of solution deployment.
[0225] Figure 8 An exemplary system architecture 800 is shown to which a distributed task scheduling method or a distributed task scheduling system according to an embodiment of the present invention can be applied.
[0226] like Figure 8 As shown, system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. Network 804 is used to provide a medium for communication links between terminal devices 801, 802, 803 and server 805. Network 804 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0227] Users can use terminal devices 801, 802, and 803 to interact with server 805 via network 804 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 801, 802, and 803, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0228] The terminal devices 801 , 802 , and 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0229] Server 805 may be a server that provides various services, such as a backend management server (for example only) that supports shopping websites browsed by users using terminal devices 801, 802, and 803. The backend management server may analyze and process received data such as product information query requests, and feed back processing results (for example, target push information and product information—for example only) to the terminal device.
[0230] It should be noted that the distributed task scheduling method provided in the embodiment of the present invention is generally executed by the server 805 , and accordingly, the distributed task scheduling system is generally set in the server 805 .
[0231] It should be understood that Figure 8 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0232] Reference below Figure 9 , which shows a schematic structural diagram of a computer system 900 of a terminal device suitable for implementing an embodiment of the present invention. Figure 9 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0233] like Figure 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the system 900 are also stored in the RAM 903. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0234] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, and the like; an output section 907 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 908 including a hard disk and the like; and a communication section 909 including a network interface card such as a LAN card or a modem. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 910 as needed, so that computer programs read therefrom can be installed into the storage section 908 as needed.
[0235] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the system of the present invention are performed.
[0236] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0237] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0238] The modules described in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be provided in a processor. For example, they may be described as follows: a processor comprising a query request receiving module, a lock object query module, a lock object processing module, an update receiving module, an update execution module, and an auxiliary checking module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the lock object query module may also be described as a module for querying the database for the lock object of the task.
[0239] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently and not be assembled into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the device includes: receiving one or more query requests for a task lock object of a task from one or more servers; querying the task lock object in a database; in response to querying the task lock object, reading a record related to the task lock object from the database and sending the record to the one or more servers, the record including a task lock status and a specific version number; receiving first update data for the task from a first server among the one or more servers, the first update data including a first version number; in response to determining that the first version number is the same as the specific version number, performing a first update on the database using the first update data; and when an error event occurs on a second server among the one or more servers, causing other servers among the one or more servers to trigger a monitoring event, wherein the first server and the second server are the same or different.
[0240] One embodiment of the above invention has the following advantages or beneficial effects: because database optimistic locking is used to schedule tasks in a distributed system while Zookeeper is used for auxiliary inspection, the technical problems of lock failure caused by improper key expiration time setting and strong dependence on Zookeeper are overcome, thereby achieving the technical effect of improving the reliability of the distributed system and reducing the complexity of solution deployment.
[0241] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A distributed task scheduling method, characterized in that: include: receiving one or more query requests for a task lock object of a task from one or more servers; Query the task lock object in the database; In response to querying the task lock object, reading a record related to the task lock object from the database and sending the record to the one or more servers, the record including the task lock state and a specific version number, wherein the optimistic locking mechanism is implemented by introducing a version number field in the database table; receiving first update data for the task from a first server among the one or more servers, the first update data including a first version number; In response to determining that the first version number is the same as the specific version number, performing a first update on the database using the first update data; as well as When an error event occurs on a second server among the one or more servers, other servers among the one or more servers are triggered to trigger a monitoring event, wherein the first server and the second server are the same or different.
2. The method according to claim 1, characterized in that in, Before receiving one or more query requests for the task lock object of the task from one or more servers, the method further includes: Creating a parent node corresponding to the check server; and One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
3. The method according to claim 2, characterized in that in, The one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
4. The method according to claim 2, characterized in that in, The parent node and the one or more child nodes are EPHEMERAL type nodes.
5. The method according to claim 2, characterized in that in, Enabling other servers among the one or more servers to trigger a monitoring event further includes: Deleting the second child node corresponding to the second server from the list to obtain a new list; Sending the new list to the child nodes in the new list; receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock; querying the database according to the lock query request; and In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
6. The method according to claim 5, characterized in that in, When the second child node holds an unreleased lock, the task lock state of the task is 1.
7. The method according to claim 5, characterized in that in, Releasing the lock further includes: setting the task lock state of the task to 0.
8. The method according to claim 1, characterized in that in, Before receiving one or more query requests for the task lock object of the task from one or more servers, the method further includes: receiving IP addresses of the one or more servers; According to the IP address, query the database for tasks with a task lock status of 1; and In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
9. The method according to claim 1, characterized in that in, After querying the database for the task lock object, the following steps are further included: In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
10. The method according to claim 1 or 9, characterized in that in, The record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
11. A distributed task scheduling system, characterized in that: include: A query request receiving module, configured to receive one or more query requests for a task lock object of a task from one or more servers; A lock object query module, used for querying the task lock object in the database; a lock object processing module, configured to, in response to querying the task lock object, read a record related to the task lock object from the database and send the record to the one or more servers, wherein the record includes a task lock state and a specific version number, and the optimistic locking mechanism is implemented by introducing a version number field in the database table; an update receiving module, configured to receive first update data for the task from a first server among the one or more servers, wherein the first update data includes a first version number; an update execution module, configured to, in response to determining that the first version number is the same as the specific version number, execute a first update on the database using the first update data; as well as An auxiliary checking module is used to enable other servers among the one or more servers to trigger a monitoring event when an error event occurs in the second server among the one or more servers, wherein the first server and the second server are the same or different.
12. The system according to claim 11, wherein: in, The auxiliary inspection module is further used to: Creating a parent node corresponding to the check server; and One or more child nodes corresponding to the one or more servers are created under the parent node, wherein the parent node maintains all currently surviving child nodes under the parent node in a list.
13. The system according to claim 12, wherein: in, The one or more sub-nodes are a list of IP addresses of the one or more servers corresponding thereto.
14. The system according to claim 12, wherein: in, The parent node and the one or more child nodes are EPHEMERAL type nodes.
15. The system according to claim 12, wherein: The auxiliary inspection module is further used for: Deleting the second child node corresponding to the second server from the list to obtain a new list; Sending the new list to the child nodes in the new list; receiving a lock query request from a server corresponding to a child node in the new list, wherein the lock query request is about whether the second child node holds an unreleased lock; querying the database according to the lock query request; as well as In response to querying that the second child node holds an unreleased lock, releasing the lock, wherein the lock is an optimistic lock for the task.
16. The system according to claim 15, wherein: in, When the second child node holds an unreleased lock, the task lock state of the task is 1.
17. The system according to claim 15, wherein: in, The auxiliary checking module is further used to: set the task lock state of the task to 0.
18. The system according to claim 11, wherein: The system further comprises: A server startup module, configured to receive the IP addresses of the one or more servers; According to the IP address, query the database for tasks with a task lock status of 1; and In response to querying the task whose task lock state is 1, setting the task lock state of the task to 0.
19. The system according to claim 11, wherein: The locked object processing module is further configured to: In response to not finding the task lock object in the query, a record related to the task lock object is written into the database.
20. The system according to claim 11 or 19, characterized in that in, The record includes at least the following fields: a task type field, a task description field, a task lock status field, and a version number field.
21. A distributed task scheduling electronic device, characterized in that: include: one or more processors; a storage system for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.
22. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Domain name system DNS server query method and apparatus
CN107613040A
Task scheduling method and device
CN108563502A