Task processing method and device, electronic equipment and storage medium

By electing a target service node and dynamically updating the set of valid service nodes in the asynchronous task processing system, the problem of task interruption caused by service node failure is solved, thus improving the stability and efficiency of the system.

CN120935176APending Publication Date: 2025-11-11INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511259708.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing asynchronous task processing systems lack self-healing capabilities when service nodes fail, leading to task execution interruptions and impacting stability and efficiency.

Method used

The target service node is elected by multiple service nodes competing for a distributed lock. The running status of all service nodes in the cluster is obtained in real time. Normal nodes are classified into the set of valid service nodes, and tasks are assigned by the target service node to achieve dynamic task redirection.

Benefits of technology

It improves the stability and efficiency of asynchronous task processing, avoids task stagnation caused by service node downtime, and realizes the system's self-healing capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935176A_ABST
    Figure CN120935176A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method and device, electronic equipment and a storage medium, and relates to the field of distributed technologies. The method comprises the following steps: competing for a distributed lock through a plurality of service nodes, and determining a target service node used for coordinating tasks from the plurality of service nodes; running states of all service nodes in the distributed service cluster are obtained through the target service node, the service nodes with the normal running states are classified as an effective service node set, and the effective service node set comprises at least one effective service node; and distributing the plurality of to-be-processed tasks to the effective service node through the target service node, so that the effective service node executes the distributed to-be-processed tasks. According to the technical scheme provided by the invention, the asynchronous task processing stability and efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed technology, and in particular to a task processing method, apparatus, electronic device and storage medium. Background Technology

[0002] In asynchronous task processing scenarios, existing systems store information about each service node (such as internet protocol addresses and service ports) in a maintenance service cluster table. When a new task is added, the task parameters, type, and execution service node information are stored in the asynchronous task table. Each service node's background thread then periodically retrieves and executes the corresponding task, meeting basic task allocation requirements. However, this architecture has a technical flaw that urgently needs to be addressed: insufficient robustness. If a service node suddenly crashes, all assigned and newly assigned tasks to that node will stall until the node recovers, resulting in task execution interruption. The existing system lacks self-healing capabilities, severely impacting the stability and efficiency of asynchronous task processing. Therefore, there is an urgent need to design a task processing method that improves the stability and efficiency of asynchronous task processing. Summary of the Invention

[0003] This application provides a task processing method, apparatus, electronic device, and storage medium that can improve the stability and efficiency of asynchronous task processing.

[0004] Firstly, this application provides a task processing method applied to a distributed service cluster, the method comprising:

[0005] By having multiple service nodes compete for a distributed lock, a target service node for coordinating tasks is determined from these multiple service nodes.

[0006] The target service node is used to obtain the running status of all service nodes in the distributed service cluster, and the service nodes with normal running status are classified into a set of effective service nodes, which includes at least one effective service node.

[0007] The target service node assigns multiple pending tasks to the effective service node, so that the effective service node executes the assigned pending tasks.

[0008] Furthermore, the distributed lock includes a Redis remote dictionary service lock and a database lock; the step of determining the target service node for coordinating tasks from among the multiple service nodes through competition for the distributed lock includes: competing for the Redis lock among the multiple service nodes; determining the service node that wins the competition for the Redis lock among the multiple service nodes as the target service node; if the Redis cluster cannot provide lock services, competing for the database lock among the multiple service nodes; and determining the service node that wins the competition for the database lock among the multiple service nodes as the target service node.

[0009] Furthermore, after determining the service node that has won the competition for the Redis lock among the plurality of service nodes as the target service node, the method further includes: determining the lock holding time of the Redis lock; starting a background daemon thread, and determining the running status of the target service node based on the background daemon thread; if the running status is "task processing in progress", then renewing the lock holding time at preset intervals, wherein the duration of the preset interval is less than the duration of the lock holding time; if the running status is "task processing completed" or "abnormal", then releasing the Redis lock when the lock holding time ends.

[0010] Furthermore, the Redis cluster is determined to be unable to provide lock services by at least one of the following methods: when multiple service nodes perform lock operations on the Redis cluster, a preset number of operation failures occur consecutively, the operation failures include at least one of connection timeout, response timeout, and operation rejection; the Redis cluster is determined to be in an abnormal state, the abnormal state includes at least one of cluster master and slave nodes disconnecting, a majority of cluster nodes going offline, and the cluster being unable to respond to status query requests.

[0011] Furthermore, the task status of the pending tasks includes at least two categories: pending service node assignment and pending assignment to a default service node, wherein the default service node is one of the plurality of service nodes. Assigning the plurality of pending tasks to the effective service node through the target service node includes: obtaining a first task from the plurality of pending tasks, wherein the first task is characterized by a task status of pending assignment to a default service node and the default service node's running status being abnormal; resetting the task status of the first task to pending assignment to a service node; and assigning the task with the pending assignment status to the effective service node based on the resource load status of the effective service node.

[0012] Furthermore, the task status also includes being processed by a first service node, where the first service node is one of the effective service nodes in the set; the step of allocating multiple pending tasks to the effective service nodes through the target service node includes: obtaining a second task from the multiple pending tasks, the second task being characterized by a task status of being processed by a first service node and the duration of being in this task status exceeding a preset time threshold; determining that the running status of the first service node is abnormal, removing the first service node from the set of effective service nodes, and resetting the task status of the second task to be assigned to a service node; and allocating tasks with a task status of being assigned to a service node from the multiple pending tasks to the effective service nodes according to the resource load status of the effective service nodes.

[0013] Furthermore, the task status of the pending task includes at least processing completion and processing failure; the method further includes: obtaining the processing result of the effective service node on the pending task through the target service node; if the processing result is successful, setting the task status of the pending task to processing completion; if the processing result is failed, obtaining the failure policy corresponding to the pending task, the failure policy including a reprocessing task policy or a direct alarm triggering policy; based on the reprocessing task policy, reassigning the pending task to the corresponding effective service node within the maximum number of failures, or generating task failure prompt information based on the direct alarm triggering policy.

[0014] Secondly, this application provides a task processing apparatus integrated into a distributed service cluster, the apparatus comprising:

[0015] The coordination node determination module is used to determine the target service node for the coordination task from multiple service nodes competing for a distributed lock.

[0016] The effective node determination module is used to obtain the running status of all service nodes in the distributed service cluster through the target service node, and classify the service nodes with normal running status into the effective service node set, wherein the effective service node set includes at least one effective service node.

[0017] The task processing module is used to assign multiple pending tasks to the effective service node through the target service node, so that the effective service node executes the assigned pending tasks.

[0018] Thirdly, this application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the task processing method described in any embodiment of this application.

[0019] Fourthly, this application provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the task processing method described in any embodiment of this application.

[0020] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the task processing method described in any embodiment of this application.

[0021] To address the shortcomings of existing technologies, this application provides a task processing method. This method offers the following advantages: First, it elects a target service node through a competitive distributed lock among multiple service nodes, avoiding the drawback of a lack of unified coordinator in traditional architectures. Subsequent task allocation is then coordinated by the target service node. Second, the target service node obtains the real-time running status of all service nodes in the cluster, adding those in normal running status to the set of valid service nodes. When a service node suddenly crashes, it is removed from the set of valid service nodes, and pending tasks are no longer assigned to that node, resolving the problem of task stagnation caused by a crashed service node. Finally, the target service node allocates tasks to valid service nodes. The dynamic updating of the set of valid service nodes ensures that tasks are always allocated to available service nodes, fundamentally compensating for the system's lack of self-healing capabilities and improving the stability and efficiency of asynchronous task processing.

[0022] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the task processing device, or it may be packaged separately from the processor of the task processing device; this application does not impose any limitations on this.

[0023] The descriptions of the second, third, and fourth aspects in this application can be referenced to the detailed description of the first aspect; and the beneficial effects described in the second, third, and fourth aspects can be referenced to the analysis of the beneficial effects of the first aspect, which will not be repeated here.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A flowchart illustrating a task processing method provided in an embodiment of this application;

[0028] Figure 2 This is a schematic diagram of the task processing framework provided in the embodiments of this application;

[0029] Figure 3 This is a schematic diagram of the structure of a task processing device provided in an embodiment of this application;

[0030] Figure 4 This is a block diagram of an electronic device used to implement a task processing method according to an embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0032] It should be noted that the terms "first," "second," "target," and "original," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein. Furthermore, the terms "comprising," "having," and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] Figure 1 This is a flowchart illustrating a task processing method provided in an embodiment of this application. This embodiment is applicable to scenarios where asynchronous tasks are processed in parallel based on a distributed service cluster. The task processing method provided in this embodiment can be executed by the task processing device provided in this embodiment. This device can be implemented in software and / or hardware and integrated into the electronic device executing this method.

[0034] See Figure 1 The method in this embodiment includes, but is not limited to, the following steps:

[0035] S110. By having multiple service nodes compete for a distributed lock, the target service node for coordinating the task is determined from among the multiple service nodes.

[0036] In this context, a service node can be a server or process within a distributed cluster that possesses independent processing capabilities and can participate in task execution or coordination. A distributed lock can be a synchronization mechanism used to resolve resource contention in a distributed system, achieving exclusive control across service nodes through shared storage media (such as a remote dictionary server (Redis) or a database). The target service node is the service node that wins the competition for the distributed lock; it undertakes task coordination and state management, ensuring orderly task execution, and acts as a temporary coordinator in the distributed cluster; the target service node can also become a coordinating service node.

[0037] In this embodiment, multiple service nodes in a distributed service cluster are orchestrated and deployed using an open-source container orchestration platform. Each service node automatically registers its own information (e.g., internet protocol address, port, service type) with an open-source distributed key-value store system upon startup and automatically unregisters upon shutdown. Each service node periodically (e.g., every 30 seconds) sends heartbeat packets to this distributed key-value store system to update its running status. If a service node loses heartbeats for a preset number of consecutive times (e.g., 3 times), its running status is marked as abnormal. This application, based on the container orchestration platform and distributed key-value store system, ensures the horizontal scalability of the service nodes themselves, abandons various configuration tables in existing frameworks that are difficult to extend, completely replaces static configuration tables, and supports second-level scaling.

[0038] A unique lock key (e.g., a Redis key-value pair with the key "coordinator_lock") is preset in the storage medium of the distributed lock (e.g., Redis). This lock is exclusive, meaning that only one service node is allowed to hold it at a time. When the distributed system starts, the existing coordination service node fails, or a periodic coordinator election occurs, all service nodes in the cluster automatically trigger the lock contention process.

[0039] Multiple service nodes simultaneously send lock acquisition requests to the distributed lock storage medium. Each request includes a unique identifier for the service node (e.g., internet address, port) and the lock holding time (to prevent permanent lock occupation due to node failure). The distributed lock storage medium ensures that only one service node successfully acquires the distributed lock through atomic operations. The service node that successfully acquires the distributed lock is marked as the target service node and is responsible for subsequent task coordination (e.g., task dispatch); service nodes that fail to acquire the distributed lock enter a waiting state, periodically retrying or listening for lock release events.

[0040] Specifically, distributed locks include Redis locks for remote dictionary services and database locks. The process involves multiple service nodes competing for the distributed lock, and then identifying the target service node for coordinating tasks. This includes: multiple service nodes competing for the Redis lock; identifying the service node that successfully acquires the Redis lock as the target service node; and immediately triggering a fault degradation process if the Redis cluster cannot provide lock services, involving multiple service nodes competing for the database lock; and identifying the service node that successfully acquires the database lock as the target service node.

[0041] During the period when the target service node holds the distributed lock, service nodes that have not acquired the distributed lock (i.e., non-coordinator service nodes) will not compete for the distributed lock to avoid concurrency storms, and will receive and process the asynchronous tasks assigned to them by the target service node (i.e., the coordinator). Non-coordinator service nodes can listen for lock release events and re-compete for the distributed lock when the target service node releases it. Optionally, the lock holding time of the target service node can be set according to the actual application, such as 60 seconds.

[0042] The advantage of this setup is that, through a two-layer degradation locking strategy using Redis locks and database locks, with the database lock serving as a fallback alternative to the Redis lock, a fault tolerance mechanism can be implemented. This addresses the problem in existing systems that rely on a single lock or configuration component (e.g., Redis), where the system cannot automatically switch to the backup solution when that component fails, leading to a complete failure of the coordination function. This solution in this application avoids a large number of service nodes operating on related tables simultaneously and ensures that the coordinator election is never interrupted.

[0043] If the lock holding time is set too long, a network failure during coordination by the target service node could prevent the lock from being released. If the lock holding time is set too short, the lock might expire before the business process is complete. For example, node A acquires the lock and starts executing a 20-second task, but the lock holding time is only 10 seconds. After 10 seconds, the lock is automatically released, and node B can then acquire the lock and operate on the same resource, leading to data conflicts. To address this issue, after identifying the service node that has won the Redis lock among multiple service nodes as the target service node, the following steps are taken: determining the Redis lock holding time; starting a background daemon thread (e.g., a watchdog thread), and determining the target service node's running status based on the background daemon thread; if the running status is "task processing in progress," renewing the lock holding time at preset intervals, where the preset interval is shorter than the lock holding time; if the running status is "task processing completed" or "an error occurred," releasing the Redis lock when the lock holding time expires.

[0044] Similarly, after identifying the service node that has won the database lock among the multiple service nodes as the target service node, the process further includes: determining the lock holding time of the database lock; starting a background daemon thread and determining the running status of the target service node based on the background daemon thread; if the running status is "task processing", then renewing the lock holding time at preset intervals, where the duration of the preset interval is less than the duration of the lock holding time; if the running status is "task processing completed" or "abnormal", then releasing the database lock when the lock holding time ends.

[0045] For example, renewing the lock holding time could be done as follows: if the lock holding time is 15 seconds, the watchdog thread would send a command to the lock's storage medium (such as Redis or a database) every 10 seconds to reset the lock holding time to 15 seconds.

[0046] Furthermore, the Redis cluster is determined to be unable to provide lock services by at least one of the following methods: when multiple service nodes perform lock operations on the Redis cluster, the operation fails for a preset number of consecutive times (e.g., 3 times), and the operation failure includes at least one of connection timeout, response timeout, and operation rejection; the Redis cluster is determined to be in an abnormal state, and the abnormal state includes at least one of cluster master and slave nodes disconnecting, a majority of cluster nodes going offline, and the cluster being unable to respond to status query requests.

[0047] S120. Obtain the running status of all service nodes in the distributed service cluster through the target service node, and classify the service nodes with normal running status into the set of valid service nodes. The set of valid service nodes includes at least one valid service node.

[0048] The operational status can be multi-dimensional data used to measure whether a service node is functioning properly, including network connectivity, hardware resource usage (e.g., Central Processing Unit (CPU), memory, disk), software process status, and service response performance. It serves as the basis for determining whether a service node is functioning correctly. The effective service node set refers to the collection of service nodes selected from the target service node that are in normal operational status. It can include the unique identifier of each service node, basic information, and real-time status data, and is the core basis for subsequent cluster task allocation and resource scheduling.

[0049] In this embodiment, the target service node obtains the running status of all service nodes from the aforementioned distributed key-value storage system. The target service node performs timestamp verification on the obtained running status data, retaining only real-time data within a preset time (e.g., 10 seconds) to avoid invalid data interfering with subsequent filtering. The target service node selects service nodes with normal running status (i.e., valid service nodes) from all service nodes and stores these valid service nodes in a valid service node set.

[0050] The selection of valid service nodes can be achieved by pre-defining a quantitative standard for normal operation. This standard should include at least network connectivity, resource consumption, and service availability, and may also include other performance standards. The target service node verifies the operational status data of each node based on this standard. Only when all indicators meet the acceptable criteria is the node considered to be operating normally and thus a valid service node. If any indicator exceeds the standard (e.g., CPU utilization reaches 90%), the node is considered to be operating abnormally and excluded from the selection process.

[0051] For example, resource usage standards can be: CPU utilization ≤ 80%, memory utilization ≤ 85%, and disk space ≥ 20GB; service availability standards can be: business process survival and business response latency ≤ 1000ms.

[0052] Ideally, if the selected set of effective service nodes is empty according to the above quantitative criteria, meaning all service nodes are in an abnormal running state, then the thresholds for non-core indicators in the above quantitative criteria should be relaxed. For example, the CPU utilization rate could be relaxed from no more than 80% to no more than 90%. Then, the selection should be carried out again according to the relaxed quantitative criteria to ensure that the final set of effective service nodes contains at least one effective service node, avoiding the extreme case of no available service nodes in the distributed service cluster.

[0053] S130. Distribute multiple pending tasks to valid service nodes through the target service node, so that the valid service nodes can execute the assigned pending tasks.

[0054] In this embodiment, the target service node obtains multiple pending tasks from the asynchronous task table and acquires the resource load status of each valid service node. Based on the resource load status of each valid service node, the target service node allocates the pending tasks to the corresponding valid service nodes according to a load balancing strategy. Each valid service node periodically (e.g., every 5 seconds) pulls (e.g., up to 10 records at a time) the allocated pending tasks from the asynchronous task table in the database and executes them. The load balancing strategy can be such that the lower the load of a valid service node, the higher its weight in allocating pending tasks.

[0055] Optionally, if the resource load status of each valid service node cannot be obtained (e.g., load data is unavailable), the target service node will randomly assign the task to be processed to any valid service node according to a random strategy.

[0056] Optionally, if the target service node fails to assign the task to be processed (e.g., the service node is unable to receive the task due to network interruption), the target service node resets the task status of the task to be processed to the pending service node and waits for the next round of allocation.

[0057] In one optional embodiment, the task status of the task to be processed includes at least a service node to be assigned and a default service node to be assigned, wherein the default service node is one of a plurality of service nodes;

[0058] When a user submits an asynchronous task (i.e., a task awaiting processing), the system writes the relevant task parameters into the asynchronous task table. Optionally, a service node is initially assigned to the task awaiting processing, designated as the default service node, and this default service node handles the task. The task status of this task is then marked as "assigned to a default service node and awaiting processing," where the node information of the default service node can be explicitly indicated in the task status. Alternatively, for cases where no default service node is initially assigned to a task awaiting processing, the task status of such tasks is marked as "awaiting a service node assignment."

[0059] The process involves assigning multiple pending tasks to valid service nodes through a target service node, including: obtaining a first task from multiple pending tasks, characterized by a task status of being assigned to a default service node and the default service node being in an abnormal running state; resetting the task status of the first task to be assigned to a service node, and reassigning a valid service node to the first task; and, based on the resource load of the valid service nodes, assigning tasks with a task status of being assigned to a service node (including the first task and tasks whose initial task status is itself being assigned to a service node) from among the multiple pending tasks to the valid service nodes, so that the valid service nodes can execute the assigned pending tasks.

[0060] Optionally, the execution process of the first task can be recorded in the repair log, such as recording the task number and repair time of the first task.

[0061] In another optional embodiment, the task status also includes being assigned to a first service node for processing, where the first service node is one of the valid service nodes in the set of valid service nodes. Based on the method of the above optional embodiments, after assigning a task with a pending service node status to a valid service node, then for a specific task to be processed, this task is assigned a first service node (i.e., one of the valid service nodes in the set of valid service nodes). Subsequently, if this first service node fails to meet the above quantitative criteria, it becomes a service node with an abnormal operating status.

[0062] To address this situation, multiple pending tasks are assigned to valid service nodes via the target service node. This includes: obtaining a second task from the multiple pending tasks, characterized by its task status being "allocated to a first service node for processing" and the duration of this task status exceeding a preset time threshold; indicating that the second task is being processed by the first service node, but the processing time has expired, suggesting that the first service node may be down. Therefore, if the first service node's operating status is determined to be abnormal, it is removed from the set of valid service nodes, and the task status of the second task is reset to "pending allocation to a service node," and a valid service node is reassigned to the second task. Based on the resource load of the valid service nodes, tasks with a "pending allocation to a service node" status (which also includes the second task and tasks whose initial task status was "pending allocation to a service node") are assigned to valid service nodes, enabling the valid service nodes to execute their assigned pending tasks.

[0063] In another optional embodiment, the task status of the task to be processed includes at least processing completion and processing failure; the task processing method of this application further includes: obtaining the processing result of the effective service node for the task to be processed through the target service node; if the processing result is successful, setting the task status of the task to be processed to processing completion; if the processing result is failed, obtaining the failure policy corresponding to the task to be processed, the failure policy including a reprocessing task policy or a direct alarm triggering policy; based on the reprocessing task policy, reassigning the task to be processed to the corresponding effective service node within the maximum number of failures, or generating task failure prompt information based on the direct alarm triggering policy.

[0064] The technical solution provided in this embodiment determines the target service node for coordinating tasks by having multiple service nodes compete for a distributed lock. The target service node then obtains the running status of all service nodes in the distributed service cluster, classifying those running normally into a set of valid service nodes, which includes at least one valid service node. The target service node then assigns multiple tasks to these valid service nodes, enabling them to execute their assigned tasks. This application first uses a distributed lock competition among multiple service nodes to elect the target service node, avoiding the drawback of lacking a unified coordinator in traditional architectures. Subsequent task allocation is then coordinated by the target service node. Secondly, the target service node obtains the running status of all service nodes in the cluster in real time, classifying those running normally into the set of valid service nodes. When a service node suddenly crashes, it is removed from the set of valid service nodes, and tasks are no longer assigned to that node, thus resolving the problem of task stagnation caused by a crashed service node. Finally, the target service node assigns tasks to available service nodes. The dynamic updating of the set of available service nodes ensures that tasks are always assigned to available service nodes, fundamentally compensating for the system's lack of self-healing capabilities and improving the stability and efficiency of asynchronous task processing.

[0065] Figure 2 This diagram illustrates the task processing framework provided in this application embodiment. Multiple service nodes in a distributed service cluster are orchestrated and deployed through an open-source container orchestration platform. The diagram shows three service nodes. Each service node automatically registers its information with an open-source distributed key-value store system upon startup. Each service node periodically sends heartbeat packets to this distributed key-value store system to update its running status. Multiple service nodes simultaneously compete for a Redis lock; the service node that successfully acquires the Redis lock is determined as the target service node. If the Redis cluster cannot provide lock services, a fault degradation process is immediately triggered, involving multiple service nodes competing for a database lock; the service node that successfully acquires the database lock is then determined as the target service node.

[0066] The target service node retrieves the running status of all service nodes from the distributed key-value storage system. The target service node scans multiple pending tasks in the asynchronous task table and repairs abnormal tasks (i.e., the first and second tasks in the above embodiment), that is, resetting the task status of the abnormal tasks to "awaiting allocation to a service node." The target service node assigns the multiple pending tasks to valid service nodes. Each valid service node pulls its assigned pending task from the asynchronous task table in the database every 5 seconds and executes the pending task to obtain the processing result. If the processing result is successful, the task status of the pending task is set to "processing completed"; if the processing result is unsuccessful, a task failure alarm message is generated through the alarm system.

[0067] Figure 3 This is a schematic diagram of the structure of a task processing device provided in an embodiment of this application, as shown below. Figure 3 As shown, the device 300 is integrated into a distributed service cluster and may include:

[0068] The coordination node determination module 310 is used to determine the target service node for the coordination task from the multiple service nodes through competition for a distributed lock.

[0069] The effective node determination module 320 is used to obtain the running status of all service nodes in the distributed service cluster through the target service node, and classify the service nodes with normal running status into the effective service node set, wherein the effective service node set includes at least one effective service node.

[0070] The task processing module 330 is used to assign multiple pending tasks to the effective service node through the target service node, so that the effective service node executes the assigned pending tasks.

[0071] In one embodiment, the distributed lock includes a Redis lock and a database lock;

[0072] The aforementioned coordination node determination module 310 can be specifically used for: competing for the Redis lock among the multiple service nodes; determining the service node that wins the competition for the Redis lock among the multiple service nodes as the target service node; when the Redis cluster cannot provide lock services, competing for the database lock among the multiple service nodes; and determining the service node that wins the competition for the database lock among the multiple service nodes as the target service node.

[0073] In one embodiment, the coordination node determination module 310 can be specifically used to: after determining the service node that has won the competition for the Redis lock among the plurality of service nodes as the target service node, determine the lock holding time of the Redis lock; start a background daemon thread, and determine the running status of the target service node based on the background daemon thread; if the running status is in the process of task processing, then perform a renewal operation on the lock holding time at preset intervals, the duration of the preset interval being less than the duration of the lock holding time; if the running status is the end of task processing or an exception, then release the Redis lock when the lock holding time ends.

[0074] In one embodiment, the Redis cluster is determined to be unable to provide lock services by at least one of the following methods: when multiple service nodes perform lock operations on the Redis cluster, a preset number of operation failures occur consecutively, the operation failures include at least one of connection timeout, response timeout, and operation rejection; the Redis cluster is determined to be in an abnormal state, the abnormal state includes at least one of cluster master-slave node disconnection, a majority of cluster nodes being offline, and the cluster being unable to respond to status query requests.

[0075] In one embodiment, the task status of the task to be processed includes at least a service node to be assigned and a default service node to be assigned, wherein the default service node is one of the plurality of service nodes;

[0076] The task processing module 330 described above can be specifically used to: obtain a first task from the plurality of pending tasks, wherein the first task is characterized by a task status of being assigned to a default service node and the running status of the default service node being abnormal; reset the task status of the first task to be assigned to a service node; and assign the task with the task status of being assigned to a service node from the plurality of pending tasks to the effective service node according to the resource load status of the effective service node.

[0077] In one embodiment, the task status further includes being processed by a first service node, where the first service node is one of the valid service node sets;

[0078] The task processing module 330 described above can be specifically used to: obtain a second task from the plurality of pending tasks, wherein the second task is characterized by a task status of being processed by a first service node and the duration of being in this task status exceeds a preset time threshold; determine that the running status of the first service node is abnormal, remove the first service node from the set of effective service nodes, and reset the task status of the second task to be assigned to a service node; and assign the tasks with the task status of being assigned to a service node from the plurality of pending tasks to the effective service node according to the resource load status of the effective service node.

[0079] In one embodiment, the task status of the task to be processed includes at least processing completed and processing failed;

[0080] The task processing module 330 described above can be specifically used to: obtain the processing result of the effective service node for the task to be processed through the target service node; if the processing result is successful, set the task status of the task to be processed to be completed; if the processing result is failed, obtain the failure strategy corresponding to the task to be processed, the failure strategy including a reprocessing strategy or a direct alarm triggering strategy; based on the reprocessing strategy, reassign the task to be processed to the corresponding effective service node within the maximum number of failures, or generate task failure prompt information based on the direct alarm triggering strategy.

[0081] The task processing device provided in this embodiment can be applied to the task processing method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0082] Figure 4 This is a block diagram of an electronic device used to implement a task processing method according to an embodiment of this application. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0083] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or loaded from storage unit 18 into the random access memory 13. The random access memory 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output interface 15 is also connected to the bus 14.

[0084] Multiple components in electronic device 10 are connected to input / output interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0085] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as task processing methods.

[0086] In some embodiments, the task processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the task processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the task processing method by any other suitable means (e.g., by means of firmware).

[0087] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), system-on-a-chip (SoC) systems, payload programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0088] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0089] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0091] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0092] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to address the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.

[0093] Note that the above are merely preferred embodiments and technical principles applied in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. For example, those skilled in the art can use the various forms of processes shown above to reorder, add, or delete steps; the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this application can be achieved, and no limitations are imposed herein.

[0094] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A task processing method, characterized in that, Applied to a distributed service cluster, the method includes: By having multiple service nodes compete for a distributed lock, a target service node for coordinating tasks is determined from these multiple service nodes. The target service node is used to obtain the running status of all service nodes in the distributed service cluster, and the service nodes with normal running status are classified into a set of effective service nodes, which includes at least one effective service node. The target service node assigns multiple pending tasks to the effective service node, so that the effective service node executes the assigned pending tasks.

2. The task processing method according to claim 1, characterized in that, The distributed lock includes a remote dictionary service Redis lock and a database lock; The step of determining the target service node for coordinating the task from among multiple service nodes through competition for a distributed lock includes: The multiple service nodes compete for the Redis lock; The service node that wins the competition for the Redis lock among the plurality of service nodes is determined as the target service node; In cases where the Redis cluster cannot provide lock services, the database lock is contested by the multiple service nodes. The service node that wins the competition for the database lock among the multiple service nodes is determined as the target service node.

3. The task processing method according to claim 2, characterized in that, After determining the service node that wins the competition for the Redis lock among the plurality of service nodes as the target service node, the method further includes: Determine the lock holding time of the Redis lock; Start a background daemon thread and determine the running status of the target service node based on the background daemon thread; If the running status is "task processing", the lock holding time is renewed every preset period, and the duration of the preset period is less than the duration of the lock holding time. If the running status is "task processing completed" or "abnormal", then the Redis lock is released when the lock holding time ends.

4. The task processing method according to claim 2, characterized in that, Determine that the Redis cluster cannot provide locking services through at least one of the following methods: When the multiple service nodes perform lock operations on the Redis cluster, a preset number of operation failures occur consecutively. The operation failures include at least one of connection timeout, response timeout, and operation rejection. The Redis cluster is determined to be in an abnormal state, which includes at least one of the following: the cluster master and slave nodes are disconnected, a majority of cluster nodes are offline, and the cluster is unable to respond to status query requests.

5. The task processing method according to claim 1, characterized in that, The task status of the pending task includes at least two service nodes: one to be assigned and the other to be assigned to a default service node. The default service node is one of the plurality of service nodes. The process of assigning multiple pending tasks to the valid service nodes through the target service node includes: Obtain a first task from the plurality of pending tasks, wherein the first task is characterized in that its task status is that it has been assigned to a default service node and is pending processing, and the running status of the default service node is abnormal. Reset the task status of the first task to pending service node allocation; Based on the resource load status of the effective service node, tasks with a pending service node status are assigned to the effective service node from among the multiple pending tasks.

6. The task processing method according to claim 5, characterized in that, The task status also includes being processed by a first service node, where the first service node is one of the valid service nodes in the set; the step of assigning multiple tasks to be processed through the target service node to the valid service nodes includes: Obtain a second task from the plurality of pending tasks. The second task is characterized in that its task status is that it has been assigned to a first service node for processing and the duration of this task status exceeds a preset time threshold. If the running status of the first service node is determined to be abnormal, the first service node is removed from the set of valid service nodes, and the task status of the second task is reset to a service node to be assigned. Based on the resource load status of the effective service node, tasks with a pending service node status are assigned to the effective service node from among the multiple pending tasks.

7. The task processing method according to claim 1, characterized in that, The task status of the task to be processed includes at least two states: processing completed and processing failed; the method further includes: The processing result of the effective service node for the task to be processed is obtained through the target service node; If the processing result is successful, then the task status of the task to be processed is set to processing completed; If the processing result is a processing failure, then the failure strategy corresponding to the task to be processed is obtained, and the failure strategy includes a reprocessing strategy or a direct alarm triggering strategy. Based on the reprocessing task strategy, the task to be processed is reassigned to the corresponding valid service node within the maximum number of failures, or a task failure prompt message is generated based on the direct triggering alarm strategy.

8. A task processing device, characterized in that, Integrated into a distributed service cluster, the device includes: The coordination node determination module is used to determine the target service node for the coordination task from multiple service nodes competing for a distributed lock. The effective node determination module is used to obtain the running status of all service nodes in the distributed service cluster through the target service node, and classify the service nodes with normal running status into the effective service node set, wherein the effective service node set includes at least one effective service node. The task processing module is used to assign multiple pending tasks to the effective service node through the target service node, so that the effective service node executes the assigned pending tasks.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that is executed by the at least one processor, which enables the at least one processor to perform the task processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the task processing method according to any one of claims 1 to 7.