Deadlock detection method, device and storage medium

By deploying deadlock detection components on each node of the distributed system, using the method of timeout triggering and lock waiting map construction, the problem of low centralized detection efficiency is solved and efficient and accurate distributed deadlock detection is achieved.

WO2025141482A1PCT designated stage expired Publication Date: 2025-07-03CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Patent Information

Application Number
PCT/IB2024/063163
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In distributed database systems, existing deadlock detection methods usually adopt the detection idea of ​​a centralized system, resulting in poor detection efficiency. As the system scale expands, the detection pressure increases, affecting system performance.

Method used

A distributed deadlock detection method is proposed. By deploying deadlock detection components on each node in the distributed system, the deadlock detection task is started using the timeout trigger mechanism, and a global lock waiting map is built for detection, avoiding repeated detection and realizing distributed execution.

Benefits of technology

It improves the efficiency and accuracy of deadlock detection in distributed systems, adapts to system scalability, avoids concurrent conflicts in detection tasks, and reduces the need for serial detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024063163_03072025_PF_FP_ABST
    Figure IB2024063163_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a deadlock detection method, a device and a storage medium. In the embodiments of the present disclosure, a detection mechanism in which a triggering entity is also responsible for execution is provided, and each node in a distributed system can serve as an executor for deadlock detection, so that deadlock detection tasks in the distributed system can be executed in a distributed manner, rather than being concentrated on a single node, thereby effectively accommodating the scalability of the distributed system and preventing the detection efficiency from being affected as the system scales up. Moreover, a detection mechanism similar to optimistic locking is further provided, so as to prevent problems such as redundant detection that may occur between the deadlock detection tasks distributed on the nodes, thereby ensuring the accuracy of deadlock detection. In addition, a single deadlock detection task can detect all deadlocks in the system in one step, eliminating the need for serial detection and thus further improving the efficiency of deadlock detection in the distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A Deadlock Detection Method, Device, and Storage Medium. This disclosure relates to the field of communications technology, and more particularly to a deadlock detection method, device, and storage medium. Background: In distributed database systems, many resources are accessed concurrently by multiple nodes. Therefore, lock services are required to ensure system correctness. When two or more nodes, while already occupying resources through lock requests, request resources already held by the other, they enter a waiting stalemate, resulting in a deadlock. Once a deadlock occurs, transactions involved in the deadlock cannot proceed, impacting system performance. Currently, deadlock detection in distributed database systems typically follows the same approach used in centralized systems. This involves centralizing deadlock detection on a single node within the distributed database system, which then performs deadlock detection on each process link in the distributed database system in a serial manner. As the system scale continues to grow, the detection pressure on this node also increases, resulting in poor deadlock detection efficiency. Therefore, how to more efficiently detect deadlock in distributed database systems has become an urgent issue to be addressed. SUMMARY OF THE INVENTION Various aspects of the present disclosure provide a deadlock detection method, device, and storage medium for improving the efficiency of deadlock detection in a distributed system. An embodiment of the present disclosure provides a deadlock detection method applicable to various nodes in a distributed system. For any node, the method includes: upon detecting a lock wait timeout event on the node, initiating a target deadlock detection task corresponding to the lock wait timeout event; monitoring whether any node in the distributed system is currently executing other deadlock detection tasks; if no such node exists, constructing a lock wait graph for the distributed system, the lock wait graph describing the lock wait relationships in the distributed system; and performing deadlock detection based on the lock wait graph to detect deadlocks in the distributed system. An embodiment of the present disclosure also provides a node device, comprising a memory, a processor, and a communication component. The node device is any node in the distributed system. The memory is configured to store one or more computer instructions. The processor is coupled to the memory and the communication component to execute the one or more computer instructions to perform the aforementioned deadlock detection method. An embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the aforementioned deadlock detection method.In the embodiments of the present disclosure, a trigger-and-execute detection mechanism is proposed. Each node in a distributed system can serve as the executor of deadlock detection. This allows the deadlock detection task in the distributed system to be executed in a distributed manner, rather than being concentrated on a single node. This effectively adapts to the scalability of the distributed system and prevents detection efficiency from being affected by system expansion. Furthermore, a detection mechanism similar to optimistic locking is proposed to avoid potential duplicate detection issues between deadlock detection tasks distributed across various nodes, thereby ensuring the accuracy of deadlock detection. Furthermore, a single deadlock detection task can detect all deadlocks in the system at once, eliminating the need for serial detection. This further improves deadlock detection efficiency in distributed systems. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are provided to provide a further understanding of the present disclosure and constitute a part of this disclosure. The illustrative embodiments of this disclosure and their description are provided to explain the present disclosure and are not intended to unduly limit it. In the accompanying drawings: Figure 1 is a schematic diagram of the structure of a distributed system provided by an exemplary embodiment of the present disclosure; Figure 2 is a flowchart of a deadlock detection method provided by an exemplary embodiment of the present disclosure; Figure 3 is a logical diagram of a monitoring solution provided by an exemplary embodiment of the present disclosure; Figure 4 is a flowchart of another deadlock detection method provided by an exemplary embodiment of the present disclosure; Figure 5 is a flowchart of yet another deadlock detection method provided by an exemplary embodiment of the present disclosure; Figure 6 is a flowchart of yet another deadlock detection method provided by an exemplary embodiment of the present disclosure; Figure 7a is a logical diagram of an exemplary solution for selecting a target process provided by an exemplary embodiment of the present disclosure; Figure 7b is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure; and Figure 8 is a schematic diagram of the structure of a node device provided by another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. Before we begin a detailed description of the technical solutions provided by various embodiments of this disclosure, we'll briefly explain several technical concepts involved. Distributed locks: In a distributed system, many resources are accessed concurrently by multiple nodes, requiring a global lock service to ensure system correctness. To ensure scalability and availability of the lock service, the lock itself is also distributed, providing services simultaneously on multiple nodes. This type of lock service is called a distributed lock.Lock wait: Node 1 holds a lock on a resource, while node 2 requests the same lock (mutual exclusion). Node 2 must wait for node 1 to release the lock. Deadlock: Deadlock occurs when two or more transactions, while simultaneously occupying resources, continue to request resources already held by the other, resulting in a wait-for-resource deadlock. Without external intervention, all transactions trapped in the deadlock cannot proceed. Currently, deadlock detection schemes in distributed database systems typically directly apply those used in centralized systems. That is, processes on different nodes in a distributed database system are treated as if they were on the same physical machine in a centralized system. Specifically, a deadlock detection program is deployed on the control node in the distributed database system. The control node communicates with all nodes in the system through the network to query each node for lock wait information. Based on this information, the control node performs deadlock detection. As can be seen, existing deadlock detection schemes are completely dependent on the control node. For control nodes, as the system scale continues to expand, the number of deadlocks in the system will also continue to increase, and the performance overhead caused by deadlock detection will also increase proportionally. Furthermore, control nodes often have other control tasks that also consume performance overhead. As a result, deadlock detection tasks on the control nodes may accumulate significantly, resulting in the inability to process deadlock detection tasks in a timely manner, which in turn affects the efficiency of deadlock detection in distributed database systems. To address this issue, embodiments of the present disclosure propose a new deadlock detection method to improve the efficiency of deadlock detection in distributed systems. The technical solutions provided by various embodiments of the present disclosure are described in detail below, in conjunction with the accompanying drawings. Figure 1 is a schematic diagram of the structure of a distributed system provided by an exemplary embodiment of the present disclosure. As shown in Figure 1, the system may include multiple nodes. The distributed system in this embodiment may be a distributed database system, a distributed storage system, or a distributed computing system. This embodiment does not limit the type of service provided by the distributed system. For different types of distributed systems, the types of resources that each node expects to occupy through distributed lock services may also vary. For example, in a distributed database system, the resources that each node expects to occupy through distributed lock services are typically data resources. For another example, in a distributed storage system, the resources that each node expects to occupy through a distributed lock service are typically storage resources. Further examples are not provided here, and in this embodiment, the types of resources that a node expects to occupy through a distributed lock service are not limited. Referring to Figure 1 , in this embodiment, a deadlock detection component can be deployed on each node in the distributed system.The deadlock detection method provided in this embodiment can be executed by these deadlock detection components. The deadlock detection components in this embodiment can be implemented as software, hardware, or a combination of software and hardware, without limitation herein. Considering the logical type of the deadlock detection method implemented on each node in a distributed system, for ease of description, the following detailed description of the deadlock detection method will be provided from the perspective of a single node. Figure 2 is a flowchart of a deadlock detection method provided by an exemplary embodiment of the present disclosure. Referring to Figure 2, for any node in a distributed system, the method may include: Step 200: Upon detecting a lock wait timeout event on the node, initiating a target deadlock detection task corresponding to the lock wait timeout event; Step 201: Monitoring whether there are nodes in the distributed system that are currently executing other deadlock detection tasks; Step 202: If no such nodes exist, constructing a lock wait graph for the distributed system. The lock wait graph is used to describe the lock wait relationships in the distributed system; Step 203: Performing deadlock detection based on the lock wait graph to detect deadlocks in the distributed system. Referring to Figures 1 and 2 , this embodiment uses a timeout triggering approach to trigger the deadlock detection task. A lock wait timeout refers to when the wait time after a lock request is issued exceeds a preset threshold. As mentioned above, this embodiment does not limit the types of services provided by the distributed system. In different types of distributed systems, the lock requesting entities may vary. For example, in a distributed database system, the lock requesting entity may be a lock group. A lock group appears to be a single entity requesting / releasing locks. A lock group typically includes one or more processes, and all processes within the same lock group are not affected by lock exclusivity. Lock groups can support tasks such as parallel query and two-phase commit in distributed database systems. For another example, in a distributed computing system, the lock requesting entity may be a single process. It is worth noting that the above examples only provide several examples of lock requesting entities that commonly appear in different types of distributed systems. This does not limit a single distributed system to only one type of lock requesting entity. This embodiment supports the simultaneous existence of multiple types of lock requesting entities in a distributed system, and this does not affect the implementation of the deadlock detection method in this embodiment. Based on this, referring to Figure 1 , if a lock wait timeout event occurs for lock request subject A on node 1, this will trigger the initiation of a deadlock detection task corresponding to the lock wait timeout event that occurred for lock request subject A on node 1. If a lock wait timeout event occurs for lock request subject B on node 2, this will trigger the initiation of a deadlock detection task corresponding to the lock wait timeout event that occurred for lock request subject B on node 2.For ease of description, in step 200, the deadlock detection task initiated on this node due to the current lock wait timeout event is described as the target deadlock detection task. This description will be used later in the description of the deadlock detection method in this section. During research, the inventors discovered that deadlocks in distributed systems are low-frequency, meaning that they do not occur very frequently. While the deadlock detection method provided in this embodiment requires a relatively short execution time of a single deadlock detection task, reaching milliseconds, deadlock detection tasks triggered by a lock wait timeout event in a distributed system may still experience concurrency issues. Specifically, multiple deadlock detection tasks may be initiated simultaneously in a distributed system. During research, the inventors discovered that this concurrency issue with deadlock detection tasks can lead to duplication in deadlock detection results. However, in this embodiment, it is desirable to use the deadlock detection results as the basis for deadlock resolution, and to maintain a distributed deadlock resolution process. Therefore, it is desirable for each node to continue to perform deadlock resolution work after completing deadlock detection. As a result, the deadlock detection results will directly affect the relevant operations performed by nodes during the deadlock resolution process. The duplication of deadlock detection results caused by the aforementioned deadlock detection task concurrency issue can lead to multiple nodes performing repeated deadlock resolution operations for the same deadlock. For example, referring to Figure 2, if a deadlock forms between lock requester A and lock requester B in Figure 2, and this deadlock causes lock wait timeout events for both lock requesters A and B to occur almost simultaneously, it is possible that nodes 1 and 2 will initiate deadlock detection tasks for the same deadlock at almost the same time, thus encountering the aforementioned concurrency issue. In this case, upon detecting the deadlock, node 1 will inevitably assume that lock requester B should release the lock and will therefore control lock requester B to release the lock. Node 2, upon detecting the deadlock, will inevitably assume that lock requester A should release the lock and will therefore control lock requester A to also release the lock. However, it's clear that in this case, the deadlock can be resolved by ensuring that only one of the lock requesting entities releases the lock; both lock requesting entities do not need to release the lock. As can be seen, the duplication of deadlock detection results caused by the concurrency of deadlock detection tasks can lead to excessive operations in the subsequent deadlock resolution process, which in turn may affect the efficiency of the mainline in the distributed system. Therefore, this embodiment also proposes a solution to this concurrency issue of deadlock detection tasks.Referring to Figure 1 , in step 201, after initiating the target deadlock detection task, the current node may first enter the monitoring phase. Specifically, this phase involves monitoring whether there are nodes in the distributed system currently executing other deadlock detection tasks. Other deadlock detection tasks are those triggered by other lock wait timeout events. Furthermore, the monitoring here encompasses the entire distributed system, encompassing both the current node and other nodes. "Executing" can be understood as "initiated." In actual applications, there may be some special cases, such as two deadlock detection tasks simultaneously in the aforementioned monitoring phase. In such cases, rules can be set to ensure exclusivity, such as requiring that the node that initiates monitoring first terminates the deadlock detection task first. Of course, this is merely illustrative; other rules can be employed in this embodiment to ensure exclusivity in these special cases, and further examples are not provided here. That is, in this embodiment, concurrent deadlock detection tasks are allowed in a distributed system. However, after initiating a deadlock detection task, a single node will first determine whether other deadlock detection tasks are already executing in the distributed system. If so, the node will immediately cease executing the current deadlock detection task. It should be understood that in this embodiment, the purpose of executing the monitoring phase after initiating the deadlock detection task is to schedule the deadlock detection tasks in the distributed system, ensuring that only one deadlock detection task is currently executing in the phases following the monitoring phase. This avoids the possibility of duplicate deadlock detection results caused by concurrent deadlock detection tasks. Based on this, this embodiment proposes that during the execution of the deadlock detection task, at least a lock wait graph construction phase and a deadlock detection phase may be set after the monitoring phase. Referring to Figure 2, in step 202, if the monitoring phase determines that no nodes in the distributed system are currently executing other deadlock detection tasks, the lock wait graph construction phase may be initiated. In this embodiment, the lock-wait graph is used to describe the lock-wait relationships within a distributed system. Specifically, the lock-wait graph in this embodiment is a global graph, not a single-node graph. To support the construction of the lock-wait graph, in this embodiment, the node can query each node in the distributed system for lock-wait information, which serves as the data foundation for constructing the lock-wait graph. In a preferred implementation, the node can concurrently issue lock-wait information collection instructions to each node in the distributed system (including the node itself), requesting each node to return lock-wait information. In this way, a lock-wait graph for the distributed system can be constructed based on the lock-wait information returned by each node. The parallel query approach proposed in this embodiment effectively improves the efficiency of constructing the lock-wait graph, thereby reducing the time required for a single deadlock detection task.During their research, the inventors discovered that to obtain lock wait information for each node, the node needs to establish network communication with other nodes in the distributed system. Furthermore, during the aforementioned monitoring phase, the node also needs to query each node whether it is currently executing a deadlock detection task. To this end, this embodiment further proposes moving the process of obtaining lock wait messages to the monitoring phase and communicating the deadlock detection task status on the node by specifying a response format for lock wait information collection instructions. Figure 3 is a logical diagram of a monitoring phase solution provided by an exemplary embodiment of the present disclosure. Referring to Figure 3, in this further solution, after initiating the target deadlock detection task, the node can concurrently initiate lock wait information collection instructions to each node in the distributed system. If no node returns a retry notification, the node determines that no nodes in the distributed system are currently executing other deadlock detection tasks. Upon receiving the lock wait message collection instruction, the node currently executing the deadlock detection task returns a retry notification as a response to the lock wait message collection instruction. Referring to Figure 3, in this further solution, node 1 can issue lock-wait message collection instructions to other nodes in the distributed system during the monitoring phase. Node 2 is currently executing a deadlock detection task (the black bar below node 2 in Figure 3 represents the duration of its deadlock detection task). Therefore, node 2 will return a retry notification in response to the lock-wait message collection instruction issued by node 1. Upon receiving this retry notification, node 1 can confirm that there are nodes in the distributed system currently executing other deadlock detection tasks. For nodes not currently executing deadlock detection tasks, such as node 3, a lock-wait message will be returned in response to the lock-wait message collection instruction issued by node 1. This way, if no nodes in the distributed system are currently executing other deadlock detection tasks, this node will not receive any retry notifications during the monitoring phase and can collect lock-wait information from all nodes in the distributed system. Thus, after entering the phase of constructing a lock-wait graph, this node can construct a lock-wait graph for the distributed system based on the lock-wait information collected during the monitoring phase. It is worth noting that the aforementioned process of moving the lock wait message acquisition process to the monitoring phase and conveying the deadlock detection task status on the node by agreeing on a response format for the lock wait information collection instruction is only a preferred solution in this embodiment. In this embodiment, other methods can also be used to implement the monitoring phase.For example, during the monitoring phase, the node can initiate a task execution status query command to determine whether each node is executing the deadlock detection task. The process of obtaining lock wait messages can be placed during the construction of the lock wait graph. During this process, the node can initiate another round of network communication to obtain lock wait messages from each node. Further example solutions are not provided here. Continuing with Figure 2, after completing step 202, the process proceeds to step 203, where deadlock detection is performed based on the lock wait graph constructed in step 202 to detect deadlocks in the distributed system. As mentioned above, the lock wait graph in this embodiment is a global graph. Therefore, in step 203, the node can detect all deadlocks in the distributed system based on the lock wait graph, rather than just those involved in the node itself. This single-node global detection concept can further improve the efficiency of deadlock detection in distributed systems. It's also worth noting that for each node in a distributed system, the execution of the deadlock detection task can be implemented in a bypass manner. That is, the deadlock detection task and the mainline work task in the node are completely asynchronous and do not interfere with each other. This ensures that the deadlock detection method in this embodiment has no impact on the mainline work performance of the distributed system. In summary, this embodiment proposes a trigger-and-execute detection mechanism, allowing each node in the distributed system to serve as the executor of deadlock detection. This allows the deadlock detection task in the distributed system to be executed in a distributed manner, rather than being concentrated on a single node. This effectively adapts to the scalability of the distributed system and eliminates the impact of detection efficiency due to system expansion. Furthermore, a detection mechanism similar to optimistic locking is proposed to avoid potential duplicate detection issues between deadlock detection tasks distributed across various nodes, thereby ensuring the accuracy of deadlock detection. Furthermore, a single deadlock detection task can detect all deadlocks in the system at once, eliminating the need for serial detection, further improving deadlock detection efficiency in the distributed system. FIG4 is a flow chart of another deadlock detection method provided by an exemplary embodiment of the present disclosure.4 , the deadlock detection method may include: step 400, when a lock wait timeout event is detected on the current node, starting a target deadlock detection task corresponding to the lock wait timeout event; step 401, monitoring whether there are nodes in the distributed system that are executing other deadlock detection tasks; if not, executing step 402; if so, executing step 404; step 402, constructing a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships in the distributed system, and continuing with step 403; step 403, performing deadlock detection based on the lock wait graph to detect deadlocks in the distributed system; step 404, stopping execution of the target deadlock detection task; using the reaching of a specified wait time as a trigger condition, triggering the restart of the target deadlock detection task, and returning to execution of step 401, waiting until there are no nodes in the distributed system that are executing other deadlock detection tasks, and completing the target deadlock detection task. Steps 400, 402, and 403 of this embodiment can be referred to the descriptions in the previous embodiments and will not be repeated here. Based on steps 401 and 404, this embodiment provides a solution for the situation where, after the target deadlock detection task is initiated, it is detected that a node in the distributed system is already executing another deadlock detection task. As mentioned in the previous embodiments, in this situation, execution of the target deadlock detection task is stopped. This embodiment further proposes that, in this situation, after a specified waiting time has elapsed, the target deadlock detection task is restarted. After the restart, it is again determined whether there are already executing deadlock detection tasks in the distributed system. This cycle continues until a suitable execution opportunity is found for the target deadlock detection task, and the target deadlock detection task is completed at the found execution opportunity. The specified waiting time in this embodiment is configurable as needed. In a preferred configuration, the waiting time can be a random value within a certain time interval to reduce the probability of task conflicts recurring. For example, in some extreme cases, two nodes might initiate deadlock detection tasks at the same time. Both nodes might detect a task conflict and stop their own deadlock detection tasks. If both nodes follow the same wait time, they would both restart at the same time, making it impossible to resolve the task conflict. Setting the wait time to a random value within a certain time interval can effectively reduce the probability of this happening.Referring to Figure 3 , for node 1, the target deadlock detection task on node 1 was stopped due to the presence of a deadlock detection task on node 2. However, after a specified waiting period, the target deadlock detection task was restarted on node 1 (the white bar below the node in Figure 3 represents the restarted target deadlock detection task). The restart mechanism proposed in this embodiment effectively prevents the omission of deadlock detection tasks in a distributed system. Deadlock detection tasks that were stopped due to the exclusivity of deadlock detection tasks in this embodiment can be completed at other times and will not be missed. During research, the inventors discovered that concurrency issues may exist between deadlock detection tasks on different nodes, and concurrency issues may also occur within the same node. Regarding the latter type of concurrency issue, again using this node as an example, it can be understood that a new lock wait timeout event occurred on this node during the execution of the target deadlock detection task. To address the latter type of concurrency issue, this embodiment can follow the execution flow in Figure 4 and treat new lock wait timeout events occurring on the node during the execution of the target deadlock detection task as ordinary lock wait timeout events. Specifically, upon detecting such a lock wait timeout event, a deadlock detection task is initiated for it. Through the aforementioned monitoring process, it is detected that the target deadlock detection task is currently being executed on the node, thereby stopping the deadlock detection task caused by such a lock wait timeout event and restarting it at an appropriate time. To address the latter type of concurrency issue, this embodiment proposes a further modification solution. Instead of treating new lock wait timeout events occurring on the node during the execution of the target deadlock detection task as ordinary lock wait timeout events, they are treated as special lock wait timeout events. Upon detecting such special lock wait timeout events, a waiting task is created for them. The waiting task will need to be passively initiated, rather than being immediately initiated upon creation like ordinary deadlock detection tasks. To this end, this further modification solution proposes setting up a waiting queue in each node of the distributed system to store waiting tasks. In this way, the node can place the waiting task created for the special lock wait timeout event into its waiting queue. Based on this, after completing the target deadlock detection task, the node can continue to control itself to maintain the state of executing the deadlock detection task and sequentially consume the waiting tasks in its waiting queue. After sequentially completing the deadlock detection tasks triggered by the waiting tasks in the waiting queue, the node can be controlled to end the state of executing the deadlock detection task.Specifically, in this further modification, for the aforementioned special lock wait timeout events, the corresponding deadlock detection task is no longer immediately initiated. Instead, deadlock detection tasks caused by these lock wait timeout events are queued within the node through a waiting queue. If there are waiting tasks in the waiting queue within the node, after completing the target deadlock detection task, the node will proceed to execute the deadlock detection task in the waiting queue without relinquishing the exclusive right to the deadlock detection task. It should be understood that if there are multiple waiting tasks in the waiting queue, the node will still adhere to the rule of executing only a single deadlock detection task at a time. After completing the deadlock detection task corresponding to a waiting task, it will consume the next waiting task in the waiting queue. Thus, with this further modification, for the aforementioned special lock wait timeout events, the monitoring process no longer needs to detect that the node is currently executing a deadlock detection task. Instead, the task is automatically queued within the node, eliminating numerous restarts and reducing node performance overhead. Furthermore, during research, the inventors discovered that, because the deadlock detection task provided by this embodiment executes very quickly, the aforementioned special lock wait timeout event may have already been resolved by the target deadlock detection task. Repeatedly executing the deadlock detection task for the lock wait timeout event would waste the performance of the node. Therefore, in the aforementioned further modification scheme, it is further proposed that: if it is detected that the lock requesting entity associated with the lock wait timeout event corresponding to the target waiting task in the waiting queue has died, the target waiting task is deleted from the waiting queue; and the lock held by the lock requesting entity associated with the lock wait timeout event corresponding to the target waiting task is released. The waiting tasks in the waiting queue are created based on the lock wait timeout event, and the lock wait timeout event corresponds to the lock requesting entity. Therefore, if the lock requesting entity has died, the lock held by the lock requesting entity can be directly released, and there is no need to execute the deadlock detection task for the lock requesting entity. This can further reduce the performance overhead of the node. During research, the inventors discovered that while this embodiment proposes that deadlock detection be performed by working nodes in a distributed system, the lock requesting entity that causes the deadlock is running on a working node. Based on this, this embodiment also proposes that, when the node shuts down, the locks held by each lock requesting entity on the node are released. Furthermore, after the node restarts, any unfinished deadlock detection tasks before the shutdown are no longer executed. In this way, node failures in a distributed system are positively utilized in this embodiment, turning the negative node failure issue into a deadlock resolution method.Based on the principle of "who triggers, who detects" in this embodiment, any deadlocks caused by this node will be resolved after the node shuts down, eliminating the need for rollback detection upon restart. Furthermore, each node is responsible for its own deadlock detection task, so it does not need to worry about whether deadlock detection tasks on other nodes need to be rolled back. In summary, in this embodiment, after starting the target deadlock detection task, if the node detects that a node in the distributed system is already executing other deadlock detection tasks, it can stop the target deadlock detection task and restart it to continuously test whether there is a deadlock detection window in the distributed system. This allows for the scheduling of deadlock detection tasks in the distributed system at a temporal level and prevents the omission of deadlock detection tasks. Furthermore, for special lock wait timeout events that occur during the execution of the node's deadlock detection task, deadlock detection tasks can be scheduled through intra-node queuing to prevent omissions, thereby reducing performance overhead for each node. In addition, node failures in a distributed system can be leveraged as a means of resolving deadlocks, further reducing the performance overhead of each node. Figure 5 is a flowchart illustrating another deadlock detection method provided by an exemplary embodiment of the present disclosure. Referring to Figure 5 , the deadlock detection method may include: Step 500: Upon detecting a lock wait timeout event on the current node, initiating a target deadlock detection task corresponding to the lock wait timeout event; Step 501: Monitoring whether there are nodes in the distributed system currently executing other deadlock detection tasks; Step 502: If not, constructing a lock wait graph using the lock request entities contained in each node in the distributed system as nodes and the lock wait relationships between the lock request entities as edges; Step 503: Annotating the out-degree and in-degree of the lock request entities in the lock wait graph; Step 504: Performing deadlock detection based on the lock wait graph containing the out-degree and in-degree to detect deadlocks in the distributed system. Steps 500 and 501 may be referred to the relevant descriptions in the previous embodiments and are not repeated here. In this embodiment, an optional implementation scheme for constructing a lock-wait graph and an optional implementation scheme for performing deadlock detection based on the lock-wait graph can be provided based on steps 502-504. Referring to Figure 5, in step 502, a lock-wait graph can be constructed using the lock request entities included in each node in the distributed system as nodes and the lock-wait relationships between lock request entities as edges. The lock-wait relationships can be extracted based on the lock-wait messages provided by each node in the aforementioned embodiment. For more information on this, reference can be made to existing lock-wait relationship extraction techniques and will not be elaborated upon here.Unlike traditional lock-wait graphs, this embodiment proposes annotating the out-degree and in-degree of lock requesting entities in the lock-wait graph. The out-degree indicates the number of locks the lock requesting entity is waiting for, while the in-degree indicates the number of locks held by the lock requesting entity that are being waited for by other lock requesting entities. When the out-degree is 0, the number of locks the lock requesting entity is waiting for is 0, meaning that the lock requesting entity does not need to wait for any other locks. When the in-degree is 0, the number of locks held by the lock requesting entity is not being waited for by any other lock requesting entities. The inventors have discovered through research that lock requesting entities with an out-degree or in-degree of 0 are not necessarily in any deadlock. Therefore, in this embodiment, in step 504, deadlock detection can be performed based on the lock-wait graph containing the out-degree and in-degree to detect deadlocks in the distributed system. That is, in step 504, by analyzing the out-degree and in-degree of each lock requesting entity, it is determined whether each lock requesting entity is in a deadlock. Ultimately, the deadlock can be detected from the lock waiting graph. As can be seen, this embodiment no longer uses the traditional depth-first search algorithm to utilize the lock waiting graph. Instead, a new deadlock detection solution based on the lock waiting graph is proposed. Specifically, deadlock detection is performed based on the out-degree and in-degree annotated for the lock requesting entity. In an exemplary detection scheme, a target point with an out-degree or in-degree of 0 is selected from the lock-wait graph; the target point and its associated edges are pruned from the lock-wait graph; the out-degree and in-degree corresponding to the remaining points in the lock-wait graph are updated; the target point selection, pruning, and updating operations are repeated until no more points with an out-degree or in-degree of 0 remain in the lock-wait graph, at which point the loop ends. If unpruned points and edges still exist in the lock-wait graph after the loop ends, a deadlock in the distributed system is determined based on the ring structure formed by the unpruned points and edges. This exemplary detection scheme employs pruning for deadlock detection. It is worth noting that the unit of pruning in this detection scheme is the lock request subject. This is because, in the lock-wait relationship, the lock request subject is the unit of lock request and, therefore, the lock request subject is the fundamental unit that causes deadlock. It is understandable that in this exemplary detection scheme, by cyclically executing the target point selection, pruning, and update operations, points with zero out-degree or in-degree and their associated edges in the lock-waiting graph are pruned. Thus, after the loop completes, the remaining points in the lock-waiting graph must have both non-zero out-degree and non-zero in-degree. According to the definitions of out-degree and in-degree, these remaining points hold locks, are being held by other points, and are themselves waiting for locks from other points. Therefore, these remaining points meet the characteristics of lock requesting entities in a deadlock and are definitely points in the deadlock.All points in a deadlock will remain, and the edges between the points can be used to connect the points in the same deadlock into a ring structure. Therefore, if a ring structure remains in the lock wait graph after pruning, the presence of a deadlock in the distributed system can be determined based on this ring structure. In other words, the lock request entities corresponding to the points contained in the remaining single link structure in the lock wait graph have caused a deadlock. Thus, in this embodiment, based on the out-degree and in-degree annotated for the lock request entities, deadlocks in the distributed system can be efficiently and comprehensively detected from the lock wait graph. During research, the inventors discovered that the lock state of a worker process in a distributed system is constantly changing (locking / releasing the lock), while the deadlock detection task in this embodiment is completely asynchronous with the worker processes in the distributed system. Therefore, timing issues may occur in the lock wait graph in this embodiment, resulting in discrepancies in the deadlock detected after deadlock detection based on the lock request graph. For example:

[0002] At time T1, the lock waiting information of node 1 is collected: lock request subject A is waiting for lock request subject B;

[0003] At time T2, the lock waiting information of node 2 is collected: lock request subject B is waiting for lock request subject A;

[0004] At T3, lock requester B releases the lock, and lock requester A successfully acquires the lock, no longer needing to wait.

[0005] Time T4: Create a lock wait graph:

[0006] Time T5: Deadlock detection is performed based on the lock wait graph. As can be seen, since the state change at T3 is not detected when the lock wait graph is constructed at T4, the deadlock detection operation at T5 may detect a deadlock between lock requesters A and B. This may result in a deadlock that does not actually exist, leading to skewed detection results. To ensure accurate deadlock detection, the worker process needs to be paused during the deadlock detection process to maintain the lock state of the worker process. However, this solution will inevitably have a negative impact on the performance of the distributed system. To address this timing issue, this embodiment proposes an exemplary solution: For each remaining ring structure in the lock request graph, the node can send a lock wait status confirmation request to the node containing each lock request subject in each link structure. The node then receives a response to the issued lock wait status confirmation request. If the response determines that the lock wait status of any lock request subject in the target ring structure has changed, the target ring structure is pruned from the lock wait graph. The remaining ring structures in the lock wait graph are then identified as deadlocks in the distributed system. This exemplary solution employs a secondary check mechanism. For deadlocks initially detected through pruning, lock wait status confirmation requests are only sent to the nodes involved in the deadlock to determine whether the lock wait status of each lock request subject in the deadlock has changed. If no changes have occurred, the deadlock detection is confirmed to be correct. If the lock wait state of any lock requesting entity in a deadlock has changed, this indicates that the aforementioned timing issue has caused an inaccurate lock wait graph. Such deadlocks can be promptly removed from the lock wait graph. This secondary check mechanism effectively improves the accuracy of deadlock detection. In summary, this embodiment proposes a new deadlock detection method based on the lock wait graph. This method labels the out-degree and in-degree of lock requesting entities and, based on this, performs vertex and edge pruning in the lock wait graph. This method discovers ring structures in the lock wait graph and detects deadlocks in distributed systems. Furthermore, a secondary check mechanism is proposed. By communicating with deadlock-related nodes, initially detected deadlocks can be reconfirmed to further ensure the accuracy of deadlock detection.FIG6 is a flow diagram of another deadlock detection method provided by an exemplary embodiment of the present disclosure. Referring to FIG6 , the method may include: Step 600: Upon detecting a lock wait timeout event on a local node, initiating a target deadlock detection task corresponding to the lock wait timeout event; Step 601: Monitoring whether a node in the distributed system is currently executing other deadlock detection tasks; Step 602: If no node exists, constructing a lock wait graph for the distributed system. The lock wait graph describes the lock wait relationships in the distributed system; Step 603: Performing deadlock detection based on the lock wait graph to detect deadlocks in the distributed system; Step 604: Upon detecting a deadlock in the distributed system, selecting a target process from among the processes involved in the target deadlock according to preset rules; and Step 605: Controlling the target process to release its held lock to resolve the target deadlock. For details about Steps 600 to 603, refer to the relevant descriptions in the previous embodiments and are not repeated here to save space. Based on steps 604 and 605, this embodiment proposes adding a deadlock resolution step during the execution of the deadlock detection task. This allows the deadlock resolution task to be performed by each node in the distributed system, enabling distributed execution of deadlock resolution within the distributed system. Referring to Figure 6 , in step 604, if a deadlock is detected in the distributed system, a target process is selected from the processes involved in the target deadlock according to preset rules. This ensures that, for a single deadlock detected, only the single process involved is addressed, effectively reducing the cost of deadlock resolution. This embodiment proposes an exemplary scheme for selecting a target process. Figure 7a is a logical diagram of an exemplary scheme for selecting a target process, provided in an exemplary embodiment of the present disclosure. Referring to Figure 7a, in this exemplary solution, each process involved in the target deadlock can be labeled with an in-degree in the lock-wait graph. Based on this, the process with the highest in-degree among the processes involved in the target deadlock can be selected as the target process. As mentioned above, in-degree represents the number of locks held by a lock requesting entity that are waited for by other lock requesting entities. Therefore, the in-degree labeled for a process here represents the number of locks held by a process that are waited for by other processes. Referring to Figure 7a, in this exemplary solution, a lock-wait relationship exists between nodes 1, 2, and 3. Specifically, processes 1 and 3 in node 3 have lock-wait relationships with nodes 1 and 2, and each process is labeled with an in-degree in the lock-wait graph.It should be understood that the in-degrees assigned to processes in Figure 7a are merely exemplary and the present embodiment is not limited thereto. In-degree can be used to represent the number of locks held by a process that are being waited for by other processes. Thus, the process with the highest in-degree can be selected from the deadlock as the target process for resolution. Continuing with Figure 6, in step 605, the target process can be controlled to release its held locks to resolve the target deadlock. This is from the perspective of a single deadlock. During the unlocking process, the locks held by the process with the highest in-degree in the deadlock can be released to resolve the deadlock. The inventors discovered during their research that processes with an in-degree exceeding 1 can often cause multiple deadlocks. Resolving such processes can simultaneously resolve multiple deadlocks. In the actual application of the above exemplary solution, it is possible that multiple processes with the highest in-degree exist among the processes involved in the target deadlock. In this case, the process with the highest in-degree can be analyzed to determine whether the target process is present. If so, the target process can be selected as the target process. This can further reduce network communication overhead. If no local processes exist, the target process can be selected through random selection or other methods, without further limitation. Of course, the above exemplary solution selects the target process from the perspective of a single deadlock. This embodiment also supports selecting processes to be handled from a global perspective. For example, based on the pruned lock wait graph in the aforementioned embodiment, the in-degree of each process in the final detected ring structure can be assigned, and the process with the highest in-degree can be selected. After releasing the locks held by this process, the out-degree and in-degree of the remaining lock waiters in the lock wait graph can be updated, and a pruning operation performed. The in-degree of the processes in the remaining lock waiters can then be updated, and the process with the highest in-degree can be selected again as the process to be handled. This cycle continues until all lock waiters in the lock wait graph have been pruned, terminating the loop. This approach of selecting processes to be handled from a global perspective can further reduce the number of processes to be handled, thereby further reducing the cost of resolving deadlocks. In summary, this embodiment proposes labeling the processes included in the lock request subject in the lock wait graph with an in-degree to represent the number of locks held by the process that are being waited for by other processes. This allows for more rational selection of processes for resolution based on the in-degree of the processes, thereby reducing the cost of resolving deadlocks. Figure 7b is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure. Referring to Figure 7b, taking node 1 in a distributed database system as an example, after a lock wait timeout event occurs in node 1, a deadlock detection task can be initiated in node 1.Referring to Figure 7b, after the deadlock detection task is initiated, node 1 can issue a lock wait message collection command (fetch lock status) to each node in the distributed database system. Each node that receives this command, if not currently executing the deadlock detection task, can return a lock wait message (reply lock status). Referring to Figure 7b, node 1 successfully collects lock wait messages from all nodes. Node 1 then constructs a global lock wait graph and performs deadlock detection by pruning the global lock wait graph. For the detection results obtained through pruning, node 1 further implements a secondary check mechanism, namely, issuing a lock wait status confirmation command (recheck) to each node involved in the deadlock. Referring to Figure 7b, if all relevant nodes confirm that the lock wait status has not changed (recheck succ), node 1 confirms the deadlock. Node 1 then resolves the deadlock according to the deadlock resolution solution provided in this embodiment and terminates the local deadlock detection task. As can be seen, node 1 uses a mechanism similar to optimistic locking to ensure that its deadlock detection task does not conflict with other deadlock detection tasks. Furthermore, a secondary check mechanism is employed to ensure the accuracy of deadlock detection. Furthermore, node 1 further performs deadlock resolution operations based on the deadlock detection results, thereby improving the efficiency of deadlock resolution. It should be noted that while some of the processes described in the above embodiments and accompanying figures include multiple operations that appear in a specific order, it should be understood that these operations may be executed in a different order or in parallel. Operation sequence numbers, such as 201 and 202, are merely used to distinguish between different operations and do not represent any specific execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. Figure 8 is a schematic diagram of the structure of a node device provided in another exemplary embodiment of the present disclosure. As shown in Figure 8, the node device may include a memory 80, a processor 81, and a communication component 82. The processor 81 is coupled to the memory 80 and the communication component 82, and is configured to execute a computer program in the memory 80, to: upon detecting a lock wait timeout event on a local node, initiate a target deadlock detection task corresponding to the lock wait timeout event; monitor whether there are any nodes in the distributed system that are executing other deadlock detection tasks; if no such nodes exist, construct a lock wait graph for the distributed system, the lock wait graph being used to describe lock wait relationships in the distributed system; and perform deadlock detection based on the lock wait graph to detect deadlocks in the distributed system.The node device provided in this embodiment can be any node device in a distributed system. In an optional embodiment, the processor 81 may further be configured to: after initiating the target deadlock detection task, if it is detected that a node in the distributed system is currently executing other deadlock detection tasks, stop executing the target deadlock detection task; trigger the restart of the target deadlock detection task upon reaching a specified waiting time, and wait until no nodes in the distributed system are currently executing other deadlock detection tasks, completing the target deadlock detection task. In an optional embodiment, when monitoring whether there are nodes in the distributed system currently executing other deadlock detection tasks, the processor 81 may be configured to: after initiating the target deadlock detection task, concurrently initiate a lock wait information collection instruction to each node in the distributed system; and monitor whether there are nodes in the distributed system currently executing other deadlock detection tasks based on whether any nodes return a retry notification. Upon receiving the lock wait message collection instruction, the node currently executing the deadlock detection task returns a retry notification as a response to the lock wait message collection instruction. In an optional embodiment, when constructing a lock waiting graph for the distributed system, the processor 81 may be specifically configured to: construct the lock waiting graph using lock requesting entities included in each node in the distributed system as nodes and lock waiting relationships between lock requesting entities as edges; and annotate the lock requesting entities with out-degree and in-degree in the lock waiting graph; wherein a single lock requesting entity includes one or more processes, the out-degree represents the number of locks the lock requesting entity is waiting for; and the in-degree represents the number of locks held by the lock requesting entity that are being waited for by other lock requesting entities. In an optional embodiment, when performing deadlock detection based on the lock wait graph, the processor 81 may be specifically configured to: select a target point with an out-degree or in-degree of 0 in the lock wait graph; prune the target point and its associated edges from the lock wait graph; update the out-degree and in-degree corresponding to the remaining points in the lock wait graph; loop through the target point selection, pruning, and updating operations until no more points with an out-degree or in-degree of 0 exist in the lock wait graph, and then terminate the loop; and if unpruned points and edges still exist in the lock wait graph obtained after the loop terminates, determine a deadlock in the distributed system based on a ring structure formed by the unpruned points and edges.In an optional embodiment, when the processor 81 determines a deadlock in the distributed system based on the ring structure formed by the unpruned points and edges, it may be specifically configured to: send a lock wait status confirmation request to each node containing each lock request subject contained in each ring structure; receive response information in response to the issued lock wait status confirmation request; if the lock wait status of any lock request subject in the target ring structure is determined to have changed based on the response information, prune the target ring structure from the lock wait graph; and determine the remaining ring structure in the lock wait graph as a deadlock in the distributed system. In an optional embodiment, the processor 81 may also be configured to: upon detecting a deadlock in the distributed system, select a target process from the processes involved in the target deadlock according to a preset rule; and control the target process to release the lock it holds to resolve the target deadlock; wherein the target deadlock is any deadlock in the distributed system. In an optional embodiment, the processor 81 may be further configured to: in the lock wait graph, mark the in-degree of each process involved in the target deadlock; and when selecting a target process from the processes involved in the target deadlock according to a preset rule, the processor 81 may be specifically configured to: select the process with the highest in-degree from the processes involved in the target deadlock as the target process; wherein the in-degree represents the number of locks held by the lock requesting subject that are being waited for by other lock requesting subjects. In an optional embodiment, the processor 81 may be further configured to: if there are multiple processes with the highest in-degree among the processes involved in the target deadlock and the current node process is one of them, select the current node process as the target process. In an optional embodiment, the processor 81 may further be configured to: if a new lock wait timeout event is detected on the local node during execution of the target deadlock detection task, create a waiting task for the new lock wait timeout event; place the waiting task in the local node's waiting queue; after completing the target deadlock detection task, continue to control the local node to maintain a state of executing the deadlock detection task and sequentially consume the waiting tasks in the local node's waiting queue; and after sequentially completing the deadlock detection tasks triggered by the waiting tasks in the waiting queue, control the local node to terminate the state of executing the deadlock detection task. In an optional embodiment, the processor 81 may further be configured to: if it is detected that the lock request subject to which the lock wait timeout event corresponding to the target waiting task in the waiting queue belongs has died, delete the target waiting task from the waiting queue; and release the lock held by the lock request subject to which the lock wait timeout event corresponding to the target waiting task belongs.In an optional embodiment, processor 81 may also be configured to: release the locks held by each lock requesting entity in the node upon shutdown; and, after the node restarts, discontinue deadlock detection tasks that were not completed before shutdown. Furthermore, as shown in FIG8 , the node device also includes other components, such as a power supply component 83. FIG8 only schematically illustrates some components and does not imply that the node device only includes the components shown in FIG8 . It is worth noting that the technical details of the above-mentioned node device embodiments can be found in the relevant descriptions of the aforementioned method embodiments. To save space, these details will not be repeated here, but this should not compromise the scope of protection of the present disclosure. Accordingly, embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program. When executed, the computer program can implement the steps that can be performed by the node device in the above-mentioned method embodiments. The memory in FIG8 is used to store the computer program and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, images, videos, and the like. The memory can be implemented using any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. The communication component in FIG8 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G, or other mobile communication networks, or a combination thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID), infrared data association (IrDA), ultra-wideband (UWB), Bluetooth (BT), and other technologies. The power supply assembly in Figure 8 provides power to various components of the device in which the power supply assembly resides. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly resides.Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code. The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that execution of the instructions by the processor of the computer or other programmable data processing device generates means for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more flow charts and / or one or more blocks in a block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device, causing the computer or other programmable device to execute a series of operational steps to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flow charts and / or one or more blocks in a block diagram. It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or device comprising a series of elements may include not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product, or device. In the absence of further restrictions, an element defined by the phrase "comprising a ..." does not exclude the existence of other identical elements in the process, method, product or apparatus comprising the element.It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. The above description is merely an embodiment of this disclosure and is not intended to limit this disclosure. It will be apparent to those skilled in the art that various modifications and variations of this disclosure are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure are intended to be included within the scope of protection of this disclosure.

Claims

Claims 1. A deadlock detection method, applicable to each node in a distributed system. For any one of the nodes, the method includes: When a lock wait timeout event occurs on this node is detected, start a target deadlock detection task corresponding to the lock wait timeout event; Monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks; if not, construct a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships existing in the distributed system; based on the lock wait graph, perform deadlock detection to detect deadlocks existing in the distributed system.

2. The method according to claim 1, wherein It further includes: After starting the target deadlock detection task, if it is detected that there are nodes in the distributed system that are executing other deadlock detection tasks, stop executing the target deadlock detection task; Use reaching a specified waiting duration as a trigger condition to trigger a restart of the target deadlock detection task, and wait until there are no nodes in the distributed system that are executing other deadlock detection tasks, and then complete the target deadlock detection task.

3. The method according to claim 1, wherein The monitoring of whether there are nodes in the distributed system that are executing other deadlock detection tasks includes: after starting the target deadlock detection task, parallelly send lock wait information collection instructions to each node in the distributed system, and based on whether a node returns a retry notice, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks; among them, a node that is executing a deadlock detection task returns a retry notice as a response to the lock wait message collection instruction after receiving the lock wait message collection instruction.

4. The method according to claim 1, wherein Constructing a lock wait graph for the distributed system includes: using the lock request entities included in each node in the distributed system as points, and using the lock wait relationships between the lock request entities as edges to construct the lock wait graph; in the lock wait graph, label the out-degree and in-degree for the lock request entities; among them, a single lock request entity includes one or more processes, the out-degree represents the number of locks that the lock request entity is waiting for; the in-degree represents the number of locks held by the lock request entity that are being waited for by other lock request entities.

5. The method according to claim 4, wherein Performing deadlock detection based on the lock wait graph includes: in the lock wait graph, select a target point with an out-degree or in-degree of 0; from the lock wait graph, cut off the target point and its associated edges; update the out-degree and in-degree corresponding to the remaining points in the lock wait graph; loop and execute the operations of selecting the target point, the cutting, and the updating until there are no points with an out-degree or in-degree of 0 in the lock wait graph, and then end the loop; if there are still uncut points and edges in the lock wait graph obtained after ending the loop, then based on the uncut points and edges The formed cyclic structure is used to determine the deadlocks existing in the distributed system.

6. The method according to claim 5, wherein Determining the deadlocks existing in the distributed system based on the circular structures formed by the uncropped points and edges, including: sending lock waiting status confirmation requests to the nodes where each lock request entity included in each circular structure is located; receiving response information for the sent lock waiting status determination requests; if it is determined according to the response information that the lock waiting status of any lock request entity in the target circular structure has changed, then cropping the target circular structure from the lock waiting graph; determining the remaining circular structures in the lock waiting graph as the deadlocks existing in the distributed system.

7. The method according to claim 1, wherein Further including: In the case of detecting that there are deadlocks in the distributed system, selecting a target process from each process involved in the target deadlock according to a preset rule; Controlling the target process to release the locks held by it to unlock the target deadlock; wherein, the target deadlock is any deadlock existing in the distributed system.

8. The method according to claim 7, wherein Further including: in the lock waiting graph, respectively injecting in-degrees for each process involved in the target deadlock; the selecting a target process from each process involved in the target deadlock according to a preset rule includes: selecting the process with the highest in-degree from each process involved in the target deadlock as the target process; wherein, the in-degree represents the number of locks held by the lock request entity that are waited for by other lock request entities.

9. The method according to claim 8, wherein, Further including: If there are multiple processes with the highest in-degree among the processes involved in the target deadlock and there is a local node process among them, then selecting the local node process as the target process.

10. The method according to claim 1, wherein, Further including: if a new lock waiting timeout event occurs on the local node during the execution of the target deadlock detection task, then creating a waiting task for the new lock waiting timeout event; putting the waiting task into the local node waiting queue; after completing the target deadlock detection task, continuing to control the local node to maintain the state of being in the process of executing the deadlock detection task, and sequentially consuming the waiting tasks in the local node waiting queue; after sequentially completing the deadlock detection tasks triggered by the waiting tasks in the waiting queue, controlling the local node to end the state of being in the process of executing the deadlock detection task.

11. The method according to claim 10, wherein Further including: if it is monitored that the lock request entity to which the lock waiting timeout event corresponding to the target waiting task in the waiting queue belongs has died, then deleting the target waiting task from the waiting queue; releasing the locks held by the lock request entity to which the lock waiting timeout event corresponding to the target waiting task belongs.

12. The method according to claim 1, wherein, Further including: In the case of the local node shutting down, releasing the locks held by each lock request entity in the local node; After the local node restarts, not executing the deadlock detection tasks that were not completed before the shutdown. 19 13. The method according to claim 1, wherein Perform deadlock detection based on the lock wait graph to detect deadlocks existing in the distributed system, including: perform deadlock detection on each node in the distributed system according to the lock wait graph to detect deadlocks existing in each node in the distributed system, wherein the lock wait graph is a global graph.

14. A node device, comprising a memory, a processor, and a communication component, where the node device is any node in a distributed system; the memory is used to store one or more computer instructions; the processor is coupled to the memory and the communication component and is used to execute the one or more computer instructions to perform the deadlock detection method according to any one of claims 1-13.

15. A computer-readable storage medium storing computer instructions, which when executed by one or more processors, cause the one or more processors to perform the deadlock detection method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Deadlock detection method based on side tracing for distributed system

    CN106557371A

  • Database deadlock detection method and device

    CN112256442A

  • Deadlock detection in distributed databases

    US20220350677A1

Cited By

  • Method and device for controlling consistency of medical insurance settlement transactions

    CN121935040A