Deadlock detection method and device and storage medium

By deploying deadlock detection components on each node of the distributed database system, using the lock waiting timeout event and the lock waiting map for global deadlock detection, the problem of poor deadlock detection efficiency in the existing technology is solved, and efficient and accurate deadlock detection is achieved.

CN120216212APending Publication Date: 2025-06-27HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311823861.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In distributed database systems, the prior art is not efficient for deadlock detection, especially as the system scale expands, the detection pressure of a single node increases, resulting in a decrease in deadlock detection efficiency.

Method used

A distributed deadlock detection method is proposed. By deploying deadlock detection components on each node in the distributed system, using the lock waiting timeout event to trigger the deadlock detection task, and building a lock waiting map for global deadlock detection.

Benefits of technology

This method can effectively adapt to the scalability of the distributed system, avoiding the reduction in detection efficiency due to the expansion of the system scale, and improving the accuracy and efficiency of deadlock detection through distributed execution and optimistic lock mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216212A_ABST
    Figure CN120216212A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a deadlock detection method and device and a storage medium. In the embodiment of the invention, a detection mechanism that who triggers execution is provided, and each node in the distributed system can serve as an execution main body of deadlock detection, so that the deadlock detection tasks in the distributed system can be executed in a distributed manner and are not concentrated at a single node any more, the expandability of the distributed system can be effectively adapted, and the expandability of the distributed system is improved. And the detection efficiency is not influenced by the scale expansion of the system. Moreover, the invention also provides a detection mechanism similar to an optimistic lock, so as to avoid the problems of repeated detection and the like possibly occurring between deadlock detection tasks distributed on each node, and further ensure the accuracy of deadlock detection. Besides, a single deadlock detection task can detect all deadlocks in the system at one time, and serial detection is not needed any more, so that the deadlock detection efficiency in the distributed system can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a deadlock detection method, device, and storage medium. Background Art

[0002] In a distributed database system, many resources are accessed concurrently by multiple nodes. Therefore, a lock service is required to ensure the correctness of the system. When two or more nodes, while already occupying resources through lock requests on their own, also request resources occupied by the other party, they will enter a deadlock situation of mutual waiting, resulting in the occurrence of the deadlock phenomenon. After the occurrence of the deadlock, the transactions involved in the deadlock cannot proceed forward, which will affect the relevant performance of the system.

[0003] Currently, when performing deadlock detection in a distributed database system, the detection idea in a centralized system is usually followed. That is, the deadlock detection work is concentrated on a single node in the distributed database system, and this node uses a serial method to perform deadlock detection on each process link in the distributed database system. Moreover, as the scale of the system continues to increase, the detection pressure on this node also continues to increase, resulting in poor deadlock detection efficiency.

[0004] Therefore, how to more efficiently perform deadlock detection in a distributed database system has become an urgent problem to be solved. Summary of the Invention

[0005] Multiple aspects of this application provide a deadlock detection method, device, and storage medium to improve the efficiency of deadlock detection in a distributed system.

[0006] An embodiment of this application provides a deadlock detection method, which is applicable to each node in a distributed system. For any one of the nodes, the method includes:

[0007] When a lock wait timeout event occurs on the local node, start a target deadlock detection task corresponding to the lock wait timeout event;

[0008] Monitor whether there is a node in the distributed system that is executing other deadlock detection tasks;

[0009] If not, construct a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships existing in the distributed system;

[0010] Perform deadlock detection according to the lock wait graph to detect the deadlocks existing in the distributed system.

[0011] An embodiment of this application also provides a node device, including a memory, a processor, and a communication component. The node device is any one of the nodes in the distributed system;

[0012] The memory is used to store one or more computer instructions;

[0013] The processor is coupled to the memory and the communication component, and is configured to execute the one or more computer instructions to perform the foregoing deadlock detection method.

[0014] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the foregoing deadlock detection method.

[0015] In the embodiments of the present application, a detection mechanism of "whoever triggers executes" is proposed. Each node in the distributed system can be used as the execution entity for deadlock detection. In this way, the deadlock detection tasks in the distributed system can be executed distributively, rather than being concentrated on a single node. Therefore, the scalability of the distributed system can be effectively adapted, and the detection efficiency will no longer be affected by the expansion of the system scale. Moreover, a detection mechanism similar to an optimistic lock is also proposed to avoid problems such as duplicate detection that may occur between the deadlock detection tasks distributed on each node, thereby ensuring the accuracy of deadlock detection. In addition, a single deadlock detection task can detect all deadlocks in the system at one time, rather than requiring serial detection, which can further improve the deadlock detection efficiency in the distributed system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 is a schematic structural diagram of a distributed system provided by an exemplary embodiment of the present application;

[0018] Figure 2 is a schematic flowchart of a deadlock detection method provided by an exemplary embodiment of the present application;

[0019] Figure 3 is a schematic logical diagram of a solution for a monitoring link provided by an exemplary embodiment of the present application;

[0020] Figure 4 is a schematic flowchart of another deadlock detection method provided by an exemplary embodiment of the present application;

[0021] Figure 5 is a schematic flowchart of yet another deadlock detection method provided by an exemplary embodiment of the present application;

[0022] Figure 6Schematic flowchart of another deadlock detection method provided by an exemplary embodiment of the present application;

[0023] Figure 7a Schematic logical diagram of an exemplary solution for selecting a target process provided by an exemplary embodiment of the present application;

[0024] Figure 7b Schematic diagram of an application scenario provided by an exemplary embodiment of the present application;

[0025] Figure 8 Schematic structural diagram of a node device provided by another exemplary embodiment of the present application. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0027] Before starting to elaborate on the technical solutions provided by the embodiments of the present application in detail, several technical concepts involved in the present application are briefly explained as follows.

[0028] Distributed lock: In a distributed system, many resources will be concurrently accessed by multiple nodes, and a global lock service is required to ensure the correctness of the system. For the scalability and availability of the lock service, the lock itself is also distributed and provides services on multiple nodes simultaneously. Such a lock service is called a distributed lock.

[0029] Lock waiting: Node 1 holds the lock of a certain resource, while another node 2 requests the lock of the same resource (mutual exclusion). At this time, node 2 needs to wait for node 1 to release the lock.

[0030] Deadlock: A deadlock refers to a situation where two or more transactions, while occupying resources, continue to request resources that are already occupied by the other party, thus entering a stalemate of mutual waiting. If there is no external system intervention, all transactions trapped in the deadlock cannot move forward.

[0031] Currently, the deadlock detection solutions in distributed database systems usually directly apply the deadlock detection solutions used in centralized systems. That is, the processes located on different nodes in a distributed database system are regarded as the processes located on the same physical machine in a centralized system. Specifically: a deadlock detection program is deployed on the control node in the distributed database system. The control node needs to communicate with all nodes in the system through the network to query the lock waiting information in each node; based on this, the control node performs deadlock detection based on the queried information. It can be seen that the existing deadlock detection solution completely depends on the control node. For the control node, as the scale of the system continues to expand, the number of deadlocks in the system will also continue to increase, and the performance overhead brought by deadlock detection will increase proportionally. Moreover, there are usually other control tasks on the control node that also need to occupy performance overhead. Therefore, there will be a serious backlog problem of deadlock detection tasks on the control node, resulting in the deadlock detection tasks not being processed in time, which will in turn affect the efficiency of deadlock detection in the distributed database system.

[0032] For this reason, a new deadlock detection method is proposed in the embodiments of the present application to improve the efficiency of deadlock detection in a distributed system.

[0033] The following will detail the technical solutions provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0034] Figure 1 It is a schematic structural diagram of a distributed system provided for an exemplary embodiment of the present application. As Figure 1 shown, the system may include multiple nodes. The distributed system in this embodiment may be a distributed database system, a distributed storage system, or a distributed computing system, etc. The type of service provided by the distributed system in this embodiment is not limited. For different types of distributed systems, the types of resources that each node expects to occupy through the distributed lock service may also be different. For example, in a distributed database system, the resources that each node expects to occupy through the distributed lock service are usually data resources. Another example is that in a distributed storage system, the resources that each node expects to occupy through the distributed lock service are usually storage resources. No more examples are given here. In this embodiment, the type of resources that each node expects to occupy through the distributed lock service is not limited either.

[0035] Referring to Figure 1 , in this embodiment, deadlock detection components may be respectively deployed on each node in the distributed system. The deadlock detection method provided in this embodiment may be executed by these deadlock detection components. The deadlock detection components in this embodiment may be implemented as software, hardware, or a combination of software and hardware, which is not limited here.

[0036] Considering the logical types of deadlock detection methods implemented on each node in a distributed system, for the sake of convenience in description, the following will elaborate on the deadlock detection method from the perspective of a single node.

[0037] Figure 2 It is a schematic flowchart of a deadlock detection method provided by an exemplary embodiment of the present application. Refer to Figure 2 , for any node in the distributed system, the method may include:

[0038] Step 200, when a lock wait timeout event occurs on this node, start a target deadlock detection task corresponding to the lock wait timeout event;

[0039] Step 201, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks;

[0040] Step 202, if not, construct a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships existing in the distributed system;

[0041] Step 203, perform deadlock detection according to the lock wait graph to detect the deadlocks existing in the distributed system.

[0042] Refer to Figure 1 and Figure 2 , in this embodiment, a timeout-triggered method is adopted to trigger the deadlock detection task. Among them, lock wait timeout means that the time waited by the lock request entity after sending the lock request has exceeded the preset duration threshold. As mentioned above, this embodiment does not limit the service types provided by the distributed system, and in different types of distributed systems, the lock request entities may vary. For example, in a distributed database system, the lock request entity may be a lock group. The lock group appears as a whole to request / release locks externally. Usually, a lock group contains one or more processes, and all processes within the same lock group are not affected by lock exclusivity. Based on the lock group, it can support work tasks such as parallel queries and two-phase commits in the distributed database system. Another example is that in a distributed computing system, the lock request entity may be a single process. It should be noted that the above only provides several exemplary lock request entities that often appear in different types of distributed systems, but does not limit that there can only be one type of lock request entity in a distributed system. In this embodiment, it is supported that there are multiple types of lock request entities in the distributed system, which will not affect the implementation of the deadlock detection method in this embodiment.

[0043] Based on this, refer to Figure 1, if a lock wait timeout event occurs for lock request entity A on node 1, a deadlock detection task corresponding to the lock wait timeout event that occurred for lock request entity A will be triggered to start on node 1. If a lock wait timeout event occurs for lock request entity B on node 2, a deadlock detection task corresponding to the lock wait timeout event that occurred for lock request entity B will be triggered to start on node 2. For ease of description, in step 200, the deadlock detection task started on the current node for the current lock wait timeout event will be described as the target deadlock detection task, and this description will be used throughout the subsequent elaboration of the deadlock detection method in this section.

[0044] The inventors found during the research process that the deadlock problem in a distributed system has low frequency, that is, the frequency of deadlocks occurring in the system is not very high. In the deadlock detection method provided in this embodiment, the execution time required for a single deadlock detection task is relatively short, reaching the millisecond level. However, in a distributed system, deadlock detection tasks triggered by lock wait timeout events may still have concurrency issues. That is, multiple deadlock detection tasks may be started in the distributed system at the same time.

[0045] The inventors found during the research process that such concurrency issues in deadlock detection tasks may lead to duplicate issues in deadlock detection results. However, in this embodiment, it is expected that the deadlock detection results will be used as the basis for resolving deadlocks, and it is also expected that the process of resolving deadlocks can still be maintained in a distributed manner. Therefore, it is expected that each node can continue to undertake the work of resolving deadlocks after completing deadlock detection. In this way, the deadlock detection results will directly affect the relevant operations in the process of nodes resolving deadlocks. The duplicate issues in deadlock detection results caused by the above-mentioned concurrency issues in deadlock detection tasks may lead to multiple nodes performing duplicate deadlock resolution operations on the same deadlock. For example, refer to Figure 2 , if Figure 2 a deadlock is formed between lock request entity A and lock request entity B, and this deadlock causes lock request entity A and lock request entity B to almost simultaneously have lock wait timeout events, then it is possible that node 1 and node 2 almost simultaneously start deadlock detection tasks for the same deadlock, that is, the aforementioned concurrency issue; in this case, for node 1, after detecting this deadlock, it will surely think that lock request entity B should release the lock, so it will control lock request entity B to release the lock, and for node 2, after detecting this deadlock, it will surely think that lock request entity A should release the lock, so it will control lock request entity A to also release the lock. However, obviously, in this case, only one of the lock request entities needs to release the lock to resolve this deadlock, and both lock request entities do not need to release the lock. It can be seen that the duplicate issues in deadlock detection results caused by the concurrency issues in deadlock detection tasks may lead to over-operation issues in the subsequent process of resolving deadlocks, and may thus affect the main work efficiency in the distributed system.

[0046] Therefore, in this embodiment, a solution is also proposed for the concurrency problem of such deadlock detection tasks.

[0047] Referring to Figure 1 , in step 201, for this node, after starting the target deadlock detection task, it can first enter the monitoring phase, that is, monitor whether there are other nodes in the distributed system that are executing other deadlock detection tasks. Among them, other deadlock detection tasks refer to the deadlock detection tasks triggered by other lock wait timeout events. In addition, the monitoring here is for the entire distributed system, including this node itself and other nodes. Being executed can be understood as having been started. In practical applications, there may be some special cases. For example, two deadlock detection tasks are both in the aforementioned monitoring phase at the same time. In this case, some rules that can ensure exclusivity can be set, such as whoever initiates the monitoring first stops the deadlock detection task first. Of course, this is only exemplary. In this embodiment, other rules can also be used to ensure exclusivity in these special cases, and no more examples are given here.

[0048] That is, in this embodiment, the deadlock detection tasks in the distributed system are allowed to be concurrent. However, for a single node, after it starts the deadlock detection task, it will first confirm whether there are already other deadlock detection tasks being executed in the distributed system. If so, this node will immediately stop continuing to execute the current deadlock detection task. It should be understood that in this embodiment, the intention of performing the monitoring phase first after starting the deadlock detection task is to schedule the deadlock detection tasks in the distributed system to ensure that only one deadlock detection task is in other phases after the monitoring phase at the same time, thereby avoiding the problem of duplicate deadlock detection results that may be caused by the concurrency problem of deadlock detection tasks.

[0049] On this basis, this embodiment proposes that during the execution of the deadlock detection task, at least a lock wait graph construction phase and a deadlock detection phase can be set after the monitoring phase.

[0050] Referring to Figure 2 , in step 202, if in the monitoring phase, it is determined that there are no other nodes in the distributed system that are executing other deadlock detection tasks, then the lock wait graph construction phase can be entered. The lock wait graph in this embodiment is used to describe the lock wait relationships existing in the distributed system. That is, the lock wait graph in this embodiment is a global graph, rather than a single-node graph.

[0051] To support the construction of a lock wait graph, in this embodiment, this node can query lock wait information from each node in the distributed system as the data basis for constructing the lock wait graph. In a preferred implementation, this node can concurrently send lock wait information collection instructions to each node (including this node) in the distributed system to request that each node return lock wait information. In this way, based on the lock wait information returned by each node, a lock wait graph of the distributed system can be constructed. The parallel query method proposed in this embodiment can effectively improve the construction efficiency of the lock wait graph, thereby reducing the time consumed for a single deadlock detection task.

[0052] The inventors found during the research process that in order to obtain the lock wait information of each node, this node needs to communicate with other nodes in the distributed system, and in the aforementioned monitoring process, this node also needs to query each node to determine whether it is performing a deadlock detection task. Therefore, in this embodiment, it is further proposed that the process of obtaining the lock wait message can be advanced to the monitoring process, and the response format of the lock wait information collection instruction is agreed upon to convey the deadlock detection task status on the node. Figure 3 It is a schematic diagram of the solution logic for a monitoring process provided in an exemplary embodiment of this application. Refer to Figure 3 , in this further solution: After starting the target deadlock detection task, this node can concurrently send lock wait information collection instructions to each node in the distributed system; if no node returns a retry notice, it is determined that there is no node in the distributed system that is performing other deadlock detection tasks; among them, the node that is performing the deadlock detection task returns a retry notice as a response to the lock wait message collection instruction after receiving the lock wait message collection instruction.

[0053] Refer to Figure 3 , in this further solution, Node 1 can send lock wait message collection instructions to other nodes in the distributed system during the monitoring process. Among them, Node 2 is performing a deadlock detection task ( Figure 3 the black bar part shown under Node 2 in represents the duration of the deadlock detection task it is performing), so Node 2 will return a retry notice in response to the lock wait message collection instruction sent by Node 1. After receiving this retry notice, Node 1 can confirm that there is a node in the distributed system that is currently performing other deadlock detection tasks. For nodes that are not in the process of performing the deadlock detection task, such as Node 3, it will return the lock wait message in response to the lock wait message collection instruction sent by Node 1.

[0054] In this way, if there is no node in the distributed system that is executing other deadlock detection tasks, this node will not receive any retry notifications during the monitoring phase and can collect the lock waiting information of each node in the distributed system. In this way, after entering the phase of constructing the lock waiting graph, this node can construct the lock waiting graph for the distributed system based on the lock waiting information collected during the monitoring phase.

[0055] It should be noted that the process of obtaining the lock waiting messages is advanced to the monitoring phase above, and the response format of the lock waiting information collection instruction is agreed upon to convey the deadlock detection task status on the node. This is only a preferred solution in this embodiment. In this embodiment, other methods can also be used to implement the monitoring phase. For example, this node can initiate a task execution status query instruction during the monitoring phase to monitor whether each node is executing the deadlock detection task, and the process of obtaining the lock waiting messages can be placed in the phase of constructing the lock waiting graph. This node can initiate another round of network communication during the process of constructing the lock waiting graph to obtain the lock waiting messages of each node. No more solution examples are provided here.

[0056] Continue to refer to Figure 2 , after completing step 202, it can continue to enter step 203. According to the lock waiting graph constructed in step 202, deadlock detection is performed to detect the deadlocks existing in the distributed system.

[0057] As mentioned above, the lock waiting graph in this embodiment is a global graph. Therefore, in step 203, this node can detect all the deadlocks in the distributed system based on the lock waiting graph, rather than only detecting the deadlocks related to this node itself. This concept of single-node global detection can further improve the efficiency of deadlock detection in the distributed system.

[0058] In addition, it should be noted that for each node in the distributed system, the execution process of the deadlock detection task can be implemented in a bypass manner, that is, the deadlock detection task is completely asynchronous with the main work task in the node and does not interfere with each other. This ensures that the deadlock detection method in this embodiment has no impact on the performance of the main work in the distributed system.

[0059] In summary, in this embodiment, a detection mechanism of "whoever triggers executes" is proposed. Each node in the distributed system can be used as the execution entity for deadlock detection. In this way, the deadlock detection tasks in the distributed system can be executed distributively, rather than being concentrated on a single node. Therefore, the scalability of the distributed system can be effectively adapted, and the detection efficiency will no longer be affected by the expansion of the system scale. Moreover, a detection mechanism similar to optimistic locking is also proposed to avoid problems such as duplicate detections that may occur between deadlock detection tasks distributed on each node, thereby ensuring the accuracy of deadlock detection. In addition, a single deadlock detection task can detect all deadlocks in the system at one time, rather than requiring serial detection, which can further improve the deadlock detection efficiency in the distributed system.

[0060] Figure 4 It is a schematic flowchart of another deadlock detection method provided by an exemplary embodiment of the present application. Refer to Figure 4 , this deadlock detection method may include:

[0061] Step 400: When a lock wait timeout event occurs on this node is monitored, start a target deadlock detection task corresponding to the lock wait timeout event;

[0062] Step 401: Monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks. If not, execute Step 402. If so, execute Step 404;

[0063] Step 402: Construct a lock wait graph for the distributed system. The lock wait graph is used to describe the lock wait relationships existing in the distributed system, and continue to execute Step 403;

[0064] Step 403: Perform deadlock detection according to the lock wait graph to detect the deadlocks existing in the distributed system;

[0065] Step 404: Stop executing the target deadlock detection task; use reaching a specified waiting duration as a trigger condition to trigger the restart of the target deadlock detection task, and return to execute Step 401 to wait until there are no nodes in the distributed system that are executing other deadlock detection tasks, and then complete the target deadlock detection task.

[0066] Steps 400, 402, and 403 in this embodiment can refer to the descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, based on Steps 401 and 404, a handling solution is provided in the case where after the target deadlock detection task is started, it is monitored that there are already nodes in the distributed system that are executing other deadlock detection tasks.

[0067] As mentioned in the foregoing embodiments, in this case, the execution of the target deadlock detection task will be stopped. In this embodiment, it is further proposed that, in this case, after reaching the specified waiting duration, the target deadlock detection task will be restarted. After the restart, it will be determined again whether there is a deadlock detection task being executed in the distributed system. This process will loop until a suitable execution opportunity is found for the target deadlock detection task, and the target deadlock detection task will be completed through the found execution opportunity. The specified waiting duration in this embodiment can be configured as needed. In a preferred configuration scheme: the waiting duration can be a random value within a certain time interval to reduce the probability of task conflicts occurring again. For example, in some extreme cases, it is possible that two nodes start the deadlock detection task at the same time. Both sides may detect task conflicts and stop their own deadlock detection tasks. If both sides follow the same waiting duration, they will initiate a restart at the same time, resulting in the task conflict between them not being resolved. By setting the waiting duration as a random value within a certain time interval, the probability of this situation occurring can be effectively reduced.

[0068] Reference Figure 3 , for node 1, its target deadlock detection task on node 1 is stopped because there is a deadlock detection task being executed on node 2. However, after the specified waiting duration, the target deadlock detection task on node 1 is restarted ( Figure 3 the white long bar part under the node in

[0069] represents the restarted target deadlock detection task).

[0070] During the research process, the inventors found that there may be concurrency issues between deadlock detection tasks of different nodes, and concurrency issues may also occur within the same node for deadlock detection tasks. For the latter type of concurrency issue, still taking this node as an example, it can be understood that during the execution of the target deadlock detection task, a new lock wait timeout event occurs on this node. For the latter type of concurrency issue, in this embodiment, it can be processed according to the Figure 4 execution process, and the new lock wait timeout event that occurs on this node during the execution of the target deadlock detection task is treated as an ordinary lock wait timeout event. That is, after such a lock wait timeout event is discovered, a deadlock detection task is started for it, and it is determined whether there is a target deadlock detection task being executed on this node through the foregoing monitoring link, so as to stop the deadlock detection task caused by such a lock wait timeout event and restart it at an opportune time.

[0071] For the latter type of concurrency problems described above, in this embodiment, a further improvement scheme is proposed. Instead of treating new lock wait timeout events that occur on this node during the execution of the target deadlock detection task as ordinary lock wait timeout events, they are treated as special lock wait timeout events. After detecting such special lock wait timeout events, a waiting task is created for them. Among them, the waiting task needs to be started passively, rather than being started immediately after being created like an ordinary deadlock detection task.

[0072] Therefore, in this further improvement scheme, it is proposed that waiting queues be respectively set up in each node of the distributed system, and the waiting queues store waiting tasks. In this way, for this node, the waiting tasks created for special lock wait timeout events can be put into the waiting queue of this node. Based on this, this node can continue to control this node to maintain the state of being in the process of executing the deadlock detection task after completing the target deadlock detection task, and sequentially consume the waiting tasks in the waiting queue of this node; after sequentially completing the deadlock detection tasks triggered by the waiting tasks in the waiting queue, control this node to end the state of being in the process of executing the deadlock detection task.

[0073] That is to say, in this further improvement scheme, for the aforementioned special lock wait timeout events, the corresponding deadlock detection tasks are not immediately started, but can be queued within the node for the deadlock detection tasks caused by such lock wait timeout events through the waiting queue. If there are waiting tasks in the waiting queue of this node, after this node finishes executing the target deadlock detection task, it will continue to execute the deadlock detection tasks involved in the waiting queue, and will not relinquish the exclusive right to the deadlock detection task. It should be understood that if there are multiple waiting tasks in the waiting queue, this node will still follow the rule of only executing a single deadlock detection task at the same time. After completing the deadlock detection task corresponding to one waiting task, it will then consume the next waiting task in the waiting queue.

[0074] In this way, through the further improvement scheme, for the aforementioned special lock wait timeout events, there is no need to monitor and find that this node is executing the deadlock detection task, but directly queue them automatically within this node, which can avoid many restart processes, thereby saving the performance overhead of this node.

[0075] In addition, during the research process, the inventors also found that since the execution time of the deadlock detection task provided in this embodiment is very short, the aforementioned special lock wait timeout event may have been resolved in the target deadlock detection task, and repeating the execution of the deadlock detection task for the lock wait timeout event is a waste of the performance of this node. Therefore, in the aforementioned further improvement solution, it is also proposed that if it is monitored that the lock request entity to which the lock wait timeout event corresponding to the target wait task in the wait queue belongs has died, then the target wait task is deleted from the wait queue; the lock held by the lock request entity to which the lock wait timeout event corresponding to the target wait task belongs is released. Among them, the wait tasks in the wait queue are created based on the lock wait timeout event, and the lock wait timeout event corresponds to the lock request entity. Therefore, if the lock request entity has died, the lock held by the lock request entity can be directly released, and there is no need to execute the deadlock detection task for the lock request entity anymore. This can further save the performance overhead of this node.

[0076] During the research process, the inventors found that in this embodiment, it is proposed that the deadlock detection work is undertaken by the working nodes in the distributed system, and the lock request entities that cause deadlocks run on the working nodes. Based on this, in this embodiment, it is also proposed that in the case of the shutdown of this node, the locks held by each lock request entity in this node are released; moreover, after this node restarts, the deadlock detection tasks that were not completed before shutdown are not executed again.

[0077] In this way, the node failures that occur in the distributed system will be positively utilized in this embodiment, borrowing the negative node failure problem as a way to solve deadlocks. It is precisely based on the concept of "who triggers, who detects" in this embodiment that after this node shuts down, the deadlocks caused by this node will all be resolved, and there is no need to roll back the detection after this node restarts. Moreover, each node can be responsible for its own deadlock detection tasks. Therefore, this node does not need to care whether the deadlock detection tasks on other nodes need to be rolled back.

[0078] In summary, in this embodiment, after this node starts the target deadlock detection task, if it is monitored that there is already a node in the distributed system that is executing other deadlock detection tasks, the target deadlock detection task can be stopped, and the target deadlock detection task can be restarted to continuously probe whether there is a window period for the deadlock detection task in the distributed system. In this way, the scheduling of the deadlock detection task in the distributed system can be realized from the time level, and the omission of the deadlock detection task can be avoided. Moreover, for the special lock wait timeout event that occurs during the execution of the deadlock detection task on this node, the scheduling of the deadlock detection task can be realized by queuing within the node and the omission can be avoided, thereby saving the performance overhead of each node. In addition, the node failures in the distributed system are also borrowed as a way to solve deadlocks, thereby further saving the performance overhead of each node.

[0079] Figure 5 The flowchart of another deadlock detection method provided for an exemplary embodiment of this application. Refer to Figure 5 , the deadlock detection method may include:

[0080] Step 500, when a lock wait timeout event occurs on this node, start a target deadlock detection task corresponding to the lock wait timeout event;

[0081] Step 501, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks;

[0082] Step 502, if not, use the lock request entities included in each node in the distributed system as points, and use the lock wait relationships between the lock request entities as edges to construct a lock wait graph;

[0083] Step 503, in the lock wait graph, label the out-degree and in-degree for the lock request entities

[0084] Step 504, based on the lock wait graph containing the out-degree and in-degree, perform deadlock detection to detect the deadlocks existing in the distributed system.

[0085] Among them, Step 500 and Step 501 can refer to the relevant descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, an optional implementation solution for constructing a lock wait graph and an optional implementation solution for performing deadlock detection based on the lock wait graph can be provided based on Steps 502 - 504.

[0086] Refer to Figure 5 , in Step 502, the lock request entities included in each node in the distributed system can be used as points, and the lock wait relationships between the lock request entities can be used as edges to construct a lock wait graph. Among them, the lock wait relationships can be extracted based on the lock wait messages provided by each node in the foregoing embodiments. Regarding this part, reference can be made to the existing lock wait relationship extraction technologies and will not be elaborated here.

[0087] Different from the traditional lock waiting graph, in this embodiment, it is proposed that in the lock waiting graph, the out-degree and in-degree are marked for the lock request entity. Among them, the out-degree represents the number of locks waited for by the lock request entity; the in-degree represents the number of locks held by the lock request entity that are waited for by other lock request entities. When the out-degree is 0, it means that the number of locks waited for by the lock request entity is 0, that is, the lock request entity does not need to wait for other locks. When the in-degree is 0, it means that none of the locks held by the lock request entity are waited for by other lock request entities. The inventor found through research that a lock request entity with an out-degree or in-degree of 0 will definitely not be in any deadlock. Therefore, in this embodiment, in step 504, deadlock detection can be performed according to the lock waiting graph including the out-degree and in-degree to detect the deadlock existing in the distributed system. That is, in step 504, it is possible to analyze whether each lock request entity is in a deadlock by analyzing the out-degree and in-degree of each lock request entity. Finally, the deadlock can be analyzed from the lock waiting graph.

[0088] It can be seen that in this embodiment, instead of using the traditional depth-first search algorithm to use the lock waiting graph, a new deadlock detection scheme based on the lock waiting graph is proposed. That is, deadlock detection is performed based on the out-degree and in-degree marked for the lock request entity. In an exemplary detection scheme:

[0089] In the lock waiting graph, a target point with an out-degree or in-degree of 0 can be selected;

[0090] From the lock waiting graph, the target point and its associated edges are trimmed;

[0091] The out-degree and in-degree corresponding to the remaining points in the lock waiting graph are updated;

[0092] The operations of selecting the target point, trimming, and updating are repeatedly executed until there are no points with an out-degree or in-degree of 0 in the lock waiting graph, and then the loop ends;

[0093] If there are still untrimmed points and edges in the lock waiting graph obtained after the loop ends, the deadlock existing in the distributed system is determined based on the circular structure formed by the untrimmed points and edges.

[0094] In this exemplary detection scheme, a pruning method is adopted for deadlock detection. It should be noted that in this detection scheme, the pruning unit is the lock request entity. This is because in the lock waiting relationship, the lock request entity is the unit of lock request. Therefore, the lock request entity is the basic unit that causes deadlocks. It can be understood that in this exemplary detection scheme, by repeatedly performing operations of selecting target points, pruning, and updating, the points with an out-degree or in-degree of 0 and the edges associated with them in the lock waiting graph are pruned off. In this way, after the loop ends, the remaining points in the lock waiting graph must be the points with neither an out-degree nor an in-degree of 0. According to the definitions of out-degree and in-degree, the remaining points hold locks themselves, and the locks they hold are being waited for by other points, and they are also waiting for the locks of other points. Therefore, the remaining points conform to the characteristics of the lock request entities in deadlocks and must be points in a deadlock. All the points in a deadlock will be remaining, and through the edges between the points, the points in the same deadlock can be connected into a cyclic structure. In this way, after pruning, if there is a cyclic structure remaining in the lock waiting graph, the deadlocks existing in the distributed system can be determined based on the cyclic structure. That is to say, the deadlock is caused by the lock request entities corresponding to the points included in the single cyclic structure remaining in the lock waiting graph.

[0095] In this way, in this embodiment, based on the out-degree and in-degree labeled for the lock request entities, deadlocks existing in the distributed system can be detected from the lock waiting graph efficiently and comprehensively.

[0096] The inventor found during the research process that the lock states of the working processes in the distributed system are constantly changing (locking / releasing locks), and the deadlock detection task in this embodiment is completely asynchronous with the working processes in the distributed system. Therefore, timing problems may occur in the lock waiting graph in this embodiment, which may lead to deviations in the detected deadlocks after performing deadlock detection operations based on the lock request graph. For example:

[0097] At time T1: The lock waiting information of node 1 is collected: The lock request entity A is waiting for the lock request entity B;

[0098] At time T2: The lock waiting information of node 2 is collected: The lock request entity B is waiting for the lock request entity A;

[0099] At time T3: The lock request entity B releases the lock, and the lock request entity A successfully acquires the lock and no longer needs to wait;

[0100] At time T4: A lock waiting graph is created:

[0101] At time T5: Deadlock detection is performed based on the lock waiting graph.

[0102] It can be seen that since the lock wait graph is constructed at time T4 and the state change that occurred at time T3 is not perceived, when the deadlock detection operation is performed at time T5, a deadlock formed by lock request entity A and lock request entity B may be detected. This results in the detected deadlock not actually existing, causing a deviation in the detection result. If it is desired to ensure the accuracy of deadlock detection, the working process needs to be paused during the deadlock detection process to keep the lock state of the working process unchanged. However, this solution will inevitably have a negative impact on the performance of the distributed system.

[0103] Therefore, in this embodiment, an exemplary solution is proposed for such timing problems:

[0104] For each remaining circular structure in the lock request graph, this node can send lock wait status confirmation requests to the nodes where each lock request entity included in each link structure is located respectively;

[0105] Receive response information for the sent lock wait status determination request;

[0106] If it is determined according to the response information that the lock wait status of any lock request entity in the target circular structure has changed, then cut off the target circular structure from the lock wait graph;

[0107] Determine the remaining circular structures in the lock wait graph as the deadlocks existing in the distributed system.

[0108] In this exemplary solution, the secondary check mechanism is adopted. For the deadlocks preliminarily detected by pruning, only lock wait status confirmation requests can be sent to the nodes involved in the deadlocks to determine whether the lock wait status of each lock request entity in the deadlocks has changed. If none of them have changed, it can be confirmed that the deadlock detection is correct. And if the lock wait status of any one of the lock request entities in the deadlock has changed, it means that the problem of inaccurate lock wait graph caused by the aforementioned timing problem has occurred, and such deadlocks can be further cut off from the lock wait graph in a timely manner. In this way, through the secondary check mechanism, the accuracy of deadlock detection can be effectively improved.

[0109] In summary, in this embodiment, a new deadlock detection method based on the lock wait graph is proposed, that is, the out-degree and in-degree are marked for the lock request entities, and based on this, points and edges are pruned in the lock wait graph to find the circular structures in the lock wait graph to detect the deadlocks existing in the distributed system. In addition, a secondary check mechanism is also proposed. By communicating with the nodes related to the deadlocks, the preliminarily detected deadlocks can be re-determined to further ensure the accuracy of deadlock detection.

[0110] Figure 6 This is a schematic flowchart of another deadlock detection method provided for an exemplary embodiment of the present application. Refer toFigure 6 , the method may include:

[0111] Step 600, when a lock waiting timeout event occurs on this node is detected, start a target deadlock detection task corresponding to the lock waiting timeout event;

[0112] Step 601, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks;

[0113] Step 602, if not, construct a lock waiting graph for the distributed system, where the lock waiting graph is used to describe the lock waiting relationships existing in the distributed system;

[0114] Step 603, perform deadlock detection according to the lock waiting graph to detect the deadlocks existing in the distributed system;

[0115] Step 604, when it is detected that there are deadlocks in the distributed system, select a target process from each of the processes involved in the target deadlock according to a preset rule;

[0116] Step 605, control the target process to release the held lock to unlock the target deadlock.

[0117] Among them, Steps 600 - 603 can refer to the relevant descriptions in the foregoing embodiments. For the sake of brevity, they will not be repeated here. Based on Steps 604 and 605, in this embodiment, it is proposed to add a link to resolve deadlocks during the execution of the deadlock detection task. In this way, the work of resolving deadlocks will be executed by each node in the distributed system, which enables the work of resolving deadlocks in the distributed system to be distributed and executed.

[0118] Reference Figure 6 , in Step 604, when it is detected that there are deadlocks in the distributed system, a target process can be selected from each of the processes involved in the target deadlock according to a preset rule. This makes it possible to only handle a single process in the deadlock for a single detected deadlock, thereby effectively reducing the cost of resolving deadlocks.

[0119] In this embodiment, an exemplary solution for selecting a target process is proposed. Figure 7a It is a logical schematic diagram of an exemplary solution for selecting a target process provided by an exemplary embodiment of the present application. Reference Figure 7a, in this exemplary solution: in the lock waiting graph, the in-degree can be injected for each process involved in the target deadlock; based on this, the process with the highest in-degree can be selected from each process involved in the target deadlock as the target process; as mentioned above, the in-degree represents the number of locks held by the lock request subject that are being waited for by other lock request subjects. Therefore, the in-degree marked for the process here can represent the number of locks held by the process that are being waited for by other processes.

[0120] Refer to Figure 7a , in this exemplary solution, there is a lock waiting relationship among Node 1, Node 2, and Node 3. Specifically, there is a lock waiting relationship between Process 1 and Process 3 in Node 3 and Node 1 and Node 2, and the in-degree is marked for each process in the lock waiting graph. It should be understood that Figure 7a the in-degree marked for the process in is only exemplary, and this embodiment is not limited thereto. The in-degree can represent the number of locks held by the process that are being waited for by other processes. In this way, the process with the highest in-degree can be selected from the deadlock as the target process for disposal.

[0121] Continue to refer to Figure 6 , in step 605, the target process can be controlled to release the locks it holds to unlock the target deadlock. This is from the perspective of a single deadlock. During the unlocking process, the locks held by the process with the highest in-degree in the deadlock can be released to unlock the deadlock. The inventor found during the research process that processes with an in-degree exceeding 1 usually may cause multiple deadlocks. After disposing of such processes, the effect of simultaneously resolving multiple deadlocks may be achieved.

[0122] In the actual application process of the above exemplary solution, in the case where there are multiple processes with the highest in-degree among the processes involved in the target deadlock, in this case, it can be further analyzed whether there is a local node process among these processes with the highest in-degree. If so, the local node process can be preferably selected as the target process. In this way, the network communication overhead can be further reduced. If there is no local process, the target process can be selected by means of random selection, etc., and no more limitations are made here.

[0123] Of course, the above exemplary solution selects the target process from the perspective of a single deadlock. In this embodiment, it is also possible to support selecting the processes to be disposed of from a global perspective. For example, based on the pruned lock wait graph in the foregoing embodiment, the degree of in-degree can be injected into the processes in each lock request subject in the finally detected circular structure, and the process with the highest in-degree can be selected from them. After releasing the locks held by this process, the out-degree and in-degree of the remaining lock wait subjects in the lock wait graph can be updated, and the pruning operation can be performed. After that, the in-degree of the processes in the remaining lock wait subjects can be updated again, and the process with the highest in-degree can be selected again as the process to be disposed of. This cycle continues until all lock wait subjects in the lock wait graph are pruned, ending the loop. This way of selecting the processes to be disposed of from a global perspective can further reduce the number of processes to be disposed of, thereby further reducing the cost of resolving deadlocks.

[0124] In summary, in this embodiment, the degree of in-degree is injected into the processes included in the lock request subject in the lock wait graph to represent the number of locks held by the process that are waited for by other processes. In this way, based on the in-degree of the process, the processes to be disposed of can be selected more reasonably, thereby reducing the cost required to resolve deadlocks.

[0125] Figure 7b It is a schematic diagram of an application scenario provided by an exemplary embodiment of the present application. Refer to Figure 7b , taking a node 1 in a distributed database system as an example, after a lock wait timeout event occurs in node 1, a deadlock detection task can be started in node 1.

[0126] Refer to Figure 7b , after the deadlock detection task is started, node 1 can send a lock wait message collection instruction fetch lock status to each node in the distributed database system. If each node that receives the instruction is not currently executing the deadlock detection task, it can return a lock wait message reply lock status. Refer to Figure 7b , node 1 successfully collects the lock wait messages of all nodes.

[0127] After that, node 1 constructs a global lock wait graph. And continue to perform deadlock detection by pruning the global lock wait graph.

[0128] For the detection result obtained by the pruning method, node 1 further implements a secondary check mechanism, that is, it sends a lock wait status confirmation instruction recheck to each node involved in the deadlock respectively. Refer to Figure 7b , if the relevant nodes all confirm that the lock wait status has not changed recheck succ, then node 1 can confirm the deadlock. After that, node 1 can complete the deadlock resolution according to the deadlock resolution solution provided in this embodiment and end the local deadlock detection task.

[0129] It can be seen that Node 1 ensures that the deadlock detection task on it does not conflict with other deadlock detection tasks according to a mechanism similar to an optimistic lock. Moreover, a secondary check mechanism is also adopted to ensure the accuracy of deadlock detection. In addition, Node 1 further performs a deadlock resolution operation based on the deadlock detection result, thereby improving the efficiency of deadlock resolution.

[0130] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 201 and 202 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.

[0131] Figure 8 The following is a schematic structural diagram of a node device provided by another exemplary embodiment of the present application. As Figure 8 shown, the node device may include: a memory 80, a processor 81, and a communication component 82.

[0132] The processor 81 is coupled to the memory 80 and the communication component 82 and is configured to execute a computer program in the memory 80 for:

[0133] When a lock wait timeout event occurs on this node is detected, start a target deadlock detection task corresponding to the lock wait timeout event;

[0134] Monitor whether there is a node in the distributed system that is executing other deadlock detection tasks;

[0135] If not, construct a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships existing in the distributed system;

[0136] Perform deadlock detection according to the lock wait graph to detect the deadlocks existing in the distributed system.

[0137] Among them, the node device provided in this embodiment may be any node device in the distributed system.

[0138] In an optional embodiment, the processor 81 may further be configured to:

[0139] After starting the target deadlock detection task, if it is monitored that there is a node in the distributed system that is executing other deadlock detection tasks, stop executing the target deadlock detection task;

[0140] Using the achievement of the specified waiting duration as a trigger condition, trigger the restart of the target deadlock detection task, and wait until there are no nodes in the distributed system that are executing other deadlock detection tasks, and then complete the target deadlock detection task.

[0141] In an alternative embodiment, when the processor 81 monitors whether there are nodes in the distributed system that are executing other deadlock detection tasks, it can specifically be used for:

[0142] After starting the target deadlock detection task, parallelly send lock waiting information collection instructions to each node in the distributed system;

[0143] Based on whether a node returns a retry notification, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks;

[0144] Among them, a node that is executing a deadlock detection task returns a retry notification as a response to the lock waiting message collection instruction after receiving the lock waiting message collection instruction.

[0145] In an alternative embodiment, when the processor 81 constructs a lock waiting graph for the distributed system, it can specifically be used for:

[0146] Using the lock request entities included in each node in the distributed system as points, and using the lock waiting relationships between the lock request entities as edges, construct the lock waiting graph;

[0147] In the lock waiting graph, label the out-degree and in-degree for the lock request entity;

[0148] Among them, a single lock request entity includes one or more processes. The out-degree represents the number of locks that the lock request entity is waiting for; the in-degree represents the number of locks held by the lock request entity that are being waited for by other lock request entities.

[0149] In an alternative embodiment, when the processor 81 performs deadlock detection based on the lock waiting graph, it can specifically be used for:

[0150] In the lock waiting graph, select a target point with an out-degree or in-degree of 0;

[0151] From the lock waiting graph, cut off the target point and its associated edges;

[0152] Update the out-degree and in-degree corresponding to the remaining points in the lock waiting graph;

[0153] Loop through the operations of selecting the target point, cutting, and updating until there are no points with an out-degree or in-degree of 0 in the lock waiting graph, and then end the loop;

[0154] If there are still unclipped points and edges in the lock wait graph obtained after the loop ends, a deadlock existing in the distributed system is determined based on the circular structure formed by the unclipped points and edges.

[0155] In an alternative embodiment, when determining the deadlock existing in the distributed system based on the circular structure formed by the unclipped points and edges, the processor 81 may specifically be used for:

[0156] Sending lock wait status confirmation requests to the nodes where each lock request entity included in each circular structure is located respectively;

[0157] Receiving response information for the sent lock wait status determination request;

[0158] If it is determined according to the response information that the lock wait status of any lock request entity in the target circular structure changes, the target circular structure is clipped off from the lock wait graph;

[0159] Determining the remaining circular structures in the lock wait graph as the deadlocks existing in the distributed system.

[0160] In an alternative embodiment, the processor 81 may also be used for:

[0161] When detecting that there is a deadlock in the distributed system, selecting a target process from each process involved in the target deadlock according to a preset rule;

[0162] Controlling the target process to release the held lock to unlock the target deadlock;

[0163] Wherein, the target deadlock is any deadlock existing in the distributed system.

[0164] In an alternative embodiment, the processor 81 may also be used for:

[0165] In the lock wait graph, injecting in-degrees for each process involved in the target deadlock respectively;

[0166] When the processor 81 selects a target process from each process involved in the target deadlock according to a preset rule, it may specifically be used for:

[0167] Selecting the process with the highest in-degree from each process involved in the target deadlock as the target process;

[0168] Wherein, the in-degree represents the number of locks held by the lock request entity that are waited for by other lock request entities.

[0169] In an alternative embodiment, the processor 81 may also be used for:

[0170] If there are multiple processes with the highest in-degree among the processes involved in the target deadlock and the local node process exists among them, then select the local node process as the target process.

[0171] In an alternative embodiment, the processor 81 may further be configured to:

[0172] If a new lock wait timeout event occurs on the local node during the execution of the target deadlock detection task, create a wait task for the new lock wait timeout event;

[0173] Put the wait task into the local node wait queue;

[0174] After completing the target deadlock detection task, continue to control the local node to maintain the state of being in the process of executing the deadlock detection task, and sequentially consume the wait tasks in the local node wait queue;

[0175] After sequentially completing the deadlock detection tasks triggered by the wait tasks in the wait queue, control the local node to end the state of being in the process of executing the deadlock detection task.

[0176] In an alternative embodiment, the processor 81 may further be configured to:

[0177] If it is monitored that the lock request entity to which the lock wait timeout event corresponding to the target wait task in the wait queue belongs has died, delete the target wait task from the wait queue;

[0178] Release the lock held by the lock request entity to which the lock wait timeout event corresponding to the target wait task belongs.

[0179] In an alternative embodiment, the processor 81 may further be configured to:

[0180] In the case of the local node shutting down, release the locks held by each lock request entity in the local node;

[0181] After the local node restarts, do not execute the deadlock detection tasks that were not completed before the shutdown.

[0182] Furthermore, as Figure 8 shown, the node device further includes: a power supply component 83 and other components. Figure 8 Only some components are schematically shown in Figure 8 and it does not mean that the node device only includes

[0183] It should be noted that for the technical details in the above embodiments of the node device, reference can be made to the relevant descriptions in the foregoing method embodiments. To save space, they are not repeated here, but this should not cause loss of the protection scope of this application.

[0184] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by the node device in the above method embodiment.

[0185] The above Figure 8 The memory in the above is used to store a computer program and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0186] The above Figure 8 The communication component in the above is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0187] The above Figure 8 The power supply component in the above provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0188] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0189] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0190] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0192] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.

[0193] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0194] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A deadlock detection method, characterized in that, Applicable to each node in a distributed system. For any one of the nodes, the method includes: When a lock wait timeout event occurs on this node is detected, start a target deadlock detection task corresponding to the lock wait timeout event; Monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks; If not, construct a lock wait graph for the distributed system, where the lock wait graph is used to describe the lock wait relationships existing in the distributed system; Based on the lock wait graph, perform deadlock detection to detect the deadlocks existing in the distributed system.

2. The method according to claim 1, characterized in that, It further includes: After starting the target deadlock detection task, if it is detected that there are nodes in the distributed system that are executing other deadlock detection tasks, stop executing the target deadlock detection task; Use reaching a specified waiting duration as a trigger condition to trigger a restart of the target deadlock detection task, and wait until there are no nodes in the distributed system that are executing other deadlock detection tasks, and then complete the target deadlock detection task.

3. The method according to claim 1, wherein The monitoring of whether there are nodes in the distributed system that are executing other deadlock detection tasks includes: After starting the target deadlock detection task, parallelly send lock wait information collection instructions to each node in the distributed system, and based on whether there is a node returning a retry notice, monitor whether there are nodes in the distributed system that are executing other deadlock detection tasks; Among them, a node that is executing a deadlock detection task returns a retry notice as a response to the lock wait message collection instruction after receiving the lock wait message collection instruction.

4. The method according to claim 1, characterized in that Constructing a lock wait graph for the distributed system includes: Using the lock request subjects included in each node in the distributed system as points, and using the lock wait relationships between the lock request subjects as edges, construct the lock wait graph; In the lock wait graph, label the out-degree and in-degree for the lock request subjects; Among them, a single lock request subject includes one or more processes. The out-degree represents the number of locks that the lock request subject is waiting for; the in-degree represents the number of locks held by the lock request subject that are being waited for by other lock request subjects.

5. The method according to claim 4, wherein Performing deadlock detection based on the lock wait graph includes: In the lock wait graph, select a target point with an out-degree or in-degree of 0; From the lock wait graph, cut off the target point and its associated edges; Update the out-degree and in-degree corresponding to the remaining points in the lock wait graph; Loop through the operations of selecting the target point, cutting, and updating until there are no points with an out-degree or in-degree of 0 in the lock wait graph, and then end the loop; If there are still uncut points and edges in the lock wait graph obtained after ending the loop, based on the circular structure formed by the uncut points and edges, determine the deadlocks existing in the distributed system.

6. The method according to claim 5, wherein Based on the circular structure formed by the uncut points and edges, determining the deadlocks existing in the distributed system includes: Send lock wait status confirmation requests to the nodes where each lock request subject included in each circular structure is located respectively; Receive response information for the sent lock wait status determination requests; If it is determined according to the response information that the lock waiting status of any lock request entity in the target circular structure has changed, then the target circular structure is trimmed from the lock waiting graph; The remaining circular structures in the lock waiting graph are determined as the deadlocks existing in the distributed system.

7. The method according to claim 1, characterized in that It further includes: In the case where a deadlock is detected in the distributed system, according to a preset rule, a target process is selected from each process involved in the target deadlock; Controlling the target process to release the held lock to unlock the target deadlock; Wherein, the target deadlock is any deadlock existing in the distributed system.

8. The method according to claim 7, characterized in that It further includes: In the lock waiting graph, the in-degree is respectively injected into each process involved in the target deadlock; The step of selecting a target process from each process involved in the target deadlock according to the preset rule includes: Selecting the process with the highest in-degree from each process involved in the target deadlock as the target process; Wherein, the in-degree represents the number of locks held by the lock request entity that are waited for by other lock request entities.

9. The method according to claim 8, wherein It further includes: If there are multiple processes with the highest in-degree among the processes involved in the target deadlock and there is a local node process among them, then the local node process is selected as the target process.

10. The method according to claim 1, wherein It further includes: If a new lock waiting timeout event occurs on the local node during the execution of the target deadlock detection task, then a waiting task is created for the new lock waiting timeout event; Putting the waiting task into the local node waiting queue; After completing the target deadlock detection task, continue to control the local node to maintain the state of being in the process of executing the deadlock detection task, and sequentially consume the waiting tasks in the local node waiting queue; After sequentially completing the deadlock detection tasks triggered by the waiting tasks in the waiting queue, control the local node to end the state of being in the process of executing the deadlock detection task.

11. The method according to claim 10, characterized in that, It further includes: If it is monitored that the lock request entity to which the lock waiting timeout event corresponding to the target waiting task in the waiting queue belongs has died, then the target waiting task is deleted from the waiting queue; Releasing the lock held by the lock request entity to which the lock waiting timeout event corresponding to the target waiting task belongs.

12. The method according to claim 1, wherein It further includes: In the case of a shutdown of the local node, releasing the locks held by each lock request entity in the local node; After the local node restarts, it does not execute the deadlock detection tasks that were not completed before the shutdown.

13. A node device, characterized in that, It includes a memory, a processor, and a communication component, and the node device is any node in the distributed system; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component, and is used to execute the one or more computer instructions to execute the deadlock detection method according to any one of claims 1-12.

14. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the deadlock detection method according to any one of claims 1-12.

Citation Information

Cited By

  • Database deadlock processing method and device and database deadlock processing system

    CN121597698A