Method, apparatus and distributed system for memory recovery in a distributed system

By obtaining the availability information of the work node in a distributed system and sending memory recovery commands, the problem that the work node memory recovery is not easy to occur or leads to the system unavailability is solved, and the execution efficiency and availability of the system are improved.

CN113296937BActive Publication Date: 2025-05-27ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110181739.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-10
Publication Date
2025-05-27
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

In distributed systems, automatic JVM memory recovery of work nodes is sometimes not easy to occur, or it may cause the entire distributed system to be unavailable, affecting execution efficiency.

Method used

The master node responds to the memory recovery conditions, obtains the availability information of the work node, determines the work node that needs to perform memory recovery, and sends a memory recovery command to it. After receiving the command, the worker node determines whether the execution is allowed and performs memory recovery when allowed.

Benefits of technology

It realizes active and selective triggering of memory recovery of work nodes, avoiding unavailability caused by GC in all nodes at a certain moment, and improving the execution efficiency and availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113296937B_ABST
    Figure CN113296937B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method, an apparatus, and a distributed system for memory recycling in a distributed system. The method includes: in response to meeting the memory recycling condition, obtaining the availability information of the working nodes in the distributed system for memory recycling; determining the working nodes that need to perform memory recycling according to the availability information; and sending a memory recycling command to the working nodes that need to perform memory recycling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technologies, and particularly to a method for memory recovery in a distributed system. One or more embodiments of this specification also relate to an apparatus for memory recovery in a distributed system, a distributed system, a computing device, and a computer-readable storage medium. Background Art

[0002] GC (Garbage Collection, memory recovery) is a mechanism used to periodically recover the memory space occupied by objects with no object references during idle time, which can improve the execution efficiency and stability of the system. For example, compilation execution is widely used for the optimization of OLAP databases. During the dynamic execution of a database, compilation execution generates different codes for different execution plans, so a large number of intermediate code files are generated. Specifically, in the Java language environment, compilation execution generates a large number of class bytecode files, which are stored in the MetaSpace of the JVM. Java's automatic memory recovery mechanism is triggered under different conditions to clean up these intermediate code files. In a distributed system, worker nodes are nodes for executing computations, and each worker node performs GC based on the JVM automatic memory recovery mechanism to improve the execution efficiency of the distributed system.

[0003] However, the JVM automatic memory recovery of worker nodes sometimes does not occur easily and sometimes even causes the unavailability of the entire distributed system, which has a certain impact on the execution efficiency of the distributed system. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a method for memory recovery in a distributed system. One or more embodiments of this specification also relate to an apparatus for memory recovery in a distributed system, a distributed system, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a method for memory recovery in a distributed system is provided, including: in response to meeting the memory recovery condition, obtaining availability information of worker nodes in the distributed system for memory recovery; determining worker nodes that need to perform memory recovery according to the availability information; and sending a memory recovery command to the worker nodes that need to perform memory recovery.

[0006] Optionally, the "responding to meeting the memory recovery condition" includes: responding to the memory water level of a worker node exceeding a preset memory water level threshold, or responding to entering a new round of cycle for periodically performing memory recovery.

[0007] Optionally, sending a memory recovery command to a worker node that needs to perform memory recovery includes: determining the node order of the worker nodes that need to perform memory recovery according to the available information; and sending a memory recovery command according to the node order.

[0008] Optionally, the method further includes: receiving a response from a worker node for the memory recovery command; if the received response is a rejection, abandoning the memory recovery of the worker node for which the response is received; if the response times out, retrying or abandoning the memory recovery of the worker node for which the response times out.

[0009] According to a second aspect of the embodiments of the present specification, there is provided an apparatus for distributed system memory recovery, configured in a master node of a distributed system, including: an information acquisition module configured to acquire availability information of worker nodes in the distributed system for memory recovery in response to meeting a memory recovery condition; a node determination module configured to determine worker nodes that need to perform memory recovery according to the availability information; and a command sending module configured to send a memory recovery command to worker nodes that need to perform memory recovery.

[0010] According to a third aspect of the embodiments of the present specification, there is provided a method for distributed system memory recovery, applied to a worker node, including: providing availability information for memory recovery to a master node, so that the master node determines worker nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the worker nodes that need to perform memory recovery; receiving the memory recovery command sent by the master node; determining whether the worker node allows the execution of the memory recovery command; and if allowed, executing the memory recovery command.

[0011] Optionally, the method further includes: when the JVM automatic memory recovery mechanism of the worker node is triggered, determining whether the worker node allows the execution of JVM automatic memory recovery; and if allowed, executing JVM automatic memory recovery.

[0012] Optionally, determining whether the worker node allows the execution of the memory recovery command includes: determining whether the time interval between the time of the last memory recovery of the worker node and the current time reaches a time interval range allowing memory recovery to be performed, and / or determining whether the memory water level of the worker node exceeds a preset memory water level threshold.

[0013] Optionally, the last memory recovery of the worker node includes: memory recovery performed in response to a memory recovery command sent by the master node or another master node last time, or memory recovery performed when the JVM automatic memory recovery mechanism of the worker node was triggered last time.

[0014] Optionally, it further includes: if not allowed, feedback a response message rejecting memory recovery to the master node, and / or, in the case where the memory recovery command is executed successfully, feedback a response message indicating that the memory recovery is successfully completed to the master node.

[0015] Optionally, the method further includes: when the worker node restarts, default setting the availability information of the worker node to available; when executing the memory recovery command, updating the availability information of the worker node to unavailable; in the case where the memory recovery command is executed successfully, updating the availability information of the worker node to available.

[0016] Optionally, before executing the memory recovery, it further includes: checking whether there are unfinished tasks on the worker node; if so, after the tasks on the worker node are completed, entering the step of executing the memory recovery; if not, entering the step of executing the memory recovery.

[0017] According to a fourth aspect of the embodiments of the present specification, there is provided an apparatus for distributed system memory recovery, configured on a worker node, including: an information providing module, configured to provide availability information for memory recovery to a master node, so that the master node determines, according to the availability information, the worker nodes that need to execute memory recovery, and sends a memory recovery command to the worker nodes that need to execute memory recovery. A command receiving module, configured to receive the memory recovery command sent by the master node. An execution judgment module, configured to judge whether the worker node allows the execution of the memory recovery command. A command execution module, configured to execute the memory recovery command if the execution judgment module determines that it is allowed.

[0018] According to a fifth aspect of the embodiments of the present specification, there is provided a distributed system, including: one or more master nodes and a plurality of worker nodes communicating with the master nodes. The master node is configured to, in response to meeting the memory recovery condition, obtain the availability information of the worker nodes in the distributed system for memory recovery; determine, according to the availability information, the worker nodes that need to execute memory recovery; and send a memory recovery command to the worker nodes that need to execute memory recovery. The worker node is configured to provide availability information for memory recovery to the master node, so that the master node determines, according to the availability information, the worker nodes that need to execute memory recovery, and sends a memory recovery command to the worker nodes that need to execute memory recovery; receive the memory recovery command sent by the master node; judge whether the worker node allows the execution of the memory recovery command; and if allowed, execute the memory recovery command.

[0019] According to a sixth aspect of the embodiments of the present specification, there is provided a computing device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions: in response to meeting the memory recovery condition, obtain the availability information of the working nodes in the distributed system for memory recovery; determine the working nodes that need to perform memory recovery according to the availability information; send a memory recovery command to the working nodes that need to perform memory recovery.

[0020] According to a seventh aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing computer instructions, and when the instructions are executed by a processor, the steps of the method for distributed system memory recovery applied to the master node described in any embodiment of the present specification are implemented.

[0021] According to an eighth aspect of the embodiments of the present specification, there is provided a computing device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions: provide availability information for memory recovery for the master node, so that the master node determines the working nodes that need to perform memory recovery according to the availability information, sends a memory recovery command to the working nodes that need to perform memory recovery; receive the memory recovery command sent by the master node; determine whether the working node allows the execution of the memory recovery command; if allowed, execute the memory recovery command.

[0022] According to a ninth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing computer instructions, and when the instructions are executed by a processor, the steps of the method for distributed system memory recovery applied to the working node described in any embodiment of the present specification are implemented.

[0023] In an embodiment of one aspect of the present specification, there is provided a method for distributed system memory recovery applied to the master node. Since the master node, in response to meeting the memory recovery condition, obtains the availability information of the working nodes in the distributed system for memory recovery, determines the working nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the working nodes that need to perform memory recovery, therefore, the master node determines the working nodes that need memory recovery from a global perspective, actively and selectively controls the working nodes to perform memory recovery, which can not only achieve active and effective triggering, but also avoid the problem that the entire system becomes unavailable when all nodes perform GC at a certain moment, thus improving the execution efficiency of the system;

[0024] An embodiment on the other hand of this specification provides a method for distributed system memory recycling applied to a worker node. Since the worker node provides availability information for memory recycling to the master node and receives a memory recycling command sent by the master node according to the availability information, the worker node can be controlled by the master node to recycle memory from a global perspective, and judge whether to allow the execution of the memory recycling command in combination with its own situation. If allowed, then execute the memory recycling command, so that it can effectively trigger the worker node to recycle memory, and ensure that the GC triggered by the master node will not have an adverse impact on its own state, effectively improving the execution efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of a method for distributed system memory recycling applied to a master node provided by an embodiment of this specification;

[0026] Figure 2 is a flowchart of the processing process of a method for distributed system memory recycling applied to a master node provided by an embodiment of this specification;

[0027] Figure 3 is a schematic structural diagram of a device for distributed system memory recycling configured on a master node provided by an embodiment of this specification;

[0028] Figure 4 is a schematic structural diagram of a device for distributed system memory recycling configured on a master node provided by another embodiment of this specification;

[0029] Figure 5 is a flowchart of a method for distributed system memory recycling applied to a worker node provided by an embodiment of this specification;

[0030] Figure 6 is a flowchart of the processing process of a method for distributed system memory recycling applied to a worker node provided by an embodiment of this specification;

[0031] Figure 7 is a schematic diagram of the communication interaction between a master node and a worker node provided by an embodiment of this specification.

[0032] Figure 8 is a schematic structural diagram of a device for distributed system memory recycling configured on a worker node provided by an embodiment of this specification;

[0033] Figure 9 is a schematic structural diagram of a device for distributed system memory recycling configured on a worker node provided by another embodiment of this specification;

[0034] Figure 10 is a block diagram of the structure of a distributed system provided by an embodiment of this specification;

[0035] Figure 11 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific embodiments

[0036] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0037] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0038] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0039] First, the noun terms related to one or more embodiments of this specification are explained.

[0040] Memory recycling (GC, Garbage Collection): A mechanism for recycling the memory space occupied by objects with no object references at irregular intervals during idle time.

[0041] MPP (Massively Parallel Processing), is an architecture that distributes tasks in parallel to multiple servers and nodes, and after the calculation is completed on each node, the results of each part are aggregated together to obtain the final result. A database using the MPP architecture is called an MPP database.

[0042] Availability information is information used to describe whether memory recovery of a working node is in an available state. For example, when the availability information is "ACTIVE", it indicates availability; when the availability information is "INACTIVE", it indicates unavailability.

[0043] In this specification, a method for memory recovery in a distributed system is provided. This specification also relates to an apparatus for memory recovery in a distributed system, a distributed system, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0044] Figure 1 The flowchart of a method for memory recovery in a distributed system according to an embodiment of this specification is shown. The method provided in this embodiment can be applied to a master node. The method includes steps 102 to 106.

[0045] Step 102: In response to meeting the memory recovery condition, obtain the availability information of the working nodes in the distributed system for memory recovery.

[0046] Step 104: Determine the working nodes that need to perform memory recovery according to the availability information.

[0047] Step 106: Send a memory recovery command to the working nodes that need to perform memory recovery.

[0048] Since the master node, in response to meeting the memory recovery condition, obtains the availability information of the working nodes in the distributed system for memory recovery, determines the working nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the working nodes that need to perform memory recovery, therefore, the master node can obtain the availability information of the global working nodes in the distributed system for memory recovery, determine the working nodes that need memory recovery from a global perspective, and actively and selectively control the working nodes to perform memory recovery, which can not only actively and effectively trigger memory recovery, but also avoid the problem that the entire system becomes unavailable when all nodes perform GC at a certain moment, improving the execution efficiency of the system.

[0049] It should be noted that the memory recovery conditions in the embodiments of this specification are not limited, and can be specifically set according to the needs of the implementation scenario. For example, in one or more embodiments of this specification, in order to be able to perform memory recovery actively when the memory water level of the working node is relatively high or regularly, the response to meeting the memory recovery condition may include: in response to the memory water level of the working node exceeding a preset memory water level threshold, or in response to entering a new round of cycle for regularly performing memory recovery.

[0050] In addition, to improve the availability of the distributed system, after the master node determines the worker nodes that need to perform memory recycling from a global perspective, it can also determine the node order from a global perspective with the goal of ensuring the service sustainability of the distributed system, so that each worker node can perform memory recycling in an orderly manner and ensure that there are enough worker nodes providing services. Specifically, for example, sending a memory recycling command to the worker nodes that need to perform memory recycling includes: determining the node order of the worker nodes that need to perform memory recycling according to the available information; and sending a memory recycling command according to the node order.

[0051] For example, in actual implementation, according to the node order, a memory recycling command can be first sent to the worker node ranked first; the response of the worker node to the memory recycling command is received; according to the response of the worker node to the memory recycling command, a memory recycling command is sent to the next one or more worker nodes; if it is determined according to the node order that there are still worker nodes that have not received the memory recycling command, return to the step of receiving the response of the worker node to the memory recycling command.

[0052] After the worker node receives the memory recycling command, to ensure that the GC triggered by the master node will not have an adverse impact on its own state, the worker node can also determine whether its own situation allows the execution of the memory recycling command, and if so, execute the memory recycling command. Therefore, after the master node sends a memory recycling command to the worker node, the master node can synchronously wait for the response of the worker node to the memory recycling command. Next, combined with three possible response situations, the corresponding processing of the master node is analyzed:

[0053] Response situation one: The master node receives the response that the worker node has successfully completed the GC. In this case, the master node can record the received response and start sending memory recycling commands to the subsequent worker nodes in order.

[0054] Response situation two: The master node receives the response that the worker node refuses to perform the GC. In this case, the master node can record the received response and abandon performing the distributed GC on this worker node in this round of distributed GC.

[0055] Response situation three: The response to the memory recycling command sent by the master node times out. In this case, the master node can choose to retry or abandon performing the distributed GC on this Worker in this round.

[0056] As can be seen from the above analysis, the worker node may execute the memory recovery command and successfully complete the memory recovery, or it may reject the memory recovery, or it may fail to receive the memory recovery command due to network problems. To ensure that the master node accurately and effectively controls the memory recovery and avoid resource consumption caused by repeatedly sending memory recovery commands, in one or more embodiments of this specification, the master node may also receive the response of the worker node to the memory recovery command; if the received response is a rejection, give up the memory recovery of the worker node corresponding to the response; if the response times out, retry or give up the memory recovery of the worker node with the timeout response.

[0057] The following combines the attached Figure 2 to further illustrate the method for memory recovery in a distributed system that combines the above-mentioned multiple embodiments. Figure 2 FIG. shows a flowchart of the processing procedure of a method for memory recovery in a distributed system provided by an embodiment of this specification. The specific steps include step 202 to step 218.

[0058] Step 202: In response to the memory water level of the worker node exceeding the preset memory water level threshold, or in response to entering a new round of cycle for periodically executing memory recovery, obtain the availability information of the global worker nodes in the distributed system.

[0059] For example, the availability information can be obtained by Workers regularly reporting heartbeats to the Master, or the Master can actively send information collection messages to the Workers.

[0060] Step 204: Determine N worker nodes that need to execute memory recovery according to the availability information.

[0061] Step 206: Determine whether N is greater than zero.

[0062] Step 208: If not, sleep and wait.

[0063] Step 210: If so, send a memory recovery command to one or more of the N worker nodes.

[0064] Step 212: Receive the response of the worker node to the memory recovery command.

[0065] Step 214: For the worker node with a timeout response, retry or give up the memory recovery of the worker node with the timeout response.

[0066] Step 216: For the worker node with a rejection response, give up the memory recovery of the worker node corresponding to the response.

[0067] Step 218: Determine whether there are still worker nodes among the N worker nodes that have not been sent the memory recovery command.

[0068] If so, return to step 210 and send a memory recovery command to the worker nodes that have not received the memory recovery command.

[0069] As can be seen from the above embodiments, according to the method provided by the embodiments of the present invention, the memory recovery of the distributed system is divided into two parts: the master node and the worker nodes. The master node formulates a distributed GC plan according to the availability information of the global worker nodes, including determining the worker nodes that need to perform GC and the execution order, and initiating a memory recovery command to the worker nodes that need to perform GC. After receiving the memory recovery command, the worker nodes can choose to perform GC or refuse to perform GC according to their own status and feedback to the master node. The processing of the two parts can be attached to the corresponding master node and worker nodes in the form of threads to jointly complete the distributed GC logic. Whenever the master node completes a round of distributed memory recovery, the distributed GC thread on the master node enters the sleep state and waits for a period of time to start the next round of distributed GC.

[0070] Corresponding to the above method embodiments, this specification also provides an apparatus embodiment for memory recovery of a distributed system. Figure 3 The structural schematic diagram of an apparatus for distributed memory recovery provided by an embodiment of this specification is shown. This apparatus can be configured in the master node of the distributed system. As Figure 3 shown, this apparatus includes: an information acquisition module 302, a node determination module 304, and a command sending module 306.

[0071] This information acquisition module 302 can be configured to acquire the availability information of the worker nodes in the distributed system for memory recovery in response to meeting the memory recovery conditions.

[0072] For example, this information acquisition module 302 can be configured to acquire the availability information of the worker nodes for memory recovery in response to the memory water level of the worker node exceeding a preset memory water level threshold, or in response to entering a new round of cycle for regularly performing memory recovery.

[0073] This node determination module 304 can be configured to determine the worker nodes that need to perform memory recovery according to the availability information.

[0074] This command sending module 306 can be configured to send a memory recovery command to the worker nodes that need to perform memory recovery.

[0075] Since the master node responds to meeting the memory recovery condition, obtains the availability information of the worker nodes in the distributed system for memory recovery, determines the worker nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the worker nodes that need to perform memory recovery. Therefore, the master node can obtain the availability information of all worker nodes in the distributed system for memory recovery, determine the worker nodes that need memory recovery from a global perspective, and actively and selectively control the worker nodes to perform memory recovery. This can not only trigger memory recovery actively and effectively, but also avoid the problem that the entire system becomes unavailable when all nodes perform GC at a certain moment, improving the execution efficiency of the system.

[0076] Figure 4 FIG. shows a schematic structural diagram of a distributed memory recovery device provided by another embodiment of this specification. As Figure 4 shown, the command sending module 306 of this device may include: a sequence determination sub-module 3062 and a command sending sub-module 3064.

[0077] The sequence determination sub-module 3062 may be configured to determine the node sequence of the worker nodes that need to perform memory recovery according to the available information.

[0078] The command sending sub-module 3064 may be configured to send a memory recovery command according to the node sequence.

[0079] In the above embodiment, in order to improve the availability of the distributed system, after the master node determines the worker nodes that need to perform memory recovery from a global perspective, it may also determine the node sequence from a global perspective with the goal of ensuring the service sustainability of the distributed system, so that each worker node performs memory recovery in an orderly manner, ensuring that there are enough worker nodes providing services.

[0080] It can be understood that after receiving the memory recovery command, the worker node may execute the memory recovery command and successfully complete the memory recovery, may also reject the memory recovery, or may fail to receive the memory recovery command due to network problems. In one or more embodiments of this specification, in order to ensure that the master node accurately and effectively controls memory recovery and avoid resource consumption caused by repeatedly sending memory recovery commands, as Figure 4 shown, the device may further include: a response receiving module 308 and a response processing module 310.

[0081] The response receiving module 308 may be configured to receive the response of the worker node to the memory recovery command;

[0082] The response processing module 310 can be configured to, if the received response is a rejection, abandon the memory recovery of the working node for the response; if the response times out, retry or abandon the memory recovery of the working node with a timeout response.

[0083] The above is a schematic solution of a device for distributed system memory recovery in this embodiment. It should be noted that the technical solution of the device for distributed system memory recovery belongs to the same concept as the above technical solution of the method for distributed system memory recovery. For the details not described in the technical solution of the device for distributed system memory recovery, reference can be made to the description of the technical solution of the above method for distributed system memory recovery.

[0084] Corresponding to the above method embodiment, this specification also provides a method embodiment for distributed system memory recovery applied to a working node. Figure 5 The flowchart of a method for distributed memory recovery provided by an embodiment of this specification is shown. As Figure 5 shown, the method includes steps 502 to 508.

[0085] Step 502: Provide availability information for memory recovery to the master node, so that the master node determines the working nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the working nodes that need to perform memory recovery.

[0086] Step 504: Receive the memory recovery command sent by the master node.

[0087] Step 506: Determine whether the working node allows the execution of the memory recovery command.

[0088] Step 508: If allowed, execute the memory recovery command.

[0089] Since the working node provides availability information for memory recovery to the master node and receives the memory recovery command sent by the master node according to the availability information, the working node can be controlled by the master node for memory recovery from a global perspective, and combined with its own situation, it can judge whether to allow the execution of the memory recovery command. If allowed, then execute the memory recovery command, so that it can effectively trigger the working node to perform memory recovery, and ensure that the GC triggered by the master node will not cause adverse effects on its own state, effectively improving the execution efficiency of the system.

[0090] It should be noted that in the method provided in the embodiments of this specification, the memory reclamation triggered by the master node sending a memory reclamation command and the automatic GC of the JVM on the worker node can be completely independent and jointly implemented in the same distributed system, so that the control of the GC timing of the worker node can be more refined based on the two memory reclamation mechanisms, further improving the execution efficiency of the system. Specifically, for example, the method may further include: when the automatic memory reclamation mechanism of the JVM on the worker node is triggered, determining whether the worker node is allowed to perform automatic memory reclamation of the JVM; if allowed, performing automatic memory reclamation of the JVM.

[0091] It should be noted that the method provided in the embodiments of this specification does not limit the judgment conditions for whether the worker node is allowed to execute the memory reclamation command, and can be specifically set according to the characteristics of the worker node and the needs of the implementation environment. For example, in order to avoid overly frequent memory reclamation or memory reclamation with unreasonable memory water levels, in one or more embodiments of this specification, determining whether the worker node is allowed to execute the memory reclamation command may include: determining whether the time interval between the time of the last memory reclamation on the worker node and the current time reaches the time interval range allowed for memory reclamation; and / or determining whether the memory water level of the worker node exceeds a preset memory water level threshold. Thus, when the time interval is too short and the memory water level is not high, the worker node can reject the memory reclamation command.

[0092] For example, when the availability information of the worker node is "ACTIVE", it means that memory reclamation can be normally performed. After receiving the memory reclamation command sent by the master node, the worker node determines whether the time interval from the current time point to the completion of the last GC is less than a preset interval threshold. If it is less, the worker node may consider the memory reclamation command to be too frequent and unreasonable, reject this memory reclamation command, and feedback to the master node. If it is greater than or equal to the preset interval threshold, the worker node approves this memory reclamation command and can start to perform GC.

[0093] It can be understood that in the scenario where the memory reclamation triggered by the master node and the automatic GC of the JVM on the worker node are jointly implemented, the last GC of the worker node may be the GC triggered by the master node through the memory reclamation command, or may be the memory reclamation triggered by the automatic memory reclamation mechanism of the JVM of the worker node itself. Therefore, in one or more embodiments of this specification, the last memory reclamation of the worker node includes: the memory reclamation performed in response to the memory reclamation command sent by the master node or another master node last time, or the memory reclamation performed when the automatic memory reclamation mechanism of the JVM of the worker node was triggered last time.

[0094] To facilitate the master node to accurately and effectively control memory recycling, after the worker node determines whether to execute the memory recycling command, if it is determined that the execution is not allowed, it can also feedback a response message rejecting memory recycling to the master node. In the case where the memory recycling command is executed successfully, a response message indicating that the memory recycling is successfully completed is feedback to the master node. Through the feedback of this embodiment, the master node can give up the memory recycling of the worker node corresponding to the response according to the received response being a rejection, and realize more accurate control according to the received response being successfully completed, avoiding resource consumption caused by repeatedly sending memory recycling commands.

[0095] In addition, to ensure the sustainability of memory recycling of worker nodes and the sustainability of system services, in one or more embodiments of this specification, when a worker node restarts, the availability information of the worker node can be default set to available; when executing the memory recycling command, the availability information of the worker node is updated to unavailable; in the case where the memory recycling command is executed successfully, the availability information of the worker node is updated to available. In this embodiment, even if a worker node restarts during the distributed GC process, since it is in the available state by default, it will not have an adverse impact on the tasks executed by the restarted worker node and the distributed GC, ensuring the continuous availability of memory recycling. In the case where the memory recycling command is executed successfully, updating the availability information of the worker node to available also ensures the continuous availability of memory recycling. In addition, since the availability information of the worker node in the process of memory recycling is unavailable, it will not be confirmed by the master node again as needing to execute GC, thus ensuring the precise control of the master node and avoiding the master node making all worker nodes in the cluster be in the state of memory recycling, ensuring the continuous availability of cluster services.

[0096] To avoid memory recycling interfering with the normal tasks of worker nodes and affecting the availability of the entire cluster, in one or more embodiments of this specification, before a worker node executes memory recycling, it further includes: checking whether there are unfinished tasks on the worker node; if there are, after the tasks on the worker node are completed, enter the step of executing memory recycling; if not, enter the step of executing memory recycling.

[0097] For example, a worker node can check whether the number of its tasks is 0. If it is not 0, it means there are still tasks unfinished, and then it can wait for all tasks to be completed before performing memory recycling. The waiting can be achieved by the thread entering the sleep state for a period of time and then checking whether the number of tasks is zero, or being awakened by the event of the completion of the last task.

[0098] The following combines the attached Figure 6 , and further describes the method for memory recycling in a distributed system that combines the above multiple embodiments. Among them, Figure 6The flowchart shows the processing procedure of a method for distributed system memory recycling provided by an embodiment of this specification. The specific steps include steps 602 to 622.

[0099] Step 602: After the worker node starts, the availability information provided to the master node is defaulted to "ACTIVE".

[0100] The worker node is defaulted to "ACTIVE", indicating that it can execute tasks normally.

[0101] Step 604: Receive the memory recycling command sent by the master node.

[0102] Step 606: Determine whether the time interval between the moment of the last memory recycling on the worker node and the current moment is greater than or equal to a preset time interval threshold, and whether the memory water level of the worker node exceeds the preset memory water level threshold.

[0103] It can be understood that if it is less, the worker node can consider that the memory recycling command is too frequent and unreasonable, and can reject this memory recycling command and feedback to the master node. If it is greater than or equal to the threshold, the worker node allows this memory recycling command and starts to execute GC.

[0104] Step 608: If so, update the availability information to "INACTIVE".

[0105] In this step, the worker node sets the availability information to "INACTIVE", indicating that it no longer receives and executes tasks. The task distribution thread of the master node will not continue to distribute tasks to the worker node when it finds that the availability information of the worker node is "INACTIVE". Therefore, before memory recycling, the number of tasks on the worker node only decreases and does not increase to ensure the successful trigger of memory recycling.

[0106] Step 610: Determine whether the number of tasks of the worker node is greater than zero.

[0107] The worker node checks whether the number of its tasks is 0. If it is not 0, it means there are still tasks not completed. The worker node does not expect its GC to affect the execution of tasks and further affect the availability of the entire cluster. Therefore, it can wait for all tasks to be completed. For example, it can sleep for a period of time as in the following step 612 and then check whether the task is completed. If it is completed, it can officially perform full GC in the following step 614.

[0108] Step 612: If so, sleep and wait until the number of tasks of the worker node is equal to zero.

[0109] Step 614: If the number of tasks of the worker node is equal to zero, determine whether the memory water level of the worker node exceeds the preset memory water level threshold.

[0110] Step 616: If so, update the availability information to "INACTIVE".

[0111] Step 618: Execute the memory recovery command.

[0112] Step 620: After execution, update the availability information to "ACTIVE".

[0113] In this step, after the worker node completes GC, it can set itself back to ACTIVE to indicate that it can provide services externally.

[0114] Step 622: Record the time when the memory recovery command is executed, and send a response message indicating successful completion to the master node.

[0115] In this step, recording the time when GC is completed can provide a reference for the judgment when receiving the next GC command.

[0116] As can be seen from the above embodiments, the master node and the worker node cooperate to complete the distributed GC logic, and the control of the GC timing is more refined, achieving the effects of high availability and high efficiency.

[0117] It can be understood that in a distributed scenario, the occurrence of some uncontrollable factors may lead to some abnormal situations. Based on the above embodiments, various abnormal situations in the distributed scenario can be correspondingly processed to ensure the high availability of the distributed cluster. To make the method provided in the embodiments of this specification easier to understand, hereinafter, Figure 7 using the communication interaction schematic diagram of the master node and the worker node shown in the following, the role played by the method of distributed memory recovery provided in the embodiments of this specification in solving these abnormal situations will be schematically described:

[0118] As Figure 7 shown, one or more master nodes "Master" initiate multiple "Worker gc doing" tasks. Among them, each "Worker gc doing" task is used to send a memory recovery command to the corresponding worker node "Worker", and end, retry, or abandon the task according to the response of the worker node to the memory recovery command. For example, "Worker gcdoing" sends the memory recovery command, that is, Figure 7 "Gc" shown, to the worker node according to the availability information of the worker node being in the "Active" state. The worker node judges whether the time interval and the memory water level reach the threshold requirements. When the worker node confirms that they reach, it updates the availability information to "Inactive", and judges whether the number of tasks is equal to zero, that is, judges "task == 0". If it is determined, it judges whether the memory water level reaches the threshold requirements. If so, it executes "Gc", that is,Figure 7 "System.gc()" as shown, otherwise reject the execution of "Gc". After the successful execution of "Gc" or in case of rejection, update the availability information to the "Active" state.

[0119] Example of abnormal situation 1: As Figure 7 shown, the distributed GC initiated by the master node actively and the automatic GC of the JVM on the worker node are completely independent, which may lead to duplicate GC. For example, duplicate GC may occur in two cases. Duplicate GC case 1: Just completed an automatic JVM GC when the worker node just received the GC command. Duplicate GC case 2: The worker node triggered an automatic JVM GC during the waiting for the task to complete. Through Figure 6 the embodiments shown as well as Figure 7 the communication interaction schematic diagram shown, at these two key points where the worker node may cause duplicate GC, the worker node checks its own memory water level, and only when it exceeds the threshold will it actually perform GC, thus effectively avoiding the occurrence of duplicate GC abnormal situations.

[0120] Example of abnormal situation 2: The master node has a brain split, and multiple master nodes send memory recovery commands to the worker node. Through Figure 6 the embodiments shown as well as Figure 7 the communication interaction schematic diagram shown, it can be seen that the worker node will not execute each memory recovery command, but will compare the timestamp of the received memory recovery command and the time interval since the last GC. For example, if it is less than the preset time interval threshold, it can directly reject the execution and feedback to the initiator of the memory recovery command. If a redundant memory recovery command arrives when the worker node is executing the previous memory recovery command, it can directly feedback the information that the memory recovery is being executed to the corresponding master node.

[0121] Example of abnormal situation 3: When the master node is out of contact, the standby master node becomes the new master node. Due to the previous distributed GC initiated by the original master node, some worker nodes are in the process of memory recovery execution. At this time, the new master node starts to execute the distributed GC. Through Figure 6 the embodiments shown as well as Figure 7 the communication interaction schematic diagram shown, since the availability information of these worker nodes is in the "INACTIVE" state, therefore, they will not be determined by the new master node as the worker nodes that need to execute GC. Therefore, the method provided by the embodiments of this specification can ensure that not all worker nodes in the cluster are in the memory recovery state due to the replacement of the master node, ensuring the availability of the cluster.

[0122] Example of abnormal situation 4: The working node is disconnected. The master node will not be able to collect the availability information of the disconnected working node, so it will not perform distributed GC scheduling on it, avoiding waste of computing resources and network resources.

[0123] Example of abnormal situation 5: The working node restarts during the distributed GC process. Through Figure 6 the embodiments shown and Figure 7 the communication interaction schematic diagram shown, it can be seen that the working node restarts in the "ACTIVE" state by default, and it will not have an adverse impact on the tasks executed by the restarted working node and the distributed GC.

[0124] As can be seen from the above embodiments, the method for distributed system memory recycling provided in the embodiments of this specification fully considers various abnormal situations in the distributed scenario, implements a complete exception handling mechanism and process, ensures the normal operation of distributed GC in an extreme environment, guarantees high availability, and further ensures the efficient and stable operation of the entire distributed system.

[0125] For example, the method provided in the embodiments of this specification can be applied to a distributed system with an MPP architecture, so as to implement a highly available and efficient distributed MPP database. It can be understood that the method provided in the embodiments of this specification can not only be applied to the MPP database scenario, but also be widely used in other distributed systems, playing an important role in improving the execution efficiency and stability of the distributed system.

[0126] Corresponding to the above method embodiments, this specification also provides an embodiment of a device for distributed system memory recycling. Figure 8 The structural schematic diagram of a device for distributed memory recycling provided by an embodiment of this specification is shown. This device can be configured in the working node of the distributed system. As Figure 8 shown, this device includes: an information providing module 802, a command receiving module 804, an execution judgment module 806, and a command execution module 808.

[0127] This information providing module 802 can be configured to provide the master node with availability information for memory recycling, so that the master node determines the working nodes that need to perform memory recycling according to the availability information, and sends a memory recycling command to the working nodes that need to perform memory recycling.

[0128] This command receiving module 804 can be configured to receive the memory recycling command sent by the master node.

[0129] This execution judgment module 806 can be configured to judge whether the working node is allowed to execute the memory recycling command.

[0130] For example, the execution judgment module 806 can be configured to determine whether the time interval between the time of the last memory recovery on the working node and the current time reaches the time interval range allowing memory recovery; and / or determine whether the memory water level of the working node exceeds a preset memory water level threshold.

[0131] The command execution module 808 can be configured to execute the memory recovery command if the execution judgment module determines that it is allowed.

[0132] Since the working node provides availability information for memory recovery to the master node and receives the memory recovery command sent by the master node according to the availability information, the working node can be controlled by the master node for memory recovery from a global perspective, and combine its own situation to judge whether to allow the execution of the memory recovery command. If allowed, then execute the memory recovery command, so that it can effectively trigger the working node to perform memory recovery, and ensure that the GC triggered by the master node will not have an adverse impact on its own state, effectively improving the execution efficiency of the system.

[0133] Figure 9 The structural schematic diagram of a distributed memory recovery device provided by another embodiment of this specification is shown. As Figure 9 shown, the device may further include: an automatic memory recovery judgment module 810 and an automatic memory recovery execution module 812.

[0134] The automatic memory recovery judgment module 810 can be configured to determine whether the working node allows the execution of JVM automatic memory recovery when the JVM automatic memory recovery mechanism of the working node is triggered.

[0135] The automatic memory recovery execution module 812 can be configured to execute JVM automatic memory recovery if allowed.

[0136] In the device provided in this embodiment, the memory recovery triggered by the master node sending a memory recovery command and the automatic GC of the JVM on the working node are independent of each other and are jointly implemented in the same distributed system, so that the control of the GC timing of the working node can be more refined based on the two memory recovery mechanisms, further improving the execution efficiency of the system.

[0137] For example, the last memory recovery on the working node may include: the memory recovery executed in response to the memory recovery command sent by the master node or another master node last time, or the memory recovery executed when the JVM automatic memory recovery mechanism of the working node was triggered last time.

[0138] In one or more embodiments of this specification, as Figure 9As shown, the device may further include: a feedback module 814, which may be configured to, if the execution judgment module 806 determines that it is not allowed, feedback a response message rejecting memory recovery to the master node, and / or, in the case where the memory recovery command execution is completed, feedback a response message indicating successful completion of memory recovery to the master node.

[0139] Through the feedback of this embodiment, the master node can, based on the received response being a rejection, abandon the memory recovery of the worker node corresponding to the response, and based on the received response being successful completion, thereby achieving more accurate control and avoiding resource consumption caused by repeatedly sending memory recovery commands.

[0140] In addition, to ensure the sustainability of memory recovery of worker nodes and the sustainability of system services, in one or more embodiments of this specification, the availability information is also updated accordingly according to changes in the process. In one or more embodiments of this specification, as Figure 9 shown, the device may further include: an information update module 816, which may be configured to, in the case where the worker node restarts, default the availability information of the worker node to available; when executing a memory recovery command, update the availability information of the worker node to unavailable; in the case where the memory recovery command execution is completed, update the availability information of the worker node to available.

[0141] To avoid interference from memory recovery to the normal tasks of worker nodes and affect the availability of the entire cluster, in one or more embodiments of this specification, as Figure 9 shown, the device may further include: a task volume check module 818, which may be configured to check whether there are unfinished tasks on the worker node. If so, after the tasks on the worker node are completed, allow the command execution module 808 and / or the automatic memory recovery execution module 812 to enter the step of executing memory recovery; if not, allow the command execution module 808 and / or the automatic memory recovery execution module 812 to enter the step of executing memory recovery.

[0142] The above is a schematic solution of a device for memory recovery in a distributed system according to this embodiment. It should be noted that the technical solution of the device for memory recovery in the distributed system belongs to the same concept as the technical solution of the above-mentioned method for memory recovery in the distributed system. For the details not described in detail in the technical solution of the device for memory recovery in the distributed system, reference can be made to the description of the technical solution of the above-mentioned method for memory recovery in the distributed system.

[0143] Figure 10 Shows a block diagram of a distributed system provided by an embodiment of this specification. As Figure 10As shown, the distributed system may include: one or more master nodes 1002 and multiple worker nodes 1004 that communicate with the master nodes.

[0144] The master node 1002 may be configured to, in response to meeting the memory reclamation condition, obtain the availability information of the worker nodes in the distributed system for memory reclamation; determine the worker nodes that need to perform memory reclamation according to the availability information; and send a memory reclamation command to the worker nodes that need to perform memory reclamation.

[0145] The worker node 1004 may be configured to provide the master node with the availability information for memory reclamation, so that the master node determines the worker nodes that need to perform memory reclamation according to the availability information, sends a memory reclamation command to the worker nodes that need to perform memory reclamation; receive the memory reclamation command sent by the master node; determine whether the worker node allows the execution of the memory reclamation command; and if so, execute the memory reclamation command.

[0146] Figure 11 FIG. shows a structural block diagram of a computing device 1100 according to an embodiment of the present specification. The components of the computing device 1100 include but are not limited to a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 through a bus 1130, and a database 1150 is used to store data.

[0147] The computing device 1100 further includes an access device 1140, and the access device 1140 enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interfaces (e.g., a Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0148] In an embodiment of the present specification, the above components of the computing device 1100 and Figure 11 other components not shown in Figure 11 may also be connected to each other, for example, through a bus. It should be understood that

[0149] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 1100 can also be a mobile or stationary server.

[0150] On the one hand, the processor 1120 is configured to execute the following computer-executable instructions:

[0151] In response to meeting the memory recovery condition, obtain the availability information of the worker nodes in the distributed system for memory recovery;

[0152] Determine the worker nodes that need to perform memory recovery according to the availability information;

[0153] Send a memory recovery command to the worker nodes that need to perform memory recovery.

[0154] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the method for distributed system memory recovery applied to the master node described above belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the method for distributed system memory recovery applied to the master node.

[0155] On the other hand, the processor 1120 is configured to execute the following computer-executable instructions:

[0156] Provide the master node with the availability information for memory recovery, so that the master node determines the worker nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the worker nodes that need to perform memory recovery;

[0157] Receive the memory recovery command sent by the master node;

[0158] Determine whether the worker node allows the execution of the memory recovery command;

[0159] If allowed, execute the memory recovery command.

[0160] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the method for distributed system memory recovery applied to the worker node described above belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the method for distributed system memory recovery applied to the worker node.

[0161] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions. On the one hand, when the instructions are executed by a processor, they are used for:

[0162] In response to meeting the memory recovery condition, obtain the availability information of the working nodes in the distributed system for memory recovery;

[0163] Determine the working nodes that need to perform memory recovery according to the availability information;

[0164] Send a memory recovery command to the working nodes that need to perform memory recovery.

[0165] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the method for distributed system memory recovery applied to the master node described above belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the method for distributed system memory recovery applied to the master node.

[0166] On the other hand, when the instructions are executed by a processor, they are used for:

[0167] Provide the master node with the availability information for memory recovery, so that the master node determines the working nodes that need to perform memory recovery according to the availability information, and sends a memory recovery command to the working nodes that need to perform memory recovery;

[0168] Receive the memory recovery command sent by the master node;

[0169] Judge whether the working node allows the execution of the memory recovery command;

[0170] If allowed, execute the memory recovery command.

[0171] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the method for distributed system memory recovery applied to the working node described above belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the method for distributed system memory recovery applied to the working node.

[0172] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0173] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0174] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0175] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0176] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A method for memory recovery in a distributed system, applied to the master node of the distributed system, comprising: Upon satisfying the memory recovery condition, obtaining the availability information of the worker nodes in the distributed system for memory recovery, wherein, the step of upon satisfying the memory recovery condition includes: upon the memory water level of the worker node exceeding a preset memory water level threshold, the availability information is information used to describe whether the memory recovery of the worker node is in an available state, and when the availability information is available, it means that memory recovery can be normally executed; According to the availability information, selecting worker nodes and formulating a memory recovery plan to determine the worker nodes that need to perform memory recovery, including: from a global perspective, selecting the worker nodes that need to perform memory recovery according to the availability information, and taking the goal of ensuring the service sustainability of the distributed system to determine the node order of the selected worker nodes, so that each worker node can perform memory recovery in an orderly manner; According to the node order, sending a memory recovery command to the worker nodes that need to perform memory recovery; Receiving the response of the worker node to the memory recovery command; If the received response is a rejection, giving up the memory recovery of the worker node corresponding to the response; If the response times out, retrying or giving up the memory recovery of the worker node with a timeout response.

2. The method according to claim 1, wherein the step of upon satisfying the memory recovery condition further comprises: Upon entering a new round of cycle for periodically performing memory recovery.

3. An apparatus for memory recovery in a distributed system, configured in the master node of the distributed system, comprising: An information acquisition module, configured to obtain the availability information of the worker nodes in the distributed system for memory recovery upon satisfying the memory recovery condition, wherein, the step of upon satisfying the memory recovery condition includes: upon the memory water level of the worker node exceeding a preset memory water level threshold, the availability information is information used to describe whether the memory recovery of the worker node is in an available state, and when the availability information is available, it means that memory recovery can be normally executed; A node determination module, configured to select worker nodes and formulate a memory recovery plan according to the availability information to determine the worker nodes that need to perform memory recovery, including: from a global perspective, selecting the worker nodes that need to perform memory recovery according to the availability information, and taking the goal of ensuring the service sustainability of the distributed system to determine the node order of the selected worker nodes, so that each worker node can perform memory recovery in an orderly manner; A command sending module, configured to send a memory recovery command to the worker nodes that need to perform memory recovery according to the node order; A command receiving module, configured to receive the response of the worker node to the memory recovery command; If the received response is a rejection, giving up the memory recovery of the worker node corresponding to the response; If the response times out, retrying or giving up the memory recovery of the worker node with a timeout response.

4. A method for memory recovery in a distributed system, applied to a worker node, comprising: When it is determined that the memory water level of the working node exceeds the preset memory water level threshold, provide availability information for memory recovery to the master node. The availability information is information used to describe whether the memory recovery of the working node is in an available state. When the availability information is available, it means that memory recovery can be normally executed, so that the master node, from a global perspective, selects working nodes and formulates a memory recovery plan according to the availability information to determine the working nodes that need to perform memory recovery, and with the goal of ensuring the service sustainability of the distributed system, determines the node order of the selected working nodes, so that each working node performs memory recovery orderly, and according to the node order, sends a memory recovery command to the working nodes that need to perform memory recovery; Receive the memory recovery command sent by the master node; Determine whether the working node allows the execution of the memory recovery command; If allowed, execute the memory recovery command; If not allowed, feedback a response message rejecting memory recovery to the master node.

5. The method according to claim 4, further comprising: When the JVM automatic memory recovery mechanism of the working node is triggered, determine whether the working node allows the execution of JVM automatic memory recovery; If allowed, execute JVM automatic memory recovery.

6. The method according to claim 4 or 5, wherein the determining whether the working node allows the execution of the memory recovery command comprises: Determine whether the time interval between the time of the last memory recovery of the working node and the current time reaches the time interval range allowing the execution of memory recovery; and / or, Determine whether the memory water level of the working node exceeds the preset memory water level threshold.

7. The method according to claim 6, wherein the last memory recovery of the working node comprises: Memory recovery performed in response to the memory recovery command sent by the master node or another master node last time, or memory recovery performed when the JVM automatic memory recovery mechanism of the working node was triggered last time.

8. The method according to claim 4, further comprising: When the memory recovery command is executed successfully, feedback a response message indicating that the memory recovery is successfully completed to the master node.

9. The method according to claim 4, further comprising: When the working node restarts, default the availability information of the working node to available; When executing the memory recovery command, update the availability information of the working node to unavailable; When the memory recovery command is executed successfully, update the availability information of the working node to available.

10. The method according to claim 4 or 5, before executing memory recovery, further comprising: Check whether there are unfinished tasks on the working node; If so, after the tasks on the working node are completed, enter the step of executing memory recovery; If not, enter the step of executing memory recovery.

11. A device for memory recovery of a distributed system, configured on a working node, comprising: An information providing module, configured to provide availability information for memory recovery to the master node when the memory water level exceeds a preset memory water level threshold, where the availability information is information used to describe whether the memory recovery of the worker node is in an available state, and when the availability information is available, it indicates that memory recovery can be normally performed, so that the master node, from a global perspective, selects worker nodes and formulates a memory recovery plan according to the availability information to determine the worker nodes that need to perform memory recovery, and with the goal of ensuring the service sustainability of the distributed system, determines the node order of the selected worker nodes, so that each worker node performs memory recovery in an orderly manner, and sends a memory recovery command to the worker nodes that need to perform memory recovery according to the node order; A command receiving module, configured to receive the memory recovery command sent by the master node; An execution judgment module, configured to judge whether the worker node is allowed to execute the memory recovery command; A command execution module, configured to execute the memory recovery command if the execution judgment module determines that it is allowed; If the execution judgment module determines that it is not allowed, feedback a response message rejecting memory recovery to the master node.

12. A distributed system, comprising: One or more master nodes and multiple worker nodes communicating with the master nodes; The master node is configured to, in response to meeting the memory recovery condition, obtain the availability information of the worker nodes in the distributed system for memory recovery, where in response to meeting the memory recovery condition includes: in response to the memory water level of the worker node exceeding a preset memory water level threshold, the availability information is information used to describe whether the memory recovery of the worker node is in an available state, and when the availability information is available, it indicates that memory recovery can be normally performed; select worker nodes and formulate a memory recovery plan according to the availability information to determine the worker nodes that need to perform memory recovery, including: selecting the worker nodes that need to perform memory recovery from a global perspective according to the availability information, and with the goal of ensuring the service sustainability of the distributed system, determining the node order of the selected worker nodes, so that each worker node performs memory recovery in an orderly manner; sending a memory recovery command to the worker nodes that need to perform memory recovery according to the node order; receiving the response of the worker node to the memory recovery command; if the received response is a rejection, abandon the memory recovery of the worker node corresponding to the response; if the response times out, retry or abandon the memory recovery of the worker node with the timeout response; The working node is configured to provide availability information for memory recovery to the master node when the memory water level exceeds a preset memory water level threshold, so that the master node determines the working nodes that need to perform memory recovery based on the availability information, and sends a memory recovery command to the working nodes that need to perform memory recovery; receive the memory recovery command sent by the master node; determine whether the working node allows the execution of the memory recovery command; if allowed, execute the memory recovery command; if not allowed, feedback a response message rejecting the memory recovery to the master node.

13. A computing device, comprising: a memory and a processor; the memory is used for storing computer-executable instructions, and the processor is used for executing the computer-executable instructions: in response to meeting the memory recovery condition, obtain the availability information of the working nodes in the distributed system for memory recovery, wherein, the response to meeting the memory recovery condition includes: in response to the memory water level of the working node exceeding a preset memory water level threshold, the availability information is information used to describe whether the memory recovery of the working node is in an available state, and when the availability information is available, it means that memory recovery can be normally executed; select working nodes and formulate a memory recovery plan according to the availability information to determine the working nodes that need to perform memory recovery, including: from a global perspective, select the working nodes that need to perform memory recovery according to the availability information, and with the goal of ensuring the service sustainability of the distributed system, determine the node order of the selected working nodes, so that each working node can perform memory recovery in an orderly manner; send a memory recovery command to the working nodes that need to perform memory recovery according to the node order; receive the response of the working node to the memory recovery command; if the received response is a rejection, abandon the memory recovery of the working node corresponding to the response; if the response times out, retry or abandon the memory recovery of the working node with the timeout response.

14. A computer-readable storage medium storing computer instructions, which when executed by a processor implement the steps of the method for memory recovery of the distributed system according to any one of claims 1 to 2.

15. A computing device, comprising: a memory and a processor; the memory is used for storing computer-executable instructions, and the processor is used for executing the computer-executable instructions: When it is determined that the memory water level of a worker node exceeds a preset memory water level threshold, availability information for memory reclaim is provided to the master node. The availability information is information used to describe whether memory reclaim of the worker node is in an available state. When the availability information is available, it means that memory reclaim can be normally executed, so that the master node, from a global perspective, selects worker nodes and formulates a memory reclaim plan according to the availability information to determine the worker nodes that need to execute memory reclaim, and with the goal of ensuring the service sustainability of the distributed system, determines the node order of the selected worker nodes, so that each worker node can perform memory reclaim in an orderly manner, and sends a memory reclaim command to the worker nodes that need to execute memory reclaim according to the node order; Receive a memory reclaim command sent by the master node; Determine whether the worker node allows execution of the memory reclaim command; If allowed, execute the memory reclaim command; If not allowed, feedback a response message rejecting memory reclaim to the master node.

16. A computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the steps of the method for memory reclaim of the distributed system according to any one of claims 4 to 10 are implemented.

Citation Information

Patent Citations

  • Coordinated Garbage Collection in Distributed Systems

    US20160070593A1