Management system, method, apparatus, and computing device for a distributed system
By introducing a cross-task memory mechanism and a large model into the distributed system, task processing nodes make independent decisions, while evaluation nodes make global decisions. This solves the management challenges of dynamic and complex decision-making scenarios in existing technologies, and achieves more efficient and accurate management decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing distributed system management decisions struggle to cope with dynamic changes and complex decision-making scenarios, resulting in low accuracy and efficiency. Existing solutions lack flexibility and intelligence, leading to decision failures or resource waste.
By introducing a cross-task memory mechanism and a large model, task processing nodes make independent decisions, and evaluation nodes make global decisions based on global constraints and historical decision information, simplifying the interaction logic between nodes and enhancing the system's flexibility and scalability.
By employing cross-task memory mechanisms and large-scale model-assisted decision-making, the accuracy and efficiency of distributed system management are improved, development and maintenance complexity are reduced, and the system's flexibility and scalability are enhanced.
Smart Images

Figure CN122111632A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a management system, method, apparatus and computing device for a distributed system. Background Technology
[0002] With the rapid development of computer technology, the demand for data storage, computation, and management is constantly increasing, leading to the emergence of distributed systems. A distributed system is a cluster composed of multiple cooperating nodes that, under the control of a scheduling mechanism, can simultaneously meet the diverse service requests of a massive number of users. As the system scales up, the management of distributed systems becomes increasingly complex.
[0003] In existing technologies, management decisions in distributed systems often rely on real-time decisions or static rules, which are insufficient to cope with dynamic changes and complex decision-making scenarios, resulting in low accuracy and efficiency in management decisions. Therefore, there is an urgent need for a more efficient and accurate management solution for distributed systems. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a management system for a distributed system. One or more embodiments of this specification also relate to a management method for a distributed system, a management device for a distributed system, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a management system for a distributed system is provided, including: a task processing node and an evaluation node; The task processing node is configured to respond to the target management task of the distributed system, obtain the corresponding target decision data, and generate the node decision result of the target management task based on the task processing node's own decision logic and the target decision data. The evaluation node is configured to acquire the node decision results generated by each task processing node, and determine the global decision result of the target management task based on global constraints and the decision results of each node, wherein the global decision result is used to manage the distributed system.
[0006] According to a second aspect of the embodiments of this specification, a management method for a distributed system is provided, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the method includes: In response to goal management tasks, acquire the corresponding goal decision data; Based on the task processing node's own decision-making logic and the target decision data, a node decision result for the target management task is generated. The node decision result is used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
[0007] According to a third aspect of the embodiments of this specification, a task processing apparatus is provided, applied to a task processing node in a management system of a distributed system, the management system further comprising an evaluation node; the apparatus includes: The acquisition module is configured to acquire the corresponding target decision data in response to target management tasks. The generation module is configured to generate node decision results for the target management task based on the task processing node's own decision logic and the target decision data. The node decision results are used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
[0008] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the management method of the above-described distributed system.
[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the management method of the distributed system described above.
[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the management method of the distributed system described above.
[0011] This specification provides a management system for a distributed system, including: task processing nodes and evaluation nodes; the task processing nodes are configured to respond to a target management task of the distributed system, acquire corresponding target decision data; and generate node decision results for the target management task based on the task processing node's own decision logic and the target decision data; the evaluation nodes are configured to acquire the node decision results generated by each task processing node, and determine the global decision result of the target management task based on global constraints and the node decision results, wherein the global decision result is used to manage the distributed system.
[0012] One embodiment of this specification implements a two-stage processing method that divides the complete decision-making process into two phases. Each task processing node makes decisions independently based on its own decision logic, generating node decision results. Then, the evaluation node evaluates the decision results of each node based on global constraints to obtain the global decision result of the target management task. The task processing nodes are unaware of each other and do not need to care about whether there is a conflict with other task processing nodes, thus avoiding complex inter-node interaction logic. The evaluation node does not need to care about the specific decision-making process of the node decision results, thereby reducing the complexity of development and maintenance and enhancing the flexibility and scalability of the system. Attached Figure Description
[0013] Figure 1 This is a structural block diagram of a management system for a distributed system provided in one embodiment of this specification; Figure 2 This is a task flow diagram of a target management task provided in one embodiment of this specification; Figure 3 This is a schematic diagram illustrating the execution process of a target management task according to one embodiment of this specification; Figure 4 This is a flowchart illustrating a management method for a distributed system according to one embodiment of this specification; Figure 5 This is a flowchart illustrating the processing steps of a distributed system management method provided in one embodiment of this specification. Figure 6 This is a schematic diagram of the structure of a management device for a distributed system provided in one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0019] Job: In batch processing and management tasks, a Job is a core concept, representing a complete and independent unit of execution. A Job typically consists of a series of steps, each responsible for completing a specific task. A step can be completed by a task processing node (WorkNode). The definition and execution method of a Job depend on the framework or system used. In the embodiments of this specification, a Job represents a complete management task completed by multiple task processing nodes, which is initiated and executed by a triggering event.
[0020] JobInstance: A common concept in batch processing and task management, it represents a single execution instance of a specific job. Each execution of a Job is called a JobInstance, which contains the execution status of multiple task processing nodes. Each JobInstance is typically associated with a set of parameters that can be used to distinguish the execution of different tasks.
[0021] WorkNode: A common concept in distributed computing and task scheduling systems, especially in custom workflow or batch processing frameworks. It represents a node or unit of work that executes a specific task. A WorkNode can be an independent task or a step in a workflow. Each WorkNode is typically responsible for completing a specific task and can collaborate with other WorkNodes to complete more complex logic. In the embodiments of this specification, a WorkNode refers to the basic execution unit in a Job, responsible for making decisions within the JobInstance and generating the corresponding node decision result (Proposal).
[0022] Proposal: The node decision results generated by the WorkNode, that is, the decision suggestions for the current management task. The Evaluator uses them to decide on the final management operation to be performed.
[0023] Evaluator: Also known as the evaluation node, it is responsible for evaluating the node decision results proposed by all WorkNodes to determine which node decision results correspond to which management operations will be executed.
[0024] JobContext: An information-sharing node that shares context information within the same JobInstance. All WorkNodes share this information.
[0025] JobMemory: A memory node that stores historical decision information across JobInstances, enabling WorkNodes to make more complex decisions based on this historical information.
[0026] It's important to note that as distributed systems grow in scale, their management (such as Elasticsearch clusters) becomes increasingly complex. Current management decisions often rely on immediate decisions for single tasks or static rules, making it difficult to handle dynamic changes and complex decision-making scenarios. Furthermore, manual decision-making and solutions with low levels of automation are not only inefficient but also prone to errors.
[0027] In one implementation, the managed service of the distributed system provides functions such as automatic scaling and load balancing. However, the automated decision-making is mainly based on the result of a single task. Such frameworks usually only make decisions based on the current state of the distributed system when each task runs. They lack the ability to remember and pass context for multiple task runs and cannot make better decisions in dynamic environments. Especially in the case of proposal conflicts, the managed service often cannot handle conflicts correctly, leading to decision failure or inefficiency.
[0028] Another approach is to rely on static rules and real-time data to implement managed services for distributed systems, which have automation capabilities. These static rule-driven automated decision-making systems rely on fixed threshold triggers (such as scaling up when CPU utilization exceeds a certain value). They have poor flexibility and intelligence, and cannot adaptively adjust, which can easily lead to overscaling or resource waste in complex scenarios, and lack more intelligent decision-making capabilities.
[0029] Another implementation approach provides an operations and maintenance monitoring center node, offering monitoring and alarm functions for the distributed system. However, it cannot automatically make decisions between multiple tasks in complex scenarios. This type of framework uses centralized global state management, that is, it uses centralized global management to decide the state of all task nodes. Although it can manage globally, it lacks the ability to process local nodes in real time, and may also lead to performance bottlenecks and insufficient flexibility. Furthermore, the code for the operations and maintenance monitoring center node is complex and difficult to add logic to.
[0030] This specification provides a distributed system management system that introduces a cross-task memory mechanism, enabling the management system to make more accurate decisions based on historical decision information. Furthermore, through the evaluation node (i.e., the global evaluator), each task processing node does not need to consider the global applicability of the decision results individually, thereby significantly simplifying the development and decision-making complexity of each task processing node.
[0031] This specification provides a management system for a distributed system, and also relates to a management method for a distributed system, a management device for a distributed system, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0032] See Figure 1 , Figure 1 This specification illustrates a structural block diagram of a management system for a distributed system according to one embodiment. Figure 1 As shown, the management system of the distributed system includes a task processing node 102 and an evaluation node 104. Task processing node 102 is configured to respond to the target management tasks of the distributed system, obtain the corresponding target decision data, and generate the node decision results of the target management tasks based on its own decision logic and target decision data. Evaluation node 104 is configured to obtain the node decision results generated by each task processing node 102, and determine the global decision result of the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
[0033] A distributed system is a cluster of multiple cooperating nodes that, under the control of a scheduling mechanism, can simultaneously meet the diverse service requests of a massive number of users. The management system is used to manage the cluster resources of this distributed system, such as dynamic scaling up, dynamic scaling down, load balancing, and cluster parameter adjustment.
[0034] The management system includes task processing nodes and evaluation nodes. A task processing node is a type of work node, representing a node or work unit that performs the target management task. The management system may include at least one task processing node, and each task processing node is used to complete a specific sub-task in the target management task, such as a data acquisition sub-task, a CPU decision sub-task, a memory decision sub-task, a complex balancing sub-task, etc.
[0035] An evaluation node (Evaluator) is a global evaluator that can globally evaluate the node decision results of each task processing node to determine the global decision result. The management operations corresponding to the global decision result will be executed on the distributed system to achieve the management of the distributed system.
[0036] In practice, the target management task (Job) of this distributed system can be triggered automatically or manually, such as automatically triggered on a schedule or based on the monitoring system, or manually triggered by the administrator. The target management task is used to instruct the analysis of the current cluster resources in the managed distributed system to determine whether certain management operations need to be performed. The target management task is the management task to be performed at the current time. The execution of a management task (Job) is called a task execution instance (JobInstance). A task execution instance can produce the global decision result of the current target management task based on the current resource status of the distributed system.
[0037] The task processing node responds to the target management task of the distributed system by obtaining the corresponding target decision data. This target decision data refers to the data required by the task processing node to make management decisions, such as the current cluster resources of the distributed system under management. The task processing node combines its own decision logic and the target decision data to generate the node decision result (i.e., proposal) for the target management task. The task processing node's own decision logic consists of the logical rules for making decision judgments, which can be configured based on actual project requirements. Different task processing nodes can be configured with different own decision logics to achieve different decision judgments.
[0038] In one alternative implementation, the global constraint is a constraint rule of the management strategy configured from the global perspective of the distributed system. The global constraint is configured in the evaluation node. After obtaining the decision results of each node, the decision results of each node can be inferred based on rule matching to determine which node decision results should be executed as the global decision results.
[0039] In another optional implementation, the evaluation node can also make global decisions based on a large model. In practice, historical global decision results can be obtained, and these historical global decision results, along with the decision results of each node generated in the current round and global constraints, can be input into the large model. The large model then performs global reasoning on the decision results of each node and outputs the global decision result for the target management task.
[0040] Large models refer to deep learning models with a massive number of parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of parameters. Large models are also known as foundation models. They are pre-trained on large-scale unlabeled corpora, producing pre-trained models with hundreds of millions of parameters. These models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0041] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based question answering reasoning, sentiment classification, text summarization generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0042] In practical implementation, this large model can be an acquired open-source general-purpose large model, or a large model that has been adjusted and trained for management decision-making scenarios in distributed systems. Historical global decision results provide reference information for the large model, enabling it to infer and determine the global decision result for the target management task from the decision results of each node generated in the current round based on global constraints.
[0043] Among them, the historical global decision results can be the global decision results within a set time period before the current time, that is, the management strategies issued to the distributed system in the past. For example, the set time period can be the global decision results within a time period of half an hour before the current time, 7 days before the current time, or half a month before the current time. The historical global decision results can be the expansion strategy issued in the past 5 minutes.
[0044] In practice, historical global decision results can carry decision reasons, which refer to the reasons for determining the historical global decision result from the decision results of each node in the past. Therefore, by inputting the historical global decision result, the decision results of each node generated in the current round, and the global constraints into the large model, the large model can, under the global constraints, refer to the historical global decision results and the decision reasons they carry, reason about the decision results of each node, and output the global decision result of the target management task.
[0045] Alternatively, historical global decision results may not carry the decision cause. The large model can infer whether there is a conflict between the historical global decision results and the decision results of each node generated in the current round. Then, combined with global constraints, it can infer the current global decision result. For example, if a node's decision result is a shrinkage strategy, but the historical global decision result is that an expansion strategy was issued five minutes ago, then the node's decision result is likely unreasonable, and the probability that the large model infers the node's decision result as the global decision result will be reduced.
[0046] It should be noted that large models can capture decision information from historical global decision results, infer the incentives for global decisions, and leverage the powerful reasoning capabilities of large models. By referencing historical global decision results and based on the decision results of each node generated in the current round, better global decisions can be made, which helps to improve the quality of decision-making.
[0047] In the embodiments described in this specification, each task processing node independently generates node decision results. The evaluation node performs a global analysis of the node decision results generated by each task processing node based on global constraints to determine which global decision results to execute. Task processing nodes do not need to concern themselves with whether there are conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to concern itself with the specific decision-making process of the node decision results, thereby reducing the complexity of development and maintenance and enhancing the flexibility and scalability of the system. Here, global constraints refer to constraints applied to the entire distributed system, used to determine which decisions to execute for the distributed system.
[0048] In one optional implementation of this embodiment, such as Figure 1 As shown, the management system also includes an information sharing node 106; Information sharing node 106 is configured to acquire and store shared information from each task processing node 102, wherein the shared information is information that needs to be transmitted between different task processing nodes 102.
[0049] The information sharing node is a sharing mechanism, also known as JobContext, that stores shared information that needs to be passed between different task processing nodes within a single task execution instance (JobInstance). This shared information is the context information that needs to be passed to other task processing nodes. The information sharing node can be a database, which is maintained by the management system and is unknown to the individual task processing nodes.
[0050] Specifically, this information sharing node is a custom program, an underlying persistent storage. Essentially, it's a custom storage unit defined in code to store shared information from each task processing node and provide it to those nodes for reading. In other words, the information sharing node is a storage unit configured with read / write interfaces to provide read / write services. This information sharing node can be accessed proactively by the task processing nodes, or it can be accessed by the management system and then distributed to the task processing nodes.
[0051] For example, some task processing nodes are data acquisition nodes, while others are decision-making nodes. Decision-making nodes need to acquire data collected by data acquisition nodes to make decisions. Since the data acquisition process is relatively complex, it is split into multiple task processing nodes to facilitate collection and management. The acquired data is the shared information that needs to be passed to other task processing nodes.
[0052] In practice, the management system includes at least one task processing node. If there is only one task processing node, it can make multiple decisions, obtaining multiple node decision results, which are then evaluated globally by an evaluation node. If there are two or more task processing nodes, each node is responsible for executing different decision steps within the target management task. The management system also includes an information sharing node, which can acquire and store shared information that needs to be transferred between different task processing nodes. This shared information can be data collected by the task processing nodes or node decision results generated by the task processing nodes.
[0053] It should be noted that when any task processing node receives the corresponding target decision data in response to a target management task, it can write the collected target decision data to the information sharing node via the first access interface of the information sharing node, based on the scheduler of the management system. This data can then be transmitted to other task processing nodes during the decision-making process of the target management task. Furthermore, after generating a node decision result, in addition to providing it to the evaluation node for global evaluation, any task processing node can also write it to the information sharing node via the first access interface of the information sharing node, based on the scheduler of the management system. This allows other task processing nodes to refer to the node decision result of this task processing node when making their own decisions.
[0054] Furthermore, since the information sharing node is used to transmit shared information during the decision-making process of the current target management task, after the current target management task decision is completed, that is, after the global decision result is determined by the evaluation node, the shared information stored in the information sharing node can be cleared so that the shared information that needs to be transmitted between different task processing nodes during the execution of the next round of management tasks can continue to be recorded.
[0055] In the embodiments of this specification, a target management task can be executed by at least two task processing nodes. Different task processing nodes can transmit shared information in the decision-making process of the current target management task through information sharing nodes, thereby realizing the sharing of context information between different task processing nodes, making the task processing nodes more intelligent and coherent when generating node decision results.
[0056] In one optional implementation of this embodiment, the information sharing node 106 is managed by the management system and provides a first access interface to the task processing node 102. The first access interface is the read / write interface of the information sharing node. The information sharing node 106 is further configured to obtain shared information from each task processing node 102 based on the first access interface, store the shared information, and manage the lifecycle of each stored shared information. Task processing node 102 is further configured to obtain target decision data stored in information sharing node 106 based on the first access interface, wherein the target decision data is shared information collected by other task processing nodes and stored in information sharing node.
[0057] Specifically, the first access interface is a read-write interface that allows shared information from each task processing node to be written to the information sharing node, enabling the information sharing node to store and manage the written shared information. Furthermore, the first access interface can also be used by task processing nodes to access the information sharing node and retrieve the corresponding shared information.
[0058] In practical implementations, within distributed computing or task scheduling frameworks, information-sharing nodes are used to store state information, configuration parameters, or temporary data related to specific tasks. This mechanism allows necessary shared information to be transferred between different stages (i.e., different task processing nodes). Information-sharing nodes are similar to database storage; the task processing nodes themselves are unaware of this information, and the management system injects the necessary shared information into the task processing nodes.
[0059] Specifically, the scheduler of the management system can be responsible for injecting information sharing nodes into each task processing node. This means that when a task processing node starts making a decision, it can already access the information sharing node to obtain the corresponding shared information. This injection is usually completed during the initialization of the task processing node.
[0060] The lifecycle of an information-sharing node is typically tied to a target management task; that is, the node is created when the task begins and destroyed upon completion. During this process, all task processing nodes related to the target management task can access the information-sharing node to obtain the shared information that needs to be transmitted. Since the information-sharing node may be accessed simultaneously by multiple parallel task processing nodes, the management system must ensure the consistency and thread safety of the internal data within the node. Furthermore, depending on actual needs, the information-sharing node can be used to store various types of shared information, such as collected resource data from multiple dimensions of the distributed system and generated node decision results.
[0061] As an example, the task processing node includes a data acquisition node and a decision node; the data acquisition node responds to the target management task, collects the target decision data corresponding to the target management task, and writes the target decision data into the information sharing node; the decision node responds to the target management task, and obtains the target decision data stored in the information sharing node based on the first access interface.
[0062] In the embodiments of this specification, for task processing nodes, the information sharing node is managed by the management system. The task processing node does not need to know the specific implementation details of the information sharing node, nor does it need to care about how the shared data is collected and stored, nor does it need to care about and manage the lifecycle of the shared information in the information sharing node. The task processing node only needs to declare that the decision-making process of the target management task requires the corresponding data, and access the information sharing node according to the first access interface provided by the management system to obtain the corresponding shared information. That is, the information sharing node provides a context information required for managing and sharing the target management task during execution. This information is transparent to the task processing node that directly participates in the calculation, which simplifies the communication complexity between different nodes in the management system, greatly reduces the complexity of the task processing node, and improves the maintainability and scalability of the management system.
[0063] Of course, in actual implementation, in addition to the information sharing nodes being managed by the management system, a traditional storage scheme can also be adopted, in which each task processing node actively accesses the information sharing node to obtain the corresponding shared information. However, in this case, the task processing node also needs to consider the lifecycle management of the corresponding shared information.
[0064] In one optional implementation of this embodiment, such as Figure 1 As shown, the management system also includes a memory node 108, which is configured to store historical decision information of historical management tasks. Task processing node 102 is further configured to acquire historical decision information of historical management tasks based on memory node 108; and generate node decision results of target management tasks based on the task processing node 102's own decision logic, target decision data and historical decision information.
[0065] It should be noted that the memory node is a cross-task decision information memory mechanism, also known as JobMemory. This memory node can store historical decision information of historical management tasks executed before the current time, which is used to assist the target management task being executed at the current time in generating decision results. In other words, the memory node can remember the historical decision information of cross-task execution instances (JobInstance). This historical decision information can include historical resource data of the distributed system when executing historical management tasks, such as historical CPU data, historical memory data, and historical water level data. In addition, this historical decision information can also include the historical decision results of each task processing node in executing historical management tasks, as well as the historical global evaluation result of the historical management task finally determined by the evaluation node. That is, the decision results of historical management tasks can also be saved across tasks through the memory node.
[0066] The task processing node can also obtain historical decision information of historical management tasks based on the memory node. The historical management task is a management task executed before the current time, and the historical decision information is the relevant resource data and / or decision results of the decision-making process of the historical management task.
[0067] Specifically, this memory node is a custom program, an underlying persistent storage, essentially a storage unit defined by code. It stores historical decision information from past management tasks. In other words, through code definition, each time a management task is executed, relevant decision information is written to the memory node, such as resource data, the decision results of each task processing node executing the management task, and the final global evaluation result of the management task determined by the evaluation node. When executing the current target management task, each task processing node can access this memory node to retrieve its stored cross-task decision information. In other words, the memory node is a storage unit configured with read / write interfaces to provide read / write services. This memory node can be accessed proactively by task processing nodes or accessed by the management system and then distributed to the task processing nodes.
[0068] For example, in Java, JVM memory usage exhibits a sawtooth pattern with garbage collection (GC), but becomes almost a straight line when memory pressure is high. It's difficult to make decisions based on only a single instantaneous value, and relying on external monitoring systems to track historical data compromises availability. Therefore, the embodiments in this specification can introduce a memory mechanism to remember historical decision information across tasks, helping task processing nodes perform decision-making support analysis for the target management tasks currently being executed.
[0069] In the embodiments of this specification, a memory node is introduced into the management system to store historical decision information of historical management tasks. When a task processing node makes a management decision, it can obtain the historical decision information of historical management tasks, refer to the historical decision information across tasks, and generate the node decision result of the current target management task. This allows the generated node decision result to refer to historical decision information, enabling more efficient and accurate decision-making by dealing with the dynamically changing resource environment and complex decision-making scenarios of the distributed system without human intervention.
[0070] In one optional implementation of this embodiment, the memory node 108 is managed by the management system and provides a second access interface, which is a read-write interface. Memory node 108 is configured to obtain decision information from task processing node 102 and / or evaluation node 104 during the decision-making process of target management tasks based on the second access interface, store the decision information corresponding to the target management tasks, and manage the lifecycle of the stored decision information of each management task. The task processing node is further configured to obtain historical decision information of historical management tasks stored in the memory node based on the second access interface.
[0071] The second access interface is a read-write interface that allows decision information to be written to the memory node, enabling the memory node to store and manage the written decision information. Furthermore, the second access interface can also be used by task processing nodes to access the memory node to read corresponding historical decision information.
[0072] It should be noted that memory nodes can remember historical decision information across task execution instances (i.e., across management tasks). A target management task is a management task currently being executed. Memory nodes can obtain and store the decision information corresponding to this target management task during its decision-making process. This decision information includes the current cluster resource status of the distributed system based on the task processing nodes' decisions, the node decision results generated by the task processing nodes, the resource information based on the evaluation nodes' global perspective evaluation, and the final determined global decision result, etc.
[0073] Furthermore, in practical implementation, memory nodes can manage the lifecycle of decision information for each management task they store, providing it to the corresponding task processing nodes in subsequent management tasks. Specifically, managing the lifecycle of decision information for each management task is a crucial process, encompassing the entire process from creation to final archiving or deletion. The lifecycle typically includes the following stages: creation, use, sharing, archiving, and destruction. Each stage requires appropriate management and strategies to ensure the security, integrity, and compliance of the decision information. The entire lifecycle of the decision information is managed by the management system.
[0074] In the embodiments described in this specification, the memory data, water level data, node decision results, and global decision results corresponding to the target management task can be stored through memory nodes. The management system automatically maintains the lifecycle of each decision information stored in the memory nodes, and the task processing nodes do not need to concern themselves with data issues. For task processing nodes, the memory nodes are managed by the management system. The task processing nodes do not need to know the specific implementation details of the memory nodes, nor do they need to care about or manage the lifecycle of each decision information in the memory nodes. This simplifies the communication complexity between different nodes in the management system, greatly reduces the complexity of task processing nodes, and improves the maintainability and scalability of the management system.
[0075] In practical implementations, within distributed computing or task scheduling frameworks, memory nodes are used to store historical decision information for management tasks executed up to the current time. This memory mechanism allows necessary decision information to be saved across different management tasks (i.e., different task execution instances). Memory nodes are similar to database storage; the task processing nodes themselves are unaware of this information, and the management system injects the required historical decision information into the task processing nodes.
[0076] Specifically, the management system is responsible for injecting memory nodes into each task processing node. This means that when a task processing node starts making a decision, it can already access the memory node to obtain the corresponding historical decision information. This injection is usually completed during the initialization of the task processing node.
[0077] In the embodiments of this specification, for task processing nodes, the memory node is managed by the management system. The task processing node does not need to know the specific implementation details of the memory node, nor does it need to care about how the historical decision information is stored and managed throughout its lifecycle. The task processing node only needs to declare that the decision process in response to the target management task requires corresponding historical decision information, and access the memory node through the second access interface provided by the management system to obtain the corresponding historical decision information. That is, the memory node provides a mechanism for managing and sharing the historical decision information of each management task in time sequence. This historical decision information is transparent to the task processing nodes that directly participate in the calculation, which simplifies the communication complexity between different nodes in the management system, greatly reduces the complexity of the task processing node, and improves the maintainability and scalability of the management system.
[0078] In one optional implementation of this embodiment, such as Figure 1 As shown, the management system also includes a scheduler 110; Scheduler 110 is configured to receive target management tasks from the distributed system, determine the current task processing node 102 of the target management task, and send the current task information of the target management task to the task processing node 102. Task processing node 102 is configured to obtain target decision data based on current task information, generate node decision results for target management tasks based on its own decision logic and target decision data, and feed back target decision data and node decision results to scheduler 110. The scheduler 110 is further configured to determine the next task processing node 102 to execute the target management task based on the target decision data and the node decision results, and to send the current task information of the target management task to the next task processing node 102.
[0079] It should be noted that the goal management task can consist of a series of steps, each step is responsible for completing a specific decision sub-task and analyzing decision data in the corresponding dimension. One step can be completed by one task processing node, that is, one task processing node can analyze and process decision data in one dimension.
[0080] In actual implementation, the management system also includes a scheduler, which is used to schedule various task processing nodes. Specifically, the scheduler can receive the target management task of the distributed system, determine the task processing node to execute the decision task, and send the corresponding task information to the task processing node, so that the task processing node can make a decision based on the task information and obtain the node decision result. The task processing node can feed back the decision data used and the output node decision result to the scheduler. The scheduler determines whether to let the next task processing node continue to make the decision, and the specific next task processing node, and continues to send the current task information of the target management task to the next task processing node to continue the node decision-making.
[0081] Specifically, task processing nodes can include starting task nodes. The scheduler can first schedule target management tasks to the starting task node. The starting task node can determine the node decision result based on the current resource status of the distributed system, its own decision logic, and other input data, and feed it back to the scheduler. The scheduler can determine the next decision subtask and pass it to the next task processing node for subsequent decision-making.
[0082] Example, Figure 2 This is a task flow diagram of a target management task provided in one embodiment of this specification, such as... Figure 2 As shown, the task flow of the target management task includes task processing nodes 1 through 4. The scheduler receives the target management task and schedules it to task processing node 1. Task processing node 1 collects decision data, performs decision output, and feeds back the input decision data and output decision to the scheduler. Based on the target management task and the data fed back from task processing node 1, the scheduler schedules the next decision subtask to task processing node 2. Task processing node 2 collects decision data, performs decision output, and feeds back the input decision data and output decision to the scheduler. Based on the target management task and the data fed back from task processing node 2, the scheduler schedules the next decision subtask to task processing node 3. Task processing node 3 collects decision data, performs decision output, and feeds back the input decision data and output decision to the scheduler. Based on the target management task and the data fed back from task processing node 3, the scheduler schedules the next decision subtask to task processing node 4. Figure 2 The dashed arrows between task processing nodes indicate a logical connection between them, but the task processing nodes do not actually interact with each other; instead, they interact with the scheduler.
[0083] In the embodiments of this specification, the management system also includes a scheduler, which schedules each task processing node to achieve a complete task flow. The task processing nodes do not need to interact directly with each other. The task processing nodes only need to interact with the scheduler and do not need to care about the processing logic of other task processing nodes. The task processing nodes are unaware of each other, which reduces the complexity of development and maintenance.
[0084] In one implementation, such as Figure 1 As shown, the management system includes a scheduler. The scheduler can inject shared information from information sharing nodes into task processing nodes. Specifically, the scheduler calls a first access interface to retrieve shared information from the information sharing nodes and transmits it to the corresponding task processing nodes. Furthermore, the scheduler can also write shared information to the information sharing nodes. Additionally, the scheduler can inject historical decision information from memory nodes into task processing nodes. This is achieved by the scheduler calling a second access interface to retrieve historical decision information from memory nodes, transmit it to the corresponding task processing nodes, and write decision information to the memory nodes.
[0085] Of course, in actual implementation, the task processing node can also actively call the first access interface to obtain the shared information in the information sharing node, and the task processing node can actively call the second access interface to obtain the historical decision information in the memory node. This specification does not limit this embodiment.
[0086] In one optional implementation of this embodiment, the target management task includes the triggering of timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system; Task processing node 102 is further configured to obtain cluster information of the distributed system; based on the task processing node's own decision logic, cluster information and historical decision information, it generates node decision results for the target management task, wherein the node decision results are used to indicate the management strategy of the distributed system.
[0087] Scheduled events refer to common requirements in many applications and systems, especially in scenarios where tasks need to be executed periodically. Scheduled events are triggered when a timer reaches its set duration. System monitoring events are events that are automatically triggered when specific conditions or anomalies are detected in a distributed system.
[0088] In practice, target management tasks can be triggered based on timed management events and / or system monitoring events of the distributed system. The task processing node collects and analyzes the current cluster information of the distributed system, such as CPU and memory. Based on the task processing node's own decision-making logic, cluster information, and historical decision information, the node decision result of the target management task is generated. The node decision result is used to indicate the management strategy of the distributed system. For example, if the CPU is less than a set threshold, the distributed system is scaled down.
[0089] Specifically, a distributed system can be configured with a timer that triggers a timed management event at set intervals, which can trigger the target management task for the current time; and / or, a distributed system can be configured with a monitoring policy that triggers a system monitoring event when the current cluster resources of the distributed system meet the monitoring policy, which can trigger the target management task for the current time.
[0090] Continuing with the previous example, such as Figure 2 As shown, target management tasks are triggered by scheduled management events and / or system monitoring events.
[0091] In the embodiments of this specification, target management tasks can be triggered based on timed management events and / or system monitoring events of the distributed system. This allows task processing nodes to collect cluster information of the distributed system, combine their own decision-making logic, cluster information, and historical decision information to generate node decision results for the target management task, thereby determining how to manage the distributed system. The target management task of the distributed system can be triggered in various forms, making the management of the distributed system more flexible and allowing for different application scenarios.
[0092] In one optional implementation of this embodiment, the evaluation node 104 is at least one; Evaluation node 104 is further configured to receive node decision results of the target project type; determine the node decision results to be executed from the node decision results of each node of the target project type based on the global constraints corresponding to the target project type; and use the node decision results to be executed as the global decision results under the target project type; wherein, the target project type is the project type corresponding to the evaluation node.
[0093] In one optional implementation, different types of evaluation nodes may be included. Each evaluation node is pre-configured with global constraints corresponding to its project type, which can be defined based on the project characteristics of that project type. After each task processing node outputs its node decision results, it can input the node decision results for a specific project type to the corresponding evaluation node based on the project type of those node decisions. This evaluation node then performs a global decision based on its own configured global constraints, determining which decision results to execute from the node decision results for that project type.
[0094] In another optional implementation, the decision results corresponding to global constraints can be pre-configured in the target task processing node. The target task processing node can be any task processing node. The decision results of each node are input to the evaluation node. The evaluation node determines whether there are conflicting decision results among the received decision results of each node. If not, it means that the decision results of each node can be executed in the distributed system, and the distributed system is managed. If there are conflicting decision results, the relevant decision data of the conflicting decision results can be obtained based on the global constraints. Based on the relevant decision data, it is determined which decision result among the conflicting decision results can be executed, or none of them can be executed, to obtain the target decision result. Then, the non-conflicting decision results from each node's decision results and the target decision result are used as the global decision result and executed in the distributed system to manage the distributed system.
[0095] For example, in automated cluster resource decision-making, it's necessary to determine whether scaling up or down is needed based on the current cluster usage. However, different metrics may lead to inconsistent and conflicting node decisions. For instance, task processing node A might decide to scale down based on CPU usage falling below a set threshold, while task processing node B might decide to scale up based on memory usage exceeding the threshold, and task processing node C might predict that scaling up is needed based on historical data. In this case, there's a conflict between the node decisions generated by task processing nodes A, B, and C. By leveraging global constraints, we can obtain historical and current CPU utilization, historical and current memory utilization, and the current and historical data of the distributed system. Based on pre-configured global decision logic (i.e., global constraints), we can determine whether the global decision is to scale up or down.
[0096] It should be noted that each task processing node independently generates its node decision results. The evaluation node then performs a global decision based on these node decisions for different project types, thereby obtaining the overall global decision result. The evaluation node does not need to concern itself with the specific decision-making process of the node decisions; it only needs to focus on the received node decisions. This reduces the complexity of development and maintenance, and enhances the system's flexibility and scalability.
[0097] As an example, Figure 3 This is a schematic diagram illustrating the execution process of a target management task according to one embodiment of this specification, such as... Figure 3 As shown, based on the timeline, the task to be executed at the current time is the target management task, and the management tasks executed before the current time are historical management tasks.
[0098] For target management tasks, the scheduler can pass the information to task processing node A in the management system. Task processing node A retrieves historical decision information from the memory node, and based on its own decision logic, the collected current resource information, and the historical decision information, generates node decision result A and node decision result B for the target management task. Furthermore, based on the data fed back by task processing node A, the scheduler schedules the next decision operation to task processing node B. Task processing node B collects the corresponding cluster resource information as the node decision result based on this decision operation and writes it to the information sharing node. Then, based on the data fed back by task processing node B, the scheduler schedules the next decision operation to task processing node C. Task processing node C retrieves the cluster resource information from the information sharing node, and based on its own decision logic, the cluster resource information, and the historical decision information, generates node decision result C for the target management task. Subsequently, based on the data fed back from task processing node C, the scheduler schedules the next decision operation to task processing node D. Task processing node D obtains shared information from the information sharing node and, based on its own decision logic, shared information, and historical decision information, generates the node decision result D for the target management task. The node decision results A, B, C, and D are then input to the evaluation node of the management system to obtain the global decision M.
[0099] Figure 3 The dashed arrows between the task processing nodes indicate that there is a logical connection between the task processing nodes, but the task processing nodes do not interact with each other. In the decision-making process of the above-mentioned task processing node AD for target management tasks, it only needs to interact with the scheduler. There is no awareness between the task processing nodes. Furthermore, all information that needs to be passed between different task processing nodes can be written to the information sharing node.
[0100] In addition, the decision-making information related to the decision-making process of the target management task being executed at the current time can be written into the memory node as historical decision information for reference when executing management tasks at subsequent times.
[0101] In the embodiments described in this specification, information sharing nodes enable information sharing between different task processing nodes, and memory nodes allow the retention of historical decision information across multiple tasks, thereby making the node decision results generated by task processing nodes more intelligent and consistent. Evaluation nodes assess the node decision results generated by multiple task processing nodes from a global perspective, handling global constraints between node decision results and avoiding misjudgments due to insufficient local information. Each task processing node is an atomic, independent module that only needs to focus on generating its own node decision results; complex global decisions and conflict handling are handled by the evaluation nodes, simplifying the development process of each task processing node, improving the system's scalability and the convenience of modular development, and reducing development complexity.
[0102] Figure 4 A flowchart is shown of a management method for a distributed system according to an embodiment of this specification, applied to a task processing node in a management system of the distributed system. The management system also includes an evaluation node. The method specifically includes the following steps.
[0103] Step 402: In response to the goal management task, obtain the corresponding goal decision data.
[0104] In one optional implementation of this embodiment, the target decision data is shared information collected by other task processing nodes and stored in the information sharing node; in response to the target management task, the corresponding target decision data is obtained, including: The target decision data stored in the information sharing node is obtained based on the first access interface.
[0105] The information sharing node is managed by the management system and provides a first access interface to the task processing node. The first access interface is the read and write interface of the information sharing node.
[0106] Step 404: Based on the task processing node's own decision logic and target decision data, generate node decision results for the target management task. The node decision results are used to instruct the evaluation node to determine the global decision results for the target management task based on global constraints and the decision results of each node. The global decision results are used to manage the distributed system.
[0107] In an optional implementation of this embodiment, the management system further includes memory nodes, which are used to store historical decision information of historical management tasks; the method further includes: Retrieve historical decision-making information for historical management tasks based on memory nodes; Accordingly, based on the task processing node's own decision-making logic and target decision data, the node decision results for the target management task are generated, including: Based on the task processing node's own decision-making logic, target decision data, and historical decision information, the node decision results of the target management task are generated.
[0108] In one optional implementation of this embodiment, the memory node is managed by the management system, which provides a second access interface, which is a read / write interface for the memory node; based on the memory node, historical decision information for historical management tasks is obtained, including: The historical decision information of historical management tasks stored in the memory node is obtained based on the second access interface.
[0109] In one optional implementation of this embodiment, the target management task is triggered based on the timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system; In response to goal management tasks, acquire corresponding goal decision data, including: In response to target management tasks, obtain cluster information of the distributed system; Based on the task processing node's own decision-making logic, target decision data, and historical decision information, the node decision results for the target management task are generated, including: Based on the task processing node's own decision-making logic, cluster information, and historical decision-making information, node decision results for target management tasks are generated, whereby the node decision results are used to indicate the management strategy of the distributed system.
[0110] This specification provides a distributed system management method. A memory node is introduced into the management system to store historical decision information for past management tasks. When making management decisions, task processing nodes can obtain this historical decision information and, by referencing cross-task historical decision information, generate the node decision result for the current target management task. This allows the generated node decision result to reference historical decision information, enabling more efficient and accurate decision-making in the face of dynamically changing resource environments and complex decision-making scenarios in a distributed system without human intervention. Furthermore, each task processing node independently makes decisions based on its own decision logic, generating node decision results. An evaluation node evaluates the decision results of each node based on global constraints to obtain the global decision result for the target management task. Task processing nodes do not need to worry about conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to concern itself with the specific decision-making process of the node decision results, thereby reducing development and maintenance complexity and enhancing the system's flexibility and scalability.
[0111] The above is an illustrative scheme of a distributed system management method according to this embodiment. It should be noted that the technical solution of this distributed system management method belongs to the same concept as the technical solution of the distributed system management system described above. For details not described in detail in the technical solution of the distributed system management method, please refer to the description of the technical solution of the distributed system management system described above.
[0112] The following is in conjunction with the appendix Figure 5 Taking the application of the distributed system management method provided in this manual in a cluster resource management scenario as an example, this paper further explains the distributed system management method. Figure 5 The present specification illustrates a process flowchart of a distributed system management method according to an embodiment, which specifically includes the following steps.
[0113] Step 502: The scheduler of the management system responds to the timed management events and / or system monitoring events of the distributed system, triggers the corresponding target management task, and schedules the target management task to the first task processing node in the management system.
[0114] Step 504: The first task processing node responds to the target management task of the distributed system and obtains the cluster information of the distributed system; it obtains the historical decision information of the historical management tasks stored in the memory node based on the second access interface; based on the first task processing node's own decision logic, the collected cluster information and the historical decision information, it generates the first management strategy for the target management task; it writes the shared information that needs to be passed to other task nodes into the information sharing node based on the first access interface; it transmits the first management strategy to the evaluation node; and it feeds back the collected cluster information and the generated first management strategy to the scheduler.
[0115] Step 506: Based on the data fed back by the first task processing node, the scheduler schedules the next decision operation to the second task processing node, which is the node following the first task processing node.
[0116] Step 508: The second task processing node obtains the shared information stored in the information sharing node based on the first access interface; obtains the historical decision information of historical management tasks stored in the memory node based on the second access interface; generates a second management strategy for the target management task based on the second task processing node's own decision logic, shared information, and historical decision information; writes the shared information that needs to be passed to other task nodes into the information sharing node based on the first access interface; transmits the second management strategy to the evaluation node; and feeds back the collected cluster information and the generated first management strategy to the scheduler, which then schedules the next task processing node until the last task processing node.
[0117] Step 510: The evaluation node obtains the management strategies generated by each task processing node, analyzes each management strategy based on global constraints, and determines the global decision results for the target management task to be executed.
[0118] This specification provides a distributed system management method. A memory node is introduced into the management system to store historical decision information for past management tasks. When making management decisions, task processing nodes can obtain this historical decision information and, by referencing cross-task historical decision information, generate the node decision result for the current target management task. This allows the generated node decision result to reference historical decision information, enabling more efficient and accurate decision-making in the face of dynamically changing resource environments and complex decision-making scenarios in a distributed system without human intervention. Furthermore, each task processing node independently makes decisions based on its own decision logic, generating node decision results. An evaluation node evaluates the decision results of each node based on global constraints to obtain the global decision result for the target management task. Task processing nodes do not need to worry about conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to concern itself with the specific decision-making process of the node decision results, thereby reducing development and maintenance complexity and enhancing the system's flexibility and scalability.
[0119] Corresponding to the above method embodiments, this specification also provides embodiments of a management device for a distributed system. Figure 6 This specification shows a schematic diagram of a management device for a distributed system according to an embodiment of the present specification. The device is applied to a task processing node in a management system of the distributed system. The management system also includes an evaluation node, such as... Figure 6 As shown, the device includes: The acquisition module 602 is configured to acquire the corresponding target decision data in response to a target management task. The generation module 604 is configured to generate node decision results for the target management task based on the task processing node's own decision logic and target decision data. The node decision results are used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
[0120] This specification provides a management device for a distributed system, which divides the complete decision-making process into two stages. Each task processing node makes decisions independently based on its own decision logic, generating node decision results. Then, an evaluation node evaluates the decision results of each node based on global constraints to obtain the global decision result of the target management task. The task processing nodes are unaware of each other, and the task processing nodes do not need to care about whether there is a conflict with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to care about the specific decision-making process of the node decision results, thereby reducing the complexity of development and maintenance and enhancing the flexibility and scalability of the system.
[0121] The above is an illustrative scheme of a distributed system management device according to this embodiment. It should be noted that the technical solution of this distributed system management device belongs to the same concept as the technical solutions of the distributed system management system and the distributed system management method described above. For details not described in detail in the technical solution of the distributed system management device, please refer to the description of the technical solutions of the distributed system management system and the distributed system management method described above.
[0122] Figure 7 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0123] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0124] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0125] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0126] The processor 720 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the management method of the above-described distributed system.
[0127] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the management system and the management method of the distributed system described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the management system and the management method of the distributed system described above.
[0128] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the management method for the distributed system described above.
[0129] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the management system and the management method of the distributed system described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the management system and the management method of the distributed system described above.
[0130] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the management method for the distributed system described above.
[0131] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the management system and management method of the distributed system described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the management system and management method of the distributed system described above.
[0132] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0133] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0134] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0135] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0136] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the embodiments described in this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A management system for a distributed system, comprising: Task processing nodes and evaluation nodes; The task processing node is configured to respond to the target management task of the distributed system and obtain the corresponding target decision data; Based on the task processing node's own decision-making logic and the target decision data, the node decision result of the target management task is generated; The evaluation node is configured to acquire the node decision results generated by each task processing node, and determine the global decision result of the target management task based on global constraints and the decision results of each node, wherein the global decision result is used to manage the distributed system.
2. The management system for the distributed system according to claim 1, wherein the management system further includes an information sharing node; The information sharing node is configured to acquire and store shared information from each task processing node, wherein, The shared information refers to information that needs to be transmitted between different task processing nodes.
3. The management system of the distributed system according to claim 2, wherein the information sharing node is managed by the management system and provides a first access interface to the task processing node, wherein the first access interface is the read / write interface of the information sharing node; The information sharing node is further configured to obtain shared information from each task processing node based on the first access interface, store the shared information, and manage the lifecycle of each stored shared information. The task processing node is further configured to obtain the target decision data stored in the information sharing node based on the first access interface, wherein, The target decision data is shared information collected by other task processing nodes and stored in the information sharing node.
4. The management system for the distributed system according to claim 1, wherein the management system further includes a memory node, the memory node being configured to store historical decision information of historical management tasks; The task processing node is further configured to acquire historical decision information of historical management tasks based on the memory node; and to generate node decision results of the target management task based on the task processing node's own decision logic, the target decision data, and the historical decision information.
5. The management system of the distributed system according to claim 4, wherein the memory node is managed by the management system and provides a second access interface, the second access interface being the read / write interface of the memory node; The memory node is configured to obtain decision information from the task processing node and / or the evaluation node during the decision-making process of the target management task based on the second access interface, and store the decision information corresponding to the target management task. The lifecycle of decision-making information for each management task in the management storage; The task processing node is further configured to obtain historical decision information of historical management tasks stored in the memory node based on the second access interface.
6. The management system for the distributed system according to claim 4, wherein the management system further includes a scheduler; The scheduler is configured to receive the target management task of the distributed system, determine the current task processing node of the target management task, and send the current task information of the target management task to the task processing node. The task processing node is configured to obtain the target decision data based on the current task information, and generate the node decision result of the target management task based on the task processing node's own decision logic and the target decision data. The target decision data and the node decision results are fed back to the scheduler; The scheduler is further configured to determine the next task processing node to execute the target management task based on the target decision data and the node decision results, and to send the current task information of the target management task to the next task processing node.
7. The management system for a distributed system according to claim 6, wherein the target management task is triggered based on the timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system; The task processing node is further configured to obtain cluster information of the distributed system; Based on the task processing node's own decision-making logic, the cluster information, and the historical decision information, a node decision result for the target management task is generated, wherein the node decision result is used to indicate the management strategy of the distributed system.
8. The management system for the distributed system according to claim 1, wherein the evaluation node is at least one; The evaluation node is further configured to receive node decision results for the target project type; determine the node decision results to be executed from the node decision results of the target project type based on the global constraints corresponding to the target project type; and use the node decision results to be executed as the global decision results under the target project type; wherein, The target project type is the project type corresponding to the evaluation node.
9. A management method for a distributed system, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the method comprising: In response to goal management tasks, acquire the corresponding goal decision data; Based on the task processing node's own decision-making logic and the target decision data, a node decision result for the target management task is generated. The node decision result is used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
10. A management device for a distributed system, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the device comprising: The acquisition module is configured to acquire the corresponding target decision data in response to target management tasks. The generation module is configured to generate node decision results for the target management task based on the task processing node's own decision logic and the target decision data. The node decision results are used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
11. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the management method of the distributed system as described in claim 9.
12. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the management method for a distributed system as described in claim 9.
13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the management method for a distributed system as described in claim 9.