Management system, method and apparatus for distributed system, and computing device
By introducing a two-stage decision-making mechanism involving task processing nodes and evaluation nodes, and combining large models and historical decision information, the challenges of dynamic changes and complex decisions in distributed system management are solved, achieving more efficient and accurate management.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-06-04
AI Technical Summary
Existing distributed system management decisions rely on real-time decisions or static rules, which makes it difficult to cope with dynamic changes and complex decision-making scenarios, resulting in low management accuracy and efficiency.
A two-stage decision-making mechanism is introduced, consisting of task processing nodes and evaluation nodes. The task processing nodes independently generate node decision results, while the evaluation nodes determine the global decision results based on global constraints and node decision results. The large model is used for global reasoning and reference to historical decision information.
It simplifies the development and maintenance complexity of task processing nodes, enhances the system's flexibility and scalability, and improves the accuracy and efficiency of decision-making.
Smart Images

Figure CN2025130707_04062026_PF_FP_ABST
Abstract
Description
Management system, method, device and computing device of distributed system
[0001] The present disclosure claims priority to Chinese patent application No. 202411747063.3, filed on November 29, 2024 with the Chinese Patent Office, entitled "Management system, method, device and computing device of distributed system", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to the technical field of computer, in particular to a management system, method, device and computing device of distributed system. BACKGROUND
[0003] With the rapid development of computer technology, the demand for data storage, calculation and management is increasing, and distributed system emerges as the times require. Distributed system is a cluster composed of multiple nodes working cooperatively, which can meet the service requests of a large number of users under the control of scheduling mechanism. With the expansion of system size, the management of distributed system becomes more and more complex.
[0004] In the prior art, the management decision of distributed system often depends on instant decision or static rule, which is difficult to cope with dynamic changes and complex decision scenarios, and the accuracy and efficiency of management decision are low. Therefore, there is an urgent need for a more efficient and accurate management scheme of distributed system. SUMMARY
[0005] Therefore, the embodiments of the present disclosure provide a management system of distributed system. One or more embodiments of the present disclosure also relate to a management method of distributed system, a management device of distributed system, a computing device, a computer readable storage medium and a computer program product, to solve the technical defects in the prior art.
[0006] According to a first aspect of the embodiments of the present disclosure, a management system of distributed system is provided, comprising: a task processing node and an evaluation node.
[0007] The task processing node is configured to, in response to a target management task of the distributed system, acquire corresponding target decision data, and based on the self decision logic of the task processing node and the target decision data, generate a node decision result of the target management task.
[0008] The evaluation node is configured to acquire the node decision result generated by each task processing node, and based on global constraints and each node decision result, determine a global decision result of the target management task, wherein the global decision result is used to manage the distributed system.
[0009] According to a second aspect of embodiments of the present disclosure, a management method of a distributed system is provided, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the method comprises:
[0010] in response to a target management task, obtaining corresponding target decision data;
[0011] based on the self decision logic of the task processing node and the target decision data, generating a node decision result of the target management task, wherein the node decision result is used to instruct the evaluation node to determine a global decision result of the target management task based on global constraints and each node decision result, and the global decision result is used to manage the distributed system.
[0012] According to a third aspect of embodiments of the present disclosure, a task processing apparatus is provided, applied to a task processing node in a management system of a distributed system, the management system further comprising an evaluation node; the apparatus comprises:
[0013] an obtaining module configured to, in response to a target management task, obtain corresponding target decision data;
[0014] a generating module configured to, based on the self decision logic of the task processing node and the target decision data, generate a node decision result of the target management task, wherein the node decision result is used to instruct the evaluation node to determine a global decision result of the target management task based on global constraints and each node decision result, and the global decision result is used to manage the distributed system.
[0015] According to a fourth aspect of embodiments of the present disclosure, a computing device is provided, comprising:
[0016] a memory and a processor;
[0017] the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned management method of the distributed system.
[0018] According to a fifth aspect of embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer executable instructions, which, when executed by a processor, implement the steps of the above-mentioned management method of the distributed system.
[0019] According to a sixth aspect of embodiments of the present disclosure, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned management method of the distributed system.
[0020] The embodiment of the present disclosure provides a management system of a distributed system, comprising: a task processing node and an evaluation node; the task processing node is configured to acquire corresponding target decision data in response to a target management task of the distributed system; the task processing node generates a node decision result of the target management task based on self decision logic of the task processing node and the target decision data; the evaluation node is configured to acquire the node decision result generated by each task processing node, and determine a global decision result of the target management task based on global constraints and the node decision result, wherein the global decision result is used for managing the distributed system.
[0021] The embodiment of the present disclosure realizes that the complete decision process is divided into two stages, each task processing node independently makes a decision based on self decision logic to generate a node decision result, and then the evaluation node evaluates each node decision result based on global constraints to obtain a global decision result of the target management task, each task processing node is not aware of each other, and the task processing node does not need to care whether it conflicts with other task processing nodes, thereby avoiding complex inter-node interaction logic, and the evaluation node does not need to care about the specific decision process of the node decision result, thereby reducing the complexity of development and maintenance, and enhancing the flexibility and expansibility of the system. BRIEF DESCRIPTION OF DRAWINGS
[0022] FIG. 1 is a structural block diagram of a management system of a distributed system according to an embodiment of the present disclosure;
[0023] FIG. 2 is a task flow diagram of a target management task according to an embodiment of the present disclosure;
[0024] FIG. 3 is an execution process diagram of a target management task according to an embodiment of the present disclosure;
[0025] FIG. 4 is a flowchart of a management method of a distributed system according to an embodiment of the present disclosure;
[0026] FIG. 5 is a processing process flowchart of a management method of a distributed system according to an embodiment of the present disclosure;
[0027] FIG. 6 is a structural diagram of a management device of a distributed system according to an embodiment of the present disclosure;
[0028] FIG. 7 is a structural block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many different ways than described herein, and those skilled in the art can make similar extensions without departing from the spirit of the present disclosure, so the present disclosure is not limited to the specific implementation disclosed below.
[0030] The terminology used in this disclosure one or more embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure one or more embodiments. As used in this disclosure one or more embodiments and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this disclosure one or more embodiments, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0031] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of this disclosure one or more embodiments, first can be termed second, and, similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to a determination."
[0032] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this disclosure one or more embodiments are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0033] First, the nomenclature involved in this disclosure one or more embodiments is explained.
[0034] Job: In batch processing and management tasks, Job is a core concept, which represents a complete and independent execution unit. A Job is usually composed of a series of Steps, each responsible for completing a specific task, and a Step can be completed by a WorkNode. The definition and execution of a Job depend on the framework or system used. In this disclosure embodiment, Job represents a complete management task completed by multiple WorkNodes, which is started and executed by a trigger event.
[0035] JobInstance: A common concept in batch processing and task management, it represents an execution instance of a specific job, i.e., each execution of a job is called a JobInstance, which contains the execution of multiple task processing nodes. Each JobInstance is usually associated with a set of parameters that can be used to distinguish different task executions.
[0036] WorkNode: A common concept in distributed computing and task scheduling systems, especially in custom workflow or batch processing frameworks, it represents a node or work unit that performs a specific task. WorkNode can be an independent task or a step in a workflow. Each WorkNode is usually responsible for completing a specific task and can work with other WorkNodes to complete more complex logic. In the embodiments of the present disclosure, WorkNode refers to a basic execution unit in a Job, responsible for making decisions in a JobInstance and generating corresponding node decision results (Proposal).
[0037] Proposal: Node decision results generated by WorkNode, i.e., decision suggestions for the current management task, Evaluator bases on them to decide the final management operation to be executed.
[0038] Evaluator: Evaluator, i.e., evaluation node, responsible for evaluating all node decision results proposed by WorkNode to determine which management operation corresponding to the node decision result will be executed.
[0039] JobContext: Information sharing node, context information shared within the same JobInstance, all WorkNodes share these information.
[0040] JobMemory: Memory node, saves historical decision information across JobInstance, so that WorkNode can make more complex decisions based on historical decision information.
[0041] It should be noted that as the scale of distributed systems expands, the management of distributed systems (such as Elasticsearch clusters) becomes more and more complex. Current management decisions often rely on immediate decisions of single tasks or static rules, which are difficult to cope with dynamic changes and complex decision scenarios. In addition, manual decision-making and low-automation solutions not only have low efficiency, but also are prone to errors.
[0042] In one implementation, a hosting service of a distributed system provides functions such as automatic scaling and load balancing, but the automated decision is mainly based on the result of a single task. Such a framework usually only makes decisions based on the current state of the distributed system at each task run, lacks the ability to remember multiple task runs and context transfer, and cannot make better decisions in a dynamic environment, especially in the case of proposal conflicts. The hosting service often cannot correctly handle the conflicts, resulting in failed or inefficient decisions.
[0043] In another implementation, the hosting service of the distributed system can also rely on static rules and real-time data to achieve automation. Such static rule-driven automated decision systems have poor flexibility and intelligence. They cannot adaptively adjust, leading to over-expansion or resource waste in complex scenarios, and lack more intelligent decision-making capabilities.
[0044] In yet another implementation, an operation and maintenance monitoring center node can be provided to provide monitoring and alarm functions for the distributed system, but it cannot automatically make decisions between multiple tasks in complex scenarios. Such a framework uses centralized global state management, that is, uses centralized global management to decide the state of all task nodes. Although it can be uniformly managed globally, it lacks real-time processing capability for local nodes, and may cause performance bottlenecks and lack of flexibility. In addition, the operation and maintenance monitoring center node code is complex and it is difficult to add logic.
[0045] Embodiments of the present disclosure provide a management system for a distributed system, which introduces a cross-task memory mechanism, enabling the management system to make more accurate decisions based on historical decision information. In addition, through the evaluation node (i.e., the global evaluator), each task processing node does not need to consider the global applicability of the decision result separately, thereby significantly simplifying the development and decision complexity of each task processing node.
[0046] In the present disclosure, a management system for a distributed system is provided. The present disclosure also relates to a management method for a distributed system, a management device for a distributed system, a computing device, and a computer-readable storage medium, which are described in detail in the following embodiments.
[0047] Referring to FIG. 1, FIG. 1 shows a structural block diagram of a management system for a distributed system according to one embodiment of the present disclosure. As shown in FIG. 1, the management system for a distributed system includes a task processing node 102 and an evaluation node 104.
[0048] The task processing node 102 is configured to obtain corresponding target decision data in response to a target management task of the distributed system, and generate a node decision result of the target management task based on a self decision logic of the task processing node 102 and the target decision data.
[0049] The evaluation node 104 is configured to obtain the node decision results generated by the task processing nodes 102, and determine a global decision result of the target management task based on global constraints and the node decision results, wherein the global decision result is used to manage the distributed system.
[0050] The distributed system is a cluster composed of multiple nodes working cooperatively, and can simultaneously satisfy multiple service requests of a large number of users under the control of a scheduling mechanism. The management system is used to manage the cluster resources of the distributed system, such as dynamic expansion, dynamic contraction, load balancing, cluster parameter adjustment, and the like.
[0051] The management system includes task processing nodes and evaluation nodes. The task processing node is a kind of work node (WorkNode), which represents a node or a work unit for executing a target management task. The management system can include at least one task processing node, and each task processing node is used to complete a specific subtask in the target management task, such as a data collection subtask, a CPU decision subtask, a memory decision subtask, and a complex balancing subtask.
[0052] The evaluation node (Evaluator) is a kind of global evaluator, which can globally evaluate the node decision results of the task processing nodes to determine a global decision result. The management operation corresponding to the global decision result will be executed on the distributed system to achieve the management of the distributed system.
[0053] In actual implementation, the target management task (Job) of the distributed system can be triggered automatically or manually, such as being triggered automatically at a fixed time or being triggered automatically based on a monitoring system, or being triggered manually by a management personnel, and the like. The target management task is used to instruct to analyze the current cluster resources of the managed distributed system, and determine whether to execute some management operation task. The target management task is a management task to be executed at the current time. The execution of one management task (Job) is called a task execution instance (JobInstance). A task execution instance can generate a global decision result of the current target management task based on the current resource status of the distributed system.
[0054] The task processing node obtains corresponding target decision data in response to the target management task of the distributed system, where the target decision data is data required for the task processing node to make a management decision, such as current cluster resources of the managed distributed system. The task processing node generates a node decision result (also referred to as a Proposal) of the target management task in combination with self decision logic and the target decision data, where the self decision logic is a logical rule for decision judgment of the task processing node and can be configured based on actual project requirements. Different task processing nodes can be configured with different self decision logic for different decision judgments.
[0055] In an optional implementation, the global constraint condition is a constraint rule of a management policy configured from a global perspective of the distributed system. The global constraint condition is configured in the evaluation node. After the node decision results are obtained, the node decision results can be inferred based on rule matching to determine which node decision results are to be executed as the global decision result.
[0056] In another optional implementation, the evaluation node can also implement global decision based on a large model. In actual implementation, historical global decision results can be obtained. The historical global decision results, the node decision results generated in the current round, and the global constraint are input into the large model. The large model performs global inference on the node decision results to output the global decision result of the target management task.
[0057] The large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, billions, tens of billions, hundreds of billions, or even tens of billions of model parameters. The large model can also be referred to as a foundation model. A large-scale unlabeled corpus is used for pre-training of the large model to output a pre-trained model with hundreds of millions of parameters. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, a large language model (LLM) and a multi-modal pre-training model.
[0058] In actual application, the large model only needs a small amount of samples to fine-tune the pre-trained model and can be applied to different tasks. The large model can be widely applied to natural language processing (NLP, Natural Language Processing) and computer vision, and can be applied to computer vision tasks such as visual question answering (VQA, Visual Question Answering), image captioning (IC, Image Caption), image generation, and natural language processing tasks such as text-based question answering reasoning, sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, and the like.
[0059] In a specific implementation, the large model can be an open-source general large model obtained, or a large model adjusted after training in a management and decision-making scenario of a distributed system. The historical global decision result is provided as reference information for the large model, so that the large model can refer to the historical global decision result and infer and determine the global decision result of the target management task from the node decision results generated in the current round based on the global constraint condition.
[0060] The historical global decision result can be a global decision result within a set time period before the current time, that is, a management strategy issued to the distributed system in the past time. For example, the set time period can be a global decision result within a time period of half an hour before the current time, 7 days before the current time, or half a month before the current time. The historical global decision result can be a capacity expansion strategy issued 5 minutes ago.
[0061] In actual implementation, the historical global decision result can carry a decision reason, which refers to the reason for determining the historical global decision result from the historical node decision results. Thus, the historical global decision result, the node decision results generated in the current round, and the global constraint condition are input into the large model, which can infer the node decision results under the condition of the global constraint, refer to the historical global decision result and the decision reason carried by the historical global decision result, and output the global decision result of the target management task.
[0062] Alternatively, the historical global decision result can not carry a decision reason, and the large model can infer whether there is a conflict between the historical global decision result and the node decision results generated in the current round, and then obtain the current global decision result in combination with the global constraint condition. For example, a node decision result is a capacity reduction strategy, but the historical global decision result is a capacity expansion strategy issued five minutes ago, so the node decision result is probably unreasonable, and the probability of the large model inferring the node decision result as the global decision result is reduced.
[0063] It should be noted that the large model can capture the decision information of the historical global decision result, infer the encouragement of the global decision, utilize the powerful inference capability of the large model, refer to the historical global decision result, and generate more optimal global decision based on the node decision result generated in the current round, which helps to improve the quality of the decision.
[0064] In the embodiments of the present disclosure, each task processing node independently generates a node decision result, and the evaluation node performs global analysis on the node decision results generated by each task processing node based on global constraints to determine which global decision results are to be executed. The task processing node does not need to care whether it conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to care about the specific decision process of the node decision result, thereby reducing the complexity of development and maintenance, and enhancing the flexibility and scalability of the system. The global constraint refers to a constraint condition applied to the entire distributed system, which is used to determine which decisions are executed for the distributed system.
[0065] In an optional embodiment of the present embodiment, as shown in FIG. 1, the management system further includes an information sharing node 106.
[0066] The information sharing node 106 is configured to obtain and store shared information of each task processing node 102, wherein the shared information is information that needs to be transmitted between different task processing nodes 102.
[0067] The information sharing node is a sharing mechanism for remembering shared information that needs to be transmitted between different task processing nodes in a job instance, that is, JobContext. The shared information is context information that needs to be transmitted to other task processing nodes. The information sharing node can be a database maintained by the management system, and each task processing node is unaware.
[0068] Specifically, the information sharing node is a custom program, and the underlying persistent storage is equivalent to a storage unit customized by code, which is used to store the shared information of each task processing node and provide it for each task processing node to read. That is, the information sharing node is a storage unit configured with read / write interfaces to provide read / write services. The information sharing node can be actively accessed by the task processing node or accessed by the management system and then distributed to the task processing node.
[0069] For example, some task processing nodes are data collection nodes, and some task processing nodes are decision nodes. The decision node needs to obtain the data collected by the data collection node to make a decision. Since the data collection process is relatively complex, it is convenient to manage by splitting into multiple task processing nodes to collect data. The collected data is the shared information that needs to be transmitted to other task processing nodes.
[0070] In actual implementation, the task processing nodes included in the management system are at least one. If the task processing nodes are one, the one task processing node can make multiple decisions to obtain multiple node decision results, and then the global evaluation node performs global evaluation. If the task processing nodes are two or more, different task processing nodes are responsible for performing different decision steps in the target management task. The management system further includes an information sharing node. The information sharing node can obtain and store shared information that needs to be transmitted between different task processing nodes. The shared information can be data collected by the task processing nodes or node decision results generated by the task processing nodes.
[0071] It should be noted that, in response to the fact that any task processing node collects corresponding target decision data in response to the target management task, the collected target decision data can be written into the information sharing node based on the scheduler of the management system through the first access interface of the information sharing node, so as to be transmitted to other task processing nodes in the decision process of the target management task. In addition, after any task processing node generates a node decision result, in addition to providing the evaluation node for global perspective evaluation, the first access interface of the information sharing node can also be called based on the scheduler of the management system to write into the information sharing node, so that other task processing nodes can refer to the node decision result of the task processing node for decision making.
[0072] Further, since the information sharing node is used to transmit shared information in the decision process of the current target management task, after the decision of the current target management task is completed, that is, after the global decision result is determined by the evaluation node, the shared information stored in the information sharing node can be emptied, so that in the process of the next round of management task execution, the shared information that needs to be transmitted between different task processing nodes in the execution process of the round of management task can continue to be recorded.
[0073] In the embodiments of the present disclosure, one target management task can be executed by at least two task processing nodes. Different task processing nodes can transmit shared information in the decision process of the current target management task through the information sharing node to realize context information sharing between different task processing nodes, so that the task processing nodes are more intelligent and coherent when generating node decision results.
[0074] In an optional embodiment of the present embodiment, the information sharing node 106 is managed by the management system and provides a first access interface to the task processing node 102. The first access interface is a read-write interface of the information sharing node.
[0075] The information sharing node 106 is further configured to obtain shared information of each task processing node 102 based on the first access interface, store the shared information, and manage the life cycle of each stored shared information.
[0076] The task processing node 102 is further configured to acquire target decision data stored in the information sharing node 106 based on the first access interface, wherein the target decision data is shared information collected by other task processing nodes and stored in the information sharing node.
[0077] Specifically, the first access interface is a read-write interface, based on which the shared information of each task processing node can be written into the information sharing node, so that the information sharing node can store and manage the written shared information. In addition, the first access interface can also be used for the task processing node to access the information sharing node and acquire corresponding shared information.
[0078] In actual implementation, in a distributed computing or task scheduling framework, the information sharing node is used to store state information, configuration parameters or temporary data related to a specific task, and this mechanism allows necessary shared information to be passed between different stages (i.e., different task processing nodes). The information sharing node is similar to database storage, and the task processing node itself is not aware, and the management system injects the required shared information into the task processing node.
[0079] Specifically, the scheduler of the management system can be responsible for injecting the information sharing node into each task processing node, which means that when a task processing node starts to make a decision, it can already access the information sharing node to acquire corresponding shared information. This injection is usually completed in the process of initializing the task processing node.
[0080] The life cycle of the information sharing node is usually bound to the target management task, that is, the information sharing node is created when the target management task starts and is destroyed after the target management task is completed. In this process, all task processing nodes related to the target management task can access the information sharing node to acquire the shared information that needs to be passed. Since the information sharing node can be accessed by multiple parallel running task processing nodes at the same time, the management system must ensure the consistency and thread safety of the internal data of the information sharing node. In addition, according to actual needs, the information sharing node can be used to save various types of shared information, such as collected resource data of multiple dimensions of the distributed system, generated node decision results, etc.
[0081] As an example, the task processing node includes a data collection node and a decision node; the data collection node collects target decision data corresponding to the target management task in response to the target management task, and writes the target decision data into the information sharing node; the decision node acquires the target decision data stored in the information sharing node based on the first access interface in response to the target management task.
[0082] In the embodiments of the present disclosure, for the task processing node, the information sharing node is managed by the management system, the task processing node does not need to know the specific implementation details of the information sharing node, and does not need to care about how the shared data is collected and stored, and does not need to care about and manage the life cycle of the shared information in the information sharing node. The task processing node only needs to declare that the corresponding data is needed for the decision-making process in response to the target management task, access the information sharing node according to the first access interface provided by the management system, and obtain the corresponding shared information. That is, the information sharing node provides a kind of management and sharing of context information needed during execution of the target management task, and these information is transparent to the task processing node directly participating in the operation, which simplifies the communication complexity between different nodes in the management system, greatly reduces the complexity of the task processing node, and improves the maintainability and expansibility of the management system.
[0083] Of course, in actual implementation, in addition to the above-mentioned management of the information sharing node by the management system, a traditional storage scheme can also be used, in which each task processing node actively accesses the information sharing node to obtain the corresponding shared information. However, at this time, the task processing node also needs to consider the life cycle management of the corresponding shared information.
[0084] In an optional embodiment of the present embodiment, as shown in FIG. 1, the management system further includes a memory node 108, which is configured to store historical decision information of historical management tasks.
[0085] The task processing node 102 is further configured to obtain the historical decision information of the historical management tasks based on the memory node 108, and generate the node decision result of the target management task based on the self decision logic of the task processing node 102, the target decision data and the historical decision information.
[0086] It should be noted that the memory node is a cross-task decision information memory mechanism, that is, JobMemory. The memory node can store historical decision information of historical management tasks executed before the current time, which is used to assist the target management task executed at the current time to generate a decision result. That is, the memory node can remember the historical decision information of cross-task execution instances (JobInstance). The historical decision information can include historical resource data of the distributed system when the historical management task is executed, such as historical CPU data, historical memory data, historical water level data, etc. In addition, the historical decision information can also include the historical decision result of each task processing node when executing the historical management task, and the historical global evaluation result of the historical management task finally determined by the evaluation node, that is, the decision result of the historical management task can also be saved across tasks through the memory node.
[0087] The task processing node can also obtain historical decision information of a historical management task executed before the current time from the memory node, the historical decision information being related resource data and / or a decision result of a decision process of the historical management task.
[0088] Specifically, the memory node is a self-defined program and a persistent storage at the bottom, which is equivalent to defining a storage unit through code, and is used to store the historical decision information of the historical management task. That is, through code definition, each time a management task is executed, the related decision information can be written into the memory node, such as resource data, decision results of the management task executed by each task processing node, and the global evaluation result of the management task finally determined by the evaluation node. When each task processing node executes the current target management task, the memory node can be accessed to obtain the cross-task decision information stored therein. That is, the memory node is a storage unit configured with a read / write interface to provide read / write services. The memory node can be actively accessed by the task processing node or accessed by the management system and then distributed to the task processing node.
[0089] For example, in Java, the memory of jvm will change in a jagged manner with gc, but when the memory pressure is large, it will be almost a straight line. If only a momentary value is referred to, it is difficult to make a decision. If the historical situation is monitored by relying on an external monitoring system, the availability is difficult to guarantee. Therefore, in the embodiments of the present disclosure, a memory mechanism can be introduced to remember the historical decision information of the cross-task to help the task processing node to make decision auxiliary analysis on the target management task executed at the current time.
[0090] In the embodiments of the present disclosure, the memory node is introduced in the management system to store the historical decision information of the historical management task. When making a management decision, the task processing node can obtain the historical decision information of the historical management task, refer to the historical decision information of the cross-task, and generate the node decision result of the current target management task. Therefore, the generated node decision result can refer to the historical decision information, and the human intervention is not required to cope with the dynamically changing resource environment and complex decision scenarios of the distributed system, and more efficient and more accurate decisions can be made.
[0091] In an optional embodiment of the present embodiment, the memory node 108 is managed by the management system and provides a second access interface, which is a read / write interface.
[0092] The memory node 108 is configured to obtain decision information of the target management task in the decision process of the target management task by the task processing node 102 and / or the evaluation node 104 based on the second access interface, store the decision information corresponding to the target management task, and manage the life cycle of the stored decision information of each management task.
[0093] The task processing node is further configured to obtain historical decision information of historical management tasks stored in the memory node based on the second access interface.
[0094] The second access interface is a read-write interface that allows decision information to be written to the memory node, enabling the memory node to store and manage the written decision information. Furthermore, the second access interface can also be used by task processing nodes to access the memory node to read corresponding historical decision information.
[0095] It should be noted that memory nodes can remember historical decision information across task execution instances (i.e., across management tasks). The target management task is the management task currently being executed. Memory nodes can obtain and store the decision information corresponding to this target management task during its decision-making process. This decision information includes the current cluster resource status of the distributed system based on the task processing nodes' decisions, the node decision results generated by the task processing nodes, the resource information based on the evaluation nodes' global perspective evaluation, and the final determined global decision result, etc.
[0096] Furthermore, in practical implementation, memory nodes can manage the lifecycle of decision information for each management task they store, providing it to the corresponding task processing nodes in subsequent management tasks. Specifically, managing the lifecycle of decision information for each management task is a crucial process, encompassing the entire process from creation to final archiving or deletion. The lifecycle typically includes the following stages: creation, use, sharing, archiving, and destruction. Each stage requires appropriate management and strategies to ensure the security, integrity, and compliance of the decision information. The entire lifecycle of the decision information is managed by the management system.
[0097] In this embodiment, the memory data, water level data, node decision results, and global decision results corresponding to the target management task can be stored through memory nodes. The management system automatically maintains the lifecycle of each decision information stored in the memory nodes, eliminating the need for task processing nodes to concern themselves with data issues. For task processing nodes, the memory nodes are managed by the management system; they do not need to know the specific implementation details of the memory nodes, nor do they need to care about or manage the lifecycle of each decision information within them. This simplifies the communication complexity between different nodes in the management system, significantly reduces the complexity of task processing nodes, and improves the maintainability and scalability of the management system.
[0098] In practical implementations, within distributed computing or task scheduling frameworks, memory nodes are used to store historical decision information for management tasks executed up to the current time. This memory mechanism allows necessary decision information to be saved across different management tasks (i.e., different task execution instances). Memory nodes are similar to database storage; the task processing nodes themselves are unaware of this information, and the management system injects the required historical decision information into the task processing nodes.
[0099] Specifically, the management system is responsible for injecting memory nodes into each task processing node. This means that when a task processing node starts making a decision, it can already access the memory node to obtain the corresponding historical decision information. This injection is usually completed during the initialization of the task processing node.
[0100] In this embodiment, the memory node is managed by the management system. The task processing node does not need to know the specific implementation details of the memory node, nor does it need to care about how the historical decision information is stored and managed throughout its lifecycle. The task processing node only needs to declare that the decision process in response to the target management task requires corresponding historical decision information, and access the memory node through the second access interface provided by the management system to obtain the corresponding historical decision information. That is, the memory node provides a mechanism for managing and sharing the historical decision information of each management task in time sequence. This historical decision information is transparent to the task processing node that directly participates in the calculation, which simplifies the communication complexity between different nodes in the management system, greatly reduces the complexity of the task processing node, and improves the maintainability and scalability of the management system.
[0101] In one optional implementation of this embodiment, as shown in FIG1, the management system further includes a scheduler 110;
[0102] Scheduler 110 is configured to receive target management tasks from the distributed system, determine the current task processing node 102 of the target management task, and send the current task information of the target management task to the task processing node 102.
[0103] Task processing node 102 is configured to obtain target decision data based on current task information, generate node decision results for target management tasks based on its own decision logic and target decision data, and feed back target decision data and node decision results to scheduler 110.
[0104] The scheduler 110 is further configured to determine the next task processing node 102 to execute the target management task based on the target decision data and the node decision results, and to send the current task information of the target management task to the next task processing node 102.
[0105] It should be noted that the goal management task can consist of a series of steps, each step is responsible for completing a specific decision sub-task and analyzing decision data in the corresponding dimension. One step can be completed by one task processing node, that is, one task processing node can analyze and process decision data in one dimension.
[0106] In actual implementation, the management system also includes a scheduler, which is used to schedule various task processing nodes. Specifically, the scheduler can receive the target management task of the distributed system, determine the task processing node to execute the decision task, and send the corresponding task information to the task processing node, so that the task processing node can make a decision based on the task information and obtain the node decision result. The task processing node can feed back the decision data used and the output node decision result to the scheduler. The scheduler determines whether to let the next task processing node continue to make the decision, and the specific next task processing node, and continues to send the current task information of the target management task to the next task processing node to continue the node decision-making.
[0107] Specifically, task processing nodes can include starting task nodes. The scheduler can first schedule target management tasks to the starting task node. The starting task node can determine the node decision result based on the current resource status of the distributed system, its own decision logic, and other input data, and feed it back to the scheduler. The scheduler can determine the next decision subtask and pass it to the next task processing node for subsequent decision-making.
[0108] For example, Figure 2 is a schematic diagram of the task flow of a target management task provided in an embodiment of this disclosure. As shown in Figure 2, the task flow of the target management task includes task processing nodes 1 to 4. The scheduler receives the target management task and schedules it to task processing node 1. Task processing node 1 collects decision data, performs decision output, and feeds back the input decision data and the output decision to the scheduler. Based on the target management task and the data fed back by task processing node 1, the scheduler schedules the next decision sub-task to task processing node 2. Task processing node 2 collects decision data, performs decision output, and feeds back the input decision data and the output decision to the scheduler. Based on the target management task and the data fed back by task processing node 2, the scheduler schedules the next decision sub-task to task processing node 3. Task processing node 3 collects decision data, performs decision output, and feeds back the input decision data and the output decision to the scheduler. Based on the target management task and the data fed back by task processing node 3, the scheduler schedules the next decision sub-task to task processing node 4. In Figure 2, the dashed arrows between task processing nodes indicate that there is a logical connection between the task processing nodes, but the task processing nodes do not actually interact with each other. The task processing nodes interact with the scheduler.
[0109] In this embodiment of the disclosure, the management system also includes a scheduler, which schedules each task processing node to achieve a complete task flow. The task processing nodes do not need to interact directly with each other. The task processing nodes only need to interact with the scheduler and do not need to care about the processing logic of other task processing nodes. The task processing nodes are unaware of each other, which reduces the complexity of development and maintenance.
[0110] In one implementation, as shown in Figure 1, the management system includes a scheduler. The scheduler can inject shared information from information sharing nodes into task processing nodes. Specifically, the scheduler calls a first access interface to retrieve the shared information from the information sharing nodes and transmits it to the corresponding task processing nodes. Furthermore, the scheduler can also write shared information to the information sharing nodes. Additionally, the scheduler can inject historical decision information from memory nodes into task processing nodes. This is achieved by the scheduler calling a second access interface to retrieve historical decision information from memory nodes and transmit it to the corresponding task processing nodes. Finally, the scheduler can also write decision information to the memory nodes.
[0111] Of course, in actual implementation, the task processing node can also actively call the first access interface to obtain the shared information in the information sharing node, and the task processing node can actively call the second access interface to obtain the historical decision information in the memory node. This embodiment does not limit this.
[0112] In one optional implementation of this embodiment, the target management task includes the triggering of timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system;
[0113] Task processing node 102 is further configured to obtain cluster information of the distributed system; based on the task processing node's own decision logic, cluster information and historical decision information, it generates node decision results for the target management task, wherein the node decision results are used to indicate the management strategy of the distributed system.
[0114] Scheduled events refer to common requirements in many applications and systems, especially in scenarios where tasks need to be executed periodically. Scheduled events are triggered when a timer reaches its set duration. System monitoring events are events that are automatically triggered when specific conditions or anomalies are detected in a distributed system.
[0115] In practice, target management tasks can be triggered based on timed management events and / or system monitoring events of the distributed system. The task processing node collects and analyzes the current cluster information of the distributed system, such as CPU and memory. Based on the task processing node's own decision-making logic, cluster information, and historical decision information, the node decision result of the target management task is generated. The node decision result is used to indicate the management strategy of the distributed system. For example, if the CPU is less than a set threshold, the distributed system is scaled down.
[0116] Specifically, a distributed system can be configured with a timer that triggers a timed management event at set intervals, which can trigger the target management task for the current time; and / or, a distributed system can be configured with a monitoring policy that triggers a system monitoring event when the current cluster resources of the distributed system meet the monitoring policy, which can trigger the target management task for the current time.
[0117] Continuing with the previous example, as shown in Figure 2, target management tasks are triggered by timed management events and / or system monitoring events.
[0118] In this embodiment of the disclosure, a target management task can be triggered based on the timed management events and / or system monitoring events of the distributed system. This allows the task processing node to collect cluster information of the distributed system, combine its own decision-making logic, cluster information, and historical decision information to generate the node decision result of the target management task, and then determine how to manage the distributed system. The target management task of the distributed system can be triggered in various forms, making the management of the distributed system more flexible and allowing for different application scenarios.
[0119] In one optional implementation of this embodiment, the evaluation node 104 is at least one;
[0120] Evaluation node 104 is further configured to receive node decision results of the target project type; determine the node decision results to be executed from the node decision results of each node of the target project type based on the global constraints corresponding to the target project type; and use the node decision results to be executed as the global decision results under the target project type; wherein, the target project type is the project type corresponding to the evaluation node.
[0121] In one optional implementation, different types of evaluation nodes may be included. Each evaluation node is pre-configured with global constraints corresponding to its project type, which can be defined based on the project characteristics of that project type. After each task processing node outputs its node decision results, it can input the node decision results for a specific project type to the corresponding evaluation node based on the project type of those node decisions. This evaluation node then performs a global decision based on its own configured global constraints, determining which decision results to execute from the node decision results for that project type.
[0122] In another optional implementation, the decision results corresponding to global constraints can be pre-configured in the target task processing node. The target task processing node can be any task processing node. The decision results of each node are input to the evaluation node. The evaluation node determines whether there are conflicting decision results among the received decision results of each node. If not, it means that the decision results of each node can be executed in the distributed system, and the distributed system is managed. If there are conflicting decision results, the relevant decision data of the conflicting decision results can be obtained based on the global constraints. Based on the relevant decision data, it is determined which decision result among the conflicting decision results can be executed, or none of them can be executed, to obtain the target decision result. Then, the non-conflicting decision results from each node's decision results and the target decision result are used as the global decision result and executed in the distributed system to manage the distributed system.
[0123] For example, in automated cluster resource decision-making, it's necessary to determine whether scaling up or down is needed based on the current cluster usage. However, different metrics may lead to inconsistent and conflicting node decisions. For instance, task processing node A might decide to scale down based on CPU usage falling below a set threshold, while task processing node B might decide to scale up based on memory usage exceeding the threshold, and task processing node C might predict that scaling up is needed based on historical data. In this case, there's a conflict between the node decisions generated by task processing nodes A, B, and C. By leveraging global constraints, we can obtain historical and current CPU utilization, historical and current memory utilization, and the current and historical data of the distributed system. Based on pre-configured global decision logic (i.e., global constraints), we can determine whether the global decision is to scale up or down.
[0124] It should be noted that each task processing node independently generates its node decision results. The evaluation node then performs a global decision based on these node decisions for different project types, thereby obtaining the overall global decision result. The evaluation node does not need to concern itself with the specific decision-making process of the node decisions; it only needs to focus on the received node decisions. This reduces the complexity of development and maintenance, and enhances the system's flexibility and scalability.
[0125] As an example, Figure 3 is a schematic diagram of the execution process of a target management task provided in an embodiment of this disclosure. As shown in Figure 3, based on the timeline, the target management task to be executed at the current time is the target management task, and the management tasks executed before the current time are historical management tasks.
[0126] For target management tasks, the scheduler can pass the information to task processing node A in the management system. Task processing node A retrieves historical decision information from the memory node, and based on its own decision logic, the collected current resource information, and the historical decision information, generates node decision result A and node decision result B for the target management task. Furthermore, based on the data fed back by task processing node A, the scheduler schedules the next decision operation to task processing node B. Task processing node B collects the corresponding cluster resource information as the node decision result based on this decision operation and writes it to the information sharing node. Then, based on the data fed back by task processing node B, the scheduler schedules the next decision operation to task processing node C. Task processing node C retrieves the cluster resource information from the information sharing node, and based on its own decision logic, the cluster resource information, and the historical decision information, generates node decision result C for the target management task. Subsequently, based on the data fed back from task processing node C, the scheduler schedules the next decision operation to task processing node D. Task processing node D obtains shared information from the information sharing node and, based on its own decision logic, shared information, and historical decision information, generates the node decision result D for the target management task. The node decision results A, B, C, and D are then input to the evaluation node of the management system to obtain the global decision M.
[0127] In Figure 3, the dashed arrows between task processing nodes indicate that there is a logical connection between the task processing nodes, but the task processing nodes do not interact with each other. In the decision-making process of the above-mentioned task processing node AD for target management tasks, it only needs to interact with the scheduler. There is no awareness between the task processing nodes, and all information that needs to be transmitted between different task processing nodes can be written to the information sharing node.
[0128] In addition, the decision-making information related to the decision-making process of the target management task being executed at the current time can be written into the memory node as historical decision information for reference when executing management tasks at subsequent times.
[0129] In this embodiment, information sharing nodes enable information sharing between different task processing nodes, and memory nodes allow the retention of historical decision information across multiple tasks, thereby making the node decision results generated by task processing nodes more intelligent and consistent. Evaluation nodes assess the node decision results generated by multiple task processing nodes from a global perspective, handling global constraints between node decision results and avoiding misjudgments due to insufficient local information. Each task processing node is an atomic, independent module that only needs to focus on generating its own node decision results; complex global decisions and conflict handling are handled by the evaluation nodes, simplifying the development process of each task processing node, improving the system's scalability and modular development convenience, and reducing development complexity.
[0130] Figure 4 shows a flowchart of a management method for a distributed system according to an embodiment of the present disclosure, applied to a task processing node in a management system of the distributed system. The management system also includes an evaluation node. The method specifically includes the following steps.
[0131] Step 402: In response to the goal management task, obtain the corresponding goal decision data.
[0132] In one optional implementation of this embodiment, the target decision data is shared information collected by other task processing nodes and stored in the information sharing node; in response to the target management task, the corresponding target decision data is obtained, including:
[0133] The target decision data stored in the information sharing node is obtained based on the first access interface.
[0134] The information sharing node is managed by the management system and provides a first access interface to the task processing node. The first access interface is the read and write interface of the information sharing node.
[0135] Step 404: Based on the task processing node's own decision logic and target decision data, generate node decision results for the target management task. The node decision results are used to instruct the evaluation node to determine the global decision results for the target management task based on global constraints and the decision results of each node. The global decision results are used to manage the distributed system.
[0136] In an optional implementation of this embodiment, the management system further includes memory nodes, which are used to store historical decision information of historical management tasks; the method further includes:
[0137] Retrieve historical decision-making information for historical management tasks based on memory nodes;
[0138] Accordingly, based on the task processing node's own decision-making logic and target decision data, the node decision results for the target management task are generated, including:
[0139] Based on the task processing node's own decision-making logic, target decision data, and historical decision information, the node decision results of the target management task are generated.
[0140] In one optional implementation of this embodiment, the memory node is managed by the management system, which provides a second access interface, which is the read / write interface for the memory node; based on the memory node, historical decision information of historical management tasks is obtained, including:
[0141] The historical decision information of historical management tasks stored in the memory node is obtained based on the second access interface.
[0142] In one optional implementation of this embodiment, the target management task is triggered based on the timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system;
[0143] In response to goal management tasks, acquire corresponding goal decision data, including:
[0144] In response to target management tasks, obtain cluster information of the distributed system;
[0145] Based on the task processing node's own decision-making logic, target decision data, and historical decision information, the node decision results for the target management task are generated, including:
[0146] Based on the task processing node's own decision-making logic, cluster information, and historical decision-making information, node decision results for target management tasks are generated, whereby the node decision results are used to indicate the management strategy of the distributed system.
[0147] This disclosure provides a management method for a distributed system. A memory node is introduced into the management system to store historical decision information for past management tasks. When making management decisions, task processing nodes can obtain this historical decision information and, by referencing cross-task historical decision information, generate node decision results for the current target management task. This allows the generated node decision results to reference historical decision information, enabling more efficient and accurate decision-making in the face of dynamically changing resource environments and complex decision-making scenarios without human intervention. Furthermore, each task processing node independently makes decisions based on its own decision logic, generating node decision results. An evaluation node then evaluates the decision results of each node based on global constraints to obtain the global decision result for the target management task. Task processing nodes do not need to worry about conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to concern itself with the specific decision-making process of the node decision results, thereby reducing development and maintenance complexity and enhancing the system's flexibility and scalability.
[0148] The above is an illustrative scheme of a distributed system management method according to this embodiment. It should be noted that the technical solution of this distributed system management method belongs to the same concept as the technical solution of the distributed system management system described above. For details not described in detail in the technical solution of the distributed system management method, please refer to the description of the technical solution of the distributed system management system described above.
[0149] The following description, in conjunction with Figure 5, uses the application of the distributed system management method provided in this disclosure in a cluster resource management scenario as an example to further illustrate the distributed system management method. Figure 5 shows a flowchart of the processing procedure of a distributed system management method provided in an embodiment of this disclosure, specifically including the following steps.
[0150] Step 502: The scheduler of the management system responds to the timed management events and / or system monitoring events of the distributed system, triggers the corresponding target management task, and schedules the target management task to the first task processing node in the management system.
[0151] Step 504: The first task processing node responds to the target management task of the distributed system and obtains the cluster information of the distributed system; it obtains the historical decision information of the historical management tasks stored in the memory node based on the second access interface; based on the first task processing node's own decision logic, the collected cluster information and the historical decision information, it generates the first management strategy for the target management task; it writes the shared information that needs to be passed to other task nodes into the information sharing node based on the first access interface; it transmits the first management strategy to the evaluation node; and it feeds back the collected cluster information and the generated first management strategy to the scheduler.
[0152] Step 506: Based on the data fed back by the first task processing node, the scheduler schedules the next decision operation to the second task processing node, which is the node following the first task processing node.
[0153] Step 508: The second task processing node obtains the shared information stored in the information sharing node based on the first access interface; obtains the historical decision information of historical management tasks stored in the memory node based on the second access interface; generates a second management strategy for the target management task based on the second task processing node's own decision logic, shared information, and historical decision information; writes the shared information that needs to be passed to other task nodes into the information sharing node based on the first access interface; transmits the second management strategy to the evaluation node; and feeds back the collected cluster information and the generated first management strategy to the scheduler, which then schedules the next task processing node until the last task processing node.
[0154] Step 510: The evaluation node obtains the management strategies generated by each task processing node, analyzes each management strategy based on global constraints, and determines the global decision results for the target management task to be executed.
[0155] This disclosure provides a management method for a distributed system. A memory node is introduced into the management system to store historical decision information for past management tasks. When making management decisions, task processing nodes can obtain this historical decision information and, by referencing cross-task historical decision information, generate node decision results for the current target management task. This allows the generated node decision results to reference historical decision information, enabling more efficient and accurate decision-making in the face of dynamically changing resource environments and complex decision-making scenarios without human intervention. Furthermore, each task processing node independently makes decisions based on its own decision logic, generating node decision results. An evaluation node then evaluates the decision results of each node based on global constraints to obtain the global decision result for the target management task. Task processing nodes do not need to worry about conflicts with other task processing nodes, avoiding complex inter-node interaction logic. The evaluation node does not need to concern itself with the specific decision-making process of the node decision results, thereby reducing development and maintenance complexity and enhancing the system's flexibility and scalability.
[0156] Corresponding to the above method embodiments, this disclosure also provides an embodiment of a management device for a distributed system. Figure 6 shows a schematic diagram of the structure of a management device for a distributed system provided in one embodiment of this disclosure. The device is applied to a task processing node in a management system of the distributed system. The management system also includes an evaluation node. As shown in Figure 6, the device includes:
[0157] The acquisition module 602 is configured to acquire the corresponding target decision data in response to a target management task.
[0158] The generation module 604 is configured to generate node decision results for the target management task based on the task processing node's own decision logic and target decision data. The node decision results are used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
[0159] This disclosure provides a management device for a distributed system, which divides the complete decision-making process into two stages. Each task processing node makes decisions independently based on its own decision logic, generating node decision results. Then, an evaluation node evaluates the decision results of each node based on global constraints to obtain the global decision result of the target management task. The task processing nodes are unaware of each other, and the task processing nodes do not need to care about whether there is a conflict with other task processing nodes, thus avoiding complex inter-node interaction logic. The evaluation node does not need to care about the specific decision-making process of the node decision results, thereby reducing the complexity of development and maintenance and enhancing the flexibility and scalability of the system.
[0160] The above is an illustrative scheme of a distributed system management device according to this embodiment. It should be noted that the technical solution of this distributed system management device belongs to the same concept as the technical solutions of the distributed system management system and the distributed system management method described above. For details not described in detail in the technical solution of the distributed system management device, please refer to the description of the technical solutions of the distributed system management system and the distributed system management method described above.
[0161] Figure 7 shows a structural block diagram of a computing device according to an embodiment of the present disclosure. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0162] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0163] In one embodiment of this disclosure, the aforementioned components of the computing device 700, as well as other components not shown in FIG. 7, may be interconnected, for example, via a bus. It should be understood that the computing device block diagram shown in FIG. 7 is merely for illustrative purposes and is not intended to limit the scope of this disclosure. Those skilled in the art can add or replace other components as needed.
[0164] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0165] The processor 720 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the management method of the above-described distributed system.
[0166] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the management system and the management method of the distributed system described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the management system and the management method of the distributed system described above.
[0167] An embodiment of this disclosure also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the management method for the distributed system described above.
[0168] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the management system and the management method of the distributed system described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the management system and the management method of the distributed system described above.
[0169] An embodiment of this disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described distributed system management method.
[0170] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the management system and management method of the distributed system described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the management system and management method of the distributed system described above.
[0171] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0173] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.
[0174] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0175] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. A management system for a distributed system, comprising: Task processing nodes and evaluation nodes; The task processing node is configured to respond to the target management task of the distributed system and obtain the corresponding target decision data; Based on the task processing node's own decision-making logic and the target decision data, the node decision result of the target management task is generated; The evaluation node is configured to acquire the node decision results generated by each task processing node, and determine the global decision result of the target management task based on global constraints and the decision results of each node, wherein the global decision result is used to manage the distributed system.
2. The management system for the distributed system according to claim 1, wherein the management system further includes an information sharing node; The information sharing node is configured to acquire and store shared information from each task processing node, wherein, The shared information refers to information that needs to be transmitted between different task processing nodes.
3. The management system of the distributed system according to claim 2, wherein the information sharing node is managed by the management system and provides a first access interface to the task processing node, wherein the first access interface is the read / write interface of the information sharing node; The information sharing node is further configured to obtain shared information from each task processing node based on the first access interface, store the shared information, and manage the lifecycle of each stored shared information. The task processing node is further configured to obtain the target decision data stored in the information sharing node based on the first access interface, wherein, The target decision data is shared information collected by other task processing nodes and stored in the information sharing node.
4. The management system for the distributed system according to claim 1, wherein the management system further includes a memory node, the memory node being configured to store historical decision information of historical management tasks; The task processing node is further configured to acquire historical decision information of historical management tasks based on the memory node; and to generate node decision results of the target management task based on the task processing node's own decision logic, the target decision data, and the historical decision information.
5. The management system of the distributed system according to claim 4, wherein the memory node is managed by the management system and provides a second access interface, the second access interface being the read / write interface of the memory node; The memory node is configured to obtain decision information from the task processing node and / or the evaluation node during the decision-making process of the target management task based on the second access interface, and store the decision information corresponding to the target management task. The lifecycle of decision-making information for each management task in the management storage; The task processing node is further configured to obtain historical decision information of historical management tasks stored in the memory node based on the second access interface.
6. The management system for a distributed system according to claim 4 or 5, wherein the management system further includes a scheduler; The scheduler is configured to receive the target management task of the distributed system, determine the current task processing node of the target management task, and send the current task information of the target management task to the task processing node. The task processing node is configured to obtain the target decision data based on the current task information, and generate the node decision result of the target management task based on the task processing node's own decision logic and the target decision data. The target decision data and the node decision results are fed back to the scheduler; The scheduler is further configured to determine the next task processing node to execute the target management task based on the target decision data and the node decision results, and to send the current task information of the target management task to the next task processing node.
7. The management system for a distributed system according to claim 6, wherein the target management task is triggered based on the timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system; The task processing node is further configured to obtain cluster information of the distributed system; Based on the task processing node's own decision-making logic, the cluster information, and the historical decision information, a node decision result for the target management task is generated, wherein the node decision result is used to indicate the management strategy of the distributed system.
8. The management system for the distributed system according to any one of claims 1-7, wherein the evaluation node is at least one; The evaluation node is further configured to receive node decision results for the target project type; determine the node decision results to be executed from the node decision results of the target project type based on the global constraints corresponding to the target project type; and use the node decision results to be executed as the global decision results under the target project type; wherein, The target project type is the project type corresponding to the evaluation node.
9. A management method for a distributed system, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the method comprising: In response to goal management tasks, acquire the corresponding goal decision data; Based on the task processing node's own decision-making logic and the target decision data, a node decision result for the target management task is generated. The node decision result is used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
10. The method according to claim 9, characterized in that, The management system further includes a memory node, which is used to store historical decision information of historical management tasks; the method further includes: Based on the memory nodes, obtain the historical decision information of the historical management task; Accordingly, the step of generating the node decision result for the target management task based on the task processing node's own decision logic and the target decision data includes: Based on the task processing node's own decision-making logic, the target decision data, and the historical decision information, the node decision result of the target management task is generated.
11. The method according to claim 10, characterized in that, The memory node is managed by the management system and provides a second access interface, which is the read / write interface of the memory node; obtaining historical decision information of the historical management task based on the memory node includes: The historical decision information of the historical management task stored in the memory node is obtained based on the second access interface.
12. The method according to any one of claims 9 to 11, characterized in that, The target decision data is shared information collected by other task processing nodes and stored in the information sharing node; the step of obtaining the corresponding target decision data in response to the target management task includes: The target decision data stored in the information sharing node is obtained based on the first access interface; The information sharing node is managed by the management system and provides the first access interface to the task processing node. The first access interface is the read / write interface of the information sharing node.
13. The method according to any one of claims 9 to 12, characterized in that, The target management task is triggered based on the timed management events and / or system monitoring events of the distributed system, and the target decision data is the cluster information of the distributed system. The process of responding to a target management task and acquiring corresponding target decision data includes: In response to the target management task, obtain the cluster information of the distributed system; The generation of the node decision result includes: Based on the task processing node's own decision-making logic, the cluster information, and historical decision information, a node decision result for the target management task is generated, wherein the node decision result is used to indicate the management strategy of the distributed system.
14. The method according to claim 13, characterized in that, The method further includes: Shared information that needs to be passed to other task processing nodes is written to the information sharing node through the first access interface, so that subsequent task processing nodes can read and use it.
15. The method according to claim 13 or 14, characterized in that, The method further includes: The collected cluster information and the generated node decision results are fed back to the scheduler of the management system, so that the scheduler can schedule the next task processing node to perform subsequent decision operations based on the execution status of the current task processing node.
16. The method according to claim 15, characterized in that, Multiple task processing nodes are scheduled to execute decision operations in a preset order. The preceding task processing node writes shared information into the information sharing node, and the subsequent task processing node continues to generate node decision results based on the shared information and the historical decision information, forming a chain decision process until all relevant task processing nodes complete the decision.
17. A management device for a distributed system, applied to a task processing node in a management system of the distributed system, the management system further comprising an evaluation node; the device comprising: The acquisition module is configured to acquire the corresponding target decision data in response to target management tasks. The generation module is configured to generate node decision results for the target management task based on the task processing node's own decision logic and the target decision data. The node decision results are used to instruct the evaluation node to determine the global decision result for the target management task based on global constraints and the decision results of each node. The global decision result is used to manage the distributed system.
18. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the management method of the distributed system as described in claim 9.
19. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the management method for a distributed system according to any one of claims 9-16.
20. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the management method for a distributed system according to any one of claims 9-16.