Operation management method and system of server cluster

By building a scheduling adaptation factor model and a deep self-tuning engine for offline reinforcement learning, optimizing resource scheduling of server clusters, solving the problem of Cluster Autoscaler's preference in unconfigured areas, improving service quality and resource utilization efficiency, and being suitable for multi-availability zone cloud computing platforms.

CN120416256AActive Publication Date: 2025-08-01BEIJING HOLYSTONE TECH CO LTD

Patent Information

Application Number
CN202510919059.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-01
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

In the multi-AZ cluster of SaaS platforms deployed in AWS, Cluster Autoscaler may choose a resource-intensive or higher-priced AZ without the preference for zones, resulting in increased operating costs and may cause increased network latency and data consistency issues, affecting service stability and user experience.

Method used

Build the operating environment of the server cluster, identify the operating situation type of the functional module copy, build a scheduling adaptation factor model, and use a deep self-tuning engine to perform offline reinforcement learning training, generate the optimal action map, dynamically judge computing power expansion operations, and combine the expansion structure map and feedback evaluation function to update the strategy, optimize resource scheduling.

Benefits of technology

It realizes the high intelligence and dynamic adaptability of cluster resource scheduling, improves service quality and resource utilization efficiency, and avoids scheduling performance degradation caused by model aging or environmental mutation. It is suitable for large-scale, multi-tenant, and multi-availability zone cloud computing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416256A_ABST
    Figure CN120416256A_ABST
Patent Text Reader

Abstract

The invention discloses an operation management method and system for a server cluster, and belongs to the technical field of server operation management, and the method comprises the steps: constructing a computing power unit pool containing a plurality of service blocks, combining the resource state of a function module with application label information, recognizing the operation situation of the function module, and generating a scheduling adaptation factor model, link capacity, position distribution and communication delay among service blocks are further fused to construct an expansion structure map, a deep self-adjusting engine is adopted to carry out offline reinforcement learning training, and an optimal capacity expansion strategy under multiple situations is generated; the system can dynamically judge the capacity expansion behavior based on the current state, and performs strategy updating by continuously collecting capacity expansion feedback, so that intelligent, self-adaptive and high-efficiency control of resource scheduling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server operation management, and particularly relates to a method and system for operating and managing a server cluster. Background Art

[0002] A server cluster refers to a system in which multiple servers are connected through a network to work collaboratively and provide services as a whole. It improves the processing capacity, stability, and reliability of the system through mechanisms such as load balancing and failover, and is commonly used in scenarios that require high availability and high concurrent processing capabilities, such as large websites, cloud computing platforms, and database services.

[0003] The prior art has the following deficiencies: In a multi-availability zone cluster where the SaaS platform is deployed on AWS, when Cluster Autoscaler automatically scales out without configuring zone preferences, it may select an availability zone with resource constraints or higher prices, resulting in increased operating costs. At the same time, if the new node is located in a region far from the core database or major users, it may also cause an increase in network latency and affect the service response speed. More seriously, cross-availability zone data master-slave synchronization is restricted, which may cause service anomalies or even data consistency problems, seriously affecting system stability and user experience. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for operating and managing a server cluster to solve the deficiencies in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for operating and managing a server cluster, comprising: Constructing an operating environment for the server cluster, including a computing power unit pool distributed in multiple service blocks and function module replicas deployed on the computing power units; Based on the current resource usage status and application label information of the function module replicas, identifying the type of operating scenario they are in, and constructing a corresponding scheduling adaptation factor model; According to the scheduling adaptation factor model, defining the state input set and feedback evaluation function of the elastic control strategy; Modeling the link capacity, location distribution, and interaction delay between each service block as an extended structure graph, and introducing it as a feature parameter into the policy model; Using a deep self-tuning engine to perform offline reinforcement learning training on the policy model to generate an optimal action mapping for different operating scenarios and structure graphs; Based on the current state input of the cluster and the trained policy model, dynamically determining whether to perform a computing power expansion operation and the expansion target service block and unit quantity; Return the operation feedback information of the computing power expansion behavior and various monitoring metrics to the policy model, and update the policy based on the continuous tuning mechanism.

[0006] Preferably, the running environment for building the server cluster includes: For each computing power unit, measure the network round-trip time between it and the dependent nodes; calculate the weighted average of the communication delays of multiple dependent nodes, where the weights of each node are allocated according to its call frequency with the functional module; use the weighted average communication time as the communication overhead metric for the computing power unit; based on the communication overhead metric, preferentially select the computing power unit with low latency for deploying the functional module replicas.

[0007] Preferably, identifying the type of running scenario and constructing the corresponding scheduling adaptation factor model includes: Collect the resource usage status data of the target functional module replicas within a predetermined monitoring period, including the average CPU occupancy rate, memory utilization rate, request response time, and call frequency; Parse the application label information carried by the functional module, where the labels include structured identification fields such as the service type to which the module belongs, the business priority level, and whether it is a service with state retention; Based on the preset running scenario discrimination rules, use rule matching to map the current state of the module to multiple types of running scenarios, including high-concurrency processing type, compute-intensive type, low-latency interaction type, or resource-saving type; For the identified type of running scenario, construct the corresponding scheduling adaptation factor model, which is a set of parameter sets that affect the selection of scheduling policies, including region selection weight, start priority, latency tolerance threshold, and expected deployment period.

[0008] Preferably, extract service characteristic parameters from the scheduling adaptation factor model, including region elasticity weight, deployment priority, resource sensitivity threshold, and latency tolerance level; based on the service characteristic parameters and the current resource status of the cluster, construct a state input set including node idle rate, network latency, historical deployment behavior, and failure rate, and use it as the input of the elastic control policy model; define multi-objective feedback items at least including deployment success rate, startup time, resource utilization rate, and deployment cost, and assign weights to each feedback item to construct a weighted combined feedback evaluation function.

[0009] Preferably, modeling the link capacity, location distribution, and interaction latency between service blocks as an extended structure graph includes: Construct a directed graph structure of service block sets, where each node in the graph corresponds to a service block, and the edges represent data transmission paths between two blocks; collect the corresponding bandwidth capacity, average network latency, and cross-region distance factor for each edge, and map them to the connection weight value of the edge; perform graph normalization processing on the structure graph, and use the normalized structure graph as the input of the elastic control strategy model for the topological embedding feature to optimize the node selection and scheduling area matching process.

[0010] Preferably, the offline reinforcement learning training of the policy model using the deep self-tuning engine includes: Initialize a dual-structure model including a policy network and a value network. The policy network is used to generate the probability distribution of scheduling actions, and the value network is used to estimate the cumulative expected return of the current state. Construct an offline training sample set including a state input set, actual deployment actions, and a feedback evaluation function. Use the pre-collected system operation history records as training samples to perform behavior cloning warm-up training on the policy network to stabilize the initial output of the model. Based on the offline reinforcement learning algorithm, perform joint optimization of the policy-value network, minimize the action prediction error and the expected return gap, and generate a basic policy mapping function.

[0011] Preferably, the offline training process further includes: For each state input sample used for training, extract the corresponding service block structure graph fragment, including the embedding vectors of the target block node and its adjacent nodes. Use the graph neural network algorithm to perform message passing and aggregation processing on the graph fragment to generate a high-dimensional structure-aware embedding vector, and jointly construct a complete state input vector with the context embedding vector generated based on the running context type and the resource state vector generated by the resource usage status. During the policy training process, guide the model to learn the scheduling policy performance under the influence of the topological structure, so as to realize the optimal expansion behavior mapping for the structure graph.

[0012] Preferably, the dynamic determination of whether to perform a computing power expansion operation based on the current state input of the cluster and the trained policy model includes: Receive the resource state input set of the current cluster, including the remaining resources of each service block, network latency, service load indicators, and the running context features provided by the scheduling adaptation factor model. Input the input set into the trained policy model, trigger the policy model to perform an inference operation, and generate an output set including expansion recommendation actions. The output set includes a Boolean decision value for whether to expand capacity, an identification of the target service block recommended for expansion, and the number of computing power units recommended for expansion. Determine whether to execute the capacity expansion action according to the output result of the policy model. If the answer is no, maintain the existing deployment status.

[0013] Preferably, the policy update based on the continuous tuning mechanism includes: After each computing power expansion operation is executed, collect operation feedback information including deployment success rate, startup time consumption, resource occupancy stability, cluster load change, and service response time, and structurally associate and file it with the corresponding capacity expansion action. Based on the feedback information, calculate the comprehensive behavior score of the current capacity expansion action, and compare the difference with the expected return value of the policy model before decision-making. Set a fixed time window or policy behavior round as the model update period, collect the feedback data set and construct an incremental training sample set. On the premise of ensuring the stability of online inference, perform incremental parameter adjustment on the policy network, and select the updated model through the version comparison verification mechanism to complete the continuous policy tuning process.

[0014] The present invention also provides an operation management system for a server cluster, including: An environment construction module for constructing the operation environment of the server cluster, including a computing power unit pool configured in multiple service blocks, and function module replicas deployed on the computing power units. A situation recognition module for identifying the type of operation situation it is in based on the current resource usage status and application label information of the function module replicas, and constructing a corresponding scheduling adaptation factor model. A policy modeling module for defining the state input set and feedback evaluation function of the elastic control policy according to the scheduling adaptation factor model. A structure graph construction module for modeling the link capacity, location distribution, and interaction delay between service blocks as an extended structure graph, and inputting the graph as a feature parameter into the policy model. A policy training module, including a deep self-tuning engine, for performing offline reinforcement learning training on the policy model to generate an optimal scheduling action mapping under different operation situations and structure graph conditions. A policy inference module for dynamically determining whether to execute a computing power expansion operation based on the current cluster state input and the trained policy model, and determining the target service block for expansion and the number of computing power units required. A feedback update module for collecting operation feedback information of the capacity expansion behavior and various system monitoring indicators, and structurally feedbacking them to the policy model to perform incremental update on the policy model through the continuous tuning mechanism.

[0015] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. By introducing a scheduling adaptation factor modeling, expanding the structure graph construction, strengthening the reinforcement learning strategy training, and a continuous tuning mechanism driven by operation feedback, the present invention realizes a high degree of intelligence and dynamic adaptability of cluster resource scheduling. Compared with the traditional method that relies on static rules or single-dimensional index scheduling, the present invention can make more context-aware and performance-optimal scheduling decisions based on the actual running situation of functional modules and the network structure relationship between service blocks, greatly improving the service quality and resource utilization efficiency.

[0016] 2. The present invention constructs a closed-loop evolution path for the policy model, which can continuously optimize the decision-making logic through the actual feedback after deployment, avoiding the problem of scheduling performance degradation caused by model aging or running environment variation. The system has strong robustness and sustainable learning ability, and is applicable to large-scale, multi-tenant, multi-availability zone cloud computing platforms and high-elasticity business scenarios, significantly improving the cluster's expansion response speed, cross-region collaboration ability, and scheduling autonomy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0018] Figure 1 It is a flowchart of the method of the present invention.

[0019] Figure 2 It is a flowchart of the system module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0021] Example 1, please refer to Figure 1 As shown, a method for running and managing a server cluster in this embodiment includes: Constructing the running environment of the server cluster, including a computing power unit pool distributed in multiple service blocks and function module replicas deployed on the computing power units; Based on the current resource usage status and application label information of the function module replicas, identify the type of running scenario they are in, and construct a corresponding scheduling adaptation factor model; According to the scheduling adaptation factor model, define the state input set and feedback evaluation function of the elastic control strategy; Model the link capacity, location distribution, and interaction delay between each service block as an extended structure graph, and introduce it as a feature parameter into the policy model; Use a deep self-tuning engine to perform offline reinforcement learning training on the policy model to generate optimal action mappings for different running scenarios and structure graphs; Based on the current state input of the cluster and the trained policy model, dynamically determine whether to perform a computing power expansion operation and the number of target service blocks and units to be expanded; Return the operation feedback information and various monitoring indicators of the computing power expansion behavior to the policy model, and perform policy updates based on the continuous tuning mechanism.

[0022] When constructing the initial running environment of the server cluster, to prevent the function module replicas from being deployed in a few service blocks with rich resources, resulting in overloading of a single area and uneven resource scheduling in the later stage, this embodiment adopts a deployment and scheduling strategy based on the calculation of multiple resource weights. The specific process is as follows: Before deployment, the system first collects the current resource usage status in each service block, including indicators such as CPU idle ratio, memory usage rate, disk IO load, and network bandwidth occupancy rate; After standardizing the above indicators, calculate a resource availability score value for each service block; Sort each block according to the score value, and determine the initial distribution ratio of the function module replicas in a weighted manner; Set a deployment skew threshold (for example, not exceeding 30% of the total number of replicas deployed in any service block). When the number of replicas in any block exceeds the threshold, automatically switch the deployment target to the second-best ranked resource block.

[0023] Through the above method, this embodiment can achieve balanced control of resource heat, effectively prevent the rapid depletion of resources in a certain block, and improve the overall load balancing ability of the system.

[0024] In view of the characteristics that some function modules are highly communication-sensitive to dependent nodes such as data sources and configuration services in the initial startup stage, this embodiment provides a deployment optimization process with network overhead as a constraint. The specific steps are as follows: Before deployment, identify the key dependent nodes of the function module, including database services, cache services, authentication services, etc.; Collect the network latency values between each candidate computing power unit and the dependent nodes, obtained using a network probe or internal RTT (round-trip time) statistical method; Combine the communication latency values of multiple nodes into a comprehensive metric to measure the communication affinity of the computing power unit for the current functional module; Preferentially select the computing power unit with the smallest latency metric for deployment, forming a "low latency first" deployment strategy.

[0025] This deployment strategy can significantly reduce the data interaction time-consuming during the service cold start phase, improve the deployment response speed and availability, and is particularly suitable for distributed application scenarios with high requirements for initialization performance.

[0026] To strengthen the life cycle management and scheduling stability of functional modules in computing power units, the present invention proposes a container binding mechanism based on a hierarchical structure, including the following operation steps: Perform multi-layer containerization encapsulation on the functional module replicas. The typical structure includes: a business logic container layer, a scheduling proxy layer, and a health detection module; After each container is bound to the target computing power unit, establish a binding relationship among the "computing power unit - module replica - deployment label"; Each container registers a running state probe interface for the cluster scheduler to monitor its survival state and performance metrics in real time; In scenarios such as expansion, migration, or fault recovery, this binding mechanism ensures that the original deployment characteristics of functional modules can be traced back, avoiding incorrect scheduling and service interruption.

[0027] This mechanism improves the stability and scheduling controllability of functional modules after deployment and is an important means to ensure service continuity.

[0028] When deploying a large number of functional modules in a distributed environment, if there is a lack of a mechanism to describe deployment strategies and constraints, it is very easy to cause the loss of replica characteristics during subsequent expansion or fault migration. For this reason, this embodiment proposes a "deployment marking system", and the specific process is as follows: When each functional module replica is generated, create a set of deployment marking metadata, the content including: functional use code, resource preference weight (such as preferring high memory or low latency), dependency path reference (for example, depending on the database type or message middleware identifier); Bind the deployment marking as a replica life cycle attribute to the computing power unit and write it into the scheduling metadata warehouse at the same time; During subsequent system expansion or module rescheduling, reconstruct the initial deployment conditions according to this marking, so as to maintain the interpretability and consistency of system scheduling.

[0029] The deployment marking mechanism can be regarded as the "identity fingerprint of functional modules", which provides structural memory capabilities during the operation of complex clusters and has great engineering practical value.

[0030] To optimize the communication efficiency of module deployment, a communication overhead calculation mechanism with dependency awareness and weighted average characteristics is proposed. The detailed steps are as follows: For each functional module replica, the system identifies several basic service nodes it depends on according to the deployment description or operation record; For each candidate computing power unit, the detection mechanism is used to measure the data transmission delay between it and each of the above dependent nodes respectively, and the unit can be in milliseconds; For each dependent node, a communication weight value is assigned according to its call frequency during the module execution process; Using the weighted average method, multiply the measured communication delay values by the corresponding weights and sum them to finally obtain the average communication overhead index of the computing power unit; Use this index as one of the reference criteria for evaluating whether the unit is suitable for deploying the target module and participate in the final scheduling and sorting.

[0031] Through this communication metric model, the deployer not only considers the resource quantity and load, but also incorporates "communication quality" into the evaluation system, effectively improving the collaborative efficiency of distributed services in the topological structure.

[0032] Collect the resource usage status data of functional module replicas. First, during the operation of the cluster, perform periodic performance monitoring on the functional module replicas deployed on each computing power unit, and collect their resource status data within a specified time window (such as the last 30 seconds or 5 minutes). The collected metrics include but are not limited to the following categories: Average CPU occupancy rate; peak memory occupancy rate; network bandwidth consumption rate; average request response latency; number of calls or QPS (request processing volume) per unit time; IO waiting time (used to identify disk or storage-intensive modules); These metrics are obtained through container-level monitoring components (such as Prometheus, cAdvisor, or kubelet), normalized by the time series processing component, and stored in the context recognition module for subsequent use.

[0033] Secondly, the system reads the application label information registered by each functional module replica during the deployment phase. These labels can be defined by developers in the deployment configuration or automatically generated by the platform according to the service meta-information.

[0034] Typical labels include but are not limited to: serviceType: identifies the business type of the module (e.g., api-gateway, data-ingest, cache-agent); priorityLevel: specifies the service priority; stateful: Indicates whether the module is a state-retaining service; scalingPattern: indicates whether fast scaling is supported; latencySensitivity: latency sensitivity level; These labels and monitoring indicators together constitute the context information set for judging the operating situation.

[0035] After obtaining resource indicators and label information, the system calls and runs the context recognition engine to classify the current functional module copies. The classification results are used for the differentiated generation of subsequent scheduling strategies.

[0036] The identification of situation types can be based on one or a combination of the following two strategies: Classification strategy based on rule matching: The system pre-defines multiple judgment rules, each of which consists of multiple indicators and label conditions. For example, if the CPU usage exceeds 80%, the request QPS exceeds 1000, and it is marked as a stateless module, it will be classified as "high concurrency processing type"; If latency sensitivity is high and response time is less than 30ms → classified as "Low Latency Interactive"; If the memory usage is always below 30% and the priority is marked as low → classify as "resource-saving"; If the state retention is true and the call frequency is very low → classified as "task type / cold start type"; Lightweight model-based classification strategy: If the system deployment has machine learning capabilities, a lightweight classification model (such as a decision tree, logistic regression, or a small neural network) can be used to input the above indicators and train a classifier to automatically determine the context type. The input feature vector includes CPU, latency, bandwidth, and label embedding codes, and the output is a discrete context label.

[0037] The results of a situation type are generally one of the following: High concurrent processing type: suitable for stateless services, high QPS, and easy scalability; Computation-intensive: heavy CPU load and high response tolerance; Low-latency interactive: requires extremely low response time and needs to be deployed close to the user; Resource-saving: It has stable load and is suitable for deployment in edge areas or reserved resource pools.

[0038] After obtaining the context type, the system constructs a set of scheduling adaptation factors, which are used to be input into subsequent scheduling control policy models (such as reinforcement learning decision networks, deployment template selectors, etc.) to affect key behaviors such as the expansion area, the number of replicas, and the deployment order.

[0039] The scheduling adaptation factors at least include the following parameter fields: Regional elasticity weight: indicating the sensitivity of the replicas of this module to the available zone location (e.g., high concurrency type → high elasticity; state type → low elasticity); Startup priority coefficient: the priority of this module during concurrent startup (which can be bound to the business priority); Delay tolerance threshold: the maximum acceptable average communication delay, used to filter deployment blocks; Expected deployment period: the longest tolerated time from scheduling to service readiness, used for resource cold start evaluation; Replica scalability label: whether it supports dynamic changes in the number of replicas, affecting the HPA policy generation logic.

[0040] This set of factors is bound to the scheduling metadata of the function module replicas in the form of structured parameters and serves as an input reference during each policy model decision or service rescheduling.

[0041] The "state input set" of the policy model is a set of vectors describing the current cluster running state, service deployment requirements, and resource supply status. Different from the fixed input template, the present invention generates a dynamic input structure based on the scheduling adaptation factor model, with high business sensitivity and resource adaptation capabilities.

[0042] Its construction process includes the following steps: Parameter extraction stage: Extract the key parameters of the current function module from the scheduling adaptation factor model, including: Regional elasticity weight (indicating whether cross-zone deployment is allowed); deployment priority (affecting the degree of competition with other tasks); resource sensitivity threshold (such as the maximum acceptable deployment delay); Resource status collection stage: Real-time collect the resource usage status of each service block in the cluster, including but not limited to: The current CPU idle rate and remaining memory of each node; Network throughput rate and IO load; Node availability index (based on outage probability and response delay statistics); Feature encoding stage: Combine the above parameters to construct a state vector, for example: The CPU idle rate in Zone A is 45% and the delay is 20ms; The function module requires low-latency deployment; Then, the weight of the state vector in Region A is decreased. If Region B has better metrics, it will be given priority.

[0043] Preprocessing stage: Normalize the state vector to ensure that parameters in different dimensions have a consistent numerical range in the policy model, facilitating the convergence and generalization of the policy model.

[0044] Through the above process, the system provides a set of dynamic input sets that can perceive the operating environment for the policy model, supporting flexible decision-making.

[0045] Another key element in the training of the policy model is the feedback evaluation function, which is used to guide the model to optimize its behavior. To adapt to complex business requirements, the present invention introduces a multi-objective fusion feedback mechanism to avoid the problem that a single performance objective dominates the scheduling result in traditional systems.

[0046] The implementation process is as follows: Objective definition: Extract the current performance requirement weights of the task from the scheduling adaptation factor model, such as the preference for "low latency, high stability"; Construction of the objective item scoring function: Deployment success rate scoring: By counting the success rate in past deployment attempts and accumulating it within a certain time window, a stability score is formed; Cold start latency scoring: By recording the time required for the module to report the first healthy signal after deployment and calculating its average value; Resource utilization scoring: Calculate the difference between the resources consumed per replica after actual deployment and the ideal resource ratio; Deployment cost scoring: Estimate the resource cost of each replica based on the price of unit resources in the target area; Weighted fusion: According to the attention of the task to each indicator, set the weighting coefficients of each objective item, and perform weighted average or weighted sum to form a comprehensive score.

[0047] For example, the current feedback function value of a certain functional module can be expressed in the following way (converted to text expression): Feedback value = (Deployment success rate score × 0.3) + (Startup delay score × 0.2) + (Resource utilization score × 0.3) + (Cost score × 0.2); The above scoring results will be used as the reward signal in the reinforcement learning model or as the reference result of policy inference for the next round of optimization.

[0048] To enhance the sensitivity of the policy model to long-term system evolution and service stability, the present invention further introduces the "historical behavior trajectory" feature into the state input set. This mechanism can more accurately reflect the scheduling risk, resource adaptability, and stability requirements of a specific service.

[0049] The specific steps are as follows: Behavior collection: The recording module records the relevant behavior data during the past several scheduling processes, including: Average deployment interval time; Scheduling failure rate (rejected due to insufficient resources, for example); Number of migrations and average migration success time; Number of resource rescheduling times and contention rate (i.e., the frequency of being interrupted in a high-concurrency scenario); Feature modeling: Statistically analyze the above historical behaviors, and output labels such as stability index, failure penalty term, and scheduling frequency level; Fusion processing: Input these types of features together with the current resource metrics and factor model parameters into the policy model to enable it to have "memory-based" reasoning ability.

[0050] This mechanism enables the model to perceive which services are sensitive to changes in the deployment environment, thereby improving scheduling robustness.

[0051] In traditional cluster systems, scheduling evaluation mainly relies on system-level metrics (such as resource allocation efficiency), while the present invention realizes the closed-loop control of the policy model by introducing business-level feedback signals, enabling it to have the ability to learn and respond to "service quality" itself.

[0052] The implementation method is as follows: Business metric collection: After the functional module is deployed, the system receives its business monitoring metrics, such as: Request response time jitter (Jitter); Error request rate (such as HTTP 5xx); Business completion rate (such as payment success rate, processing completion rate); Causal analysis mechanism: Determine whether business fluctuations are affected by the resource scheduling strategy through regression analysis or causal relationship modeling (for example, requests time out due to being deployed in an area with poor network quality); Feedback quantization: Convert the business fluctuations with relatively high correlation into positive rewards or negative penalty terms and include them in the policy training rewards; Feedback closed-loop: The above business feedback signals participate in the iterative update of the model, gradually improving the adaptability of the system policy to business objectives.

[0053] The main goal of constructing the structure graph is to express the resource reachability and communication cost between service blocks. For this purpose, the graph structure modeling method is adopted in this embodiment, which specifically includes the following steps: Link parameter collection: The system periodically collects the communication link metrics between service blocks through the embedded network monitoring module or service mesh probe component, mainly including: Link bandwidth upper limit: Represents the theoretical maximum transmission capacity of the network interface between two blocks; Real-time utilization rate: It refers to the proportion of current bandwidth usage, reflecting the network congestion status; Packet loss rate and retransmission times: Used to evaluate the link stability.

[0054] Block location modeling: Obtain the logical topology number (such as Kubernetes zone ID or subnet identifier) of each service block, or its geographical coordinate information (such as the region label provided by the cloud platform). This information is used to express the physical or topological "proximity".

[0055] Communication delay measurement: By periodically sending heartbeat packets or test requests, measure the round-trip time (RTT) between each service block, and calculate the average value within multiple sampling periods. This delay metric serves as the key weight basis for the edges in the graph.

[0056] Graph structure construction: Model each service block as a node in the graph, and connect an edge between any two service blocks with communication capabilities. The initial weight of each edge is a combined metric representing the communication cost from block A to block B. Its calculation method is as follows: Combine communication delay, packet loss rate, and bandwidth congestion rate with certain weights, and obtain the communication cost through weighted average.

[0057] For example: Communication cost = (delay × 0.5) + (packet loss rate × 0.3) + (bandwidth congestion rate × 0.2). The result assigns a communication cost score to each edge.

[0058] Finally, a structure graph with weighted edges and node attributes is formed, which is used for the policy model to perceive the network relationship and deployment preference between service blocks.

[0059] To further improve the expression ability of the graph, in this implementation, structured attributes are added to the nodes and edges of the graph respectively, and the communication metrics are fused and optimized as follows: Node attribute vector design: Attach an attribute vector to each service block node, and this vector includes: Resource richness (such as the reverse index of average CPU utilization); Historical average response time; Deployment success rate and fault recovery time in the last 30 minutes; Network stability coefficient.

[0060] The above attributes are stored in a numerical way, which is used to assist the policy model in evaluating the schedulability of service blocks.

[0061] Edge Weight Fusion Function Design: To ensure the global comparability of the communication cost weights of edges, the original bandwidth, latency, and stability metrics are standardized (such as Z-score or Min-Max normalization). After unifying the dimensions, they are weighted and aggregated according to the priorities defined by the policy model. The aggregation method can be expressed based on the following formula: Edge weight = α × Normalized Latency + β × Normalized Bandwidth + γ × Normalized Packet Loss Rate. Here, α, β, and γ are the weight coefficients given by the policy model or business configuration.

[0062] Graph Structure Update Mechanism: To ensure the adaptability of the model to the dynamic changes of the network topology, this system sets a graph update period (for example, update once every 10 minutes) and performs immediate reconstruction when the cluster topology changes (such as a new node joining or a block failure).

[0063] The above attribute enhancement mechanism ensures that the graph data received by the policy model not only has structural information but also integrates performance and reliability elements, constituting a complete input for the deployment environment.

[0064] To enable the policy model to directly utilize the above graph information to participate in scheduling decisions, the present invention adopts graph embedding technology (Graph Embedding) to convert the graph structure into a vector representation that can be input into the policy model. The main steps are as follows: Application of Graph Embedding Technology: Use graph neural network (GNN) or graph convolutional neural network (GCN) methods to encode each node in the graph. This process considers the attributes of the node itself and the feature information of adjacent nodes to form a high-dimensional vector representation (for example, a vector with a length of 64 or 128 dimensions), representing the "scheduling features" of the service block.

[0065] Input Fusion of Policy Model: Jointly input the embedding vector of each service block and the scheduling adaptation factor of the functional module into the elastic control policy model, enabling the model to consider the requirements of the functional module and the network characteristics of the service block simultaneously when evaluating the deployment area.

[0066] Policy Inference and Deployment Recommendation: During the policy inference process, the model sorts the target blocks according to the graph embedding values and preferentially recommends deployment areas with low communication costs and rich resources, thereby optimizing the expansion path and response time.

[0067] Reverse Update Mechanism of Graph Embedding: During the training of the policy model or the process of reinforcement learning, if the feedback effect is poor after a certain deployment (such as a sharp increase in response latency after deployment), the edge weights or node embeddings in the graph are adjusted through the backpropagation mechanism to achieve the co-optimization of the structure diagram and policy behavior.

[0068] In the resource scheduling scenario of a server cluster, the scheduling behavior decision-making can be regarded as a Markov decision process (MDP). Each expansion or deployment operation is an "action", the current state of the system is the "state input", the scheduling result is the "environmental feedback", and the goal of the scheduling system is to learn to select the optimal action in different states to maximize the long-term benefit.

[0069] To this end, the present invention adopts a policy-value dual-network structure as the core framework of the policy model, including: Policy Network: It is used to generate scheduling actions or action probability distributions, such as selecting whether to expand, selecting a deployment area, setting the number of replicas, etc.

[0070] Value Network: It is used to estimate the long-term return expectation in the current state, that is, the estimation of the pros and cons of the future overall system performance after executing a certain action in this state.

[0071] This dual-network structure can preferably be optimized using mainstream reinforcement learning algorithms such as DDPG (Deep Deterministic Policy Gradient) or PPO (Proximal Policy Optimization).

[0072] In terms of constructing training data, an off-policy reinforcement learning training mechanism is adopted, that is, it does not rely on real-time execution of operations in the online environment, but uses the system's historical deployment records to construct training samples, including: State input vectors (such as resource status, historical behaviors, structural topologies); Executed deployment actions; The system feedback brought by this action (such as startup time, cost, service success rate); Optional action space (that is, the set of deployable strategies in this state).

[0073] At the initial stage of training, behavior cloning is used for warm-up training, that is, the policy network first learns to imitate the scheduling actions with better effects in history to help the model stably output; then it switches to the reinforcement learning process to optimize the policy model through the reward function.

[0074] The reward function can be: Reward value = (deployment success rate × weight 1) + (resource utilization rate × weight 2) - (startup delay × weight 3) - (deployment cost × weight 4); where the weight values are dynamically set according to business goals.

[0075] Since different functional modules have significantly different resource preferences and performance requirements during operation (such as high-concurrency services requiring fast startup, low-latency services focusing on communication quality, etc.), the present invention introduces an operating context embedding mechanism to enable the policy model to have "context awareness ability".

[0076] The implementation process is as follows: Scenario classification and label generation: Before deployment, determine the type of operating scenario to which the functional module belongs according to the scheduling adaptation factor model and application labels. For example: concurrent processing type; state retention type; real-time interaction type; energy-saving steady state type.

[0077] Scenario label embedding processing: Encode each scenario label, convert it into a discrete or continuous embedding vector, and use it as part of the state input vector.

[0078] Training sample annotation: When constructing training samples, add corresponding scenario embedding fields to each sample state, so that the model can learn "how to make decisions in a certain type of scenario" during the training process.

[0079] Policy diversity modeling: During the training process, the model adjusts the output policy based on the scenario embedding. For example: in the "low-latency" scenario, the policy is more inclined to select a service block with a faster network; in the "energy-saving" scenario, it is more inclined to select a node with low resource occupancy.

[0080] This mechanism greatly enhances the behavior diversity and scheduling adaptability of the model.

[0081] There are complex topological structure relationships between service blocks in the server cluster, such as cross-region communication costs and node-to-node latency differences. If the policy model cannot perceive these structures, it will affect the deployment efficiency and service quality. Therefore, the present invention further introduces the structure graph embedding feature into the policy model training, specifically as follows: Graph fragment extraction: Each training sample is associated with a service block structure graph fragment, which represents the network structure and attribute relationship between the target deployment area and its adjacent nodes, including: Node resource attributes (such as resource abundance); Edge weights (such as bandwidth, latency); Historical scheduling performance (such as stability score).

[0082] Graph neural network encoding: Use GCN (Graph Convolutional Network) or GAT (Graph Attention Network) to process the graph fragment, perform node information aggregation and feature propagation, and output the structure embedding vector.

[0083] State input fusion: The state input vector consists of three parts: the structure embedding vector: generated by the graph neural network extracting features from the service block structure graph; the scenario embedding vector: generated by the scenario classification model according to the module operation label; the resource state vector: generated according to the state parameters such as CPU, memory, and latency collected by monitoring; the combination of the three is used to describe the environment where the current module is deployed, and is used as the input feature set of the policy model and input into the policy network and value network.

[0084] Policy Optimization and Graph Feedback Mechanism: The model automatically identifies which topology combinations can bring better scheduling effects through policy training. In reinforcement learning, if the deployment corresponding to a certain structure leads to negative feedback (such as high latency or failure), the graph embedding weights are adjusted through backpropagation to enhance the model's adaptability to "spatial structures".

[0085] Based on the input set of the current state of the cluster, generate suggestions for scaling-up actions through the trained policy model.

[0086] The input set of states includes but is not limited to the following: The overall load situation of the cluster (such as the total CPU occupancy rate, memory consumption rate); The current resource utilization rate and resource remaining rate of each service block; Network communication latency (such as the round-trip time between service blocks); The scheduling adaptation factors of functional modules (such as the tolerance to latency, elasticity weight, operating context label); The current scalable capacity of the system (such as whether each service block has the ability to add new replicas); These data are processed through unified normalization and used as input vectors to be fed into the policy model for inference.

[0087] After the policy model makes inferences, the output usually includes the following fields: Scaling-up decision value: A boolean value indicating whether to perform scaling-up (True / False); Target service block identifier: The target block recommended for scaling-up (such as zone-A, zone-B); Recommended scaling-up scale: The number of computing power units to be increased, such as adding 2 Pod replicas; Optional: Suggestions for the scaling-up timing (execute immediately / observe with delay), etc.

[0088] The above output can be regarded as the preliminary judgment result of the model on "whether to scale up" and "how to scale up".

[0089] To prevent the policy model from outputting low-confidence suggestions in a state of high uncertainty, the present invention introduces a confidence score mechanism as a safety boundary control method for the execution of scaling-up actions.

[0090] Each time the policy model outputs a scaling-up action, it generates a confidence score for this action, and the value range is usually 0~1, indicating its credibility in the current state. The confidence can be calculated based on the following factors: The sharpness of the action probability distribution (the maximum value in the softmax output); The average feedback of the same behavior in historical similar states; Uncertainty indicators (such as entropy or standard deviation) within the policy model.

[0091] The system dynamically adjusts the confidence threshold according to the current business level and the risk level of cluster resources. For example: Under normal operating conditions, the standard threshold is set to 0.7; If the system resources are tight or during the critical business operation period, the threshold is increased to 0.85; If in a test environment, it can be reduced to 0.5 to allow the model to play freely.

[0092] If the confidence of the scaling action output by the model is lower than the current threshold, the system enters the delayed observation mode, that is, the scaling action is not executed temporarily and will be re-evaluated in the next cycle; If the confidence is higher than the threshold, the recommended scaling plan is immediately executed, including the target block selection and replica deployment operations.

[0093] Through this mechanism, the risk of incorrect scaling caused by model errors or unstable outputs is effectively reduced, and the system robustness is improved.

[0094] Even if the policy model makes a high-confidence scaling decision, it may deviate due to factors such as changes in business scenarios and inconsistent historical behaviors. Therefore, the present invention introduces a context backtracking mechanism to perform multi-dimensional verification on the policy output results.

[0095] The system calls the scaling behavior records of the function module in the past several scheduling cycles and extracts the following information: The target block and scaling scale of the last scaling; The time taken for the service to start after scaling; The service performance metrics after deployment (such as average response time); Whether there is a record of scaling failure or deployment interruption.

[0096] The system compares the current model-recommended action with the historical successful scaling behavior in terms of vector similarity. Common methods include cosine similarity or Euclidean distance calculation, and a matching score (such as 0-1) is output.

[0097] If the matching degree between the current recommended action and the historical efficient behavior is lower than the threshold (such as 0.6), and the model confidence is in the "vicinity of the critical value" (such as 0.65-0.75), the policy fallback mechanism is triggered; The fallback mechanism may include: Enabling the expert policy (rule system first); Using the parameters of the last successful scaling action; Postponing the scaling plan and waiting for further observation; If the matching degree is high and the confidence is high, the action is directly executed, and the system will automatically record the execution effect for the next training sample iteration.

[0098] This mechanism forms a behavioral closed-loop control path of "policy model → decision recommendation → historical behavior backtracking → action confirmation", ensuring that the model decisions are not only scientific but also traceable and business-consistent.

[0099] When the policy model in the cluster outputs a scaling decision and completes the deployment of computing units, the system will start the feedback collection sub-module to structurally observe and record the scaling results, including the following: Collection of deployment-level metrics: Record whether the computing units for this scaling are successfully scheduled, the cold start time (from scheduling trigger to the first heartbeat signal), the startup resource ratio (the ratio of initial CPU and memory allocation to actual occupancy), whether the replica health status continuously meets the standards, etc.

[0100] Collection of system-level metrics: Record the change trend of the overall system performance after the scaling is completed, including: Changes in the total CPU and memory utilization of the cluster; Changes in the service response time (P95 / P99); Fluctuations in network IO bandwidth and latency; Whether the failure rate has improved, etc.

[0101] Structured archiving method: Encode the above data into a structured JSON or Protobuf format, and the fields include: Scaling behavior identifier (including service module ID, timestamp, target block ID, number of scaled replicas, etc.); The collected performance metrics; Summary of the current running scenario and scheduling policy.

[0102] This feedback information will be written into the "policy learning data buffer" for batch loading during model tuning.

[0103] To achieve targeted learning and optimization of the policy model, the present invention constructs a feedback evaluation function mechanism to quantify the execution effect of each scheduling behavior and compare it with the original policy's expected effect to guide policy adjustment.

[0104] Definition of the evaluation function: Preset a set of evaluation index items and weighted combination functions, usually including: Deployment success rate score; Startup time score (the shorter the time, the higher the score); Resource utilization rate of a single replica (stability indicator); Improvement value of the system performance after deployment (such as the reduction amplitude of the response time); Deployment cost control score (such as whether it enters a high-cost block).

[0105] Behavior score output: After each system expansion, a normalized score (e.g., in the range 0–1, where a higher score represents better behavior) is calculated using the feedback metrics.

[0106] Difference calculation and loss function update: The policy model has predicted the expected return in this state before execution; The system calculates the difference between the actual score and the expected return; This difference is used as part of the policy loss function to participate in the backpropagation to update the policy network parameters.

[0107] This "expected - actual" difference - driven learning mechanism can guide the model to correct misestimations and improve the accuracy of behavior selection.

[0108] To achieve the continuous evolution ability of the model rather than one - time training, the present invention introduces a continuous tuning mechanism based on time windows and policy rounds to achieve incremental learning and version verification of the policy model. The process is as follows: Tuning trigger mechanism: Set a tuning window, such as "every 10 minutes" or "every 20 expansion behaviors are executed"; After the end of each window, collect all expansion records and scoring information to form a batch training sample set.

[0109] Incremental model update: Do not retrain the entire policy model to avoid interruption of online services; Instead, only update the "shallow weight group" or "sub - policy branch" in the policy network to keep the overall model structure stable; Use mini - batch gradient descent or a low - rate optimizer (such as RMSProp) to complete the update.

[0110] Model version control and verification mechanism: Generate a new model version after the update and save it to the "candidate policy pool"; In subsequent expansion behaviors, use the A / B test method to deploy the new and old models on different service modules simultaneously; Compare their differences in response time, deployment success rate, and resource consumption, and select the winner to update as the main model.

[0111] This mechanism effectively avoids the risk of performance degradation during the policy update process and allows the model to automatically learn and evolve with the operating environment.

[0112] Example 2, please refer to Figure 2As shown in the figure, the operation management system of a server cluster described in this embodiment includes a data acquisition module, a sample scoring module, a generation model training module, a sample enhancement and management module, and a distributed recognition module; An environment construction module, which is used to construct the operation environment of the server cluster, including a computing power unit pool configured in multiple service blocks, and function module replicas deployed on the computing power units; A situation recognition module, which is used to identify the type of running situation it is in based on the current resource usage status and application label information of the function module replicas, and construct a corresponding scheduling adaptation factor model; A policy modeling module, which is used to define the state input set and feedback evaluation function of the elastic control policy according to the scheduling adaptation factor model; A structure graph construction module, which is used to model the link capacity, location distribution, and interaction delay between service blocks as an extended structure graph, and input this graph as a feature parameter into the policy model; A policy training module, including a deep self-tuning engine, which is used to perform offline reinforcement learning training on the policy model to generate an optimal scheduling action mapping under different running situations and structure graph conditions; A policy inference module, which is used to dynamically determine whether to perform a computing power expansion operation based on the current cluster state input and the trained policy model, and determine the target service block for expansion and the number of computing power units required; A feedback update module, which is used to collect the operation feedback information of the expansion behavior and various system monitoring indicators, and structurally feedback them to the policy model, and perform incremental update on the policy model through a continuous tuning mechanism.

[0113] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0114] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. Specifically, it can be understood by referring to the context before and after.

[0115] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0116] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for running and managing a server cluster, characterized in that: Including: Construct the operating environment of the server cluster, including a computing power unit pool distributed in multiple service blocks and replicas of functional modules deployed on the computing power units; Based on the current resource usage status and application label information of the functional module replicas, identify the types of operating scenarios they are in, and construct corresponding scheduling adaptation factor models; According to the scheduling adaptation factor models, define the state input set and feedback evaluation function of the elastic control strategy; Model the link capacity, location distribution, and interaction delay between each service block as an extended structure graph, and introduce it as a feature parameter into the policy model; Use a deep self-tuning engine to perform offline reinforcement learning training on the policy model to generate optimal action mappings for different operating scenarios and structure graphs; Based on the current state input of the cluster and the trained policy model, dynamically determine whether to perform a computing power expansion operation and the target service block and unit quantity for expansion; Return the operation feedback information and various monitoring metrics of the computing power expansion behavior to the policy model, and perform policy updates based on the continuous tuning mechanism.

2. The operation management method of a server cluster according to claim 1, wherein: The construction of the operating environment of the server cluster includes: For each computing power unit, measure the network round-trip time between it and the dependent nodes; calculate the weighted average of the communication delays of multiple dependent nodes, where the weight of each node is allocated according to its call frequency with the functional module; use the weighted average communication time as the communication overhead metric of the computing power unit; based on the communication overhead metric, preferentially select the computing power unit with low delay for deploying the functional module replicas.

3. A method for operating and managing a server cluster according to claim 1, characterized in that: The identification of the operating scenario type and the construction of the corresponding scheduling adaptation factor model include: Collect the resource usage status data of the target functional module replicas within a predetermined monitoring period, including the average CPU occupancy rate, memory utilization rate, request response time, and call frequency; Parse the application label information carried by the functional module, and the label includes structured identification fields such as the service type to which the module belongs, the business priority level, and whether it is a state-preserving service; Based on the preset operating scenario discrimination rules, use rule matching to map the current state of the module to multiple types of operating scenarios, including high-concurrency processing type, compute-intensive type, low-latency interaction type, or resource-saving type; For the identified operating scenario types, construct corresponding scheduling adaptation factor models, and the models are a set of parameter sets that affect the selection of scheduling strategies, including regional selection weights, startup priorities, latency tolerance thresholds, and expected deployment periods.

4. A method for operating and managing a server cluster according to claim 1, wherein: Extract service characteristic parameters from the scheduling adaptation factor models, including regional elasticity weights, deployment priorities, resource sensitivity thresholds, and latency tolerance levels; based on the service characteristic parameters and the current resource status of the cluster, construct a state input set including node idle rates, network delays, historical deployment behaviors, and failure rates, and use it as the input of the elastic control policy model; define multi-objective feedback items including at least deployment success rate, startup time, resource utilization rate, and deployment cost, and assign weights to each feedback item to construct a weighted combined feedback evaluation function.

5. The operation management method of a server cluster according to claim 1, wherein: Modeling the link capacity, location distribution, and interaction delay between service blocks as an extended structure graph includes: Construct a directed graph structure of the service block set, where each node in the graph corresponds to a service block, and the edge represents the existence of a data transmission path between two blocks; collect the corresponding bandwidth capacity, average network latency, and cross-region distance factor for each edge, and map them to the connection weight value of the edge; perform graph normalization processing on the structure graph, and use the normalized structure graph as the input of the elastic control strategy model to optimize the node selection and scheduling area matching process.

6. The operation management method of a server cluster according to claim 5, wherein: The offline reinforcement learning training of the policy model using the deep self-tuning engine includes: Initialize a dual-structure model including a policy network and a value network. The policy network is used to generate the probability distribution of scheduling actions, and the value network is used to estimate the cumulative expected return of the current state; Construct an offline training sample set including a state input set, actual deployment actions, and a feedback evaluation function; Use the pre-collected system operation history records as training samples to perform behavior cloning warm-up training on the policy network to stabilize the initial output of the model; Based on the offline reinforcement learning algorithm, perform joint optimization of the policy-value network, minimize the action prediction error and the expected return gap, and generate a basic policy mapping function.

7. A method for operating and managing a server cluster according to claim 6, characterized in that: The offline training process further includes: For each state input sample used for training, extract the corresponding service block structure graph fragment, including the embedding vectors of the target block node and its adjacent nodes; Use the graph neural network algorithm to perform message passing and aggregation processing on the graph fragment to generate a high-dimensional structure-aware embedding vector, and jointly construct a complete state input vector with the context embedding vector generated based on the running context type and the resource state vector generated by the resource usage status; During the policy training process, guide the model to learn the scheduling policy performance under the influence of the topological structure, so as to realize the optimal capacity expansion behavior mapping for the structure graph.

8. A method for operating and managing a server cluster according to claim 1, characterized in that: The dynamic judgment of whether to perform the computing power expansion operation based on the current state input of the cluster and the trained policy model includes: Receive the resource state input set of the current cluster, including the remaining resources of each service block, network latency, service load indicators, and the running context characteristics provided by the scheduling adaptation factor model; Input the input set into the trained policy model, trigger the policy model to perform inference operations, and generate an output set including capacity expansion recommendation actions; The output set includes a boolean determination value of whether to expand capacity, the target service block identifier recommended for expansion, and the number of computing power units recommended for expansion; Judge whether to perform the capacity expansion action according to the output result of the policy model. If not, maintain the existing deployment state.

9. A method for operating and managing a server cluster according to claim 1, characterized in that: The policy update based on the continuous tuning mechanism includes: After each computing power expansion operation is executed, collect the operation feedback information including deployment success rate, startup time, resource occupancy stability, cluster load change, and service response time, and structurally associate and file it with the corresponding capacity expansion action; Based on the feedback information, calculate the comprehensive behavior score of the current capacity expansion action, and compare it with the expected return value of the policy model before decision-making; Set a fixed time window or policy behavior round as the model update period, collect a feedback data set, and construct an incremental training sample set; On the premise of ensuring the stability of online inference, perform incremental parameter adjustment on the policy network, and select the updated model through a version comparison verification mechanism to complete the continuous policy tuning process.

10. An operation management system for a server cluster, which is used to implement an operation management method for a server cluster described in any one of claims 1-9, and is characterized in that, Including: An environment construction module for constructing the operating environment of the server cluster, including a computing power unit pool configured in multiple service blocks and function module replicas deployed on the computing power units; A situation recognition module for identifying the type of running situation it is in based on the current resource usage status and application label information of the function module replicas, and constructing a corresponding scheduling adaptation factor model; A policy modeling module for defining the state input set and feedback evaluation function of the elastic control policy according to the scheduling adaptation factor model; A structure graph construction module for modeling the link capacity, location distribution, and interaction delay between service blocks as an extended structure graph, and inputting the graph as a feature parameter into the policy model; A policy training module, including a deep self-tuning engine, for performing offline reinforcement learning training on the policy model to generate optimal scheduling action mappings under different running situations and structure graph conditions; A policy inference module for dynamically determining whether to perform a computing power expansion operation based on the current cluster state input and the trained policy model, and determining the target service block for expansion and the number of computing power units required; A feedback update module for collecting the operation feedback information of the expansion behavior and various system monitoring indicators, and structurally feeding them back to the policy model to incrementally update the policy model through a continuous tuning mechanism.

Citation Information

Patent Citations

  • Cloud and multi-edge network node collaborative micro-service deployment method

    CN118337640A

  • Computing power scheduling strategy optimization system based on reinforcement learning

    CN119356824A

  • AI-based big data distributed computing task automatic optimization method and system

    CN119576507A

  • Cloud platform computing power resource performance monitoring and real-time scheduling optimization method

    CN119883651A

  • Strong-adaptation distributed data distribution method supporting dynamic expansion

    CN119960991A

Cited By

  • Server resource dynamic scheduling method for dealing with video stream high concurrent access

    CN120751207A

  • A server resource dynamic scheduling method for high concurrency access of video stream

    CN120751207B

  • Server assembly method and device, medium and program product

    CN120851816A

  • Motor cluster remote monitoring and energy efficiency optimization method and system based on cloud platform

    CN120872502A

  • Method and system for evaluating operation efficiency of high-performance computing cluster

    CN122220194A