Node scheduling method and scheduling node

By collecting user request characteristics and cluster operation indicators, the number of P nodes and D nodes is dynamically adjusted, which solves the problem of resource idleness or overload in large model deployment and improves resource utilization and management efficiency.

CN122053700APending Publication Date: 2026-05-15XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XFUSION DIGITAL TECH CO LTD
Filing Date
2026-01-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional deployment models for large-scale models lead to idle or overloaded resources, failing to effectively adapt to demand and resulting in low resource utilization.

Method used

By collecting user request characteristics and cluster operation metrics, the number of P nodes and D nodes is dynamically adjusted to achieve precise scaling up and down operations, and resource allocation is optimized by combining scenario characteristics and decision rules.

Benefits of technology

It improved resource utilization and management efficiency, ensured service stability and business adaptability, and enabled rapid cluster response and resource optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053700A_ABST
    Figure CN122053700A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computing, and particularly provides a node scheduling method and a scheduling node, and the node scheduling method comprises the steps: obtaining a request feature corresponding to at least one user request of a target cluster; obtaining an operation index of the target cluster, wherein the operation index is used for representing an operation state of the target cluster; determining an adjustment strategy of the target cluster according to the request feature and the operation index; and according to the adjustment strategy, carrying out capacity reduction or capacity expansion on the target cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computing, and in particular to a node scheduling method and a scheduling node. Background Technology

[0002] In the application of large models in fields such as natural language processing and computer vision, inference services are generally involved. Inference services refer to the process where users send inference requests to nodes that deploy large models, the nodes process the inference requests based on the large models, obtain the inference results, and then return the results to the users.

[0003] Currently, the traditional deployment mode for large models is to deploy large models using a fixed number of nodes. This can easily lead to a mismatch between the required number of nodes and the configured number of nodes, resulting in idle or overloaded resources. Summary of the Invention

[0004] This application provides a node scheduling method and a scheduling node, which dynamically adjusts the target cluster based on two dimensions: user needs and system operation status, in order to improve the resource utilization of the target cluster.

[0005] According to a first aspect of the embodiments of this application, a node scheduling method is provided, the method comprising: Collect request features corresponding to at least one user request in the target cluster. The request features are used to represent the overall characteristics of at least one user request. Collect operational metrics of the target cluster; these metrics are used to represent the operational status of the target cluster. Based on the request characteristics and operational metrics, determine the adjustment strategy for the target cluster. The adjustment strategy refers to the strategy for adjusting the nodes in the target cluster. Based on the adjustment strategy, the target cluster is scaled down or expanded.

[0006] In this embodiment, request characteristics corresponding to at least one user request from the target cluster, as well as the target cluster's operational metrics, are collected. Request characteristics represent the overall features of at least one user request, reflecting the overall service characteristics of the target cluster from the perspective of user needs. Operational metrics represent the operational status of the target cluster, analyzing its overall carrying capacity. Combining request characteristics and operational metrics achieves multi-dimensional information fusion, ensuring that adjustment strategies are adapted to the actual business needs and real operational status of the target cluster. This enables precise scaling up and down operations, ensuring service stability while improving resource utilization, and effectively enhancing cluster management efficiency.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the target cluster includes: pre-populated P nodes and encoded D nodes, and the adjustment strategy includes at least one of the following: Increase or decrease the number of P nodes; Increase or decrease the number of D nodes.

[0008] In this embodiment, by flexibly adjusting the number of P nodes and / or D nodes, the management efficiency of nodes is improved, resource elastic scaling and cost optimization are achieved, and the overall operating efficiency and business adaptability of the cluster are guaranteed.

[0009] In conjunction with the first aspect, some implementations of the first aspect also include: If it is necessary to increase the number of P nodes and reduce the number of D nodes, switch the D node to be switched to a P node. Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P node to be switched can be switched to a D node.

[0010] In this embodiment, when adjusting nodes in the target cluster, if it is necessary to increase the number of P nodes and reduce the number of D nodes, the D nodes to be switched can be switched to P nodes. Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P nodes to be switched can be switched to D nodes. In scenarios with high computing demands, redundant D nodes are converted to P nodes to supplement computing power; in scenarios with high storage demands, idle P nodes are converted to D nodes to expand storage. This achieves functional reuse of existing nodes, ensuring the continuity of cluster operation and data security, while also improving the cluster's adaptability to dynamic changes in business scenarios, thus achieving dual optimization of resource utilization and business adaptability.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the operational metrics include scenario metrics, which refer to the operational metrics that participate in scenario decision-making. Based on the request characteristics and operational metrics, determine the adjustment strategy for the target cluster, including: Based on request characteristics and operational metrics, obtain the scenario characteristics corresponding to the target cluster; Based on scene characteristics, determine the target scene category from at least one scene category; Based on the decision rules associated with the target scenario category, determine the adjustment strategy for the target cluster.

[0012] In this embodiment, the operational metrics include scenario metrics, which are the metrics involved in scenario decision-making within the operational metrics. Therefore, scenario metrics are core metrics that comprehensively reflect and influence scenario classification. Request features and scenario metrics are fused to obtain richer and more comprehensive scenario features. Scenario classification is performed using the fused scenario features to obtain the target scenario category to which the system state belongs. Finally, the adjustment strategy for the target cluster is mapped using pre-set decision rules for the target scenario category, forming a corresponding strategy decision-making mechanism. This allows the system to make targeted adjustments based on its actual operating conditions and user needs, significantly improving the system's accuracy and automation level, and enhancing the efficiency and accuracy of converting multi-dimensional monitoring data into executable adjustment strategies.

[0013] In conjunction with the first aspect, in certain implementations of the first aspect, the adjustment strategy for the target cluster is determined based on the decision rules associated with the target scenario category, including: Obtain the decision parameters associated with the target scene category and the corresponding parameter thresholds; Determine the decision indicators corresponding to the decision parameters from the operational indicators and request characteristics; The decision indicators and parameter thresholds corresponding to the decision parameters are compared to obtain the comparison results, and different comparison results are associated with different action strategies. The action strategy associated with the comparison results is determined as the adjustment strategy for the target cluster.

[0014] In this embodiment, to obtain the adjustment strategy, the decision parameters associated with the target scene category and the corresponding parameter thresholds can be acquired first, thereby extracting data dimensions and key decision features. Then, decision indicators corresponding to the decision parameters are determined from the operational indicators. These decision indicators serve as the basis for decision-making and possess strong real-time and accuracy. Therefore, a simple comparison operation is performed by comparing the decision indicators corresponding to the decision parameters with the parameter thresholds to obtain the comparison result. The comparison result is directly mapped to the action strategy, which becomes the adjustment strategy for the target cluster. This achieves seamless acquisition of the action strategy from the awareness of the comparison state. The entire process only requires automatic parameter collection and comparison, without overly complex actions and processing, effectively improving the response speed and capability from decision rules to adjustment strategies.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, the decision parameters include concurrency parameters and length parameters; Determine the decision indicators corresponding to the decision parameters from the operational indicators, including: Determine the concurrency related to the concurrency parameter and the data length related to the length parameter from the operational metrics; Also includes: If the comparison result shows that the concurrency is greater than the concurrency threshold and the data length is greater than the length threshold, then the action strategy associated with the comparison result is the first action strategy. Alternatively, if the comparison result shows that the concurrency is greater than the concurrency threshold and the data length is less than or equal to the length threshold, then the action strategy associated with the comparison result is the second action strategy. Alternatively, if the comparison result shows that the concurrency is less than or equal to the concurrency threshold and the data length is greater than the length threshold, then the action strategy associated with the comparison result is the third action strategy. Alternatively, if the comparison result shows that the concurrency is less than or equal to the concurrency threshold and the data length is less than or equal to the length threshold, then the action strategy associated with the comparison result is the fourth action strategy.

[0016] In this embodiment, when the decision parameters include concurrency parameters and length parameters, the concurrency level related to the concurrency parameter and the data length related to the length parameter are first obtained from the operational metrics to achieve real-time acquisition of the decision basis. Thus, different comparison results will be generated when comparing the concurrency level with the concurrency threshold and the data length with the length threshold. Different comparison results are associated with different action strategies. Therefore, after obtaining the comparison results, the corresponding adjustment strategy can be automatically and quickly selected and triggered based on the obtained comparison results, greatly improving the speed and accuracy of obtaining the adjustment strategy.

[0017] In conjunction with the first aspect, some implementations of the first aspect also include: Identify multiple training requests, where a training request refers to a user request that participates in request classification training. By using a clustering algorithm, multiple training requests are clustered into scenarios to obtain at least one scenario category.

[0018] In this embodiment, at least one scenario category is obtained through multiple training requests and a clustering algorithm. At least one scenario category obtained through clustering via training requests is more accurately aligned with actual business scenarios, providing an analytical basis for subsequent scenario-based strategy formulation.

[0019] In conjunction with the first aspect, in some implementations of the first aspect, the target cluster is scaled down or expanded according to the adjustment strategy, including: Based on the adjustment strategy, generate operation instructions; Send operation commands to the target cluster to expand or shrink the target cluster.

[0020] In this embodiment, abstract adjustment strategies are transformed into standardized operation instructions, achieving a shift from strategy to execution quality. These instructions are then sent to the target cluster to scale it up or down, enhancing node control. This streamlined process from strategy formulation to resource adjustment ensures the cluster can quickly respond to changes in business needs while maintaining stability and security during the adjustment process.

[0021] In conjunction with the first aspect, in certain implementations of the first aspect, operation instructions are sent to the target cluster to expand or shrink the target cluster, including: When the operation command adds a command to a node, the node in the target cluster is added according to the preset node template and the node's resource configuration information; If the operation command is a node reduction command, then the node in the target cluster will be shut down.

[0022] In this embodiment, for the node addition command, a preset node template and node resource configuration information can be used to add nodes to the target cluster. The node addition relies on the preset node template and resource configuration information, avoiding parameter confusion in new nodes and ensuring that new nodes can quickly adapt to the existing cluster architecture, integrating and undertaking tasks without repeated debugging. For the node reduction command, a simple logic of direct shutdown can quickly release resources occupied by idle nodes, avoiding long-term resource waste. Furthermore, both types of operations are uniformly executed by the scheduling node, ensuring standardized implementation of expansion or contraction operations and reducing potential errors caused by manual intervention. Ultimately, this achieves rapid adaptation, resource controllability, and operational security for cluster node adjustments, allowing the cluster to flexibly respond to changes in business load.

[0023] In conjunction with the first aspect, some implementations of the first aspect also include: Establish a shared task queue for the target cluster, which includes at least one computation task. When the target cluster is expanded, control the newly added nodes in the target cluster to obtain computing tasks from the shared task queue; When the target cluster is scaled down, control the nodes in the target cluster that need to be shut down to stop obtaining computing tasks from the shared task queue and go offline after completing the remaining computing tasks.

[0024] In this embodiment, a shared task queue can be established for the target cluster during task execution. This shared task queue aggregates all computing tasks, enabling centralized management. When the target cluster expands, newly added nodes are controlled to obtain computing tasks from the shared task queue and execute them. When node changes occur in the target cluster, newly added nodes can directly schedule computing tasks from the shared task queue, avoiding scheduling chaos caused by scattered storage of computing tasks and improving the utilization rate of cluster computing resources. Conversely, when the target cluster shrinks, nodes that need to be shut down are controlled to stop obtaining computing tasks from the shared task queue, preventing the accumulation of new computing tasks on these nodes. Furthermore, nodes are taken offline after completing remaining computing tasks, avoiding task interruptions and data loss due to forced node shutdown, ensuring a smooth and orderly shrinking process. This ensures that the cluster maintains the continuity and stability of task processing during both shrinking and expansion.

[0025] In conjunction with the first aspect, some implementations of the first aspect, after scaling down or scaling up the target cluster according to the adjustment strategy, also include: Monitor the node status of each node in the target cluster, including running status or idle status; If the idle time of a node in an idle state is greater than or equal to the idle time threshold, the node in the idle state shall be shut down.

[0026] In this embodiment, after scaling up or down the target cluster, the node status of each node in the target cluster can be monitored, including running or idle status. Therefore, by accurately reflecting the specific operating status of each node through its status, idle nodes that are not carrying computing tasks but are occupying resources can be identified. If the idle time of a node in the idle state is greater than or equal to an idle time threshold, the node in the idle state is shut down. By setting the idle time threshold, the frequent start-stop losses caused by blindly shutting down nodes due to short periods of idleness can be prevented, while timely shutdown can be performed when a node is indeed idle for a long period, quickly releasing the resources it occupies and reallocating them to business scenarios with demand. Overall, while ensuring the normal operation of business, the cost of idle resources is minimized, improving the overall resource utilization efficiency and economy of the cluster.

[0027] According to a second aspect of the embodiments of this application, a node scheduling apparatus is provided, the apparatus comprising: The first acquisition unit is used to acquire request features corresponding to at least one user request in the target cluster. The request features are used to represent the overall characteristics of at least one user request.

[0028] The second acquisition unit is used to collect the operational metrics of the target cluster, which are used to represent the operational status of the target cluster.

[0029] The strategy acquisition unit is used to determine the adjustment strategy for the target cluster based on request characteristics and operational metrics. The adjustment strategy refers to the strategy for adjusting the nodes in the target cluster.

[0030] The cluster adjustment unit is used to scale down or up the target cluster according to the adjustment strategy.

[0031] According to a third aspect of the embodiments of this application, a scheduling node is provided, comprising: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to implement any of the above-described node scheduling methods.

[0032] According to a fourth aspect of the embodiments of this application, a communication device is provided, including a transceiver and a processor, wherein the transceiver is used to receive or send data, and the processor is used to execute any of the node scheduling methods of the embodiments of this application.

[0033] According to a fifth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements any node scheduling method.

[0034] According to a sixth aspect of the embodiments of this application, a computer product is provided, comprising: a computer program that, when executed by a processor, implements the steps of any node scheduling method.

[0035] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description

[0036] The above and other objects, features, and advantages of the embodiments of this application will become more apparent from the more detailed description of the embodiments in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the embodiments of this application and do not constitute a limitation thereof. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0037] Figure 1a The figure shows an application example of a node scheduling system according to an embodiment of this application; Figure 1b The figure shows another application example of a node scheduling system according to an embodiment of this application; Figure 2 The figure shows a flowchart of a node scheduling method according to an embodiment of this application; Figure 3 The figure shows another flowchart of a node scheduling method according to an embodiment of this application; Figure 4 The figure shows an example diagram of a decision rule according to an embodiment of this application; Figure 5 The illustration shows an example of scene classification according to an embodiment of this application; Figure 6 The figure shows an example diagram of node adjustment according to an embodiment of this application; Figure 7 The figure shows an example diagram of a node management system according to an embodiment of this application; Figure 8 The figure shows a schematic diagram of a node scheduling device according to an embodiment of this application; Figure 9 The figure shows a hardware block diagram of a scheduling node according to an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application more apparent, exemplary embodiments according to the embodiments of this application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the embodiments of this application, and not all embodiments of the embodiments of this application. It should be understood that the embodiments of this application are not limited to the exemplary embodiments described herein.

[0039] The technical solution of this application embodiment can be applied to the computing field. Based on information from both the user and the cluster itself, the target cluster is adjusted so that the adjustment strategy is adapted to the actual business needs and real operating status of the target cluster, thereby achieving precise scaling up and down operations, ensuring service stability while improving resource utilization, and effectively improving the management efficiency of the cluster.

[0040] For large models, model training and model usage are generally separated. Based on different service functions, the architecture of a large model may include a P (Prefill) module and a D (Decode) module. The P module can refer to the inference module, which is the training / inference module for the large model and can be used to train and obtain the large model. The D module can refer to the application module, which can obtain the parameters required by the large model from the P module and use the large model corresponding to these parameters to process user requests and obtain inference results. These inference results can be fed back to the user, forming a complete request processing chain. Nodes used to run the P module can be called P nodes, and nodes used to run the D module can be called D nodes. Since large models have many parameters and complex calculations, there are generally multiple P nodes and multiple D nodes. A cluster consisting of multiple P nodes and multiple D nodes can be the target cluster in the embodiments of this application.

[0041] In this embodiment, data is collected from two dimensions: request characteristics of user requests and operational metrics of the target cluster. Request characteristics represent the overall features of at least one user request, reflecting the services the target cluster needs to provide from the perspective of user needs. Operational metrics represent the operational status of the target cluster, analyzing its overall carrying capacity. Combining request characteristics and operational metrics achieves multi-dimensional information fusion, ensuring that adjustment strategies are adapted to the actual business needs and real operational status of the target cluster. This enables precise scaling up and down operations, ensuring service stability while improving resource utilization, and effectively enhancing cluster management efficiency.

[0042] The technical solutions of the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0043] like Figure 1a The diagram shown is an application example of a node scheduling system provided in an embodiment of this application. The time processing system may include a scheduling node 10, a target cluster 20, and a user terminal 30.

[0044] It is understood that the target cluster 20 can be a cluster for training large models and providing inference services corresponding to large models, and may include at least one P node 20a and at least one D node 20b.

[0045] Specifically, scheduling node 10 can execute step 101, receiving user requests, which can be sent by user terminals 30. Of course, there can be multiple user terminals 30, and multiple user requests. Scheduling node 10 performs node scheduling on the target cluster 20 based on the received user requests. Specifically, it can execute: 102. Data collection, which may specifically include: collection of request features and collection of operational metrics.

[0046] 103. Adjustment strategy analysis: The adjustment strategy is obtained. Specifically, the adjustment strategy of the target cluster can be determined based on the request characteristics and operating indicators. The adjustment strategy refers to the strategy for adjusting the nodes in the target cluster.

[0047] 104. Node scheduling, specifically, can be used to scale down or up the target cluster based on adjustment strategies.

[0048] Understandably, after scaling down or up the target cluster, step 105, task allocation, can be executed, which means allocating user requests to node D of the scaled-down or scaled-up target cluster. Node D then executes the inference task corresponding to the allocated user request. After obtaining the inference result, node D in the target cluster can execute step 106, feeding back the inference result to scheduling node 10. Scheduling node 10 can then aggregate the inference results and execute step 107, feeding back the result to the user terminal 30, whereby the user terminal outputs the inference result for the user to view.

[0049] Understandable, Figure 1a The scheduling node 10 shown is a node different from the target cluster 20. In practical applications, the scheduling node 10 can also be a node in the target cluster 20, such as the central node or any node in the target cluster 20.

[0050] like Figure 1b The diagram shown is another application example of a node scheduling system provided in this application embodiment.

[0051] Taking a target cluster 20 comprising at least one P node 20a and at least one D node 20b as an example. The P node can store P modules to perform model training and obtain model parameters. The D node can send a parameter retrieval request to the parameter distribution interface of the P node, and the P node can send model parameters to the D node through the parameter distribution interface. After receiving the model parameters sent by the P node through the parameter distribution interface, the D node obtains the model parameters and loads them into the model to obtain the D model.

[0052] In scheduling node 10, the request scheduler 20c assigns tasks, that is, sends user requests to the D node. The D node uses the D model to perform inference calculations on the user requests sent by the request scheduler 20c, and obtains inference results or calculation results. The P node can also update the P module based on the inference results or calculation results.

[0053] In practical applications, the number and frequency of user requests are constantly changing. Therefore, the computational tasks of P nodes and D nodes are constantly changing. To improve system operating efficiency, scheduling node 10 executes the node scheduling method of this application embodiment to expand or shrink the capacity of P nodes and D nodes. Specifically, expanding or shrinking the capacity of P nodes and D nodes can refer to adjusting the number and ratio of P nodes and D nodes.

[0054] like Figure 2 The diagram shown is a flowchart of a node scheduling method provided in an embodiment of this application. The node scheduling method may include the following steps: S201. Obtain the request characteristics corresponding to at least one user request in the target cluster.

[0055] Optionally, the user request can be a request sent by the user terminal, such as an image processing request, text analysis request, content query request, or any other request. At least one user request can be a request sent to one or more models. Request features can be user-scenario-related features extracted from HTTP (Hypertext Transfer Protocol) request headers or API (Application Programming Interface) parameters.

[0056] Request characteristics can refer to the overall features that describe at least one user request. For example, request characteristics may include at least one of the following: model type, input data size, input scale, response time requirement, request frequency, number of requests, etc.

[0057] Here, "model type" can refer to the type of model, such as image processing model, text analysis model, natural language processing model, question answering system, multilingual joint inference, etc. "Input data size or scale" can refer to the length of the input data. "Input scale" can refer to the number of input requests, such as high concurrency or low concurrency, batch input or single input. "Response time requirement" can refer to the maximum time required to process a user request, such as high latency or low latency. Latency means that exceeding the maximum time is considered a failure to process the user request. "Request frequency" can refer to the number of user requests received per unit of time (e.g., per millisecond, per second). "Request quantity" refers to the number of user requests received per unit of time, at a specific moment, or within a specific time period.

[0058] In one possible design, the type of metric for the request feature can be specified by the user. Prior to S201, the process also includes: detecting a first triggering operation performed by the user. The first triggering operation can be setting or selecting the metric type of the request feature, and obtaining the request feature corresponding to the first triggering operation. Based on the metric type of the request feature, request features corresponding to at least one user request in the target cluster are collected.

[0059] S202. Obtain the operating metrics of the target cluster. The operating metrics are used to represent the operating status of the target cluster.

[0060] As mentioned above, the target cluster can include at least one P node and at least one D node. The operational metrics of the target cluster refer to the metrics when at least one P node and at least one D node in the target cluster are evaluated as a whole. Specifically, the operational metrics of P nodes and D nodes can be evaluated first at the node level, and then the node-level metrics can be combined.

[0061] Optionally, the performance metrics include at least one of the following: CPU (Central Processing Unit) utilization, memory usage, GPU (Graphics Processing Unit) memory usage, queue length, average response time, task completion rate, and throughput.

[0062] CPU utilization refers to the proportion of time the CPU actually spends processing tasks within a unit of time, reflecting the CPU's workload. For example, a utilization rate of 80% means that the CPU spends 80% of its time processing data and 20% is idle. A utilization rate that is too high (such as consistently above 90%) may lead to slower task response.

[0063] Taking CPU utilization as an example, the CPU utilization of the target cluster can be calculated by the ratio of the used CPU resources of all nodes (including P nodes and D nodes) to the total CPU resources of all nodes.

[0064] Memory usage refers to the proportion or specific value of the memory currently used by all nodes in a target cluster relative to the total available memory of all nodes. It can be used to indicate resource consumption. For example, if 10GB of 16GB of memory is used, it indicates the degree of memory usage. Excessive usage may lead to frequent use of virtual memory by the system, reducing operating efficiency.

[0065] GPU memory utilization can be defined as the ratio of the total used GPU resources across all nodes in a target cluster to the total available GPU resources across all nodes. For example, consider 2 P nodes and 3 D nodes. Each P node has 8 logical GPUs, with 6 and 5 currently in use respectively. Each D node has 16 logical GPUs, with 10, 8, and 12 currently in use respectively. The total used GPU resources are: 6 + 5 + 10 + 8 + 12 = 41 logical GPUs. The total available GPU resources across all nodes = (2 × 8) + (3 × 16) = 16 + 48 = 64 logical GPUs. The GPU utilization of the target cluster = 41 / 64 × 100% ≈ 64.06%.

[0066] Queue length refers to the number of computational tasks waiting to be processed in the shared task queue of the target cluster. Queue length reflects the degree of task backlog. Average response time refers to the average time elapsed from when a computational task is submitted by the scheduling node to when the processing result is returned by the D node, reflecting the system's response speed to tasks. For example, an average response time of 200 milliseconds represents the average waiting time from task initiation to completion. Excessive time will affect user experience or business process efficiency.

[0067] Task completion rate refers to the proportion of computing tasks successfully executed within a certain period of time out of the total number of submitted computing tasks, reflecting the reliability of the target cluster in processing tasks. For example, if 98 out of 100 submitted computing tasks are completed, the completion rate is 98%. A rate that is too low may indicate system malfunctions, insufficient resources, or task compatibility issues.

[0068] Throughput can refer to the number of requests completed or the amount of data transmitted per unit of time.

[0069] In this embodiment, the computational tasks can be established based on user requests provided by the scheduling node. A shared task queue can be established in the target cluster, which can include multiple computational tasks, each corresponding to a user request.

[0070] Optionally, after obtaining the target cluster's operational metrics, the method further includes: displaying the target cluster's operational metrics to achieve a visual representation of the metrics for easy viewing by the user. This user can refer to the cluster administrator.

[0071] Optionally, the target cluster's operational metrics can also be specified by the user. Before S202, the process also includes: detecting a second triggering operation performed by the user. The second triggering operation can be setting or selecting a metric type for the operational metrics, and obtaining the metric type corresponding to the second triggering operation. Specifically, S201 may include: collecting at least one operational metric requested by a user according to the metric type of the operational metrics.

[0072] S203. Based on the request characteristics and operational metrics, determine the adjustment strategy for the target cluster. The adjustment strategy refers to the strategy for adjusting the nodes in the target cluster.

[0073] The adjustment strategy can modify the number and / or proportion of nodes in the target cluster. The node proportion refers to the ratio of the number of P nodes to the number of D nodes in the target cluster. For example, if the target cluster includes 3 P nodes and 7 D nodes, the node proportion is 3:7.

[0074] S204. Based on the adjustment strategy, scale down or expand the target cluster.

[0075] Optionally, the target cluster may include at least one P node and at least one D node.

[0076] The adjustment strategy may include at least one of the following: adjustment category; the total number of nodes to be adjusted; and the ratio of P nodes to D nodes. Adjustment categories may include: adding nodes and / or reducing nodes.

[0077] Adjustment strategies may include, for example, the following implementations: Example 1: The adjustment strategy includes: the adjustment category of P nodes and / or the adjustment category of D nodes, as well as the change amount of P nodes and / or the change amount of D nodes in the target cluster.

[0078] If the adjustment category of node P is to add a node, then the number of nodes P will be increased according to the amount of change in node P.

[0079] If the adjustment category for node P is "reducing nodes", then the number of nodes P is reduced according to the amount of change in nodes P. If the adjustment category for node D is "adding nodes", then the number of nodes D is increased according to the amount of change in nodes D.

[0080] If the adjustment category for node D is to reduce nodes, then the number of nodes D will be reduced according to the amount of change in node D.

[0081] For example, assuming the number of P nodes is 2 and the number of D nodes is 5, if the adjustment category is to add nodes, then two P nodes and five D nodes can be added to the target cluster.

[0082] Example 2: The adjustment strategy includes the node ratio of P nodes and D nodes.

[0083] According to the adjustment strategy, scaling up or down the target cluster may include: determining the total number of nodes in the target cluster; calculating the product of the node ratio of P nodes and the total number of nodes to obtain the number of P nodes, and adjusting the P nodes in the target cluster according to the number of P nodes; calculating the product of the node ratio of D nodes and the total number of nodes to obtain the number of D nodes, and adjusting the D nodes in the target cluster according to the number of D nodes.

[0084] Example 3: The adjustment strategy includes: the total number of nodes in the target cluster, and the node ratio of P nodes to D nodes.

[0085] Specifically, you can first adjust the nodes in the target cluster according to the total number of nodes. Then, determine the number of P nodes and the number of D nodes based on the total number of nodes and the ratio of P nodes to D nodes. After that, adjust the P nodes according to the number of P nodes and adjust the D nodes according to the number of D nodes.

[0086] Optionally, after obtaining the adjustment strategy, the adjustment strategy for the target cluster can also be displayed.

[0087] If a user triggers a confirmation operation to execute the adjustment policy for the target cluster, the target cluster will be scaled down or expanded according to the adjustment policy.

[0088] If a user triggers a modification operation on the adjustment policy of the target cluster, the system retrieves the new adjustment policy provided by the user. The target cluster is then scaled up or down according to the new adjustment policy.

[0089] In this embodiment, request characteristics corresponding to at least one user request from the target cluster, along with the target cluster's operational metrics, are collected. Request characteristics represent the overall features of at least one user request, reflecting the service characteristics of the target cluster from the perspective of user needs. Operational metrics represent the operational status of the target cluster, analyzing its overall carrying capacity. Combining request characteristics and operational metrics achieves multi-dimensional information fusion, ensuring that adjustment strategies are adapted to the actual business needs and real operational status of the target cluster. This enables precise scaling up and down operations, ensuring service stability while improving resource utilization, and effectively enhancing cluster management efficiency.

[0090] As an example, the target cluster includes: P nodes and D nodes, and the adjustment strategy includes at least one of the following: Increase or decrease the number of P nodes; Increase or decrease the number of D nodes.

[0091] In this embodiment, by flexibly adjusting the number of P nodes and / or D nodes, the management efficiency of nodes is improved, resource elastic scaling and cost optimization are achieved, and the overall operating efficiency and business adaptability of the cluster are guaranteed.

[0092] like Figure 3 The diagram shown is another flowchart of a node scheduling method provided in this application embodiment. This node scheduling method, along with... Figure 2 The difference in the illustrated embodiment is that the operational metrics include scenario metrics, which refer to the metrics involved in scenario decision-making within the operational metrics. S203 determines the adjustment strategy for the target cluster based on the request characteristics and operational metrics, which may include the following steps: S301. Based on the request characteristics and the scenario indicators in the operation indicators, obtain the scenario characteristics corresponding to the target cluster.

[0093] Optionally, S301 may include: concatenating request features and operational metrics to obtain the scenario features of the target cluster.

[0094] Operational metrics may include scenario-based metrics. Scenario-based metrics refer to those operational metrics related to user scenarios. For example, scenario-based metrics may include: concurrency, input length, output length, and throughput. Scenario-based metrics can be pre-set or obtained by pre-categorizing metrics. This application embodiment does not impose excessive limitations on the specific types of scenario-based metrics.

[0095] Operational metrics include scenario metrics, which are the metrics that participate in scenario decision-making within the operational metrics. Therefore, scenario metrics are the core metrics that can comprehensively reflect the impact on scenario classification.

[0096] Optionally, the scenario features corresponding to the target cluster can be obtained by feature fusion of request features and scenario metrics in the operational metrics. This also includes: concatenating the request features and scenario metrics in the operational metrics to obtain the scenario features corresponding to the target cluster.

[0097] S302. Based on scene characteristics, determine the target scene category from at least one scene category.

[0098] Optionally, each scene category can be pre-associated with standard features. S302 may include: calculating the feature similarity between the standard features and scene features of each scene category, obtaining the feature similarity corresponding to at least one scene category respectively. The scene category with the highest feature similarity is determined as the target scene category.

[0099] S303. Determine the adjustment strategy for the target cluster based on the decision rules associated with the target scenario category.

[0100] Understandably, each scenario category can be pre-associated with corresponding decision rules. Different scenario categories have different decision rules. The adjustment strategy for the target cluster can then be the decision rules associated with the target scenario category.

[0101] Decision rules refer to the rules for making decisions on the adjustment strategy of the target cluster. They may include decision parameters, parameter thresholds corresponding to each decision parameter, and the comparison method between decision parameters and operating indicators.

[0102] For example, in scenario 1, decision parameters might include concurrency and data length. In scenario 2, decision parameters might include batch size and throughput.

[0103] To facilitate understanding, Table 1 below uses four scenario categories—high-concurrency long input long output, high-concurrency long input short output, high-concurrency short input long output, and high-concurrency short input long output—as examples to illustrate the correspondence between decision rules and adjustment strategies. For instance, the decision parameters (i.e., the parameters involved in the decision) in the decision rules include: input data length and output data length. Each decision rule is associated with a corresponding adjustment strategy.

[0104] Table 1

[0105] As shown in Table 1, N represents the original total number of nodes in the target cluster. In the first row, the 2 in N+2 represents the change in the number of nodes. Adjustment strategies could include adding 2 nodes, resulting in a 5:5 ratio of P nodes to D nodes. In the second row, the 1 in N+1 represents the change in the number of nodes. Adjustment strategies could include adding 1 node, resulting in a 6:4 ratio of P nodes to D nodes. In the third row, N (number of nodes) remains unchanged, and the ratio of P nodes to D nodes is 3:7. In the fourth row, N remains unchanged, and the ratio of P nodes to D nodes is 4:6.

[0106] Furthermore, if two decision rules contain the same decision parameter, the parameter threshold for that decision parameter can be the same or different in different decision rules. As shown in Table 1, the parameter threshold for each decision rule is the same, which is 8K.

[0107] It is understandable that in the decision rules corresponding to different scenario categories, the comparison method between decision indicators and parameter thresholds is the same: size comparison and pre-established action decisions for each comparison branch. The action decisions for different branches can be different. The action decisions for the same branch can also be different in different decision rules.

[0108] Optionally, the decision-making rule can be made using three dimensions: state space S, planning space O, and action space A.

[0109] The state space (S) represents the current state of the system and has three key dimensions: current load, number of nodes, and user scenario category.

[0110] The planning space (O) is used to represent a detailed analysis of system requirements, which can be further divided into resource requirements based on "state space + user scenario".

[0111] The action space (A) is used to represent the final adjustment strategy obtained through the planning space. The adjustment strategy may include: adjustment category, number of nodes to be adjusted, and PD node ratio.

[0112] In this embodiment, request features and operational metrics are fused to obtain richer and more comprehensive scene features. The fused scene features are then used for scene classification to determine the target scene category to which the system state belongs. Finally, based on pre-set decision rules for the target scene category, adjustment strategies for the target cluster are mapped, forming a corresponding strategy decision-making mechanism. This allows the system to make targeted adjustments based on its actual operating conditions and user needs, significantly improving the system's accuracy and automation level. This enhances the efficiency and accuracy of converting multi-dimensional monitoring data into executable adjustment strategies.

[0113] The adjustment strategy for the target cluster can be determined based on the decision rules associated with the target scenario category. These decision rules involve decision parameters, threshold values ​​for those parameters, and specific decision-making methods. In one possible design, determining the adjustment strategy for the target cluster in step S303 based on the decision rules associated with the target scenario category may include the following steps: A1. Obtain the decision parameters associated with the target scene category and the corresponding parameter thresholds.

[0114] A2. Determine the decision indicators corresponding to the decision parameters from the operational indicators and request characteristics.

[0115] A3. Compare the decision indicators and parameter thresholds corresponding to the decision parameters.

[0116] Optionally, by comparing the decision indicators and parameter thresholds corresponding to the decision parameters, comparison results can be obtained, and different comparison results can be associated with different action strategies.

[0117] A4. Determine the action strategy for the target cluster based on the comparison results.

[0118] Understandably, the decision-making principle used by the decision-making rules is: if the load is high and the forecast shows a continuous increase, increase the number of nodes; if the load is low and the forecast shows a continuous decrease, decrease the number of nodes. For mixed scenarios, the ratio of P nodes to D nodes can be adjusted.

[0119] Taking Table 1 as an example, decision parameters can include input length and output length. The threshold for the input length parameter is 8K. The threshold for the output parameter is also 8K. The decision index and parameter thresholds are compared to obtain the comparison result. If the comparison result is: the input data is greater than 8K and the output data is less than 8K, then the corresponding adjustment strategy is "N+1, P:D=6:4".

[0120] In this embodiment, to obtain the adjustment strategy, the decision parameters associated with the target scene category and the corresponding parameter thresholds can be acquired first, thereby extracting data dimensions and key decision features. Then, decision indicators corresponding to the decision parameters are determined from the operational indicators. These indicators serve as the basis for decision-making and possess strong real-time and accuracy. Therefore, the comparison results are obtained by comparing the decision indicators corresponding to the decision parameters with the parameter thresholds. The comparison results are directly mapped to the action strategy, which becomes the adjustment strategy for the target cluster. This achieves seamless acquisition of the action strategy from the awareness of the comparison state. The entire process only requires automatic parameter collection and comparison, without overly complex actions and processing, effectively improving the response speed and capability from decision rules to adjustment strategies.

[0121] In the process of using decision rules to obtain the adjustment strategy for the target cluster, decision parameters are involved, which can include one or more. For example, if there are two decision parameters, they can be a concurrency parameter and a length parameter.

[0122] Based on this, A2 determines the decision indicators corresponding to the decision parameters from the operational indicators, including: A21. Determine the concurrency related to the concurrency parameter and the data length related to the length parameter from the operational metrics.

[0123] Performing step A3 yields the comparison results. These results can include various scenarios; therefore, the following examples illustrate the comparison results and their corresponding action strategies: Result 1: The comparison results show that the concurrency is greater than the concurrency threshold and the data length is greater than the length threshold.

[0124] The length parameter can be, for example, the length of the input data or the length of the output data.

[0125] If the comparison result is 1, then the action policy associated with the comparison result is the first action policy.

[0126] Result 2: The comparison results show that the concurrency is greater than the concurrency threshold, and the data length is less than or equal to the length threshold.

[0127] If the comparison result is result 2, then the action policy associated with the comparison result is the second action policy.

[0128] Result 3: The comparison results show that the concurrency is less than or equal to the concurrency threshold, and the data length is greater than the length threshold.

[0129] If the comparison result is 3, then the action strategy associated with the comparison result is the third action strategy.

[0130] Result 4: The comparison results show that the concurrency is less than or equal to the concurrency threshold, and the data length is less than or equal to the length threshold.

[0131] If the comparison result is 4, then the action policy associated with the comparison result is the fourth action policy.

[0132] To facilitate understanding, the following will use a high-concurrency text generation scenario as an example to explain the decision rules in detail. In this scenario, based on the decision parameters in the decision rules—concurrency parameter and length parameter, and parameter thresholds—the concurrency threshold K and length threshold L, the decision indicators corresponding to the decision parameters are: the concurrency parameter corresponds to the concurrency level; the length parameter corresponds to the data length.

[0133] like Figure 4The diagram shown is an example of a decision rule provided in an embodiment of this application. The decision rule may include the following steps: S401. Determine the concurrency level corresponding to the concurrency parameter and the data length corresponding to the length parameter from the operating indicators.

[0134] S402. Determine if the concurrency is greater than the concurrency threshold K. If yes, proceed to S403; otherwise, proceed to S406.

[0135] S403. Determine if the data length is greater than the length threshold L. If yes, execute S404; otherwise, execute S405.

[0136] S404. Determine action strategy B1 as the adjustment strategy for the target cluster.

[0137] S405. Determine action strategy B2 as the adjustment strategy for the target cluster.

[0138] S406. Determine if the data length is greater than the length threshold L. If yes, proceed to S407; otherwise, proceed to S408.

[0139] S407. Determine action strategy B3 as the adjustment strategy for the target cluster.

[0140] S408. Determine the action strategy B4 as the adjustment strategy for the target cluster.

[0141] refer to Figure 4 The comparison results of decision indicators and parameter thresholds may include, for example: concurrency is greater than the concurrency threshold and data length is greater than the length threshold; concurrency is greater than the concurrency threshold and data length is less than or equal to the length threshold; concurrency is less than or equal to the concurrency threshold and data length is greater than the length threshold; concurrency is less than or equal to the concurrency threshold and data length is less than or equal to the length threshold.

[0142] Understandably, if the concurrency is greater than the concurrency threshold and the data length is greater than the length threshold, then action strategy B1 is determined to be the adjustment strategy; if the concurrency is greater than the concurrency threshold and the data length is less than or equal to the length threshold, then action strategy B2 is determined to be the adjustment strategy; if the concurrency is less than or equal to the concurrency threshold and the data length is greater than the length threshold, then action strategy B3 is determined to be the adjustment strategy; and if the concurrency is less than or equal to the concurrency threshold and the data length is less than or equal to the length threshold, then action strategy B4 is determined to be the adjustment strategy.

[0143] Optionally, action strategy B1, action strategy B2, action strategy B3, or action strategy B4 can be used as adjustment strategies. Any action strategy can include at least one of the following: adjustment category; the amount of node change of the node to be adjusted; the node ratio of P nodes and D nodes. Adjustment category can include: adding nodes and / or reducing nodes.

[0144] For example, in the case where the concurrency is greater than the concurrency threshold K and the data length is greater than the length threshold L, which is a high-concurrency long input scenario, the action strategy B1 may include, for example, adjusting the category to add nodes, the node change amount of the nodes to be adjusted is t1, and the node ratio of P nodes and D nodes is w1, such as w1=5:5.

[0145] In scenarios where the concurrency exceeds the concurrency threshold and the data length is less than or equal to the length threshold, which fall under the category of high-concurrency short-input scenarios, action strategy B2 may include, for example, the total number of nodes in the target cluster remains unchanged, and the node ratio of P nodes to D nodes is w2, such as w2=3:7.

[0146] In scenarios where the concurrency is less than or equal to the concurrency threshold and the data length is greater than the length threshold, which fall under the category of low-concurrency long input scenarios, action strategy B3 may include, for example, adjusting the category to reduce nodes, with the node change amount of the nodes to be adjusted being t2, and the node ratio of P nodes to D nodes being w3, such as w3=4:6.

[0147] In scenarios where the concurrency is less than or equal to the concurrency threshold and the data length is less than or equal to the length threshold, which fall under the category of low-concurrency short-input scenarios, action strategy B4 may include, for example, the total number of nodes in the target cluster remains unchanged, and the node ratio of P nodes to D nodes is w4, such as w4=6:4.

[0148] Understandably, the higher the concurrency and the longer the data length, the greater the inference pressure on node D. Therefore, when determining the ratio of P nodes to D nodes, node D should account for a larger proportion. Conversely, the lower the concurrency and the shorter the data length, the less inference pressure on node D. Therefore, when determining the ratio of P nodes to D nodes, node D should account for a smaller proportion, and the proportion of node D can be appropriately reduced.

[0149] Of course, action strategies B1, B2, B3, and B4 are merely examples and do not constitute specific limitations.

[0150] In this embodiment, when the decision parameters include concurrency parameters and length parameters, the concurrency level related to the concurrency parameter and the data length related to the length parameter are first obtained from the operational metrics to achieve real-time acquisition of the decision basis. Thus, different comparison results will be generated when comparing the concurrency level with the concurrency threshold and the data length with the length threshold. Different comparison results are associated with different action strategies. Therefore, after obtaining the comparison results, the corresponding adjustment strategy can be automatically and quickly selected and triggered based on the obtained comparison results, greatly improving the speed and accuracy of obtaining the adjustment strategy.

[0151] The relevant embodiments involve at least one scene category, which can be obtained through classification. In one possible design, it also includes: Multiple training requests are identified, where a training request refers to a user request that participates in the request classification training.

[0152] By using a clustering algorithm, multiple training requests are clustered into scenarios to obtain at least one scenario category.

[0153] Optionally, by using a clustering algorithm to cluster multiple training requests into scenarios and obtain at least one scenario category, it may include: extracting scenario features corresponding to multiple training requests respectively to obtain multiple scenario features. The scenario features include at least one of the following: model identifier, data length of input data, and timeout.

[0154] Based on clustering algorithms, multiple scene features are classified to obtain at least one scene category.

[0155] Optionally, the scenario category can be any of the following: high concurrency and low latency, high throughput, mixed load, high concurrency and high latency, low concurrency and low latency, or low concurrency and high latency.

[0156] The high-concurrency, low-latency category refers to a load scenario where the number of requests per unit time exceeds a first threshold (high concurrency), and the processing time for a single request is extremely short (low latency). The first threshold is a relatively large value.

[0157] High throughput can be defined as a total throughput processed per unit time that is greater than a second threshold value, which is a specified large value.

[0158] Mixed load categories can refer to loads with multiple characteristics (such as some requests having high concurrency and low latency, and some tasks having high throughput), which need to take into account the needs of different scenarios.

[0159] The high-concurrency, high-latency category can refer to a situation where the number of requests per unit time exceeds the first threshold (high concurrency), but the processing flow of a single request is complex and time-consuming (high latency). The second threshold is a specified, relatively large value.

[0160] The low-concurrency, low-latency category refers to a situation where the number of requests per unit time is less than the third threshold (low concurrency), the processing time for a single request is short (low latency), and the system load is low. The third threshold is a specified, relatively small value.

[0161] The low-concurrency, high-latency category can refer to a situation where the number of requests per unit time is less than the third threshold (low concurrency), but a single request takes a long time due to complex processes, large amounts of data, etc. (high latency).

[0162] The first, second, and third threshold values ​​can be preset. "Shorter duration" means the duration is less than the first time interval, and "longer duration" means the duration is greater than the second time interval. The first time interval is less than the second time interval.

[0163] like Figure 5 The diagram shown is an example of scene classification provided in an embodiment of this application. The following steps can be performed during the scene classification process: S501, Identify multiple training requests.

[0164] S502. Extract the scene features corresponding to multiple training requests respectively.

[0165] Scenario features describe the characteristics related to the usage scenario of the training request. The acquisition methods are the same as those mentioned above, both obtainable through request features and runtime metrics, and will not be repeated here. Scenario features may include at least one of the following: model type, input data size, response time requirements, request frequency, number of requests, throughput, etc.

[0166] S503. Use a clustering algorithm to classify multiple request features and obtain at least one scenario category.

[0167] At least one scenario category may include, for example, at least one of the following: low-load text generation scenario, high-load low-latency text generation scenario, and batch processing high-throughput scenario.

[0168] Optionally, after obtaining at least one scenario category, the method further includes: setting decision rules for each scenario category and associating adjustment strategies with each decision rule.

[0169] List 2 below shows examples of scenario categories and decision rules.

[0170]

[0171] Table 2 Alternatively, the clustering algorithm can be, for example, the K-means algorithm.

[0172] In this embodiment, at least one scenario category is obtained through multiple training requests and a clustering algorithm. At least one scenario category obtained through clustering via training requests is more accurately aligned with actual business scenarios, providing an analytical basis for subsequent scenario-based strategy formulation.

[0173] To improve the adjustment efficiency of the target cluster, scheduling nodes can be used to adjust the nodes. Specifically, depending on the adjustment strategy, the target cluster can be scaled down or expanded, which may include: Based on the adjustment strategy, generate operation instructions; Send operation commands to the target cluster to expand or shrink the target cluster.

[0174] Optionally, the scheduling node may include a Kubernetes integration module.

[0175] The scheduling node can be configured with node templates and node resource configuration information.

[0176] Furthermore, the node's resource configuration information may include resource metrics specified by the Horizontal Pod Autoscaler (HPA) as well as custom resource metrics.

[0177] HPA is an automatic scaling mechanism in Kubernetes. Its core function is to automatically adjust the number of Pods (container groups) based on preset monitoring metrics (such as CPU utilization, memory usage, and custom business metrics). When the metrics exceed the threshold, the number of Pod replicas is increased (horizontal scaling) to distribute the load; when the metrics are below the threshold, the number of Pod replicas is decreased (horizontal scaling) to save resources.

[0178] For example, when the CPU utilization of a web service consistently exceeds 70%, HPA can automatically increase the number of Pods from 3 to 5; when the utilization drops below 30%, it can reduce it back to 2, thereby achieving dynamic matching between service capacity and business load.

[0179] When adding nodes to a target cluster, new nodes can be created using node templates, and then communication resources can be configured for them based on their resource configuration information. Node templates are standardized configuration templates used to quickly create or configure new nodes in a target cluster, containing preset information such as node hardware specifications, software environment, network settings, and operating parameters. Using node templates ensures consistency in configuration for new nodes, avoiding parameter confusion or omissions caused by manual configuration.

[0180] For example, a node template can include: the number of CPUs, the number of GPUs, the software environment, network configuration information (such as IP address and port), and running parameters (such as starting 2 computing processes by default). When expansion is needed, the node management module can directly call this template to quickly generate new nodes that meet the standards, without the need for repeated manual configuration, ensuring that the new nodes can be seamlessly integrated into the cluster and work collaboratively.

[0181] Resource configuration information may refer to predefined queue resources, GPU resources, CPU resources, etc.

[0182] Furthermore, to expand the target cluster, after adding new nodes, these new nodes can be configured as P nodes or D nodes. Specifically, when configuring a new node as a P node, P node functions can be set for it; for example, configuring a P module for the new node will result in a new P node. When configuring a new node as a D node, D node functions can be set for it; for example, configuring a D module for the new node will result in a new D node.

[0183] In this embodiment, abstract adjustment strategies are transformed into standardized operation instructions, realizing the conversion from strategy to execution quality. The operation instructions are sent to the target cluster by the scheduling node, enhancing the control over the nodes through a more specialized scheduling node. This streamlines the process from strategy formulation to resource adjustment, ensuring the cluster can quickly respond to changes in business needs while guaranteeing the stability and security of the adjustment process.

[0184] To improve resource utilization efficiency, a resource recycling mechanism can also be established. Specifically, after scaling down or up the target cluster according to the adjustment strategy, this also includes: The node recycling module reclaims idle nodes from the target cluster.

[0185] Optionally, the node reclamation module can be used to reclaim idle nodes in the target cluster, including: Monitor the node status of each node in the target cluster, including running status or idle status; If the idle time of a node in an idle state is greater than or equal to the idle time threshold, the node in the idle state shall be shut down.

[0186] In this embodiment, for the node addition command, a preset node template and node resource configuration information can be used to add nodes to the target cluster. The node addition relies on the preset node template and resource configuration information, avoiding parameter confusion in new nodes and ensuring that new nodes can quickly adapt to the existing cluster architecture, integrating and undertaking tasks without repeated debugging. For the node reduction command, a simple logic of direct shutdown can quickly release resources occupied by idle nodes, avoiding long-term resource waste. Furthermore, both types of operations are uniformly executed by the scheduling node, ensuring standardized implementation of expansion or contraction operations and reducing potential errors caused by manual intervention. Ultimately, this achieves rapid adaptation, resource controllability, and operational security for cluster node adjustments, allowing the cluster to flexibly respond to changes in business load.

[0187] Based on the different adjustment categories, adjustments can be divided into adding nodes and reducing nodes.

[0188] For scenarios involving adding nodes: when the operation command is to add a node, the node in the target cluster is added according to the preset node template and the node's resource configuration information.

[0189] For scenarios involving reducing nodes: if the operation command is a node reduction command, then the nodes in the target cluster will be shut down.

[0190] Increment and decrement instructions are different. Increment and decrement instructions can be carried in the same field of the same signaling. If the field value is 1, it indicates that the operation instruction is an increment instruction; if the field value is 0, it indicates that the operation instruction is a decrement instruction.

[0191] The scheduling node can pre-configure node templates and resource information, allocate node resources to the target cluster, and create new nodes. The scheduling node can also shut down nodes in the target cluster, such as abnormal nodes that are malfunctioning or idle nodes.

[0192] As mentioned above, the target cluster can include P nodes and D nodes. P nodes can retrieve new tasks from the P task queue. D nodes can retrieve new tasks from the D task queue.

[0193] In this embodiment, a shared task queue can be established for the target cluster during task execution. This shared task queue aggregates all computing tasks, enabling centralized management. When the target cluster expands, newly added nodes are controlled to obtain computing tasks from the shared task queue and execute them. When node changes occur in the target cluster, newly added nodes can directly schedule computing tasks from the shared task queue, avoiding scheduling chaos caused by scattered storage of computing tasks and improving the utilization rate of cluster computing resources. Conversely, when the target cluster shrinks, nodes that need to be shut down are controlled to stop obtaining computing tasks from the shared task queue, preventing the accumulation of new computing tasks on these nodes. Furthermore, nodes are taken offline after completing remaining computing tasks, avoiding task interruptions and data loss due to forced node shutdown, ensuring a smooth and orderly shrinking process. This ensures that the cluster maintains the continuity and stability of task processing during both shrinking and expansion.

[0194] For ease of understanding, Figure 6 An example diagram of node adjustment is shown.

[0195] Multiple P nodes 601 can obtain P tasks from the P task queue 602. Multiple D nodes 603 can obtain D tasks from the D task queue 604.

[0196] If it is necessary to increase the number of P nodes and reduce the number of D nodes, switch the D node to be switched to a P node.

[0197] by Figure 6 Taking node 2 as an example, it was originally node D. It can be taken offline and brought online as node P, thereby switching the original node D to node P and obtaining a new node P.

[0198] Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P node to be switched can be switched to a D node.

[0199] by Figure 6 Taking node 1 to be switched as an example, node 1 was originally a P node. It can be taken offline as a P node and brought online as a D node, thereby switching the original P node to a D node and obtaining a new D node.

[0200] In this embodiment, when adjusting nodes in the target cluster, if it is necessary to increase the number of P nodes and reduce the number of D nodes, the D nodes to be switched can be switched to P nodes. Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P nodes to be switched can be switched to D nodes. In scenarios with high computing demands, redundant D nodes are converted to P nodes to supplement computing power; in scenarios with high storage demands, idle P nodes are converted to D nodes to expand storage. This achieves functional reuse of existing nodes, ensuring the continuity of cluster operation and data security, while also improving the cluster's adaptability to dynamic changes in business scenarios, thus achieving dual optimization of resource utilization and business adaptability.

[0201] To improve the efficiency of node adjustment, hot migration can be used to adjust nodes in the target cluster. Specifically, this can include the following steps: Establish a shared task queue for the target cluster, which includes at least one computation task. When the target cluster is expanded, control the newly added nodes in the target cluster to obtain computing tasks from the shared task queue; When the target cluster is scaled down, control the nodes in the target cluster that need to be shut down to stop obtaining computing tasks from the shared task queue and go offline after completing the remaining computing tasks.

[0202] For example, if the target cluster is being expanded, refer to Figure 6 A new P node is added, which means it can retrieve and execute P tasks from the P task queue 602. A newly added D node is added, which can retrieve and execute D tasks from the D task queue 604.

[0203] If the target cluster is downsized to have node P, such as Figure 6 In the process, node 1 to be switched is a P node that needs to be offline. That is, node 1 to be switched will no longer obtain P tasks from the P task queue 602 and will go offline after completing the remaining tasks.

[0204] If the target cluster shrinks node D, such as Figure 6 In the process, node 2 to be switched is node D that needs to be taken offline. That is, node 2 to be switched will no longer obtain D tasks from task queue 604 and will go offline after completing the remaining tasks.

[0205] Optionally, the offline status of idle nodes in the target cluster can be determined using TTL (Time-to-Live), meaning that idle nodes automatically go offline after a specified time. The specified time can be a pre-set waiting time for the remaining tasks to execute. The specified time can be greater than or equal to the minimum execution time of the remaining tasks. For example, if there are 10 remaining computation tasks and the average execution time of each computation task is 10 milliseconds, then the specified time is greater than or equal to 100 milliseconds.

[0206] Optionally, when adding new nodes to the target cluster, a priority queue can be used to ensure resource allocation for critical tasks. Specifically, a priority queue assigns priority weights to different computing tasks, allowing critical computing tasks to gain priority scheduling rights in resource contention, thereby accurately guaranteeing their resource allocation. Critical computing tasks may include, for example, core transaction processes or critical business orders. Priority weights can be pre-set numerical values ​​representing the execution priority of computing tasks. For example, a priority weight of 0.8 or 0.6; the higher the priority weight, the higher the priority.

[0207] Furthermore, load balancing strategies can be used to distribute computing tasks among the nodes in the target cluster (including existing and new nodes). For example, a consistent hashing algorithm can be used to distribute computing tasks, ensuring load balancing between old and new nodes.

[0208] In this context, using consistent hashing to allocate computing tasks involves calculating a hash value for each node and distributing them in a ring-shaped space. Then, a hash value is calculated for each computing task, and the task is assigned to the nearest node moving clockwise along the ring. Consistent hashing offers advantages such as load balancing, strong fault tolerance, and high scalability.

[0209] In this embodiment, after scaling up or down the target cluster, the node status of each node in the target cluster can be monitored, including running or idle status. Therefore, by accurately reflecting the specific operating status of each node through its status, idle nodes that are not carrying computing tasks but are occupying resources can be identified. If the idle time of a node in the idle state is greater than or equal to an idle time threshold, the node in the idle state is shut down. By setting the idle time threshold, the frequent start-stop losses caused by blindly shutting down nodes due to short periods of idleness can be prevented, while timely shutdown can be performed when a node is indeed idle for a long period, quickly releasing the resources it occupies and reallocating them to business scenarios with demand. Overall, while ensuring the normal operation of business, the cost of idle resources is minimized, improving the overall resource utilization efficiency and economy of the cluster.

[0210] like Figure 7 The diagram illustrates an example of a node management system provided in this application. This system may include a scheduling node, which may comprise a node management module 701 and a resource reclamation module 703. The node management module 701 may be, for example, a Kubernetes integration module. The scheduling node can scale up or down the nodes in the target cluster 702. The resource reclamation module 703 can reclaim idle nodes in the target cluster 702 to prevent resource waste. Furthermore, the scheduling node may also include a monitoring module and a scheduling module (not shown in the diagram). The monitoring module can collect various parameters of the target cluster, such as request characteristics corresponding to at least one user request and operational metrics of the target cluster. The scheduling module can execute the node scheduling method of this application to scale up or down the target cluster according to an adjustment strategy, and to allocate resources and manage the lifecycle of nodes in the target cluster.

[0211] like Figure 8 The diagram shown is a structural schematic of a node scheduling device provided in an embodiment of this application. The node scheduling device 800 includes: The first acquisition unit 801 is used to acquire the request characteristics corresponding to at least one user request in the target cluster.

[0212] The second acquisition unit 802 is used to acquire the operating indicators of the target cluster, which are used to represent the operating status of the target cluster.

[0213] The strategy acquisition unit 803 is used to determine the adjustment strategy of the target cluster based on the request characteristics and operating metrics.

[0214] Cluster adjustment unit 804 is used to shrink or expand the target cluster according to the adjustment strategy.

[0215] Optionally, the target cluster includes: P nodes and D nodes, and the adjustment strategy includes at least one of the following: Increase or decrease the number of P nodes; Increase or decrease the number of D nodes.

[0216] Optionally, it also includes: If it is necessary to increase the number of P nodes and reduce the number of D nodes, switch the D node to be switched to a P node. Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P node to be switched can be switched to a D node.

[0217] Optionally, the policy acquisition unit 803 may include: The feature fusion module is used to obtain the scene features corresponding to the target cluster based on request features and operational metrics.

[0218] The category determination module is used to determine at least one target scene category corresponding to a user request from at least one scene category based on scene characteristics.

[0219] The strategy acquisition module is used to determine the adjustment strategy for the target cluster based on the decision rules associated with the target scenario category.

[0220] Optionally, the policy acquisition module may include: The parameter acquisition submodule is used to obtain the decision parameters associated with the target scene category and the corresponding parameter thresholds.

[0221] The indicator acquisition submodule is used to determine the decision indicators corresponding to the decision parameters from the running indicators and request features.

[0222] The parameter comparison submodule is used to compare the decision indicators and parameter thresholds corresponding to the decision parameters, and determine the action strategy of the target cluster based on the comparison results.

[0223] Optionally, the decision parameters include concurrency parameters and length parameters. The metric acquisition submodule is specifically used to determine the concurrency related to the concurrency parameter and the data length related to the length parameter from the running metrics.

[0224] Optionally, it also includes: If the comparison result shows that the concurrency is greater than the concurrency threshold and the data length is greater than the length threshold, then the action strategy associated with the comparison result is the first action strategy. Alternatively, if the comparison result shows that the concurrency is greater than the concurrency threshold and the data length is less than or equal to the length threshold, then the action strategy associated with the comparison result is the second action strategy. Alternatively, if the comparison result shows that the concurrency is less than or equal to the concurrency threshold and the data length is greater than the length threshold, then the action strategy associated with the comparison result is the third action strategy. Alternatively, if the comparison result shows that the concurrency is less than or equal to the concurrency threshold and the data length is less than or equal to the length threshold, then the action strategy associated with the comparison result is the fourth action strategy.

[0225] Optionally, it also includes: The request retrieval unit is used to identify multiple training requests, which refer to user requests that participate in request classification training.

[0226] The request classification unit is used to cluster multiple training requests into scenes using a clustering algorithm to obtain at least one scene category.

[0227] Optionally, the cluster adjustment unit 804 may include: The instruction generation module is used to generate operation instructions based on the adjustment strategy.

[0228] The cluster adjustment module is used to send operation commands to the target cluster to expand or shrink the target cluster.

[0229] Optionally, the cluster adjustment module is specifically used to: add nodes in the target cluster according to the preset node template and the node's resource configuration information when the operation instruction is to add nodes; or, shut down nodes in the target cluster when the operation instruction is to remove nodes.

[0230] Optionally, it also includes: The queue creation unit is used to create a shared task queue for the target cluster, which includes at least one computation task.

[0231] The first adjustment unit is used to control the newly added nodes in the target cluster to obtain computing tasks from the shared task queue when the target cluster is expanded.

[0232] The second adjustment unit is used to control the nodes in the target cluster that need to be shut down to stop obtaining computing tasks from the shared task queue and to go offline after completing the remaining computing tasks when the target cluster is scaled down.

[0233] Optionally, it also includes: The status monitoring unit is used to monitor the node status of each node in the target cluster, including running status or idle status.

[0234] The idle shutdown unit is used to shut down a node that is in an idle state if the idle time of the node is greater than or equal to the idle time threshold.

[0235] In the embodiments of this application, Figure 8The device shown can also be a chip or a chip system, such as a system on chip (SoC) or a baseboard management controller (BMC).

[0236] Figure 9 This is a hardware block diagram of a scheduling node provided in an embodiment of this application. The scheduling node 900 according to this embodiment includes at least a memory 901, a processor 902, and a transceiver 903. The memory 901 stores a computer program, and the transceiver 903 communicates with the target cluster to perform operations such as data acquisition and instruction transmission. The processor 902 executes the computer program to implement the node scheduling method of any of the above embodiments.

[0237] In addition, the memory 901, processor 902, and transceiver 903 are all electrically connected to the bus 904.

[0238] Furthermore, embodiments of this application also provide a computer-readable storage medium for storing a computer program. When executed by a processor, the computer program implements the node scheduling method of any of the preceding embodiments of this application.

[0239] Computer-readable storage media include, but are not limited to, volatile storage media and / or non-volatile storage media. Volatile storage media may include, for example, random access storage media (RAM) and / or cache storage media. Non-volatile storage media may include, for example, read-only storage media (ROM), hard disks, flash memory, optical disks, magnetic disks, etc.

[0240] This application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the node scheduling method of any of the preceding embodiments of this application.

[0241] The basic principles of the embodiments of this application have been described above with reference to specific examples. However, it should be noted that the advantages, benefits, and effects mentioned in the embodiments of this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the embodiments of this application from necessarily employing the aforementioned specific details.

[0242] The block diagrams of devices, apparatuses, devices, and systems involved in the embodiments of this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context explicitly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0243] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0244] It should also be noted that in the systems and methods of this application embodiment, each component or step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the embodiments of this application.

[0245] Various changes, substitutions, and modifications can be made to the technology herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the embodiments of this application is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0246] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use embodiments of this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of embodiments of this application. Therefore, embodiments of this application are not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0247] The above description has been given for illustrative and descriptive purposes. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A node scheduling method, characterized in that, include: Obtain the request characteristics corresponding to at least one user request in the target cluster; Obtain the operating metrics of the target cluster, which are used to represent the operating status of the target cluster; Based on the request characteristics and the operational metrics, determine the adjustment strategy for the target cluster; According to the adjustment strategy, the target cluster is scaled down or expanded.

2. The method according to claim 1, characterized in that, The target cluster includes: pre-populated P nodes and encoded D nodes, and the adjustment strategy includes at least one of the following: Increase or decrease the number of P nodes; Increase or decrease the number of the D nodes.

3. The method according to claim 2, characterized in that, Also includes: If it is necessary to increase the number of P nodes and reduce the number of D nodes, switch the D node to be switched to a P node. Alternatively, if it is necessary to increase the number of D nodes and reduce the number of P nodes, the P node to be switched can be switched to a D node.

4. The method according to any one of claims 1-3, characterized in that, The step of determining the adjustment strategy for the target cluster based on the request characteristics and the operational metrics includes: Based on the request characteristics and the operational metrics, the scenario characteristics corresponding to the target cluster are obtained; Based on the scene characteristics, determine the target scene category from at least one scene category; The adjustment strategy for the target cluster is determined based on the decision rules associated with the target scenario category.

5. The method according to claim 4, characterized in that, The step of determining the adjustment strategy for the target cluster based on the decision rules associated with the target scenario category includes: Obtain the decision parameters associated with the target scene category and the parameter thresholds corresponding to the decision parameters; Determine the decision indicators corresponding to the decision parameters from the operational indicators and the request characteristics; The decision indicators and parameter thresholds corresponding to the decision parameters are compared, and the action strategy of the target cluster is determined based on the comparison results.

6. The method according to any one of claims 1-5, characterized in that, The step of scaling down or scaling up the target cluster according to the adjustment strategy includes: Based on the adjustment strategy, generate operation instructions; The operation command is sent to the target cluster to expand or shrink the target cluster.

7. The method according to claim 6, characterized in that, Sending the operation command to the target cluster to expand or shrink the target cluster includes: When the operation instruction is a node addition instruction, the node is added to the target cluster according to the preset node template and the node's resource configuration information; or If the operation instruction is a node reduction instruction, then the nodes in the target cluster are shut down.

8. The method according to claim 7, characterized in that, Also includes: Establish a shared task queue for the target cluster, wherein the shared task queue includes at least one computing task; In the event of expansion of the target cluster, the newly added nodes in the target cluster are controlled to obtain computing tasks from the shared task queue; When the target cluster is scaled down, the nodes in the target cluster that need to be shut down are controlled to stop acquiring computing tasks from the shared task queue and go offline after completing the remaining computing tasks.

9. The method according to any one of claims 1-8, characterized in that, After scaling down or scaling up the target cluster according to the adjustment strategy, the process further includes: Monitor the node status of each node in the target cluster, including running status or idle status; If the idle time of a node in an idle state is greater than or equal to an idle time threshold, the node in the idle state shall be turned off.

10. A scheduling node, characterized in that, include: A processor and a memory, the memory storing a computer program that is invoked by the processor to execute the node scheduling method according to any one of claims 1-9.