Scheduling method of bastion host cluster, program product, cushion device and storage medium

By acquiring multi-dimensional operational metrics of bastion host cluster nodes to generate a comprehensive weight value, and combining this with scheduling strategies to select target nodes, the problem that traditional load scheduling methods cannot adapt to dynamic business loads is solved, achieving efficient resource utilization and low-latency response.

CN121814850APending Publication Date: 2026-04-07BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional bastion host cluster load balancing methods cannot adapt to dynamically changing business loads, resulting in low resource utilization efficiency and service delays.

Method used

By acquiring multi-dimensional operational metrics of each node in the bastion host cluster, a comprehensive weight value is generated, and the target node is selected in conjunction with the current scheduling strategy to achieve dynamic load scheduling.

Benefits of technology

It improves the resource utilization efficiency and scheduling rationality of bastion host clusters in complex and heterogeneous environments, avoids overload of a single node, and ensures low-latency response and overall availability of services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814850A_ABST
    Figure CN121814850A_ABST
Patent Text Reader

Abstract

The invention provides a bastion host cluster scheduling method, a program product, cushion equipment and a storage medium, and is applied to the technical field of operation and maintenance security, and the bastion host cluster scheduling method comprises the following steps: obtaining a multi-dimensional operation index of each node in a bastion host cluster, the multi-dimensional operation index comprises at least one of the following items: a computing resource index, a network quality index, a service load index and an environmental factor index; for each node, generating a comprehensive weight value of the node based on the multi-dimensional operation index of the node; and determining a target node based on the comprehensive weight value of each node and the current scheduling strategy so as to schedule the operation and maintenance request to the target node. According to the scheme, the multi-dimensional operation indexes such as the computing resources, the network quality, the service load and the environmental factors are obtained and fused in real time, and the comprehensive weight value of the node is generated accordingly, so that the defect of single index in a traditional load scheduling method is overcome, and the dynamically changing service load can be responded in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of operation and maintenance security technology, and more specifically, to a scheduling method, program product, padding device and storage medium for a bastion host cluster. Background Technology

[0002] As enterprises deepen their digital transformation, the scale and complexity of information technology (IT) operations and maintenance are increasing dramatically. Bastion hosts, as the core gateway for operational security, often need to be deployed in clusters to ensure high availability and processing capacity. Furthermore, as enterprise business scales up, the management and operational tasks that bastion host clusters need to handle grow exponentially, thus requiring reasonable scheduling of the bastion host clusters. However, traditional load balancing methods for bastion host clusters often employ static or simple dynamic strategies such as round-robin and least connections, which cannot adapt to dynamically changing business loads. Summary of the Invention

[0003] The purpose of this application is to provide a scheduling method, program product, padding device and storage medium for a bastion host cluster, so as to solve the technical problem that the load scheduling method in the prior art cannot adapt to dynamically changing business loads.

[0004] In a first aspect, embodiments of this application provide a scheduling method for a bastion host cluster, comprising: obtaining multi-dimensional operational metrics for each node in the bastion host cluster, wherein the multi-dimensional operational metrics include at least one of the following: computing resource metrics, network quality metrics, service load metrics, and environmental factor metrics; generating a comprehensive weight value for each node based on the multi-dimensional operational metrics of that node; and determining a target node based on the comprehensive weight values ​​of each node and the current scheduling strategy, so as to schedule maintenance requests to the target node.

[0005] In the above scheme, by acquiring and integrating multi-dimensional operational indicators such as computing resources, network quality, business load and environmental factors in real time, and generating a comprehensive weight value for nodes, the shortcomings of single indicators in traditional load scheduling methods are overcome, thereby enabling timely response to dynamically changing business loads. At the same time, by combining the above comprehensive weight value with dynamic scheduling strategies for node selection, the overall resource utilization efficiency and scheduling rationality of the bastion host cluster in complex heterogeneous environments can be improved.

[0006] In an optional implementation, determining the target node based on the comprehensive weight value of each node and the current scheduling policy includes: determining the node with the smallest comprehensive weight value as the target node. In the above scheme, by consistently scheduling maintenance requests to the node with the smallest comprehensive weight value (indicating the lightest current load or optimal state), the even distribution of cluster load can be achieved quickly and directly, avoiding overload of a single node and ensuring optimal utilization of cluster processing capacity and low-latency service response.

[0007] In an optional implementation, before determining the target node based on the comprehensive weight value of each node and the current scheduling strategy, the method further includes: identifying the task type corresponding to the maintenance request, and determining the current scheduling strategy according to the task type. In the above scheme, by introducing the identification of the task type corresponding to the maintenance request and determining the appropriate scheduling strategy accordingly, the scheduling system acquires task awareness capabilities, laying the foundation for subsequent implementation of differentiated resource scheduling and support measures that match the criticality of the task.

[0008] In an optional implementation, determining the current scheduling strategy based on the task type includes: if the task type is a core management task, then the current scheduling strategy is determined to be a stability-first strategy, wherein the stability-first strategy includes: determining the node with a comprehensive weight value greater than a stability threshold and having established a dedicated resource pool for the core task as the target node, the dedicated resource pool for the core task being formed by reserving a portion of the computing resources of the node; if the task type is a general operation and maintenance task or a batch processing task, then the current scheduling strategy is determined to be a load balancing strategy, wherein the load balancing strategy includes: determining the node with the smallest comprehensive weight value as the target node. In the above scheme, for core management tasks, the stability-first strategy is adopted, which not only requires the target node to have a higher comprehensive weight value (better state), but also requires it to have established a dedicated resource pool for the core task to ensure stable runtime latency and no interference; for general operation and maintenance tasks or batch processing tasks, an efficient load balancing strategy is adopted. This differentiated processing mechanism improves the overall throughput of the cluster while ensuring the service quality and reliability of critical services, achieving a balance between efficiency and stability.

[0009] In an optional implementation, after determining the target node based on the comprehensive weight value of each node and the current scheduling policy to schedule maintenance requests to the target node, the method further includes: for each node, generating an overload threshold for that node based on the multi-dimensional operating indicators of that node; if the node load corresponding to that node is greater than the overload threshold, then suspending the scheduling of maintenance requests for that node. In the above scheme, by generating and updating the overload threshold for each node in real time, and suspending the scheduling of new requests to that node when the node load exceeds the threshold, the system can adaptively prevent node overload crashes and trigger load balancing, thereby improving the cluster's self-protection capability, overall availability and robustness, and avoiding the spread of local faults.

[0010] In an optional implementation, the multi-dimensional operational metrics include: historical load baseline, current real-time load, and scenario correction factor; generating an overload threshold for each node based on the multi-dimensional operational metrics includes: calculating the overload threshold using the following formula: ; Where, is a node In time The overload threshold within, For nodes The historical load baseline, For nodes In time The current real-time load, For nodes The scene correction factor, , , These are the dynamic adjustment parameters for the historical load baseline, the current real-time load, and the scenario correction factor, respectively. In the above scheme, the historical load baseline (reflecting periodic patterns), the current real-time load (perceiving instantaneous states), and the scenario correction factor (adapting to specific scenarios) are integrated. By dynamically adjusting the weights of these parameters, the overload threshold can be intelligently adjusted according to business modes, real-time states, and operational scenarios. This improves the accuracy and foresight of overload judgment and reduces false alarms and missed alarms.

[0011] In an optional implementation, the multi-dimensional operational metrics include: normalized values ​​of Central Processing Unit (CPU) utilization, normalized values ​​of memory usage, network quality score, and load change trends; generating a comprehensive weight value for each node based on the multi-dimensional operational metrics includes: calculating the comprehensive weight value using the following formula: ; in, For nodes In time The comprehensive weight value within, For nodes In time The CPU utilization normalized value mentioned above, For nodes In time The normalized value of memory usage within, For nodes In time The network quality score mentioned above, For nodes In time The load change trend within, , , , These are the normalized values ​​of CPU utilization, memory usage, network quality score, and load change trends, respectively, which are dynamically adjusted parameters. In this scheme, CPU, memory, network quality, and load change trends are quantified and integrated, and dynamic parameters are used to reflect the relative importance of each indicator at different times and in different scenarios. This allows for a more accurate and sensitive depiction of the real-time health and load pressure of nodes, providing high-quality, quantifiable decision-making basis for upper-layer scheduling strategies.

[0012] Secondly, embodiments of this application provide a scheduling device for a bastion host cluster, comprising: an acquisition module, configured to acquire multi-dimensional operational metrics of each node in the bastion host cluster, wherein the multi-dimensional operational metrics include at least one of the following: computing resource metrics, network quality metrics, service load metrics, and environmental factor metrics; a first generation module, configured to generate a comprehensive weight value for each node based on the multi-dimensional operational metrics of that node; and a determination module, configured to determine a target node based on the comprehensive weight value of each node and the current scheduling strategy, so as to schedule maintenance requests to the target node.

[0013] In the above scheme, by acquiring and integrating multi-dimensional operational indicators such as computing resources, network quality, business load and environmental factors in real time, and generating a comprehensive weight value for nodes, the shortcomings of single indicators in traditional load scheduling methods are overcome, thereby enabling timely response to dynamically changing business loads. At the same time, by combining the above comprehensive weight value with dynamic scheduling strategies for node selection, the overall resource utilization efficiency and scheduling rationality of the bastion host cluster in complex heterogeneous environments can be improved.

[0014] In an optional implementation, the determining module is specifically used to: determine the node with the smallest comprehensive weight value as the target node. In the above scheme, by always scheduling maintenance requests to the node with the smallest comprehensive weight value (indicating the lightest current load or optimal state), the even distribution of cluster load can be achieved quickly and directly, avoiding overload of a single node and ensuring optimal utilization of cluster processing capacity and low-latency service response.

[0015] In an optional implementation, the scheduling device of the bastion host cluster further includes an identification module, used to identify the task type corresponding to the maintenance request and determine the current scheduling strategy based on the task type. In the above scheme, by introducing the identification of the task type corresponding to the maintenance request and determining the appropriate scheduling strategy accordingly, the scheduling system acquires task awareness capabilities, laying the foundation for subsequent implementation of differentiated resource scheduling and protection measures that match the criticality of the task.

[0016] In an optional implementation, the identification module is specifically used to: if the task type is a core management task, determine the current scheduling strategy as a stability-first strategy, wherein the stability-first strategy includes: determining the node with a comprehensive weight value greater than a stability threshold and having established a dedicated resource pool for the core task as the target node, the dedicated resource pool for the core task being formed by reserving a portion of the computing resources of the node; if the task type is a general operation and maintenance task or a batch processing task, determine the current scheduling strategy as a load balancing strategy, wherein the load balancing strategy includes: determining the node with the smallest comprehensive weight value as the target node. In the above scheme, for core management tasks, the stability-first strategy not only requires the target node to have a higher comprehensive weight value (better state), but also requires it to have established a dedicated resource pool for the core task to ensure stable runtime latency and no interference; for general operation and maintenance tasks or batch processing tasks, an efficient load balancing strategy is adopted. This differentiated processing mechanism improves the overall throughput of the cluster while ensuring the service quality and reliability of critical services, achieving a balance between efficiency and stability.

[0017] In an optional implementation, the scheduling device of the bastion host cluster further includes: a second generation module, used to generate an overload threshold for each node based on the multi-dimensional operating indicators of that node; and a pause module, used to pause the scheduling of the operation and maintenance request for that node if the node load corresponding to that node exceeds the overload threshold. In the above scheme, by generating and updating the overload threshold for each node in real time, and pausing the scheduling of new requests to that node when the node load exceeds the threshold, the system can adaptively prevent node overload crashes and trigger load balancing, thereby improving the cluster's self-protection capability, overall availability and robustness, and avoiding the spread of local faults.

[0018] In an optional implementation, the multi-dimensional operating metrics include: historical load baseline, current real-time load, and scenario correction factor; the second generation module is specifically used to: calculate the overload threshold using the following formula: ; Where, is a node In time The overload threshold within, For nodes The historical load baseline, For nodes In time The current real-time load, For nodes The scene correction factor, , , These are the dynamic adjustment parameters for the historical load baseline, the current real-time load, and the scenario correction factor, respectively. In the above scheme, the historical load baseline (reflecting periodic patterns), the current real-time load (perceiving instantaneous states), and the scenario correction factor (adapting to specific scenarios) are integrated. By dynamically adjusting the weights of these parameters, the overload threshold can be intelligently adjusted according to business modes, real-time states, and operational scenarios. This improves the accuracy and foresight of overload judgment and reduces false alarms and missed alarms.

[0019] In an optional implementation, the multi-dimensional operational metrics include: normalized CPU utilization, normalized memory usage, network quality score, and load change trend; the first generation module is specifically used to: calculate the comprehensive weight value using the following formula: ; in, For nodes In time The comprehensive weight value within, For nodes In time The CPU utilization normalized value mentioned above, For nodes In time The normalized value of memory usage within, For nodes In time The network quality score mentioned above, For nodes In time The load change trend within, , , , These are the normalized values ​​of CPU utilization, memory usage, network quality score, and load change trends, respectively, which are dynamically adjusted parameters. In this scheme, CPU, memory, network quality, and load change trends are quantified and integrated, and dynamic parameters are used to reflect the relative importance of each indicator at different times and in different scenarios. This allows for a more accurate and sensitive depiction of the real-time health and load pressure of nodes, providing high-quality, quantifiable decision-making basis for upper-layer scheduling strategies.

[0020] Thirdly, embodiments of this application provide a computer program product, including computer program instructions, which, when read and executed by a processor, perform the scheduling method for a bastion host cluster as described in the first aspect.

[0021] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a bus; the processor and the memory communicate with each other via the bus; the memory stores computer program instructions that can be executed by the processor, and the processor can execute the scheduling method of the bastion host cluster as described in the first aspect by calling the computer program instructions.

[0022] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a computer, cause the computer to perform the scheduling method for the bastion host cluster as described in the first aspect.

[0023] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, embodiments of this application are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart illustrating a scheduling method for a bastion host cluster provided in an embodiment of this application; Figure 2 A structural block diagram of a scheduling device for a bastion host cluster provided in an embodiment of this application; Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] As enterprise business scales up, the management and maintenance tasks that bastion host clusters need to handle grow exponentially. Existing technologies have two prominent technical problems: First, traditional load balancing strategies mainly use static algorithms such as round-robin and least connections, which cannot adapt to dynamically changing business loads; second, existing scheduling schemes do not fully consider differences in task characteristics and changes in the operating environment, and cannot cope with complex and ever-changing real-world business scenarios.

[0027] In view of this, embodiments of this application provide a scheduling method for a bastion host cluster. Through multi-dimensional real-time indicator collection and dynamic weight calculation, it achieves accurate load assessment and allocation. Furthermore, based on real-time fault prediction and self-healing mechanisms, it can significantly improve system reliability and operational efficiency. The technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0028] Please refer to Figure 1 , Figure 1 A flowchart illustrating a scheduling method for a bastion host cluster provided in this application embodiment is shown. This method can be, but is not limited to, executed by an electronic device. Figure 3 The possible structure of this electronic device is shown below; for details, please refer to the following section. Figure 3 The above-mentioned scheduling method for bastion host clusters can specifically include: S101: Obtain multi-dimensional operational metrics for each node in the bastion host cluster.

[0029] A bastion host cluster refers to a system collection consisting of multiple bastion host (also known as an operations and maintenance security audit system) nodes connected via a network, providing a unified operations and maintenance access point for centralized management and auditing of IT asset operations. A node in a bastion host cluster refers to an independent physical server or virtual machine with full bastion host functionality. It can be understood that a bastion host cluster typically includes multiple nodes, and the number of nodes can be adjusted appropriately according to the actual scenario.

[0030] Multi-dimensional operational metrics refer to a set of data that measures the operational status and external environment of the aforementioned nodes from different perspectives. As one implementation method, multi-dimensional operational metrics may include at least one of the following: computing resource metrics, network quality metrics, service load metrics, and environmental factor metrics. For example, multi-dimensional operational metrics may only include service load metrics, or they may simultaneously include computing resource metrics, network quality metrics, service load metrics, and environmental factor metrics.

[0031] Computational resource metrics are parameters used to quantify the hardware resource capacity of a bastion host node. They directly reflect the saturation level of the node's basic capabilities for performing computations, processing data, and running programs. For example, computational resource metrics may include at least one of the following: CPU utilization, memory utilization, and disk input / output (I / O) throughput. CPU utilization refers to the percentage of time the node's CPU is not idle within a specific time period; memory utilization refers to the proportion of the node's physical memory (RAM) that is used by the operating system and applications; and disk I / O throughput reflects the level of activity of the node reading and writing data from persistent storage.

[0032] Network quality metrics are a set of parameters used to measure the performance and stability of the communication link between a bastion host node and the external network (including the user side, the managed asset side, and other cluster nodes), and are used to evaluate the node's network responsiveness. For example, network quality metrics may include at least one of the following: network round-trip latency, jitter, and packet loss rate. Network round-trip latency refers to the time required for a data packet to travel from the source to the destination and back; jitter refers to the degree of variation in latency; and packet loss rate refers to the percentage of data packets lost during transmission out of the total number of packets sent.

[0033] Business load metrics are parameters that are directly measured at the bastion host application layer and reflect the actual operational workload it currently bears. They are used to reflect the current operational load pressure on a node. For example, business load metrics may include at least one of the following: concurrent connections, request processing time, and request processing rate. Concurrent connections refer to the total number of operational protocol connections established and maintained on the current node; request processing time refers to the average time spent processing a single typical request; and request processing rate refers to the number of various operation requests successfully processed by the node per unit time.

[0034] Environmental factor indicators refer to monitoring parameters related to the physical operating environment of the bastion host node that may affect its long-term stability and reliability, and are used to avoid hardware failures. For example, environmental factor indicators may include at least one of the following: node geographical location, data center temperature, and power status. Here, node geographical location refers to the physical data center or region where the node is located; data center temperature refers to the physical environmental parameters of the rack or data center where the node is located; and power status refers to the availability of backup power, battery level, etc.

[0035] It should be noted that the embodiments of this application do not specifically limit the specific implementation methods for obtaining multi-dimensional operational indicators, and those skilled in the art can make appropriate adjustments according to the actual situation. For example, multi-dimensional operational indicators can be received from external devices; or, multi-dimensional operational indicators can be read from local or cloud sources; or, a distributed proxy network can be established on each bastion host node to collect the multi-dimensional operational indicators of the nodes at a frequency of seconds, providing an accurate data source for subsequent weight calculation.

[0036] S102: For each node, generate a comprehensive weight value for that node based on its multi-dimensional operational metrics.

[0037] The comprehensive weight value is a numerical value calculated from multiple operational metrics to comprehensively quantify the current load and health status of a node. It is typically designed so that a lower value indicates a more idle or healthier node. For example, after obtaining the multi-dimensional operational metrics for each node, a dynamic weight calculation model can be invoked for each node. This model first normalizes each metric to eliminate the influence of dimensions, and then calculates a comprehensive weight value based on preset or dynamically adjusted weight coefficients.

[0038] As one implementation method, an adaptive weighting algorithm based on entropy weighting can be used. This algorithm calculates the weights of various multi-dimensional operational indicators using entropy weighting and combines this with dynamically adjusted parameters to achieve accurate calculation of the node's overall weight. This solves the problem of fixed indicator weights in traditional weighting calculations, which cannot adapt to dynamic load scenarios. For example, assuming that the multi-dimensional operational indicators include normalized CPU utilization, normalized memory usage, network quality score, and load change trend, then the above S102 can specifically include: S201: Calculate the overall weight value using the following formula: ; in, For nodes In time The overall weight value within, For nodes In time The normalized value of CPU utilization within the range, For nodes In time Normalized memory usage within the memory, For nodes In time Network quality rating within the network, For nodes In time The load change trend within, , , , These are the normalized values ​​for CPU utilization, normalized values ​​for memory usage, network quality score, and dynamic adjustment parameters for load change trends.

[0039] The CPU utilization normalization value refers to the linear normalization of the collected raw CPU utilization (0%-100%) to a normal value. Range; Memory usage normalization value refers to normalizing the proportion of used memory to total memory to a range. The range; network quality score refers to the score calculated based on indicators such as network latency and packet loss rate. For example, the lower the latency and the less packet loss, the higher the score (closer to 1); load change trend refers to the parameter obtained by calculating the first derivative or difference of the recent load. A positive value indicates that the load is increasing, and a negative value indicates that it is decreasing.

[0040] The dynamic adjustment parameters are the weighting coefficients corresponding to the four indicators mentioned above, and their sum is 1. For example, the above dynamic adjustment parameters can be dynamically determined based on the real-time network status or through algorithms such as the entropy weight method; for example, when the system is more concerned about computational pressure, the dynamic adjustment parameters of the normalized value of CPU utilization and the normalized value of memory usage can be increased; when the network fluctuates, the dynamic adjustment parameters of the network quality score can be increased.

[0041] S103: Determine the target node based on the comprehensive weight value of each node and the current scheduling policy, so as to schedule the operation and maintenance request to the target node.

[0042] A scheduling strategy refers to a set of predefined logical rules used to select a target node based on inputs such as a comprehensive weight value. A target node is the specific bastion host node selected to handle the current operation and maintenance request after the scheduling decision; while an operation and maintenance request is a network request initiated by a user to establish an operation and maintenance session or to perform a specific operation and maintenance operation.

[0043] For example, when a new operation and maintenance request arrives at the cluster entry point, the most suitable target node can be calculated and selected based on the latest comprehensive weight value of all nodes and the currently effective scheduling policy (i.e., the current scheduling policy). For example, if the minimum weight policy is adopted, the node with the smallest comprehensive weight value is selected. After the target node is determined, the above operation and maintenance request can be forwarded to that node, and the node will take over the subsequent authentication, authorization, session establishment and auditing work.

[0044] By periodically comparing the comprehensive weight values ​​of each node, the optimal target node can be identified. That is, the node with the lightest comprehensive load and the best state is selected after the real-time status of each node is quantitatively evaluated by the dynamic weight calculation model.

[0045] It should be noted that the embodiments of this application do not specifically limit the specific implementation method for determining the current scheduling strategy, and those skilled in the art can make appropriate adjustments according to the actual situation. For example, a fixed scheduling strategy can be pre-configured as the current scheduling strategy; or, multiple scheduling strategies can be pre-configured, and a scheduling strategy can be dynamically selected as the current scheduling strategy according to the actual situation.

[0046] In the above scheme, by acquiring and integrating multi-dimensional operational indicators such as computing resources, network quality, business load and environmental factors in real time, and generating a comprehensive weight value for nodes, the shortcomings of single indicators in traditional load scheduling methods are overcome, thereby enabling timely response to dynamically changing business loads. At the same time, by combining the above comprehensive weight value with dynamic scheduling strategies for node selection, the overall resource utilization efficiency and scheduling rationality of the bastion host cluster in complex heterogeneous environments can be improved.

[0047] Furthermore, based on the above embodiments, S103 may specifically include: S301: The node with the smallest overall weight value is determined as the target node.

[0048] For example, upon receiving an operation and maintenance request, the latest comprehensive weight value of all nodes can be iterated through. By directly comparing the values ​​of the comprehensive weight values, the node with the smallest comprehensive weight value is selected as the target node. The smallest comprehensive weight value means that, after considering factors such as CPU, memory, and network, this node is evaluated as having the lightest load, the most abundant resources, or the healthiest state. Therefore, new operation and maintenance requests can be preferentially assigned to it, thereby achieving the most effective balanced distribution of the overall cluster load.

[0049] In the above scheme, by always scheduling maintenance requests to the node with the smallest comprehensive weight value (indicating the lightest current load or the best state), the cluster load can be distributed evenly quickly and directly, avoiding overload of a single node and ensuring optimal utilization of cluster processing capacity and low-latency service response.

[0050] Furthermore, based on the above embodiments, prior to S103, the scheduling method for the bastion host cluster provided in this application embodiment may further include: S401: Identify the task type corresponding to the operation and maintenance request, and determine the current scheduling strategy based on the task type.

[0051] The task type corresponding to an operation and maintenance request refers to the nature or importance level of the operation and maintenance operation associated with the request, such as core management tasks, general operation and maintenance tasks, batch processing tasks, etc. It should be noted that the specific implementation methods for identifying task types in this application are not specifically limited, and those skilled in the art can make appropriate adjustments according to actual circumstances. For example, the task type corresponding to an operation and maintenance request can be determined by parsing the protocol, target asset Internet Protocol (IP) / port, user identity, or specific tags carried in the request; for example, a request to access the core production database can be identified as a core management task, while a request to download logs can be identified as a batch processing task.

[0052] In the above scheme, by introducing the identification of the task type corresponding to the operation and maintenance request, and determining the appropriate scheduling strategy accordingly, the scheduling system has the ability to perceive tasks, which lays the foundation for the subsequent implementation of differentiated resource scheduling and guarantee measures that match the criticality of tasks.

[0053] Furthermore, based on the above embodiments, S401 may specifically include: S501: If the task type is a core management task, then the current scheduling strategy will be set to the stability priority strategy.

[0054] The stability-first strategy involves identifying nodes with a comprehensive weight value greater than the stability threshold and which have established dedicated resource pools for core tasks as target nodes. In other words, the core objective of the stability-first strategy is not simply to find the lightest-loaded node, but to find a node that is both in good and stable condition and can provide resource guarantees for critical tasks.

[0055] The stability threshold is a preset weighted threshold (e.g., 0.7) used to select nodes with higher overall weighted values, i.e., those with excellent overall performance. The core task-specific resource pool is a portion of computing resources (e.g., 20% CPU time slice, 20% available memory) pre-allocated on the target node's operating system or virtualization layer through resource control technology. This portion of resources is locked and exclusively used by core management tasks scheduled to this node; other non-core tasks cannot occupy it.

[0056] For example, when the task type is identified as a core management task, a stability-first strategy can be implemented: First, from all nodes, select nodes whose comprehensive weight value is greater than the stability threshold (indicating that the current state of the node is good enough) to form a candidate set; second, from this candidate set, further select those nodes that have successfully established and enabled the core task-specific resource pool; finally, from the final candidate nodes, select one as the target node according to certain rules (such as the highest weight value or the most remaining resources in the resource pool).

[0057] This is because, for core management tasks, stability and reliability are more important than simple performance metrics. Nodes with high overall weight values ​​have typically been running for a long time, are stable, have robust monitoring and protection mechanisms, and have a low probability of system anomalies; while nodes with low weight values ​​may be newly deployed nodes whose stability has not been verified, or idle nodes that may have undiscovered potential problems, making their ability to cope with sudden traffic surges uncertain.

[0058] S502: If the task type is a general operation and maintenance task or a batch processing task, then the current scheduling policy will be determined as a load balancing policy.

[0059] Its load balancing strategy includes identifying the node with the lowest overall weight value as the target node. For example, when the task type is identified as a general operation and maintenance task or a batch processing task, the load balancing strategy can be directly adopted, that is, selecting the node with the lowest overall weight value to maximize the cluster's throughput and resource utilization.

[0060] In the above scheme, a stability-first strategy is adopted for core management tasks. This requires not only a higher overall weight value for the target node (better state) but also the establishment of a dedicated resource pool for the core task to ensure stable latency and prevent interference. For general maintenance tasks or batch processing tasks, an efficient load balancing strategy is used. This differentiated processing mechanism improves the overall throughput of the cluster while ensuring the service quality and reliability of critical businesses, achieving a balance between efficiency and stability.

[0061] Furthermore, based on the above embodiments, a differentiated service quality assurance mechanism can be constructed based on the three-dimensional feature labels of the task profile and the output results of the scheduling decision engine to ensure the priority execution of core tasks. That is, for core management tasks, resources are prioritized for allocation; for example, the scheduling strategy allocates the nodes with the highest weights first. For general operation and maintenance tasks, resources are allocated evenly; for example, the scheduling strategy allocates nodes with medium weights. For batch processing tasks, idle resources are allocated; for example, the scheduling strategy allocates nodes with lower weights that have not triggered the overload threshold.

[0062] Furthermore, based on the above embodiments, after S103, the scheduling method for the bastion host cluster provided in this application embodiment may further include: S601: For each node, generate the overload threshold for that node based on its multi-dimensional operational metrics.

[0063] The overload threshold is a dynamically calculated critical value used to determine whether a node is about to or has already been overloaded. For example, the overload threshold can be calculated periodically (e.g., every 5 seconds) for each node based on its multi-dimensional performance metrics.

[0064] As one implementation method, an adaptive threshold algorithm based on entropy weighting can be used. This algorithm calculates the weights of various multi-dimensional operational indicators using entropy weighting and combines this with dynamically adjusted parameters for real-time correction. This achieves accurate and dynamic generation of node overload thresholds, effectively addressing the limitations of traditional threshold settings, such as fixed thresholds and inability to adapt to business fluctuations and dynamic load scenarios, thus improving system response accuracy and robustness. For example, assuming the multi-dimensional operational indicators include historical load baselines, current real-time load, and scenario correction factors, then the aforementioned S601 could specifically include: S701: Calculate the overload threshold using the following formula: ; Where, is a node In time Overload threshold within, For nodes Historical load baseline, For nodes In time The current real-time load, For nodes Scene correction factor, , , These are the dynamic adjustment parameters for the historical load baseline, the current real-time load, and the scenario correction factor, respectively.

[0065] The historical load baseline refers to the average load calculated based on data from the same period over the past 7 days; the current real-time load refers to the real-time load of a node determined by a combination of indicators such as CPU, memory, and network; the scenario correction factor is a dynamic parameter used to quantify and adjust the impact of specific external or internal events or business models on system load thresholds or resource allocation strategies. For example, the scenario correction factor can be adjusted based on special scenarios such as business peaks or fault drills. The dynamic adjustment parameters are weighted coefficients corresponding to the above three indicators, and their sum is 1. For example, the above dynamic adjustment parameters can be dynamically determined based on operation and maintenance strategies or through algorithms such as the entropy weight method.

[0066] S602: If the node load corresponding to this node is greater than the overload threshold, then the scheduling and maintenance requests for this node will be suspended.

[0067] Node load can be a comprehensive load score calculated in real time, or it can be based on a key metric (such as CPU utilization). For example, the real-time load of each node can be continuously monitored; when the node load of a node is detected to be greater than its current overload threshold, a protection mechanism is immediately triggered.

[0068] As one implementation, the node can be marked as overloaded, and the scheduling decision module can temporarily exclude the node in subsequent scheduling decisions and no longer assign new maintenance requests to it. As another implementation, some of the non-core maintenance sessions that have been established on the node can be migrated to other lightly loaded nodes to help them deload quickly.

[0069] For example, in peak business scenarios, the overload threshold of 0.74 can be calculated based on the historical load baseline of the past 7 days (0.68), real-time CPU utilization (85%), and business peak correction factor (weight 0.2). When the node load is detected to rise to 0.78, subsequent operation and maintenance requests will not be allocated to that node.

[0070] In the above scheme, by generating and updating the overload threshold for each node in real time, and pausing the scheduling of new requests to a node when the node load exceeds the threshold, the system can adaptively prevent node overload and crash, trigger load balancing, thereby improving the cluster's self-protection capability, overall availability and robustness, and avoiding the spread of local failures.

[0071] The following describes the scheduling method for the bastion host cluster provided in the embodiments of this application. A national commercial bank's Beijing data center bastion host cluster faces the problem of insufficient processing capacity of three nodes during peak business hours.

[0072] First, multi-dimensional operational metrics for each node were collected at a sampling frequency of 500ms / time. The specific multi-dimensional operational metrics are as follows: Node A: CPU utilization 85%, memory usage 70%, network latency 15ms, concurrent connections 120; Node B: CPU utilization 45%, memory usage 50%, network latency 12ms, concurrent connections 80; Node C: CPU utilization 90%, memory usage 75%, network latency 18ms, concurrent connections 150.

[0073] Secondly, for each node, a comprehensive weight value is generated based on the multi-dimensional operational metrics of that node, using the following formula: ; The calculation result is: The overall weight value corresponding to node A: ; The overall weight value corresponding to node B: ; The overall weight value corresponding to node C: .

[0074] Finally, based on the comprehensive weight values ​​of each node, the scheduling strategy is determined as follows: new access tasks are preferentially assigned to node B (lowest weight); the load of nodes A and C is dynamically adjusted, and the subsequent 20% of asset maintenance is allocated to node B; weight changes are monitored in real time to ensure that the weight deviation of each node is less than 0.2.

[0075] Therefore, in the scheduling method of the bastion host cluster provided in this application embodiment, a dynamic perception mechanism is established, and an adaptive weight calculation model is constructed by collecting operating indicators in real time; a dynamic threshold calculation model is used to dynamically update the load threshold to ensure allocation accuracy; based on the dynamic weight calculation results, a load balancing execution mechanism for real-time decision-making and overload protection is constructed to ensure the rationality and stability of load allocation for remote clusters; and a multi-dimensional feature analysis model is designed to combine features such as task type, business criticality, and resource requirements to achieve intelligent scheduling decisions.

[0076] Please refer to Figure 2 , Figure 2 This application provides a structural block diagram of a scheduling device for a bastion host cluster. The scheduling device 800 may include: an acquisition module 801, used to acquire multi-dimensional operating indicators of each node in the bastion host cluster, wherein the multi-dimensional operating indicators include at least one of the following: computing resource indicators, network quality indicators, service load indicators, and environmental factor indicators; a first generation module 802, used to generate a comprehensive weight value for each node based on the multi-dimensional operating indicators of that node; and a determination module 803, used to determine a target node based on the comprehensive weight value of each node and the current scheduling policy, so as to schedule maintenance requests to the target node.

[0077] In the above scheme, by acquiring and integrating multi-dimensional operational indicators such as computing resources, network quality, business load and environmental factors in real time, and generating a comprehensive weight value for nodes, the shortcomings of single indicators in traditional load scheduling methods are overcome, thereby enabling timely response to dynamically changing business loads. At the same time, by combining the above comprehensive weight value with dynamic scheduling strategies for node selection, the overall resource utilization efficiency and scheduling rationality of the bastion host cluster in complex heterogeneous environments can be improved.

[0078] Furthermore, based on the above embodiments, the determining module 803 is specifically used to: determine the node with the smallest comprehensive weight value as the target node.

[0079] In the above scheme, by always scheduling maintenance requests to the node with the smallest comprehensive weight value (indicating the lightest current load or the best state), the cluster load can be distributed evenly quickly and directly, avoiding overload of a single node and ensuring optimal utilization of cluster processing capacity and low-latency service response.

[0080] Furthermore, based on the above embodiments, the scheduling device 800 of the bastion host cluster further includes: an identification module, used to identify the task type corresponding to the operation and maintenance request, and determine the current scheduling strategy according to the task type.

[0081] In the above scheme, by introducing the identification of the task type corresponding to the operation and maintenance request, and determining the appropriate scheduling strategy accordingly, the scheduling system has the ability to perceive tasks, which lays the foundation for the subsequent implementation of differentiated resource scheduling and guarantee measures that match the criticality of tasks.

[0082] Furthermore, based on the above embodiments, the identification module is specifically used for: if the task type is a core management task, then the current scheduling strategy is determined to be a stability priority strategy, wherein the stability priority strategy includes: determining the node with the comprehensive weight value greater than the stability threshold and having established a core task dedicated resource pool as the target node, the core task dedicated resource pool being formed by reserving a portion of the computing resources of the node; if the task type is a general operation and maintenance task or a batch processing task, then the current scheduling strategy is determined to be a load balancing strategy, wherein the load balancing strategy includes: determining the node with the smallest comprehensive weight value as the target node.

[0083] In the above scheme, a stability-first strategy is adopted for core management tasks. This requires not only a higher overall weight value for the target node (better state) but also the establishment of a dedicated resource pool for the core task to ensure stable latency and prevent interference. For general maintenance tasks or batch processing tasks, an efficient load balancing strategy is used. This differentiated processing mechanism improves the overall throughput of the cluster while ensuring the service quality and reliability of critical businesses, achieving a balance between efficiency and stability.

[0084] Furthermore, based on the above embodiments, the scheduling device 800 of the bastion host cluster further includes: a second generation module, used to generate an overload threshold for each node based on the multi-dimensional operating indicators of the node; and a pause module, used to pause the scheduling of the operation and maintenance request for the node if the node load corresponding to the node is greater than the overload threshold.

[0085] In the above scheme, by generating and updating the overload threshold for each node in real time, and pausing the scheduling of new requests to a node when the node load exceeds the threshold, the system can adaptively prevent node overload and crash, trigger load balancing, thereby improving the cluster's self-protection capability, overall availability and robustness, and avoiding the spread of local failures.

[0086] Furthermore, based on the above embodiments, the multi-dimensional operating indicators include: historical load baseline, current real-time load, and scenario correction factor; the second generation module is specifically used to: calculate the overload threshold using the following formula: ; Where, is a node In time The overload threshold within, For nodes The historical load baseline, For nodes In time The current real-time load, For nodes The scene correction factor, , , These are the dynamic adjustment parameters for the historical load baseline, the current real-time load, and the scene correction factor, respectively.

[0087] The above solution integrates historical load baseline (reflecting periodic patterns), current real-time load (perceiving instantaneous states), and scenario correction factors (adapting to special scenarios). By dynamically adjusting the weights through parameters, the overload threshold can be intelligently adjusted according to business mode, real-time status, and operational scenarios, thereby improving the accuracy and foresight of overload judgment and reducing false alarms and false negatives.

[0088] Furthermore, based on the above embodiments, the multi-dimensional operational indicators include: normalized CPU utilization, normalized memory usage, network quality score, and load change trend; the first generation module 802 is specifically used to: calculate the comprehensive weight value using the following formula: ; in, For nodes In time The comprehensive weight value within, For nodes In time The CPU utilization normalized value mentioned above, For nodes In time The normalized value of memory usage within, For nodes In time The network quality score mentioned above, For nodes In time The load change trend within, , , , These are the normalized CPU utilization value, the normalized memory usage value, the network quality score, and the dynamic adjustment parameters for the load change trend, respectively.

[0089] The above scheme quantifies and integrates the trends of CPU, memory, network quality, and load changes, and uses dynamic parameters to reflect the relative importance of each indicator at different times and in different scenarios. This allows for a more accurate and sensitive characterization of the real-time health and load pressure of nodes, providing a high-quality and quantifiable basis for upper-level scheduling strategies.

[0090] Please refer to Figure 3 , Figure 3 This application provides a structural block diagram of an electronic device 900, which includes at least one processor 901, at least one communication interface 902, at least one memory 903, and at least one communication bus 904. The communication bus 904 enables direct communication between these components, the communication interface 902 facilitates signaling or data communication with other node devices, and the memory 903 stores machine-readable instructions executable by the processor 901. When the electronic device 900 is running, the processor 901 communicates with the memory 903 via the communication bus 904. When the machine-readable instructions are invoked by the processor 901, the aforementioned bastion host cluster scheduling method is executed.

[0091] For example, the processor 901 in this embodiment of the application can read a computer program from the memory 903 via the communication bus 904 and execute the computer program to implement the following method: obtaining multi-dimensional operating indicators of each node in the bastion host cluster, wherein the multi-dimensional operating indicators include at least one of the following: computing resource indicators, network quality indicators, service load indicators, and environmental factor indicators; for each node, generating a comprehensive weight value for that node based on the multi-dimensional operating indicators of that node; determining a target node based on the comprehensive weight value of each node and the current scheduling policy, so as to schedule the operation and maintenance request to the target node.

[0092] The processor 901 comprises one or more, and can be an integrated circuit chip with signal processing capabilities. The processor 901 can be a general-purpose processor, including a CPU, microcontroller unit (MCU), network processor (NP), or other conventional processor; it can also be a special-purpose processor, including a neural network processing unit (NPU), graphics processing unit (GPU), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Furthermore, when there are multiple processors 901, some can be general-purpose processors, and others can be special-purpose processors.

[0093] The memory 903 includes one or more, which may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0094] Understandable. Figure 3 The structure shown is for illustrative purposes only; the electronic device 900 may also include components that are more advanced than those shown. Figure 3 The more or fewer components shown, or having the same Figure 3 The different configurations shown. Figure 3The components shown can be implemented using hardware, software, or a combination thereof. In the embodiments of this application, the electronic device 900 can be, but is not limited to, physical devices such as desktop computers, laptops, smartphones, smart wearable devices, and in-vehicle devices, or virtual devices such as virtual machines. Furthermore, the electronic device 900 is not necessarily a single device; it can be a combination of multiple devices, such as a server cluster, etc.

[0095] This application also provides a computer program product, including a computer program stored on a computer-readable storage medium. The computer program includes computer program instructions. When the computer program instructions are executed by a computer, the computer can perform the steps of the scheduling method for the bastion host cluster described in the above embodiments, such as: S101: Obtaining multi-dimensional operating indicators for each node in the bastion host cluster, wherein the multi-dimensional operating indicators include at least one of the following: computing resource indicators, network quality indicators, service load indicators, and environmental factor indicators. S102: For each node, generating a comprehensive weight value for that node based on its multi-dimensional operating indicators. S103: Determining a target node based on the comprehensive weight values ​​of each node and the current scheduling policy, so as to schedule the operation and maintenance request to the target node.

[0096] This application also provides a computer-readable storage medium that stores computer program instructions. When the computer program instructions are executed by a computer, the computer performs the scheduling method for the bastion host cluster described in the foregoing method embodiments.

[0097] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0098] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0099] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0100] It should be noted that if the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0102] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A scheduling method for a bastion host cluster, characterized in that, include: Obtain multi-dimensional operational metrics for each node in the bastion host cluster, wherein the multi-dimensional operational metrics include at least one of the following: computing resource metrics, network quality metrics, service load metrics, and environmental factor metrics; For each node, a comprehensive weight value for that node is generated based on the multi-dimensional operational metrics of that node; The target node is determined based on the comprehensive weight value of each node and the current scheduling strategy, so that the operation and maintenance request is scheduled to the target node.

2. The scheduling method for a bastion host cluster according to claim 1, characterized in that, The determination of the target node based on the comprehensive weight value of each node and the current scheduling strategy includes: The node with the smallest overall weight value is determined as the target node.

3. The scheduling method for a bastion host cluster according to claim 1, characterized in that, Before determining the target node based on the comprehensive weight values ​​of each node and the current scheduling policy, the method further includes: Identify the task type corresponding to the maintenance request, and determine the current scheduling strategy based on the task type.

4. The scheduling method for a bastion host cluster according to claim 3, characterized in that, Determining the current scheduling policy based on the task type includes: If the task type is a core management task, then the current scheduling strategy is determined to be a stability priority strategy. The stability priority strategy includes: determining the node whose comprehensive weight value is greater than the stability threshold and whose core task dedicated resource pool has been established as the target node. The core task dedicated resource pool is formed by reserving a portion of the computing resources of the node. If the task type is a general operation and maintenance task or a batch processing task, then the current scheduling strategy is determined as a load balancing strategy, wherein the load balancing strategy includes: determining the node with the smallest comprehensive weight value as the target node.

5. The scheduling method for a bastion host cluster according to any one of claims 1-4, characterized in that, After determining the target node based on the comprehensive weight value of each node and the current scheduling policy, and scheduling the operation and maintenance request to the target node, the method further includes: For each node, an overload threshold is generated based on the multi-dimensional operational metrics of that node. If the node load corresponding to the node is greater than the overload threshold, then the scheduling of the operation and maintenance request for the node will be suspended.

6. The scheduling method for a bastion host cluster according to claim 5, characterized in that, The multi-dimensional operational metrics include: historical load baseline, current real-time load, and scenario correction factor; for each node, generating the overload threshold based on the multi-dimensional operational metrics includes: The overload threshold is calculated using the following formula: ; Where, is a node In time The overload threshold within, For nodes The historical load baseline, For nodes In time The current real-time load, For nodes The scene correction factor, , , These are the dynamic adjustment parameters for the historical load baseline, the current real-time load, and the scene correction factor, respectively.

7. The scheduling method for a bastion host cluster according to any one of claims 1-4, characterized in that, The multi-dimensional operational metrics include: normalized CPU utilization, normalized memory usage, network quality score, and load change trend; for each node, a comprehensive weight value is generated based on the multi-dimensional operational metrics, including: The comprehensive weight value is calculated using the following formula: ; in, For nodes In time The comprehensive weight value within, For nodes In time The CPU utilization normalized value mentioned above, For nodes In time The normalized value of memory usage within, For nodes In time The network quality score mentioned above, For nodes In time The load change trend within, , , , These are the normalized CPU utilization value, the normalized memory usage value, the network quality score, and the dynamic adjustment parameters for the load change trend, respectively.

8. A computer program product, characterized in that, It includes computer program instructions, which are read and executed by a processor to perform the scheduling method for a bastion host cluster as described in any one of claims 1-7.

9. An electronic device, characterized in that, include: Processor, memory, and bus; The processor and the memory communicate with each other via the bus; The memory stores computer program instructions that can be executed by the processor, and the processor can execute the scheduling method of the bastion host cluster as described in any one of claims 1-7 by calling the computer program instructions.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a computer, cause the computer to perform the scheduling method for a bastion host cluster as described in any one of claims 1-7.