Service scheduling method and device, electronic equipment and storage medium

By screening similar service cases and candidate nodes in the K8S environment, combining resource status information, selecting the optimal target node and allocating resources, the resource and performance requirements in online service scheduling are resolved, and the scheduling effect and operation quality are improved.

CN120762813APending Publication Date: 2025-10-10DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510879379.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In a co-located Kubernetes online service environment, making scheduling decisions to meet the resource and performance requirements of different services is an urgent issue that needs to be addressed.

Method used

By obtaining the service type and resource requirement information, screening similar service cases, combining the resource status information of candidate nodes, selecting the optimal target node, and configuring the target resources, the scheduling decision is optimized using service indicators and interference prediction models.

Benefits of technology

It achieves scientific scheduling decisions, meets service resource requirements, avoids interference with services already running on nodes, and improves service scheduling effects and operation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762813A_ABST
    Figure CN120762813A_ABST
Patent Text Reader

Abstract

The invention provides a service scheduling method and device, electronic equipment and a storage medium, and relates to the technical field of service management. The method comprises the following steps: acquiring a first service type and first resource demand information of a to-be-scheduled first service; screening matched similar service cases from historical service cases based on the first service type and the first resource demand information; obtaining resource state information of at least one first candidate node; screening a first target node from the at least one first candidate node based on the resource state information, the first resource demand information and the similar service case; and scheduling the first service to the first target node for operation, and configuring a target resource for the first service. According to the method, intelligent scheduling decision of the service can be realized, the requirements of the service for resources and operation performance are met, and the service quality is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of service management technology, and in particular to a service scheduling method, device, electronic device, and storage medium. Background Art

[0002] Containerized colocation refers to the integration and deployment of multiple different types of online services on the same physical servers using container technology. Containers are a lightweight, portable software packaging method that encapsulates an application and all its dependencies, achieving resource isolation and efficient utilization, reducing operating costs and improving server resource utilization. In a colocation environment for Kubernetes online services, making scheduling decisions that meet the resource and performance requirements of different services is a pressing issue. Summary of the Invention

[0003] This application provides a service scheduling method, device, electronic device, and storage medium that can implement intelligent scheduling decisions for services, meet the service's requirements for resources and operating performance, and ensure service quality. The technical solution is as follows:

[0004] According to one aspect of the present application, a service scheduling method is provided, the method comprising:

[0005] Obtaining a first service type and first resource requirement information of a first service to be scheduled;

[0006] Filtering matching similar service cases from historical service cases based on the first service type and the first resource requirement information;

[0007] Obtaining resource status information of at least one first candidate node;

[0008] Filtering a first target node from the at least one first candidate node based on the resource status information, the first resource demand information, and the similar service case;

[0009] The first service is scheduled to run on the first target node, and target resources are configured for the first service.

[0010] According to another aspect of the present application, a service scheduling device is provided, the device comprising:

[0011] A first acquisition module, configured to acquire a first service type and first resource requirement information of a first service to be scheduled;

[0012] A first screening module, configured to screen matching similar service cases from historical service cases based on the first service type and the first resource requirement information;

[0013] A second acquisition module is used to obtain resource status information of at least one first candidate node;

[0014] a second screening module, configured to screen a first target node from the at least one first candidate node based on the resource status information, the first resource demand information, and the similar service case;

[0015] The first scheduling module is used to schedule the first service to run on the first target node and configure target resources for the first service.

[0016] According to one aspect of the present application, an electronic device is provided, comprising: a processor and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the service scheduling method as described above.

[0017] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the service scheduling method as described above.

[0018] According to another aspect of the present application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-mentioned service scheduling method.

[0019] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:

[0020] Target nodes are screened by comprehensively considering the service's resource requirements, the resource status of schedulable nodes, and historical similar service cases. Similar service cases provide historical references for scheduling decisions, making them more scientific. Considering both resource requirements and resource status information can meet service resource requirements while avoiding interference with existing node services, improving both service scheduling effectiveness and operational quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 A flow chart of a service scheduling method according to an exemplary embodiment of the present application is shown;

[0023] Figure 2 A flowchart of another service scheduling method according to an exemplary embodiment of the present application is shown;

[0024] Figure 3 A flowchart of another service scheduling method according to an exemplary embodiment of the present application is shown;

[0025] Figure 4 This is a structural diagram of a service scheduling device provided in an embodiment of the present application;

[0026] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application is shown. DETAILED DESCRIPTION

[0027] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0028] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0029] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; and the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifiers of "one" and "a plurality of" mentioned in this application are illustrative and not restrictive. Those skilled in the art should understand that unless the context clearly indicates otherwise, they should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0030] The following describes the solution of the present application with reference to the accompanying drawings, and illustrates the technical solution provided by the embodiments of the present application in detail through specific embodiments and their application scenarios.

[0031] Please refer to Figure 1, which shows a flow chart of a service scheduling method according to an exemplary embodiment of the present application. This method is applied to the K8s cluster control platform as an example for exemplary description. Figure 1 As shown, the method includes:

[0032] Step 101: Acquire a first service type and first resource requirement information of a first service to be scheduled.

[0033] Step 102: Filter matching similar service cases from historical service cases based on the first service type and the first resource requirement information.

[0034] Step 103: Obtain resource status information of at least one first candidate node.

[0035] Step 104 : Filter a first target node from at least one first candidate node based on the resource status information, the first resource requirement information, and similar service cases.

[0036] Step 105: Schedule the first service to run on the first target node, and configure target resources for the first service.

[0037] When a new first service needs to be scheduled, the intelligent scheduling decision module in the K8s cluster control platform first filters out similar service cases from historical service cases stored in the historical data storage platform based on the first service type and first resource requirement information of the first service (the first resource requirement information can be obtained through the resource requirement description file of the first service).

[0038] The first service type is identified by its metadata tag, and the first resource requirement information includes the first service's requested and limited amounts of resources such as CPU, memory, and network bandwidth. For example, if the new service is a big data analytics service with high CPU resource requirements, its resource requirement description file specifies a CPU request of 4 cores and a memory request of 8GB. The intelligent scheduling decision module searches historical data for similar service cases with similar CPU and memory requirements and a service type tag of big data analytics.

[0039] Optionally, an indexing mechanism may be established to establish a secondary index based on the service type and resource requirements, so as to quickly locate similar service cases related to the first service.

[0040] In step 103, at least one first candidate node for each scheduling target is obtained, so that the optimal scheduling solution can be screened from the first candidate nodes. The first candidate node can be determined based on the current resource status of each existing node. For example, a node whose remaining idle resources can meet the first resource requirement corresponding to the first service can be determined as the first candidate node; or a node whose current load is less than a load threshold can be determined as the first candidate node.

[0041] After determining at least one candidate node, in order to select the optimal scheduling solution from the first candidate nodes, one possible implementation comprehensively considers the resource status information of each first candidate node, the first resource requirement information of the first service, and the resource requirement characteristics of similar service cases. Specifically, based on the resource status information, the first resource requirement information, and similar service cases, the first target node is selected from the at least one first candidate node. This ensures that the selected optimal scheduling solution can meet the resource requirements of the first service and guarantee the operational performance of the first service.

[0042] In one possible implementation, Figure 2 As shown, step 104 may further include steps 104A to 104C.

[0043] Step 104A, for each first candidate node, based on the resource status information, the first resource demand information and similar service cases, determine the candidate resource utilization of the first service on the first candidate node, the candidate performance indicators of the first service running on the first candidate node, and the candidate interference indicators between the first service and the services already running on the first candidate node.

[0044] In order to ensure the service operation quality of the first service when it is scheduled to a specific node, for each first candidate node, the candidate resource utilization of the first service when running on the first candidate node, the candidate performance indicators of the first service running on the first candidate node, and the candidate interference indicators between the first service and the services already running on the first candidate node are predicted in advance based on the resource status information of the first candidate node, the first resource demand information of the first service and the resource demand characteristics of similar service cases. Then, based on the candidate resource utilization, the candidate performance indicators and the candidate interference indicators, the most suitable first target node is selected from each first candidate node.

[0045] Specifically, step 104A may also include steps 104A1 to 104A4.

[0046] Step 104A1 : Determine the candidate resource utilization of the first service on the first candidate node based on the resource status information and the first resource demand information.

[0047] In a possible implementation, the candidate resource utilization of the first service on the first candidate node can be predicted according to the first resource requirement information (such as CPU request amount) of the first service and the resource state information (such as the current CPU idle core number of the first candidate node) of the first candidate node.

[0048] In step 104A2, the resource state information, the first resource requirement information, and the resource usage characteristics corresponding to the similar service case are input into a service index prediction model to obtain a candidate performance index output by the service index prediction model.

[0049] The candidate performance index mainly refers to the response time and throughput of the service. In a possible implementation, the service index prediction model is pre-trained, which can predict the response time, throughput, and other business indexes of the service according to the input resource usage characteristics. The resource state information of the first candidate node, the first resource requirement information of the first service, and the resource usage characteristics of the similar service case are input into the service index prediction model to obtain the candidate performance index predicted by the service index prediction model.

[0050] In a possible implementation, the training process of the service index prediction model is as follows: the resource usage amount (such as CPU, memory, and the like) of the service is taken as an independent variable, and the response time, throughput, and other business indexes of the service are taken as a dependent variable to construct a training sample. For example, n training samples are constructed, and each training sample includes m resource usage characteristics (such as CPU usage rate, memory occupancy, and the like) and a business index value (such as response time). Let the independent variable matrix be X, where X ij represents the jth feature value of the ith training sample, and the dependent variable vector be y, where y i represents the business index value of the ith training sample. At the same time, a linear regression model is constructed, and the goal of the linear regression model is to find a weight vector θ, so that the mean square error (MSE) between the predicted value y' and the true value y is minimized, that is, is minimized. The weight vector θ is updated iteratively by a gradient descent algorithm until the MSE converges, and the trained service index prediction model is obtained. For example, the service index prediction model obtained by training can predict that when the CPU usage rate reaches 80%, the response time of the service will increase by 50%.

[0051] In step 104A3, the first resource requirement information and the first business index of the first service, and the second resource requirement information and the second business index of the running service are input into an interference prediction model to obtain a candidate interference index output by the interference prediction model.

[0052] Among them, the interference indicators include interference degree and interference type. Interference types can be divided into mild interference, moderate interference and severe interference, and interference degrees can be divided into CPU competition interference or memory competition interference, etc. In one possible implementation, an interference prediction model is pre-trained, and the interference prediction model can predict the interference indicators between multiple services based on the resource usage and business indicators of multiple input services. When predicting the interference between the first service and the service already running on the first candidate node, the first resource demand information and the first business indicator of the first service, as well as the second resource demand information and the second business indicator of the service already running on the first candidate node can be input into the interference prediction model to obtain candidate interference indicators such as the interference degree and interference type between the first service and the already running service output by the interference prediction model.

[0053] In one possible implementation, the interference prediction model is trained using a recurrent neural network (RNN) and its variant, the long short-term memory (LSTM) network, to construct a deep learning-based interference prediction model. Taking the LSTM as an example, the input layer receives data sequences of resource usage (e.g., CPU utilization, memory usage) and business metrics (e.g., request response time, throughput) for multiple services. LSTM units control the flow of information through forget gates, input gates, and output gates, effectively handling long-term dependencies in time series. Multiple LSTM units are stacked to form a deep network. The network's output layer, using a fully connected layer and a softmax function, outputs a probability distribution of the type of interference (e.g., CPU contention interference, memory contention interference, etc.) and the degree of interference (e.g., mild interference, moderate interference, severe interference) between services. During the training process, the model is iteratively trained using a large amount of historical data. The backpropagation algorithm calculates gradients and updates the model's parameters, continuously optimizing the model's prediction accuracy.

[0054] To obtain the training samples required for model training, one possible implementation involves deploying a lightweight data collection agent on each node in the Kubernetes cluster. This agent utilizes a distributed architecture and is tightly integrated with the Kubernetes Container Runtime Interface (CRI), enabling efficient and real-time collection of resource usage data for each service container running on the node. For CPU usage and memory usage, this agent interacts with the CRI to obtain operating system-level process resource usage information for the container. For example, in Linux-based Kubernetes nodes, the proc file system is leveraged to read the stat and status files in the proc directory corresponding to each container to accurately obtain data such as CPU usage time and memory resident size. For network I / O traffic data, traffic monitoring is performed at the network stack level. A custom traffic monitoring hook function is mounted on the network interface driver layer of each node. As network packets pass through the network interface, the hook function captures and counts the packet size and number, thereby calculating network I / O traffic. Furthermore, using the Kubernetes API, the kubectl top pods command is used to obtain aggregated CPU and memory usage data for each pod (an abstraction of a service container) for supplemental and verification purposes.

[0055] In order to collect key business indicators of the service, such as request response time and throughput, a distributed tracing and log collection architecture is adopted. For Web services based on the HTTP protocol, a distributed tracing agent, such as OpenTelemetry, is deployed at the front-end access layer of the Web server. When an HTTP request enters the Web service, the distributed tracing agent generates a unique tracing ID for each request and injects it into the request header. As the request flows between the various components within the service, each component records log information containing the tracing ID when processing the request, including the timestamps of the request entering and leaving the component. For throughput collection, traffic statistics are performed at the load balancer level of the Web server, and the throughput is calculated by analyzing the number of requests forwarded per unit time.

[0056] The collected raw data is transferred to the data preprocessing server cluster. On the preprocessing server, an outlier detection algorithm based on statistical methods, such as the 3σ principle, is first used. For each data dimension (such as CPU usage, memory usage, etc.), the mean μ and standard deviation σ of the data are calculated. If a data point x satisfies |x-μ|>3σ, it is determined to be an outlier and removed. For example, for a set of CPU usage data, the calculated mean is 60% and the standard deviation is 5%. If a data point is 85%, then according to the 3σ principle, this data point will be treated as an outlier and processed.

[0057] The data is then normalized using the Z-Score normalization algorithm. For each data point x, the formula for calculating the normalized new data x' is x' = (x - μ) / σ, where μ and σ are the mean and standard deviation of the data dimension, respectively. This allows data of varying magnitudes to be converted to a standard scale with a mean of 0 and a standard deviation of 1, facilitating subsequent data analysis and model training. For example, if a memory usage data point is 500MB, the mean of the memory usage data dimension is 400MB and the standard deviation is 50MB, then the normalized data is (500 - 400) / 50 = 2.

[0058] Finally, the data is aggregated at 5-minute intervals. Using a sliding window aggregation algorithm, the data within a 5-minute window is aggregated. For example, CPU usage data within a 5-minute window is averaged, and the maximum value of multiple memory usage data points is taken. This reduces the data volume and highlights data trends.

[0059] Preprocessed historical data is stored in a big data storage platform using a hybrid storage architecture combining the Hadoop Distributed File System (HDFS) and HBase. HDFS is used to store large amounts of unstructured and semi-structured data, such as raw log data and pre-aggregated data files. Structured data that requires fast random read and write speeds, such as intermediate results and model parameters generated during model training, is stored in HBase. HBase's distributed table structure efficiently supports random read and write operations on massive amounts of data. Its row key design allows for rapid data location and retrieval based on key information such as service ID and time range.

[0060] Data mining tools, such as Apache Mahout, are used to conduct in-depth analysis of historical data. Time series analysis algorithms, such as the Autoregressive Integrated Moving Average (ARIMA) model, are used to analyze the temporal patterns of resource usage peaks and troughs over time, as well as the cyclical changes in resource demand. For a data analysis service, the CPU usage time series data is first tested for stationarity. If the data is not stationary, it is differentiated to stabilize it. The autocorrelation function (ACF) and partial autocorrelation function (PACF) are then used to determine the ARIMA model parameters: p (autoregressive order), d (differentiation order), and q (moving average order). Model training and validation revealed that the CPU usage of this data analysis service peaks between 10 PM and 2 AM daily, exhibiting a clear daily cyclical pattern.

[0061] Step 104B: Determine a candidate scheduling score corresponding to scheduling the first service to the first candidate node based on the candidate resource utilization, the candidate performance indicator, and the candidate interference indicator.

[0062] To comprehensively analyze the pros and cons of each possible scheduling solution, an overall score is calculated based on candidate resource utilization, candidate performance indicators, and candidate interference indicators to comprehensively consider the impact of various factors. In one possible implementation, a candidate scheduling score for the scheduling solution of scheduling the first service to the first candidate node is determined based on the candidate resource utilization, candidate performance indicators, and candidate interference indicators. The first candidate node is then further selected based on this candidate scheduling score.

[0063] When calculating the scheduling score, the resource utilization score, the performance score and the interference degree score are calculated respectively. Specifically, step 104B may also include steps 104B1 to 104B4.

[0064] Step 104B1 : Determine a resource utilization score based on the candidate resource utilization.

[0065] In one possible implementation, the resource utilization score is determined by the candidate resource utilization and a preset full score, and the higher the resource utilization, the higher the score. For example, if the resource utilization score is 70% and the preset full score is 100, the resource utilization score is 100*70%=70 points.

[0066] Step 104B2: Determine a performance operation score based on the candidate response time in the candidate performance indicators and the performance target of the first service.

[0067] The operational performance score is calculated by comparing the predicted response time and throughput with the performance target of the service, with shorter response times and higher throughputs resulting in higher scores. In one possible implementation, the performance operational score is determined based on the candidate response time of the candidate performance indicators and the performance target of the first service.

[0068] For example, if the predicted candidate response time is 50 ms and the performance target of the first service is 80 ms, the running performance is obtained as (1-50 / 80)*100=37.5 points.

[0069] Step 104B3: Determine an interference level score based on the interference level in the candidate interference indicators.

[0070] The interference level score is calculated based on the probability distribution of the interference level output by the LSTM interference prediction model, and the lower the interference level, the higher the score. In one possible implementation, the interference level score is determined based on the interference level in the candidate interference indicator. For example, if the LSTM model predicts that the interference level is mild, the corresponding interference level score is 80 points (assuming that mild interference corresponds to 80 points, moderate interference corresponds to 50 points, and severe interference corresponds to 20 points).

[0071] Step 104B4: Determine the candidate scheduling score corresponding to scheduling the first service to the first candidate node based on the resource utilization score and the first preset weight, the performance operation score and the second preset weight, the interference level score and the third preset weight.

[0072] A comprehensive score formula is provided: Comprehensive score = 0.4 × resource utilization score + 0.4 × operation performance score - 0.2 × interference level score. In one possible implementation, a candidate scheduling score corresponding to scheduling the first service to the first candidate node is determined based on the resource utilization score and a first preset weight, the performance operation score and a second preset weight, and the interference level score and a third preset weight.

[0073] For example, if the resource utilization score is 70 points, the performance operation score is 37.5 points, and the interference level score is 80 points, the candidate scheduling score corresponding to scheduling the first service to the first candidate node is: 0.4×70+0.4×37.5-0.2×80=28+15-16=27 points.

[0074] Step 104C: Filter a first target node from at least one first candidate node based on the candidate scheduling score.

[0075] In a possible implementation, the first candidate node with the highest candidate scheduling score may be directly determined as the first target node.

[0076] Alternatively, an optimization algorithm, such as a genetic algorithm, can be used to search for the optimal solution among numerous possible scheduling solutions. Genetic algorithms iteratively optimize scheduling solutions by simulating the selection, crossover, and mutation operations used in biological evolution. First, each scheduling solution is encoded as a chromosome. The genes in the chromosome can represent information such as the node ID to which the service is assigned and resource configuration parameters. In each iteration, chromosomes are selected based on their comprehensive scores, with the scheduling solution with the highest comprehensive score chosen as the parent. For example, a roulette wheel selection method can be used, where the probability of each chromosome being selected is proportional to its comprehensive score. A crossover operation, such as a single-point crossover, is then performed on the selected parent chromosome. This randomly selects a gene position and swaps the gene segments after that position between the two parent chromosomes to generate a new daughter chromosome. Next, a mutation operation is performed on the newly generated daughter chromosome with a certain mutation probability, such as randomly changing the value of a gene (e.g., assigning the service to another node). This process continues until the optimal scheduling solution with the highest comprehensive score is found. For example, after multiple iterations, the genetic algorithm finds a solution that schedules a new service to a node and allocates a specific resource configuration to it. This solution has the highest overall score and can minimize interference with other services while meeting the service resource requirements.

[0077] After determining the first target node, the intelligent scheduling decision module in the K8s cluster control platform can send the optimal scheduling plan to the K8S scheduler. The K8S scheduler adopts an extensible architecture design and receives custom scheduling strategies and scheduling plans from the intelligent scheduling decision module. The scheduler schedules the first service to the corresponding first target node according to the optimal scheduling plan, and allocates specified target resources to the first service. During the scheduling process, the scheduler interacts with the APIServer of K8S, creates a new Pod (container) object, and schedules it to run on the first target node. At the same time, the scheduler allocates CPU, memory, network bandwidth and other resources to the Pod (container) according to the resource allocation plan to ensure that the first service can run under the optimal resource configuration while avoiding interference with other services.

[0078] In summary, the embodiments of the present application provide a service scheduling method that screens target nodes by comprehensively considering the service's resource demand information, the resource status information of schedulable nodes, and historical similar service cases. Similar service cases provide historical references for scheduling decisions, making them scientific. Considering resource demand information and resource status information can meet service resource requirements while avoiding interference with existing node services, thereby improving service scheduling effectiveness and service operation quality.

[0079] After dispatching the service to the target node, a dynamic adjustment and feedback mechanism is also set up to timely discover scheduling problems and optimize them. Figure 3 FIG. 8 shows a flowchart of another service scheduling method according to an example embodiment of the present application. The method is exemplarily illustrated by taking the application of the K8s cluster control platform as an example. As shown in FIG. 8, the method comprises the following steps. Figure 3

[0080] Step 301: obtaining a first service type and first resource requirement information of a first service to be scheduled;

[0081] Step 302: screening a similar service case matched from historical service cases based on the first service type and the first resource requirement information;

[0082] Step 303: obtaining resource state information of at least one first candidate node;

[0083] Step 304: screening a first target node from the at least one first candidate node based on the resource state information, the first resource requirement information and the similar service case;

[0084] Step 305: scheduling the first service to run on the first target node, and configuring target resources for the first service.

[0085] The implementation of steps 301 to 305 can refer to the above embodiments, and the present embodiment will not be described here.

[0086] Step 306: obtaining real-time running data of the first service during the scheduling of the first service to run on the first target node;

[0087] Step 307: determining a running deviation value of the first service based on the real-time running data and predicted running data;

[0088] Step 308: triggering a dynamic adjustment mechanism in a case where the running deviation value is greater than a preset threshold, the dynamic adjustment mechanism being used for reallocating resources for the first service and / or reallocating a running node for the first service.

[0089] During the running of the first service on the first target node, the real-time state monitoring module continuously monitors the actual running of the first service to obtain real-time running data of the first service, including resource usage, business indicators and interference between services. Then, the actual running data is compared with the predicted running data, and the deviation is measured by calculating indicators such as mean square error (MSE). The predicted running data can include predicted performance indicators output by a service indicator prediction model.

[0090] ​For example, taking the service response time in real-time operation data as an example, if the i-th monitoring data in the predicted operation data is the predicted response time Zi, and the actual response time is Zi', the operation deviation value of the first service is determined by calculating the mean square error between M predicted operation data and the actual operation data.

[0091] After calculating the operation deviation value of the first service, it is determined whether the operation deviation value exceeds an expected range, for example, whether it is greater than a preset threshold. If it is greater than the preset threshold, the dynamic adjustment mechanism is triggered.

[0092] Specifically, the dynamic adjustment method may include the following steps:

[0093] Step 1. When real-time operation data indicates that there is CPU resource competition between the first service and the second service on the first target node, determine the adjusted first occupied CPU core number of the first service and the adjusted second occupied CPU core number of the second service based on the first important weight, the first occupied CPU core number and the first resource demand information of the first service, the second important weight, the second occupied CPU core number and the second resource demand information of the second service, and the total number of CPU cores of the first target node.

[0094] In one possible implementation, if the first service and the second service on the first target node have serious interference due to CPU resource competition, the dynamic adjustment and feedback optimization module can use the following algorithm to make adjustments. First, consider dynamically adjusting resource allocation. Based on the linear programming algorithm, the CPU resources on the node are reallocated while meeting the basic resource requirements of each service. Assume that the total number of CPU cores of the node is C, and the first service and the second service currently occupy C respectively. A and C B The number of cores and interference. Through the linear programming model, the importance weight W of the service is introduced. A (importance weight of the first service) and W B (Importance weight of the second service), the importance weight is set according to business needs, and the objective function is to maximize W A ×performance A +W B ×performance B , performance A Performance indicators for the first service, performance B Performance indicators for the second service; constraints include C A +C B ≤C and the minimum resource guarantee of each service. Solve the model and get a new resource allocation plan, such as adjusting the first service to occupy C A 'Number of cores, the second service occupies C Bthe number of cores to alleviate the competition for CPU resources.

[0095] Step two, if the first service still competes for CPU resources with the second service after adjusting the CPU resources, obtain a second candidate node that the first service can migrate to;

[0096] Step three, obtain candidate migration costs and candidate migration benefits of migrating the first service to each second candidate node;

[0097] Step four, determine a second target node from the second candidate nodes based on the candidate migration costs and the candidate migration benefits.

[0098] Step five, migrate the first service to the second target node for running.

[0099] If resource reallocation cannot effectively solve the interference problem, the dynamic adjustment and feedback optimization module starts the service migration strategy. A migration algorithm based on cost-benefit analysis is used to evaluate the feasibility of migrating the interfered service or the interference source service to other nodes. The candidate migration costs are calculated, including network transmission costs (calculated according to the service data volume and the target node network bandwidth), temporary interruption costs caused by migration to the service (estimated according to the service business nature and the acceptable interruption duration), etc. At the same time, the candidate migration benefits brought by migration are calculated, such as the service performance benefits improved by avoiding interference after migration (evaluated by a prediction model). The cost-benefit analysis is performed for each possible target node, and the node with the highest benefit-cost ratio is selected as the migration target. For example, the first service produces interference with other services in node N1, and after calculation, the benefit-cost ratio of migrating to node N3 is the highest, so the dynamic adjustment and feedback optimization module generates an instruction to migrate the first service to node N3.

[0100] After each dynamic adjustment is completed, the dynamic adjustment and feedback optimization module stores the adjustment results and the real-time running data of the adjusted service as feedback data in the historical data storage platform. The historical data analysis and model training module regularly reads these feedback data from the historical data storage platform to optimize the prediction model and the scheduling strategy. For example, if a certain type of service frequently produces interference in a specific resource configuration and node environment and has been migrated and adjusted, the historical data analysis and model training module will enhance the learning of the characteristics of this type of situation in subsequent interference prediction model training, so that the model can make more accurate predictions for similar interference scenarios. At the same time, in the process of formulating the scheduling strategy, services with such characteristics will be preferentially avoided from being scheduled to environments prone to interference. Through continuous feedback optimization, the system's adaptability to the K8S online service mixed deployment environment is continuously improved, and the interference avoidance effect and overall performance are enhanced.

[0101] In this embodiment, scientific indicators such as mean square error (MSE) are used to accurately measure the deviation between the actual operation of the service and the predicted results. Compared with traditional simple threshold judgment, it can more timely and accurately detect potential interference and performance problems, and effectively trigger a dynamic adjustment mechanism.

[0102] Please refer to Figure 4 , which is a structural diagram of a service scheduling device provided in an embodiment of the present application. For example, Figure 4 As shown, the apparatus 400 includes.

[0103] A first acquisition module 401 is configured to acquire a first service type and first resource requirement information of a first service to be scheduled;

[0104] A first screening module 402 is configured to screen matching similar service cases from historical service cases based on the first service type and the first resource requirement information;

[0105] A second acquisition module 403 is configured to acquire resource status information of at least one first candidate node;

[0106] A second screening module 404 is configured to screen a first target node from the at least one first candidate node based on the resource status information, the first resource demand information, and the similar service case;

[0107] The first scheduling module 405 is configured to schedule the first service to run on the first target node and configure target resources for the first service.

[0108] Optionally, the second screening module 403 is further configured to:

[0109] determining, for each first candidate node, based on the resource status information, the first resource requirement information, and the similar service case, a candidate resource utilization rate of the first service on the first candidate node, a candidate performance indicator of the first service running on the first candidate node, and a candidate interference indicator between the first service and a service already running on the first candidate node;

[0110] Determining, based on the candidate resource utilization, the candidate performance indicator, and the candidate interference indicator, a candidate scheduling score corresponding to scheduling the first service to the first candidate node;

[0111] The first target node is screened from the at least one candidate node based on the candidate scheduling score.

[0112] Optionally, the second screening module 403 is further configured to:

[0113] Determining, based on the resource status information and the first resource demand information, a utilization rate of the candidate resource of the first service on the first candidate node;

[0114] Inputting the resource status information, the first resource demand information, and the resource usage characteristics corresponding to the similar service cases into a service indicator prediction model to obtain the candidate performance indicators output by the service indicator prediction model;

[0115] The first resource requirement information and the first business indicator of the first service, and the second resource requirement information and the second business indicator of the running service are input into an interference prediction model to obtain the candidate interference indicator output by the interference prediction model.

[0116] Optionally, the second screening module 403 is further configured to:

[0117] Determining a resource utilization score based on the candidate resource utilization;

[0118] Determine a performance operation score based on a candidate response time in the candidate performance indicator and a performance target of the first service;

[0119] Determining an interference level score based on the interference level in the candidate interference indicators;

[0120] The candidate scheduling score corresponding to scheduling the first service to the first candidate node is determined based on the resource utilization score and the first preset weight, the performance operation score and the second preset weight, the interference level score and the third preset weight.

[0121] Optionally, the device further comprises:

[0122] A third acquisition module is configured to acquire real-time operation data of the first service when the first service is dispatched to the first target node for operation;

[0123] a first determining module, configured to determine an operation deviation value of the first service based on the real-time operation data and the predicted operation data;

[0124] The trigger module is used to trigger a dynamic adjustment mechanism when the operation deviation value is greater than a preset threshold, and the dynamic adjustment mechanism is used to reallocate resources for the first service and / or reallocate operation nodes for the first service.

[0125] Optionally, the device further comprises:

[0126] The second determination module is used to determine the adjusted first number of occupied CPU cores of the first service and the adjusted second number of occupied CPU cores of the second service based on the first important weight, the first number of occupied CPU cores and the first resource requirement information of the first service, the second important weight, the second number of occupied CPU cores and the second resource requirement information of the second service, and the total number of CPU cores of the first target node, when the real-time operation data indicates that there is CPU resource competition between the first service and the second service on the first target node.

[0127] Optionally, the device further comprises:

[0128] A fourth acquisition module is configured to acquire a second candidate node to which the first service can migrate if the first service still competes with the second service for CPU resources after the CPU resources are adjusted;

[0129] a fifth acquisition module, configured to acquire a candidate migration cost and a candidate migration benefit of migrating the first service to each of the second candidate nodes;

[0130] a third determining module, configured to determine a second target node from the second candidate nodes based on the candidate migration costs and the candidate migration benefits;

[0131] The second scheduling module is used to schedule the first service to run on the second target node.

[0132] The exemplary embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, wherein the computer program, when executed by the at least one processor, causes the electronic device to perform the service scheduling method according to the embodiments of the present application.

[0133] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the service scheduling method according to an embodiment of the present application.

[0134] An exemplary embodiment of the present application further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to enable the computer to execute the service scheduling method according to the embodiment of the present application.

[0135] refer to Figure 5, a block diagram of an electronic device 500 that can serve as a server or client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0136] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0137] Multiple components within electronic device 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into electronic device 500. Input unit 506 can receive input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0138] The computing unit 501 may be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, Figure 1 、 Figure 2 、 Figure 3 The illustrated method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 may be configured to execute the computer program in any other suitable manner (e.g., by means of firmware). Figure 1 、 Figure 2 、 Figure 3 The method shown.

[0139] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0141] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0144] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A service scheduling method, characterized in that: The method comprises: Obtaining a first service type and first resource requirement information of a first service to be scheduled; Filtering matching similar service cases from historical service cases based on the first service type and the first resource requirement information; Obtaining resource status information of at least one first candidate node; Filtering a first target node from the at least one first candidate node based on the resource status information, the first resource demand information, and the similar service case; The first service is scheduled to run on the first target node, and target resources are configured for the first service.

2. The method according to claim 1, characterized in that The selecting a first target node from the at least one first candidate node based on the resource status information, the first resource requirement information, and the similar service case includes: determining, for each first candidate node, based on the resource status information, the first resource requirement information, and the similar service case, a candidate resource utilization rate of the first service on the first candidate node, a candidate performance indicator of the first service running on the first candidate node, and a candidate interference indicator between the first service and a service already running on the first candidate node; Determining, based on the candidate resource utilization, the candidate performance indicator, and the candidate interference indicator, a candidate scheduling score corresponding to scheduling the first service to the first candidate node; The first target node is screened from the at least one first candidate node based on the candidate scheduling score.

3. The method according to claim 2, characterized in that The determining, based on the resource status information, the first resource requirement information, and the similar service case, a candidate resource utilization of the first service on the first candidate node, a candidate performance indicator of the first service running on the first candidate node, and a candidate interference indicator between the first service and a service already running on the first candidate node includes: Determining, based on the resource status information and the first resource demand information, a utilization rate of the candidate resource of the first service on the first candidate node; Inputting the resource status information, the first resource demand information, and the resource usage characteristics corresponding to the similar service cases into a service indicator prediction model to obtain the candidate performance indicators output by the service indicator prediction model; The first resource requirement information and the first business indicator of the first service, and the second resource requirement information and the second business indicator of the running service are input into an interference prediction model to obtain the candidate interference indicator output by the interference prediction model.

4. The method according to claim 2, characterized in that The determining, based on the candidate resource utilization, the candidate performance indicator, and the candidate interference indicator, a candidate scheduling score corresponding to scheduling the first service to the first candidate node includes: Determining a resource utilization score based on the candidate resource utilization; Determine a performance operation score based on a candidate response time in the candidate performance indicator and a performance target of the first service; Determining an interference level score based on the interference level in the candidate interference indicators; The candidate scheduling score corresponding to scheduling the first service to the first candidate node is determined based on the resource utilization score and the first preset weight, the performance operation score and the second preset weight, the interference level score and the third preset weight.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: During the process of dispatching the first service to the first target node for operation, obtaining real-time operation data of the first service; determining an operation deviation value of the first service based on the real-time operation data and the predicted operation data; When the operation deviation value is greater than a preset threshold, a dynamic adjustment mechanism is triggered, and the dynamic adjustment mechanism is used to reallocate resources for the first service and / or reallocate operation nodes for the first service.

6. The method according to claim 5, characterized in that The method further comprises: When the real-time running data indicates that there is CPU resource competition between the first service and the second service on the first target node, based on the first importance weight, the first number of occupied CPU cores and the first resource demand information of the first service, the second importance weight, the second number of occupied CPU cores and the second resource demand information of the second service, and the total number of CPU cores of the first target node, the adjusted first number of occupied CPU cores of the first service and the adjusted second number of occupied CPU cores of the second service are determined.

7. The method according to claim 6, characterized in that The method further comprises: If the first service still competes with the second service for CPU resources after adjusting the CPU resources, obtaining a second candidate node to which the first service can migrate; Obtaining a candidate migration cost and a candidate migration benefit of migrating the first service to each of the second candidate nodes; determining a second target node from the second candidate nodes based on the candidate migration cost and the candidate migration benefit; Migrate the first service to the second target node for execution.

8. A service scheduling device, characterized in that: The device comprises: A first acquisition module, configured to acquire a first service type and first resource requirement information of a first service to be scheduled; A first screening module, configured to screen matching similar service cases from historical service cases based on the first service type and the first resource requirement information; A second acquisition module is used to obtain resource status information of at least one first candidate node; a second screening module, configured to screen a first target node from the at least one first candidate node based on the resource status information, the first resource demand information, and the similar service case; The first scheduling module is used to schedule the first service to run on the first target node and configure target resources for the first service.

9. An electronic device comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, enable the processor to perform the service scheduling method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the service scheduling method according to any one of claims 1 to 7.