Resource scheduling method and system

By obtaining the historical causal relationship of multiple indicators of the computing system and using prediction models to predict future resource requirements, the reasonable resource allocation of the computing system is achieved, the lag and irrationality of dynamic resource scheduling are solved, and the stability and resource utilization of the system are improved.

CN120276815APending Publication Date: 2025-07-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing dynamic resource scheduling methods mainly rely on changes in a single indicator, resulting in unreasonable scheduling lag and resource allocation, and the inability to fully utilize the changes in multiple indicators of the computing system, affecting resource utilization and system stability.

Method used

By obtaining the change information of multiple indicators of the computing system in the historical period and their causal relationship, using the prediction model to predict the change information in the future period, determining the resource scheduling plan, and using frequent scheduling or constant scheduling methods to reasonably allocate computing resources.

Benefits of technology

It improves the accuracy of resource scheduling and the stability of the system, avoids excessive allocation or insufficient resources, and improves resource utilization efficiency and the operation efficiency of the computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276815A_ABST
    Figure CN120276815A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a resource scheduling method and system. In the method, after change information of a plurality of indexes of a computing system in a historical time period is obtained, the change information of the plurality of indexes in a future time period is predicted based on the change information of the plurality of indexes in the historical time period and a causal relationship among the plurality of indexes. And then determining a resource scheduling scheme corresponding to the computing system based on the change information of the plurality of indexes in the future time period, and performing resource scheduling on the computing system according to the resource scheduling scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of cloud computing technology, and in particular, to a resource scheduling method and system. Background Art

[0002] Resource scheduling has wide applications in multiple fields. Taking the cloud computing field as an example, cloud service providers need to dynamically allocate computing resources to a computing system according to the load condition of the computing system. The process of allocating computing resources to the computing system as described above can be referred to as a resource scheduling process. Reasonable and accurate resource scheduling can not only ensure the normal operation of the computing system, but also improve the overall resource utilization rate and reduce the computing cost of cloud service providers.

[0003] Generally, the ways of resource scheduling include static resource scheduling and dynamic resource scheduling. Static resource scheduling generally needs to allocate fixed computing resources to a computing system according to experience. In practical applications, the load of the computing system usually changes dynamically, which makes the way of static resource scheduling have great uncertainty. For example, there may be a situation where resources are surplus in some periods and insufficient in some periods. Compared with static resource scheduling, dynamic resource scheduling can dynamically adjust resources according to the change of a certain load index (such as the number of user requests) of the computing system, so as to realize the elastic allocation and release of computing resources. For example, when it is detected that the number of user requests of the computing system increases, computing resources are added to the computing system; when it is detected that the number of user requests of the computing system decreases, computing resources are reduced from the computing system.

[0004] The content of the background art part is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the filing date of this disclosure, nor does it mean that it can become the prior art of this disclosure. Summary of the Invention

[0005] This specification provides a resource scheduling method and system, which can accurately predict the change information of multiple indicators in a future period based on the change information of multiple indicators in a historical period and the causal relationship between the indicators, so as to accurately allocate computing resources to the computing system and improve the rationality of resource scheduling.

[0006] In a first aspect, this specification provides a resource scheduling method, including: obtaining change information of multiple metrics of a computing system within a historical period, where the multiple metrics describe the operating conditions of the computing system from different dimensions; predicting change information of the multiple metrics within a future period based on the change information of the multiple metrics within the historical period and the causal relationships between the multiple metrics; determining a resource scheduling plan corresponding to the computing system based on the change information of the multiple metrics within the future period, where the resource scheduling plan at least represents the amount of computing resources to be allocated to the computing system within the future period; and performing resource scheduling on the computing system according to the resource scheduling plan.

[0007] In a second aspect, this specification further provides a resource scheduling system, including at least one storage medium and at least one processor. The at least one storage medium stores at least one instruction set for performing resource scheduling. The at least one processor is communicatively connected to the at least one storage medium. Wherein, when the at least one processor runs, it reads the at least one instruction set and executes the method described in the first aspect above according to the instructions of the at least one instruction set.

[0008] In a third aspect, this specification further provides a computer-readable non-volatile storage medium, where at least one instruction set is stored in the computer-readable non-volatile storage medium, and when the at least one instruction set is executed by at least one processor, the resource scheduling method described in any item of the first aspect above is implemented.

[0009] As can be seen from the above technical solutions, for the resource scheduling method and system provided in this specification, after obtaining the change information of multiple metrics within a historical period, it predicts the change information of the multiple metrics within a future period based on the change information of the multiple metrics within the historical period and the causal relationships between the multiple metrics. Then, based on the change information of the multiple metrics within the future period, it determines a resource scheduling plan corresponding to the computing system and performs resource scheduling on the computing system according to the resource scheduling plan. Since the method provided in this specification comprehensively considers the change information of multiple metrics within a historical period and the causal relationships between the multiple metrics when predicting the change information of the multiple metrics within a future period, therefore, the above prediction process can make more reasonable and scientific inferences based on the causal relationships between the multiple metrics, improve the accuracy of the prediction results, and further improve the accuracy of the resource scheduling plan. That is, the resource scheduling plan provided in this specification can better adapt to the operating requirements of the computing system within the future period. Thus, it can be seen that the above solution makes the resource scheduling result more reasonable, thereby ensuring that the computing system runs more stably and efficiently within the future period.

[0010] Other functions of the resource scheduling method and system provided in this specification will be partially listed in the following description. The creative aspects of the resource scheduling method and system provided in this specification can be fully explained through practice or by using the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 FIG. shows a schematic diagram of an application scenario of a resource scheduling provided according to an embodiment of this specification;

[0013] Figure 2 FIG. shows a schematic diagram of the hardware structure of a computing device provided according to some embodiments of this specification;

[0014] Figure 3 FIG. shows a schematic flowchart of a resource scheduling method provided according to an embodiment of this specification;

[0015] Figure 4 FIG. shows a schematic flowchart of a resource scheduling method provided according to another embodiment of this specification;

[0016] Figure 5 FIG. shows a schematic diagram of the structure of a prediction model provided according to an embodiment of this specification; and

[0017] Figure 6 FIG. shows a schematic diagram of frequent scheduling and constant scheduling provided according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the disclosed embodiments, but has the broadest scope consistent with the claims.

[0019] The terms used herein are for the purpose of describing particular example embodiments only and are not limiting. For example, unless the context clearly dictates otherwise, as used herein, the singular forms "a", "an" and "the" may also include the plural forms. When used in this specification, the terms "comprising", "including" and / or "containing" mean that the associated integers, steps, operations, elements and / or components are present, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups in the system / method.

[0020] In view of the following description, these and other features of the present specification, as well as the operations and functions of the relevant elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0021] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations of the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0022] In this specification, the expression "X includes at least one of A, B or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only any one of A, B, C, or may include any combination of A, B, C and other possible contents / elements at the same time. Any combination of A, B, C may be A, B, C, AB, AC, BC, or ABC.

[0023] In this specification, unless clearly stated otherwise, the association relationship generated between structures may be a direct association relationship or an indirect association relationship. For example, when describing "A is connected to B", unless it is clearly stated that A is directly connected to B, it should be understood that A may be directly connected to B or indirectly connected to B; for another example, when describing "A is above B", unless it is clearly stated that A is directly above B (A and B are adjacent and A is above B), it should be understood that A may be directly above B or A may be indirectly above B (there are other elements between A and B and A is above B). And so on.

[0024] It should be noted that the user data obtained in this specification is authorized by the user and does not involve user privacy.

[0025] With the wide application of cloud services, cloud systems, especially heterogeneous cloud systems containing different resource types, have also been widely deployed. Various resources are deployed in cloud systems, such as Central Processing Unit (CPU), memory, storage, etc. The resource scheduling system can reasonably allocate these resources according to the needs of the computing system, ensuring the normal operation of the computing system on the one hand and avoiding resource idleness or waste on the other hand.

[0026] As mentioned above, traditional dynamic resource scheduling schemes dynamically adjust resources according to the change of a certain operating index (such as the number of user requests) of the computing system. For example, when it is detected that the number of user requests of the computing system increases, computing resources are added to the computing system, and when it is detected that the number of user requests of the computing system decreases, computing resources are reduced from the computing system. However, such a scheduling method adjusts the computing resources only after detecting the change of the operating index, resulting in a lag in scheduling. In addition, the above scheduling method only characterizes the operation change of the computing system based on the change of a single index. In fact, the operation change of the computing system is usually related to multiple indexes. Therefore, the above scheduling method cannot fully utilize the change information of different indexes, which limits the available data for resource scheduling and thus affects the rationality of resource scheduling.

[0027] The resource scheduling method provided in this specification can, after obtaining the change information of multiple indexes of the computing system in the historical period, predict the change information of multiple indexes in the future period based on the change information of multiple indexes in the historical period and the causal relationship between multiple indexes. Subsequently, based on the change information of multiple indexes in the future period, a resource scheduling scheme corresponding to the computing system is determined, and the computing system is resource-scheduled according to the resource scheduling scheme. Such a resource scheduling method can obtain multiple indexes of the computing system and also take into account the causal relationship between multiple indexes. The method provided in this specification can fully explore the operation state performance of all aspects of the computing system in the historical period, rather than being limited to a single index. When predicting the resource scheduling scheme of the computing system subsequently, more reasonable and scientific predictions can be made based on the change information and causal relationship of multiple indexes of the computing system, making the decision of resource scheduling more in line with the actual situation of the system and improving the accuracy of resource scheduling. That is to say, the method provided in this specification for resource scheduling can effectively avoid the problems of over-allocation or under-allocation of resources, improve resource utilization efficiency, and ensure that the computing system runs more stably and efficiently in the future period.

[0028] It should be noted that the above description of the application scenario is only one of the multiple usage scenarios provided in this specification. Those skilled in the art should understand that when the resource scheduling method and system provided in this specification are applied to other usage scenarios, their implementation manners and technical effects are similar.

[0029] Figure 1 FIG. 100 is a schematic diagram of an application scenario of a resource scheduling provided according to an embodiment of this specification.

[0030] As Figure 1 shown, the application scenario 100 may include a resource scheduling system 130 and a computing system 110. It should be noted that the number of computing systems 110 may be one or more, and this specification does not limit this. Figure 1 Two computing systems 110 are taken as an example for illustration.

[0031] During the operation of the computing system 110, the resource scheduling system 130 may dynamically allocate computing resources to the computing system 110. For example, in combination with Figure 1 , the resource scheduling system 130 may monitor the computing system 110 to obtain the change information of multiple metrics of the computing system within a historical period. Then, the resource scheduling system 130 may predict the change information of multiple metrics of the computing system 110 within a future period based on the change information of multiple metrics within the historical period and the causal relationship between multiple metrics. Furthermore, the resource scheduling system 130 determines the resource scheduling scheme corresponding to the computing system 110 based on the change information of multiple metrics within the future period, and performs resource scheduling on the computing system 110 according to the resource scheduling scheme.

[0032] It can be understood that in the case where there are multiple computing systems 110 in the application scenario 100, the scheduling process of the resource scheduling system 130 for each computing system 110 is similar.

[0033] The resource scheduling system 130 may be an electronic device with certain computing capabilities. The resource scheduling system 130 may execute the resource scheduling method provided in this specification. The resource scheduling system 130 may store data or instructions for executing the resource scheduling method described in this specification, and may execute or be used to execute the data or instructions. The resource scheduling system 130 may include a hardware device with data information processing capabilities and necessary programs for driving the hardware device to work.

[0034] The resource scheduling system 130 may be a single computing device or a cluster system composed of multiple computing devices.

[0035] In some embodiments, the resource scheduling system 130 may adopt a periodic scheduling strategy. For example, the resource scheduling system 130 may trigger resource scheduling once every preset period. Herein, the preset period may be flexibly adjusted according to the needs of the current scenario. In some embodiments, the resource scheduling system 130 may also adopt an operator-triggered scheduling strategy. For example, when the operator observes that there is a shortage of resources in the computing system 110, or predicts that there may be a shortage of resources in a period of time in the future, the operator sends a resource scheduling request to the resource scheduling system 130 to supplement the computing resources of the computing system 110 in a timely manner.

[0036] In some embodiments, the resource scheduling system 130 may be the computing system 110 itself. In some embodiments, the resource scheduling system 130 may be another system independent of the computing system 110.

[0037] It should be understood that Figure 1 the application scenario shown is only one of the multiple application scenarios provided in this specification. Figure 1 The numbers of the resource scheduling system 130 and the computing system 110 shown are merely exemplary. According to the implementation requirements, any number can be provided.

[0038] Figure 2 FIG. shows a schematic diagram of the hardware structure of a computing device 200 provided according to some embodiments of this specification. The computing device 200 may be used as Figure 1 the resource scheduling system 130 in. In some embodiments, when the resource scheduling system 130 adopts a device cluster, the computing device 200 may be used as any device in the resource scheduling system 130.

[0039] As Figure 2 shown, the computing device 200 includes at least one storage medium 230 and at least one processor 220. In some embodiments, the computing device 200 may further include an internal communication bus 210. In some embodiments, the computing device 200 may further include a communication port 250. In some embodiments, the computing device 200 may further include an I / O component 260.

[0040] The internal communication bus 210 may connect different system components, including the storage medium 230 and the processor 220. The I / O component 260 supports input / output between the computing device 200 and other components.

[0041] The communication port 250 is used for data communication between the computing device 200 and the outside world. For example, the computing device 200 may be connected to a network through the communication port 250.

[0042] The storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. The storage medium 230 further includes at least one instruction set stored in the data storage device. The instruction set is computer program code, and the computer program code may include programs, routines, objects, components, data structures, procedures, modules, etc. for executing the resource scheduling method provided in this specification.

[0043] At least one processor 220 is communicatively connected to at least one storage medium 230 via an internal communication bus 210. The at least one processor 220 is configured to execute the above at least one instruction set. When the system 130 is running, the at least one processor 220 reads the at least one instruction set and executes the resource scheduling method provided in this specification according to the instructions of the at least one instruction set.

[0044] The processor 220 may execute all steps included in the resource scheduling method. The processor 220 may be in the form of one or more processors. The processor 220 may issue execution instructions. The processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.

[0045] For illustrative purposes only, only one processor 220 is shown in the computing device 200 in the present specification. However, it should be noted that the computing device 200 in the present specification may further include multiple processors. Therefore, the operations and / or method steps disclosed in the present specification may be executed by one processor as described in the present specification or jointly executed by multiple processors. For example, if the processor 220 of the computing device 200 in the present specification executes step A and step B, it should be understood that step A and step B may also be jointly or separately executed by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0046] Figure 3 A flowchart showing a resource scheduling method provided according to an embodiment of the present specification is shown; the resource scheduling method P300 may be performed by Figure 1It is executed by the resource scheduling system 130 in []. As Figure 3 shown, the method P300 provided in this specification may include S310 - S370, where:

[0047] S310: Obtain the change information of multiple metrics of the computing system within a historical period, where the multiple metrics describe the operating conditions of the computing system from different dimensions.

[0048] In some possible embodiments, the multiple metrics may include: at least one user - layer metric, at least one resource - layer metric, and at least one service - layer metric. Among them, the user - layer metric describes the operating conditions of the computing system from the dimension of user requests, and the user - layer metric may include: the number of user requests, the number of concurrent users, etc. The resource - layer metric describes the operating conditions of the computing system from the dimension of resource occupancy. The resource - layer metric may include: CPU occupancy rate, GPU occupancy rate, memory occupancy rate, etc. The service - layer metric describes the operating conditions of the computer from the dimension of service capabilities. The service - layer metric may include: response time, throughput, etc.

[0049] Among them, the historical period may be several scheduling cycles before the current moment. The number of scheduling cycles included in the historical period can be flexibly adjusted according to user needs.

[0050] In some embodiments, the change information of the multiple metrics within the historical period may be represented by time - series data. For example, assume that the historical period includes time 1 to time N. The time - series data corresponding to metric 1 includes: the metric values of metric 1 from time 1 to time N. The time - series data corresponding to metric 2 includes: the metric values of metric 2 from time 1 to time N. The time - series data corresponding to metric 3 includes: the metric values of metric 3 from time 1 to time N.

[0051] In the above - mentioned manner, using time - series data to characterize the change information of each metric within the historical period can more intuitively reflect the change trend of each metric itself within the historical period. It can more accurately track the performance of the causal chain in the time series, which helps to more clearly explain the causal relationship between different metrics.

[0052] In the embodiments of this specification, information on the changes of multiple metrics of a computing system within a historical period can be obtained in response to the triggering of a periodic scheduling event. Alternatively, information on the changes of multiple metrics of the computing system within a historical period can also be obtained in response to detecting a resource scheduling request triggered by an operator. That is to say, in order to ensure the normal operation of the computing system, the method provided in this specification adopts multiple scheduling strategies. In addition to periodically scheduling the resources of the computing system, there will also be staff (operators) paying attention to the actual usage of the computing system. If there is a resource shortage, the staff will manually trigger a resource scheduling request to quickly respond to emergencies, timely replenish the resources of the computing system, and avoid a decline in the performance of the computing system or service interruption.

[0053] Among them, the scheduling period of the scheduling event can be flexibly set according to user needs or the requirements of the current scenario. For example, the scheduling period of the scheduling event can be set to any time interval such as 1 hour, 2 hours, 3 hours, 6 hours, 12 hours, etc. Alternatively, the scheduling period of the scheduling event can also be flexibly adjusted according to the date. Taking the shopping scenario as an example: For example, on ordinary dates, the user request volume is relatively stable, and the scheduling period of the scheduling event can be set longer to reduce the system overhead caused by frequent scheduling. On holidays or promotional dates, the user request volume is relatively large, and the scheduling period of the scheduling event can be set shorter so that the system can respond to these changes more timely. For example, the scheduling period of the scheduling event on ordinary dates can be set to 6 hours, and the scheduling period of the scheduling event on promotional dates can be set to 2 hours. Alternatively, different scheduling periods can also be set according to different time periods of each day. For example, the scheduling period at night (when there may be fewer shoppers) can be set longer, and the scheduling period during the day (when there may be more shoppers) can be set shorter. It should be understood that the above embodiments are only for illustrative purposes, and the setting method of the scheduling period of the specific scheduling event and the specific time interval of the scheduling period setting can be flexibly adjusted according to user needs, and are not limited to those given in the above embodiments.

[0054] In practical applications, the periodic scheduling can be combined with the method of the operator triggering a resource scheduling request. It can not only allocate and adjust the resources of the computing system according to the preset scheduling period, but also trigger a resource scheduling request in the way of manual triggering by the operator when a sudden resource shortage occurs. By combining the periodic scheduling and the manual triggering of the operator, the computing system can have corresponding countermeasures when facing various complex situations. Whether it is the resource changes brought about by the daily index changes or the sudden resource crisis, the method provided in this specification can timely adjust the resources of the computing system, ensure the normal operation of the computing system, make the resource allocation of the computing system more accurate, improve the reliability and stability of the computing system, and reduce the risk of computing system failures.

[0055] S330: Predict the change information of multiple metrics in a future period based on the change information of multiple metrics in a historical period and the causal relationships between the multiple metrics.

[0056] Figure 4 The flowchart of a resource scheduling method provided according to another embodiment of this specification is shown. As Figure 4 shown, in some embodiments of this specification, the change information of multiple metrics in a future time period can be predicted based on a pre-trained prediction model. The prediction method can be: input the change information of multiple metrics in a historical period and the causal relationships between the multiple metrics into the pre-trained prediction model, so as to predict the change information of multiple metrics in a future period through the prediction model. Among them, the prediction model is trained based on multiple training samples, and each training sample includes: the change information of multiple metrics in a first period and the change information of multiple metrics in a second period, and the second period is after the first period.

[0057] To improve the accuracy of prediction, before predicting the change information of multiple future metrics in a future period, this specification also needs to mine the causal relationships between the multiple metrics from the change information of the multiple metrics in the historical period. Since the changes between metrics are often not isolated, the change of one metric may cause the change of other metrics. Therefore, by mining the causal relationships between metrics, the internal laws and logics between metrics can be indicated. Incorporating the causal relationships between multiple metrics into the consideration of prediction can enable the prediction model to capture the changes of various factors more comprehensively, thereby improving the accuracy of the prediction model for predicting the changes of future metrics.

[0058] Among them, the method for obtaining the causal relationships between multiple metrics can be: mining a causal matrix from the change information of multiple metrics in a historical period, and the causal matrix includes the causal relationships between any two of the multiple metrics. The causal relationships between any two metrics can be clearly presented in the causal matrix, forming a complete relationship network, which helps the prediction model to subsequently determine the internal structure between metrics as a whole, avoiding the situation that the prediction model only focuses on the local part and ignores the overall situation, and providing a comprehensive perspective for in-depth analysis and understanding of complex computing systems. Subsequently, the method of this specification also needs to perform denoising processing and / or dimensionality reduction processing on the causal matrix to obtain the causal relationships between multiple metrics.

[0059] In this specification, removing noise and outliers from the causal matrix can purify the causal matrix, reduce interference factors in the causal matrix, enable the causal matrix to more truly reflect the actual relationships between indicators, and improve the accuracy and credibility of causal analysis. The causal relationships in the denoised causal matrix will become more accurate and reliable, facilitating the analysis and understanding of subsequent prediction models. For a prediction model constructed based on causal relationships, its stability and generalization ability will be significantly improved.

[0060] In this specification, dimensionality reduction processing of the causal matrix can reduce the data dimension in the causal matrix, lowering subsequent computational requirements and storage needs. Additionally, performing dimensionality reduction processing on the causal matrix can extract key features and main information in the causal matrix, enhancing the pertinence and effectiveness of subsequent causal analysis.

[0061] In some other possible embodiments, singular value decomposition can be used to perform dimensionality reduction on the causal matrix. During the process of using singular value decomposition to perform dimensionality reduction on the causal matrix, due to the properties of the algorithm itself, generally, during the process of performing dimensionality reduction on the causal matrix, denoising of the causal matrix can be achieved, simplifying the processing flow and improving the processing efficiency of the causal matrix.

[0062] Figure 5 The schematic structural diagram of a prediction model provided by the embodiments of this specification is shown, as Figure 5 shown, the prediction model includes a first prediction branch and a second prediction branch. Among them, the first prediction branch can be a causal prediction branch, and the second prediction branch can be a spurious prediction branch. The change information of multiple indicators in the future time period is predicted through the first prediction branch. During the training process of the prediction model, the second prediction branch provides interference information to the first prediction branch to improve the ability of the first prediction branch to perform indicator prediction based on the causal relationships between indicators.

[0063] Continuing as Figure 5 shown, the first prediction branch includes a first aggregator and a first predictor, that is, the causal prediction branch includes a causal aggregator and a causal predictor. The second prediction branch includes a second aggregator and a second predictor, that is, the spurious prediction branch includes a spurious aggregator and a spurious predictor.

[0064] The training process of the prediction model may include: inputting the change information of multiple metrics within a first time period and the causal relationships between the multiple metrics into a first aggregator to generate a first aggregated feature. Subsequently, inputting the first aggregated feature into a first predictor to obtain a first prediction result, and generating a first loss function based on the difference between the first prediction result and the change information of the multiple metrics within a second time period. Inputting the change information of the multiple metrics within the first time period and the causal relationships between the multiple metrics into a second aggregator to generate a second aggregated feature, and inputting the second aggregated feature into a second predictor to obtain a second prediction result, and generating a second loss function based on the difference between the second prediction result and the random metric change information. Interfering with the first aggregated feature based on the second aggregated feature to obtain a third aggregated feature, and inputting the third aggregated feature into the first predictor to obtain a third prediction result, and generating a third loss function based on the difference between the third prediction result and the first prediction result. Taking minimizing the sum of the first loss function, the second loss function, and the third loss function as the training objective, updating the parameters of the prediction model.

[0065] In some possible embodiments, the first loss function, the second loss function, and the third loss function may each have a preset weight. During the training process of the prediction model, the parameters of the prediction model may be updated by taking minimizing the weighted sum of the first loss function, the second loss function, and the third loss function as the training objective.

[0066] Among them, the first aggregated feature represents the aggregated representation of the prediction model for multiple metrics and some causal relationships under normal circumstances, and is used to generate a relatively accurate first aggregated result. The first aggregated feature can intuitively reflect the accuracy of the prediction result, facilitating the optimization algorithm to perform calculations and converge. The second aggregated feature represents the false information or interference factors in the prediction model, and is used to measure the error or loss of the prediction model when processing these false information, prompting the prediction model to better identify and eliminate these interference / false factors, thereby improving the accuracy and reliability of the prediction model.

[0067] In some possible embodiments, the manner of interfering with the first aggregated feature based on the second aggregated feature may be: shuffling the second aggregated feature, and splicing the shuffled second aggregated feature and the first aggregated feature to obtain a third aggregated feature. Among them, the manner of shuffling the second aggregated feature may be random shuffling. This way of obtaining the third aggregated feature increases the diversity of the data by creating a combined feature (the third aggregated feature) different from the original features (the first aggregated feature and the second aggregated feature), avoids overfitting of the prediction model, and improves the robustness of the prediction model.

[0068] In some possible embodiments, by introducing a second aggregated feature to interfere with the first aggregated feature, various uncertainties and noise effects that data may be subject to in the real world can be simulated. Interfering with the first aggregated feature by the second aggregated feature is equivalent to creating an interference scenario for the prediction model, enabling the prediction model to still make predictions as accurately as possible (obtaining a third prediction result) when facing the interfered feature (the third aggregated feature). If the prediction model can maintain good performance under this interference, that is, the difference between the third prediction result and the first prediction result is small (the above difference is measured and constrained by the third loss function), it indicates that the prediction model has strong robustness, indicating that the prediction model can still accurately capture key information and make effective predictions under the mixed influence of two different feature representations. The prediction model obtained in this way can well handle various interferences and data changes that may occur in practical applications, and will not cause large deviations in the prediction results due to minor changes or noise in the input data.

[0069] In the embodiments of this specification, the second aggregation matrix corresponding to the second aggregator is the complementary matrix of the first aggregation matrix corresponding to the first aggregator. That is, subtracting the first aggregation matrix from the all-ones matrix (a matrix in which all elements are 1) can obtain the second aggregation matrix. Among them, the first aggregation matrix is mainly used to reflect the causal relationship between multiple indicators, and by subtracting it from the all-ones matrix to obtain a pseudo-matrix, relationships irrelevant to the causal relationship (such as secondary relationships or accidental relationships, etc.) can be highlighted.

[0070] If the difference between the first aggregated feature and the second aggregated feature is too small, the prediction model may confuse these two types of information and cannot well improve the overall prediction accuracy. To ensure that the prediction model can clearly distinguish the first aggregated feature and the second aggregated feature during the learning process, in the embodiments of this specification, the difference between the second aggregated feature generated by the second aggregator and the first aggregated feature generated by the first aggregator needs to be greater than or equal to a preset difference. Or rather, the difference between the second aggregated feature and the first aggregated feature is large. It should be understood that the specific setting of the difference between the first aggregated feature and the second aggregated feature can be flexibly adjusted according to user needs, as long as there is a certain difference between the first aggregated feature and the second aggregated feature, and it is not limited to that given in the above embodiments.

[0071] That is to say, in the embodiments of this specification, by setting the difference requirement between the first aggregated feature and the second aggregated feature, it is possible to prevent the prediction model from overfitting the first aggregated feature or underfitting the second aggregated feature. In addition, the difference requirement can also improve the generalization ability of the prediction model and its ability to process complex data. At the same time, the difference requirement can also enhance the robustness and stability of the prediction model, and the adaptability of the trained prediction model is significantly improved, enabling it to handle complex and changing data environments.

[0072] S350: Determine a resource scheduling plan corresponding to the computing system based on the change information of multiple metrics in a future time period. The resource scheduling plan at least represents the amount of computing resources to be allocated to the computing system in the future time period.

[0073] Wherein, the future time period may include S scheduling cycles, S is an integer greater than 1, and the resource scheduling plan also represents the scheduling method adopted in the future time period. The scheduling method may be frequent scheduling or constant scheduling. Among them, under frequent scheduling, the amount of computing resources corresponding to each of the S scheduling cycles may be different. Under constant scheduling, the amount of computing resources corresponding to each of the S scheduling cycles is the same.

[0074] The method provided in this specification can flexibly determine different scheduling methods according to different actual situations. For example, if the change amount of multiple metrics in the future time period is stable, the constant scheduling method is adopted to maintain a relatively fixed resource configuration, avoid waste caused by frequent resource adjustment, and improve the utilization efficiency of resources. Another example is that if the change amount of multiple metrics in the future time period fluctuates greatly, the frequent scheduling method is adopted to perform real-time and flexible allocation for different scheduling cycles to prevent resource idleness or shortage.

[0075] In the embodiments of this specification, after determining the scheduling method, the amount of computing resources allocated to each of the S scheduling cycles can be determined based on the scheduling method and the change information of multiple metrics in the future time period.

[0076] Figure 6 FIG. shows a schematic diagram of frequent scheduling and constant scheduling provided according to an embodiment of this specification, as Figure 6 shown. Taking the future S scheduling cycles including scheduling cycle i, scheduling cycle i + 1, scheduling cycle i + 2,..., scheduling cycle s as an example for illustration. Let W opt represent the amount of computing resources allocated in a scheduling cycle. Under frequent scheduling, the amounts of computing resources corresponding to scheduling cycle i, scheduling cycle i + 1,..., scheduling cycle s may be different, and the resource scheduling strategies corresponding to each scheduling cycle may be respectively: Under constant scheduling, the amounts of computing resources corresponding to scheduling cycle i, scheduling cycle i + 1,..., scheduling cycle s are the same, all being W opt .

[0077] In some possible embodiments, an index change index may be determined based on the change information of multiple metrics within a future time period. The index change index characterizes the differences of the multiple metrics between different scheduling cycles. Based on the index change index, a scheduling method to be adopted within the future time period is determined. Based on the scheduling method and the change information of the multiple metrics within the future time period, the amounts of computing resources allocated for each of the S scheduling cycles are determined.

[0078] In some possible embodiments, when the index change index is greater than a preset threshold, the scheduling method may be determined as frequent scheduling. When the index change index is greater than the preset threshold, it indicates that the change range of the current metric within the future time period is relatively large. To ensure good resource allocation effects, different amounts of computing resources may be determined for different time periods within the future time period. Alternatively, when the index change index is less than or equal to the preset threshold, the scheduling method may be determined as constant scheduling. When the index change index is less than or equal to the preset threshold, it indicates that the change range of the current metric within the future time period is relatively small. At this time, the same amount of computing resources may be determined for different time periods within the future time period to reduce the resource waste caused by determining the amount of computing resources.

[0079] Among them, the index change index includes: the change index corresponding to each of the multiple metrics. Determining the scheduling method to be adopted within the future time period may include at least one of the following: when the change index of at least one of the multiple metrics is greater than the preset threshold, determining the scheduling method as frequent scheduling; when the change indexes of all the multiple metrics are less than or equal to the preset threshold, determining the scheduling method as constant scheduling; when the average value of the change indexes of the multiple metrics is greater than the preset threshold, determining the scheduling method as frequent scheduling; or when the average value of the change indexes of the multiple metrics is less than or equal to the preset threshold, determining the scheduling method as constant scheduling. That is to say, the scheduling method may be determined according to the change index of each of the multiple metrics; or the scheduling method may be determined according to the average value of the change indexes of the multiple metrics. It should be understood that the above embodiments are only for illustrative purposes, and the specific determination method of the scheduling method may be flexibly adjusted according to user needs and is not limited to those given in the above embodiments.

[0080] In the case where the scheduling method is frequent scheduling: for the i-th scheduling cycle among the S scheduling cycles, based on the change information O of the multiple metrics within the i-th scheduling cycle i , the required computing resource amount for the i-th scheduling cycle is determined and the required computing resource amount is used as the computing resource amount allocated for the i-th scheduling cycle, where i ranges from 1 to S.

[0081] In the case where the scheduling method is constant scheduling: For the i-th scheduling period among S scheduling periods, the demand computing resource amount for the i-th scheduling period can be determined based on the change information O of multiple metrics within the i-th scheduling period i , and determine the demand computing resource amount for the i-th scheduling period The value of i ranges from 1 to S. Among the demand computing resource amounts for the S scheduling periods, select the maximum value as the computing resource amount W allocated for the S scheduling periods respectively opt .

[0082] In the embodiments of this specification, the quantity of computing resources can be described by W, that is, W is used to represent the total amount of computing resources, and its measurement unit is "piece" or "portion", etc. The value of W fluctuates within a preset range. For example, the value of W fluctuates between [W min , W max .

[0083] When determining the demand computing resource amount within each scheduling period, it is possible to first traverse within a preset range to determine multiple candidate computing resource amounts. For each candidate computing resource amount, based on the change information of multiple metrics within the i-th scheduling period, determine the computing cost generated by the candidate computing resource amount within the i-th scheduling period. The candidate computing resource amount corresponding to the minimum computing cost among the multiple candidate computing resource amounts is determined as the demand computing resource amount for the i-th scheduling period

[0084] The above method for determining the demand computing resource amount for the i-th scheduling period can, on the premise of ensuring that the operation requirements of the computing system are met, minimize the computing cost to the greatest extent. And it can avoid the problem of a large number of computing resource restrictions caused by over-allocation of resources, improving the overall utilization rate of resources. This method for determining the demand computing resource amount can flexibly adjust the allocation of computing resources according to the actual situation, ensuring that the computing system can always operate with the resource configuration of the minimum computing cost, and improving the stability and reliability of the computing system

[0085] When determining the computing cost generated by each candidate computing resource amount, based on the candidate computing resource amount and the change information of multiple metrics within the i-th scheduling period, determine the first probability that the computing system is in a resource slack state under the candidate computing resource amount, and determine the first computing cost generated by the candidate computing resource amount within the i-th scheduling period based on the first probability. Based on the candidate computing resource amount and the change information of multiple metrics within the i-th scheduling period, determine the second probability that the computing system is in a resource tight state under the candidate computing resource amount, and determine the second computing cost generated by the candidate computing resource amount within the i-th scheduling period based on the second probability. Based on the sum of the first computing cost and the second computing cost, determine the computing cost generated by the candidate computing resource amount within the i-th scheduling period

[0086] In some possible embodiments, the first computing cost and the second computing cost each have their corresponding preset weights. For example, W slack may be the preset weight of the first computing cost, and W throt may be the preset weight of the second computing cost. When determining the computing cost, the weighted sum of the first computing cost and the second computing cost can be determined based on W slack and W throt to determine the computing cost generated by the candidate computing resource amount in the i-th scheduling period. When determining the computing cost as described above, by introducing the probability of resource slack and the probability of resource tension, the uncertainty and risk of resource use can be incorporated into the cost calculation. When the probability of resource tension is relatively high, it may lead to additional risk costs, such as processing speed delays and service quality degradation. When the probability of resource slack is relatively high, there may also be some opportunity costs, such as waste caused by resource idleness. By comprehensively considering these factors, a more forward-looking and risk management-conscious decision can be made when selecting the candidate computing resource amount. So that the computing system can adjust its resource status according to changes in multiple indicators such as user requirements and task types in different scheduling periods, thereby making the resource scheduling method of this specification more adaptable to various dynamic environments.

[0087] When this specification determines each candidate computing resource, it not only comprehensively considers the first probability that the computing system is in a resource slack state and the second probability that it is in a resource tension state, but also combines the change information of the corresponding indicators in the i-th scheduling period. This specification takes into account the possible different states of the resources and the occurrence probabilities of different state situations, avoiding the problem of inaccurate determination results caused by determining the cost of the computing resource amount from only a single resource state perspective. The method provided in this specification can more accurately reflect the cost situation of the candidate computing resource amount in different states. Since the costs required for the two states of resource slack and resource tension may vary greatly (for example, when the resources are slack, only less maintenance cost may be required, while when the resources are tense, it may involve emergency deployment costs of resources, etc.), the costs required in these two states of resource slack and resource tension are calculated separately and combined with the corresponding probabilities, making the final computing cost closer to the actual cost. Avoiding the deviation caused by a single cost calculation method improves the overall performance and stability of the computing system.

[0088] For the convenience of understanding, the following content gives the calculation method of the required computing resource amount for each scheduling period in the case where the scheduling method is frequent scheduling :

[0089]

[0090] Among them, the scheduling period can be the i-th scheduling period among 1... S scheduling periods; W is the quantity of computing resources, and its measurement unit is "piece" or "share", etc. U is a unit of computing resources, which represents a benchmark quantity for measuring computing resources. For example, U may be the number of CPU cores corresponding to a unit of computing resources. A unit of computing resources may correspond to 8 CPU cores. Of course, it can also be 4, 6, or any number of CPU cores. This is only an exemplary illustration and is not limited thereto. The WU obtained by multiplying W by U represents the total computing resource quantity measured based on the standard unit U. For example, still taking the example that a unit of computing resources may correspond to 8 CPU cores, and W = 5 (indicating that there are 5 shares of such computing resources), then WU = 5 x 8 = 40 CPU cores, that is, the total computing resource quantity is the computing resource quantity corresponding to 40 CPU cores. O i is the change information of multiple indicators in the i-th scheduling period. P(O i < WU) is the first computing cost for the computing system to be in a resource slack state under the currently allocated computing resource quantity. P(O i > WU) is the second computing cost for the computing system to be in a resource tight state under the currently allocated computing resource quantity. W slack is the weight of the first computing cost; W throt is the weight of the second computing cost.

[0091] In the case where the scheduling method is constant scheduling, the determined method of the required computing resource quantity W opt for each scheduling period can be obtained based on the following algorithm:

[0092]

[0093] Among them is the required computing resource quantity for the 1st scheduling period, is the required computing resource quantity for the S-th scheduling period.

[0094] S370: Perform resource scheduling on the computing system according to the resource scheduling plan.

[0095] In the embodiments of this specification, according to the resource scheduling plan, perform resource scheduling on the computing system at each scheduling moment. Until the trigger of the next periodic scheduling event, or until the resource scheduling request triggered by the operator, then re-trigger the determination of the resource scheduling plan.

[0096] For example, the current periodic scheduling triggers resource scheduling once every 3 hours. Each scheduling predicts the resource scheduling plan for the next 4 hours. For example, if the current time is 12:00 and a resource scheduling for the computing system is triggered, based on the change information of multiple metrics in the historical period, the required computing resource amounts at 13:00, 14:00, 15:00, and 16:00 in the future are predicted. The resource scheduling at 13:00, 14:00, and 15:00 in the future can be carried out according to the predicted required computing resource amounts. At 15:00, since it has been 3 hours since the last resource scheduling trigger, a new round of resource scheduling will be triggered at 15:00. Under the new resource scheduling, the required computing resource amounts at 16:00, 17:00, 18:00, and 19:00 in the future are predicted. At this time, the required computing resource amount at 16:00 is subject to the required computing resource amount determined by the new resource scheduling triggered at 15:00.

[0097] In summary, in the resource scheduling method and system provided in this specification, after obtaining the change information of multiple metrics of the computing system in the historical period, the change information of multiple metrics in the historical period and the causal relationship between multiple metrics can be input into the prediction model to predict the change information of multiple metrics in the future period through the prediction model. Then, based on the change information of multiple metrics in the future period, the corresponding scheduling method of the computing system is determined, and the corresponding resource scheduling plan is determined under the scheduling method. The resource scheduling of the computing system is carried out according to the resource scheduling plan.

[0098] That is to say, by using the resource scheduling method provided in this specification, after obtaining the change information of multiple metrics describing the operation of the computing system from different dimensions in the historical period, the causal relationship between the metrics is mined. And a prediction model including a first prediction branch and a second prediction branch and trained specifically is used to predict the change of multiple metrics in the future period, and then the corresponding resource scheduling plan of the computing system is determined and implemented. Such a resource scheduling method can not only comprehensively grasp the state of the computing system and improve the prediction accuracy. And in the process of determining the resource scheduling plan, intelligent decision-making can be carried out according to the index change index. By distinguishing between frequent scheduling or constant scheduling, different resource scheduling strategies are used to determine the resource amount for each scheduling cycle. When determining the computing resource amount, this specification also comprehensively considers the calculation cost of the resource state probability, realizing the balance between efficient resource utilization and cost control. At the same time, it supports periodic data acquisition or request-triggered data acquisition, adapts to multiple scenarios, and multiple metrics cover user layer metrics, resource layer metrics, and service layer metrics, which can comprehensively reflect the status of the computing system and improve the scientificity, flexibility, and effectiveness of overall resource scheduling.

[0099] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for performing resource scheduling executable instructions. When the at least one instruction set is executed by a processor, the at least one instruction set directs the processor to implement the steps of the resource scheduling method P300 of this specification. In some possible implementation manners, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on the resource scheduling system 130, the program code is used to cause the resource scheduling system 130 to execute the steps of the resource scheduling method P300 described in this specification. The program product for implementing the above method can use a portable compact disc read-only memory (CD-ROM) to include the program code and can run on the resource scheduling system 130. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code included on the readable storage medium can be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages.

[0100] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0101] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0102] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0103] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art, when reading this specification, may very well mark out some of the devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as an integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.

[0104] Each patent, patent application, published patent application, and other materials cited in this disclosure, such as articles, books, specifications, publications, documents, references, etc. (excluding any historical prosecution files associated therewith), are hereby incorporated by reference for all purposes relevant to this disclosure, e.g., in the specification and claims of this disclosure. However, if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terms of the above materials and those used in this disclosure, the descriptions, definitions, and / or terms used in this disclosure shall prevail.

[0105] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely illustrative and not restrictive. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Accordingly, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A resource scheduling method, comprising: Obtaining change information of multiple metrics of a computing system within a historical period, where the multiple metrics describe the operating conditions of the computing system from different dimensions; Predicting the change information of the multiple metrics within a future period based on the change information of the multiple metrics within the historical period and the causal relationships between the multiple metrics; Determining a resource scheduling plan corresponding to the computing system based on the change information of the multiple metrics within the future period, where the resource scheduling plan at least characterizes the amount of computing resources to be allocated to the computing system within the future period; And Performing resource scheduling on the computing system according to the resource scheduling plan.

2. The method according to claim 1, wherein The predicting the change information of the multiple metrics within the future period based on the change information of the multiple metrics within the historical period and the causal relationships between the multiple metrics includes: Inputting the change information of the multiple metrics within the historical period and the causal relationships between the multiple metrics into a pre-trained prediction model to predict the change information of the multiple metrics within the future period through the prediction model, where The prediction model is trained based on multiple training samples, and each training sample includes: the change information of the multiple metrics within a first period and the change information of the multiple metrics within a second period, where the second period is after the first period.

3. The method according to claim 2, wherein, The prediction model includes a first prediction branch and a second prediction branch, where: The change information of the multiple metrics within the future period is predicted through the first prediction branch; During the training process of the prediction model, the second prediction branch provides interference information to the first prediction branch to improve the ability of the first prediction branch to perform metric prediction based on the causal relationships between metrics.

4. The method according to claim 3, wherein The first prediction branch includes a first aggregator and a first predictor, and the second prediction branch includes a second aggregator and a second predictor; the training process of the prediction model includes: Inputting the change information of the multiple metrics within the first period and the causal relationships between the multiple metrics into the first aggregator to generate a first aggregated feature, inputting the first aggregated feature into the first predictor to obtain a first prediction result, and generating a first loss function based on the difference between the first prediction result and the change information of the multiple metrics within the second period; Inputting the change information of the multiple metrics within the first period and the causal relationships between the multiple metrics into the second aggregator to generate a second aggregated feature, inputting the second aggregated feature into the second predictor to obtain a second prediction result, and generating a second loss function based on the difference between the second prediction result and the random metric change information; Interfering with the first aggregated feature based on the second aggregated feature to obtain a third aggregated feature, inputting the third aggregated feature into the first predictor to obtain a third prediction result, and generating a third loss function based on the difference between the third prediction result and the first prediction result; With the objective of minimizing the sum of the first loss function, the second loss function, and the third loss function, update the parameters of the prediction model.

5. The method according to claim 4, wherein, The second aggregation matrix corresponding to the second aggregator is the complementary matrix of the first aggregation matrix corresponding to the first aggregator.

6. The method according to claim 1, wherein Before predicting the change information of the multiple metrics in the future period based on the change information of the multiple metrics in the historical period and the causal relationship between the multiple metrics, the method further includes: Mine the causal relationship between the multiple metrics from the change information of the multiple metrics in the historical period.

7. The method according to claim 6, wherein, The mining of the causal relationship between the multiple metrics from the change information of the multiple metrics in the historical period includes: Mine a causal matrix from the change information of the multiple metrics in the historical period, where the causal matrix includes the causal relationship between any two of the multiple metrics; and Perform denoising processing and / or dimensionality reduction processing on the causal matrix to obtain the causal relationship between the multiple metrics.

8. The method according to claim 1, wherein The future period includes S scheduling cycles, where S is an integer greater than 1. The resource scheduling plan also characterizes the scheduling method adopted in the future period, and the scheduling method is frequent scheduling or constant scheduling. Among them, Under the frequent scheduling, the computing resource amounts corresponding to the S scheduling cycles are different; Under the constant scheduling, the computing resource amounts corresponding to the S scheduling cycles are the same.

9. The method according to claim 8, wherein, The determining of the resource scheduling plan corresponding to the computing system based on the change information of the multiple metrics in the future period includes: Determine an index change index based on the change information of the multiple metrics in the future period, where the index change index characterizes the difference situation of the multiple metrics between different scheduling cycles; Determine the scheduling method adopted in the future period based on the index change index; and Determine the computing resource amounts allocated in the S scheduling cycles respectively based on the scheduling method and the change information of the multiple metrics in the future period.

10. The method according to claim 9, wherein, The determining of the computing resource amounts allocated in the S scheduling cycles respectively based on the scheduling method and the change information of the multiple metrics in the future period includes: In the case where the scheduling method is frequent scheduling, for the i-th scheduling cycle among the S scheduling cycles, determine the required computing resource amount of the i-th scheduling cycle based on the change information of the multiple metrics in the i-th scheduling cycle, and use the required computing resource amount as the computing resource amount allocated in the i-th scheduling cycle, where the value of i ranges from 1 to S.

11. The method according to claim 9, wherein, The determining of the computing resource amounts allocated in the S scheduling cycles respectively based on the scheduling method and the change information of the multiple metrics in the future period includes: In the case where the scheduling method is constant scheduling, for the i-th scheduling cycle among the S scheduling cycles, determine the required computing resource amount of the i-th scheduling cycle based on the change information of the multiple metrics in the i-th scheduling cycle, where the value of i ranges from 1 to S; Among the required computing resource amounts in the S scheduling cycles, select the maximum value as the computing resource amount allocated in each of the S scheduling cycles respectively.

12. The method according to claim 10 or 11, wherein Determining the required computing resource amount in the i-th scheduling cycle based on the change information of the multiple metrics in the i-th scheduling cycle includes: Determine a plurality of candidate computing resource amounts within a preset range; For each candidate computing resource amount, determine the computing cost generated by the candidate computing resource amount in the i-th scheduling cycle based on the change information of the multiple metrics in the i-th scheduling cycle; and Determine the candidate computing resource amount corresponding to the minimum computing cost among the plurality of candidate computing resource amounts as the required computing resource amount in the i-th scheduling cycle.

13. The method according to claim 12, wherein, Determining the computing cost generated by the candidate computing resource amount in the i-th scheduling cycle based on the change information of the multiple metrics in the i-th scheduling cycle includes: Based on the candidate computing resource amount and the change information of the multiple metrics in the i-th scheduling cycle, determine the first probability that the computing system is in a resource slack state under the candidate computing resource amount, and determine the first computing cost generated by the candidate computing resource amount in the i-th scheduling cycle based on the first probability; Based on the candidate computing resource amount and the change information of the multiple metrics in the i-th scheduling cycle, determine the second probability that the computing system is in a resource tension state under the candidate computing resource amount, and determine the second computing cost generated by the candidate computing resource amount in the i-th scheduling cycle based on the second probability; and Determine the computing cost generated by the candidate computing resource amount in the i-th scheduling cycle based on the sum of the first computing cost and the second computing cost.

14. The method according to claim 9, wherein, Determining the scheduling method to be adopted in the future period based on the metric change index includes: When the metric change index is greater than a preset threshold, determine that the scheduling method is frequent scheduling; or When the metric change index is less than or equal to the preset threshold, determine that the scheduling method is constant scheduling.

15. The method according to claim 9, wherein, The metric change index includes the change index corresponding to each of the multiple metrics. Determining the scheduling method to be adopted in the future period based on the metric change index includes at least one of the following: When the change index of at least one of the multiple metrics is greater than the preset threshold, determine that the scheduling method is frequent scheduling; When the change indices of the multiple metrics are all less than or equal to the preset threshold, determine that the scheduling method is constant scheduling; When the average value of the change indices of the multiple metrics is greater than the preset threshold, determine that the scheduling method is frequent scheduling; or When the average value of the change indices of the multiple metrics is less than or equal to the preset threshold, determine that the scheduling method is constant scheduling.

16. The method according to claim 1, wherein Obtaining the change information of multiple metrics of the computing system in the historical period includes: In response to the triggering of a periodic scheduling event, obtain the change information of multiple metrics of the computing system in the historical period; or, In response to detecting a resource scheduling request triggered by an operator, obtain the change information of multiple metrics of the computing system within a historical period.

17. The method according to claim 1, wherein, The multiple metrics include: at least one user layer metric, at least one resource layer metric, and at least one service layer metric.

18. A resource scheduling system, comprising: At least one storage medium storing at least one instruction set for resource scheduling; And At least one processor communicatively connected to the at least one storage medium, wherein when the at least one processor runs, it reads the at least one instruction set and executes the method according to any one of claims 1-17 as described above according to the indication of the at least one instruction set.