Method, apparatus, and program product for managing a computing system

By applying a allocation model based on reinforcement learning in the computing system and combining manual interaction to generate training data, the problem of inefficient computing unit management in the computing system is solved, and more efficient operational allocation and computing system performance improvement is achieved.

CN114816722BActive Publication Date: 2025-06-13EMC IP HLDG CO LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110111133.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-27
Publication Date
2025-06-13
Estimated Expiration
2041-01-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage a large number of computing units in the computing system, resulting in low operational allocation efficiency and difficulty in improving the overall performance of the computing system.

Method used

Using reinforcement learning-based technology, the initial allocation model is constructed, and more effective training data is generated through the manual interaction process, combining machine learning and manual experience to optimize the allocation model.

Benefits of technology

By optimizing the allocation model, operations can be allocated to multiple computing units in the computing system more effectively, improving the operating efficiency and performance of the computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816722B_ABST
    Figure CN114816722B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods, devices, and program products for managing a computing system. In one method, a set of operations to be executed on a plurality of computing units in a computing system is obtained. Based on the set of operations, the states of the plurality of computing units, and an allocation model that describes the association relationship between the set of operations, the states of the plurality of computing units, an allocation action for allocating the set of operations to the plurality of computing units, and the reward of the allocation action, an allocation action for allocating the set of operations to the plurality of computing units and the reward of the allocation action are determined. In response to determining that the matching degree between the reward of the allocation action and the performance metric of the computing system after executing the allocation action meets a predetermined condition, an adjustment to the reward is received. Training data for updating the allocation model is generated based on the adjustment. Corresponding devices and program products are provided. With the exemplary implementation of the present disclosure, training data for the allocation model can be generated in a more efficient manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various implementations of the present disclosure relate to the management of computing systems, and more particularly, to methods, devices, and computer program products for allocating a set of operations to multiple computing units in a computing system. Background Art

[0002] With the development of computer technology, computing systems can include a large number of computing units. For example, a computing system can include one or more computing devices, and each computing device can include one or more central processing units (CPUs) and graphics processing units (GPUs), etc. Further, the CPUs and GPUs can include one or more processor cores. At this time, the computing system will include a large number of computing units, and the computing system can perform a variety of operations. At this time, how to allocate these operations among multiple computing units to improve the overall performance of the computing system has become a research hotspot. Summary of the Invention

[0003] Therefore, it is desirable to develop and implement a technical solution for managing a large number of computing units in a computer system in a more effective manner. It is desirable that this technical solution can allocate the operations to be executed to each computing unit in a more convenient and effective manner, thereby improving the operating efficiency of the computing system.

[0004] According to a first aspect of the present disclosure, there is provided a method for managing a computing system. In this method, a set of operations to be executed on multiple computing units in the computing system is obtained. Based on the set of operations, the states of the multiple computing units, and an allocation model, an allocation action for allocating the set of operations to the multiple computing units and a reward for the allocation action are determined, and the allocation model describes the association relationship between the set of operations, the states of the multiple computing units, the allocation action for allocating the set of operations to the multiple computing units, and the reward for the allocation action. In response to determining that the matching degree between the reward for the allocation action and the performance metric of the computing system after executing the allocation action meets a predetermined condition, an adjustment for the reward is received. Training data for updating the allocation model is generated based on the adjustment.

[0005] According to a second aspect of the present disclosure, there is provided an electronic device, including: at least one processor; a volatile memory; and a memory coupled to the at least one processor, the memory having instructions stored therein, and the instructions, when executed by the at least one processor, cause the device to execute the method according to the first aspect of the present disclosure.

[0006] According to a third aspect of the present disclosure, there is provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions for executing the method according to the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In conjunction with the accompanying drawings and with reference to the following detailed description, the features, advantages, and other aspects of the various implementations of the present disclosure will become more apparent. Several implementations of the present disclosure are shown herein by way of illustration and not limitation. In the drawings:

[0008] Figure 1 A block diagram schematically showing an application environment in which exemplary implementations of the present disclosure can be implemented;

[0009] Figure 2 A block diagram schematically showing a process for managing a computing system according to an exemplary implementation of the present disclosure;

[0010] Figure 3 A flowchart schematically showing a method for managing a computing system according to an exemplary implementation of the present disclosure;

[0011] Figure 4 A block diagram schematically showing a process of using an allocation model for managing a computing system according to an exemplary implementation of the present disclosure;

[0012] Figure 5 A block diagram schematically showing a process for determining a reward that needs to be adjusted according to an exemplary implementation of the present disclosure;

[0013] Figure 6 A block diagram schematically showing a process for managing a computing system according to an exemplary implementation of the present disclosure;

[0014] Figure 7 A block diagram schematically showing a process for filtering a training data set according to an exemplary implementation of the present disclosure; and

[0015] Figure 8 A block diagram schematically showing a device for managing a computing system according to an exemplary implementation of the present disclosure. DETAILED DESCRIPTION

[0016] Preferred implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the implementations set forth herein. On the contrary, these implementations are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0017] As used herein, the term "including" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless otherwise specified, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example implementation" and "an implementation" mean "at least one example implementation". The term "another implementation" means "at least one additional implementation". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0018] For ease of description, first refer to Figure 1 Describe an application environment according to an exemplary implementation of the present disclosure. Figure 1 A block diagram of an application environment 100 in which exemplary implementations of the present disclosure can be implemented is schematically shown. As Figure 1 shown, the computing system 110 may include one or more computing devices 120, and each computing device 120 may include various types of computing units. For example, the computing device 120 may include a CPU-type computing unit 130 and a GPU-type computing unit 140. These computing units may jointly serve the computing system 110 for processing a set of operations 150 executed on the computing system 110.

[0019] Currently, technical solutions for managing the allocation of workloads to individual computing units in a computing system based on machine learning techniques have been proposed. It will be understood that in the environment of workload management, people's needs and the state of the computing system are constantly changing. If new computing units are added to the computing system, the trained model needs to be updated again, resulting in a waste of time and resources. The proposed allocation model based on reinforcement learning involves a huge amount of computation and is thus difficult to use in small computing systems with limited computing power. Further, the training process may involve a large amount of manual labor, and it is difficult to combine the accumulated expert knowledge with reinforcement learning techniques. This results in the unsatisfactory effect of the existing allocation models based on reinforcement learning.

[0020] To address the above deficiencies, according to an exemplary implementation of the present disclosure, a technical solution for managing a computing system is proposed. Specifically, an initial allocation model is constructed based on reinforcement learning techniques. During the further training of the initial allocation model, an artificial interaction process is introduced to manually intervene in the training process based on the knowledge of technical experts, thereby generating training data that is more helpful for improving the performance of the computing system. In this way, the machine learning process can be combined with artificial experience to obtain a more accurate and effective training model.

[0021] Hereinafter, first refer to Figure 2 Provide an overview according to an exemplary implementation of the present disclosure.Figure 2 FIG. schematically shows a block diagram of a process 200 for managing a computing system 110 according to an exemplary implementation of the present disclosure. For ease of description, in the context of the present disclosure, a GPU will be taken as an example of a computing unit to describe the technical solution for managing multiple computing units in a computing system. According to an exemplary implementation of the present disclosure, the computing unit may include, but is not limited to, a computing device, a CPU, a GPU, and processor cores of a CPU and a GPU, and so on. As Figure 2 shown, the states 260 of multiple computing units 140 (e.g., including n computing units) can be obtained, and a set of operations 150 to be assigned (e.g., including m operations) can be obtained. A preliminary trained allocation model 210 can be obtained based on reinforcement learning techniques, and the states 260 of multiple computing units and a set of operations 150 are input to the allocation model 210. The allocation model 210 can then output an allocation action for allocating a set of operations 150 to multiple computing units 140 and a reward 220 associated with the action.

[0022] According to an exemplary implementation of the present disclosure, a filter 230 can be used to determine whether the reward 220 is consistent with the performance metric expected to be obtained in the computing system 110. If the two are inconsistent, a technical expert 270 can be consulted. An adjustment 250 from the technical expert 270 can be received to generate training data 240 for a subsequent further training process. If the two are consistent, the training data 240 can be directly generated based on the reward 220. The training data 240 generated here can be used to train 252 the allocation model 210 in a subsequent process. With the exemplary implementation of the present disclosure, the filter 230 can be implemented based on active learning techniques, thereby greatly reducing the manual labor in the training process. In this way, when the allocation model 210 fails to meet the requirements of the administrator of the computing system 110, an artificial intervention process can be initiated, and then the experience of the technical expert 270 can be fully utilized to improve the accuracy of the allocation model 210.

[0023] Hereinafter, refer to Figure 3 Describe the respective steps of a method according to an exemplary implementation of the present disclosure. Figure 3 FIG. schematically shows a flowchart of a method 300 for managing a computing system 110 according to an exemplary implementation of the present disclosure. At block 310, a set of operations 150 to be executed on multiple computing units in the computing system 110 is obtained. The computing system 110 may include n computing units, and the multiple computing units may be represented by the symbol CU: CU = (U1, U2,..., Un).

[0024] According to an exemplary implementation of the present disclosure, operations can have different granularities. For example, a set of operations can include a code segment, and in this case, each operation can be a line of code. For another example, a set of operations can include a task, and in this case, each operation can include the functions called by the task, and so on. A set of m operations can be represented by the symbol OP: OP = (O1, O2,..., Om). A set of operations 150 to be executed on multiple computing units 140 can be obtained from the task list of the computing system 110.

[0025] At block 320, based on a set of operations 150, the states 260 of the multiple computing units 140, and the allocation model 210, determine the allocation action for allocating the set of operations 150 to the multiple computing units 140 and the reward for the allocation action. Here, the state 260 can include various metrics of each computing unit. For example, processor utilization rate, the number of operations in the waiting queue, processor frequency, etc. The state of each computing unit can be represented by a multi-dimensional vector, and further form a higher-dimensional vector representing the overall state 260 of all the multiple computing units.

[0026] According to an exemplary implementation of the present disclosure, the allocation model 210 can be a machine learning model initially trained based on reinforcement learning techniques. A variety of training techniques that have been developed currently and / or will be developed in the future can be used to obtain the allocation model 210. The allocation model 210 can describe the association relationship between a set of operations, the states of multiple computing units, the allocation action of allocating a set of operations to multiple computing units, and the reward for the allocation action. The allocation action can be represented based on a vector: AC = (P1, P2,..., Pm). The i-th dimension Pi in AC can represent to which computing unit the i-th operation in a set of operations is allocated. For example, the value range of Pi can be defined as [1, n], and in this case, the i-th operation can be allocated to any of the n computing units. For example, the allocation action (1, n,..., 3) can represent: allocating the first operation in a set of operations to the first computing unit, the second operation to the n-th computing unit,..., and allocating the m-th operation in a set of operations to the third computing unit.

[0027] According to an exemplary implementation of the present disclosure, the initial action space of the allocation action can be constructed based on multiple ways. For example, an action space representing all allocation possibilities can be constructed, and in this case, the action space will include n mAn allocation action. The action space can be constructed based on a random manner. Alternatively and / or additionally, the action space can be constructed based on expert knowledge for allocating a set of operations to multiple computing units. Specifically, allocation actions that have been verified to be helpful in improving the overall performance of the computing system 110 can be selected from historical allocation actions to construct the action space. For example, operations can be preferentially allocated to computing units in an idle state, each operation can be distributed as evenly as possible to multiple computing units, and excessive operations can be avoided from being allocated to one computing unit, etc. In the case where the action space has been determined, the rewards for each action in the action space can be obtained to obtain a preliminarily trained allocation model 210.

[0028] In the following, reference will be made to Figure 4 describe more details about the allocation model 210. Figure 4 A block diagram schematically showing a usage process 400 of an allocation model for managing a computing system according to an exemplary implementation of the present disclosure is shown. A preliminarily trained allocation model 210 can be obtained based on labeled training data. Subsequently, the preliminarily trained allocation model 210 can be used to predict the rewards corresponding to the allocation actions that can be performed. As Figure 4 shown, a set of operations 410 and the states 420 of multiple computing units can be input to the allocation model 210. At this time, the allocation model 210 can predict the allocation action 430 and the corresponding reward 440.

[0029] According to an exemplary implementation of the present disclosure, the allocation action 430 can be performed in the computing system 110 to determine the performance metrics of the computing system 110 after the allocation action is performed. Alternatively and / or additionally, a simulator can be used to simulate the execution of the allocation action 430 in the computing system 110, and then a prediction of the corresponding performance metrics can be obtained. In the following, reference will be made to Figure 5 describe more details about the performance metrics. Figure 5 A block diagram schematically showing a process 500 for determining the rewards that need to be adjusted according to an exemplary implementation of the present disclosure is shown. As Figure 5 shown, the performance metrics 510 can include the waiting time 512 of the operations in a set of operations. The waiting time of a set of operations can be represented in a vector manner, and the higher the waiting time, the lower the overall performance metrics 510. The performance metrics 510 can further include the cumulative workload 514 of the computing units in multiple computing units. The cumulative workload of multiple computing units can be represented in a vector manner, and the higher the cumulative workload, the lower the overall performance metrics 510.

[0030] Further, based on the comparison between the reward 440 and the performance metrics, it can be determined whether manual intervention is required. Return Figure 3, at block 330, in response to determining that the match between the reward 440 for the allocation action 430 and the performance metric of the computing system 110 calculated after performing the allocation action 430 meets a predetermined condition, receive an adjustment 250 for the reward 440. According to an exemplary implementation of the present disclosure, a filter 230 can be used to distinguish rewards that need adjustment from those that do not.

[0031] Specifically, the filter 230 can be set based on the direction of the reward and the change direction of the performance metric. For example, the reward 440 can include a positive reward and a negative reward. The positive reward is used to indicate that the allocation action can operate in a direction that helps improve the performance of the computing system 110, and the negative reward is used to indicate that the allocation action can operate in a direction that is harmful to improving the performance of the computing system 110. According to an exemplary implementation of the present disclosure, if the reward is a positive reward and the change direction of the performance metric is "decrease", it is considered that the reward 440 needs to be adjusted. Another example is that if the reward is a negative reward and the change direction of the performance metric is "increase", it is also considered that the reward 440 needs to be adjusted. The filter 230 can determine the reward 520 that needs adjustment based on the above conditions.

[0032] According to an exemplary implementation of the present disclosure, the reward 520 that needs adjustment can be provided to the technical expert 270 so that the technical expert 270 can perform the adjustment 250 based on their own experience. Assuming that the value range of the reward is [-1, 1], the technical expert 270 can adjust the value of the received reward so that the adjusted reward can truly reflect whether the allocation action can manage the computing system 110 in a direction that improves the performance of the computing system 110. For example, if the allocation action causes a decrease in the performance metric 510 and the reward is positive, the value of the reward can be decreased (e.g., setting the reward as a negative reward). Another example is that if the allocation action causes an increase in the performance metric 510 and the reward is negative, the value of the reward can be increased (e.g., setting the reward as a positive reward).

[0033] According to an exemplary implementation of the present disclosure, the technical expert 270 can also modify the action space of the allocation action. For example, an allocation action that seriously degrades the performance of the computing system 110 (e.g., the action of allocating multiple operations to the same computing unit) can be removed from the existing action space. Another example is that a new allocation action that can improve the performance of the computing system 110 (e.g., the action of evenly distributing multiple operations to multiple computing units) can be added to the action space.

[0034] Using the exemplary implementation of the present disclosure, when needed, with the help of the experience of the technical expert 270, the labeled data can be re-provided, and then the new labeled data can be used to train the allocation model 210. The process of determining the reward 520 that needs to be adjusted has been described above. In some cases, the filter 230 can determine the reward 530 that does not need to be adjusted based on the above conditions. At this time, the reward 530 can be directly used to generate training data.

[0035] In the following, it will be returned Figure 3 to describe how to generate the training data 240 based on the adjustment 250. At Figure 3 box 340, the training data 240 for updating the allocation model 210 is generated based on the adjustment 250. Specifically, the adjusted reward, a set of operations 410, the states 420 of multiple computing units, and the allocation actions 430 can be used to generate the training data 240. For another example, the newly added allocation actions and corresponding rewards of the technical expert 270 can be combined with a set of operations 410 and the states 420 of multiple computing units to generate the training data 240.

[0036] According to an exemplary implementation of the present disclosure, the method 300 described above can be iteratively executed in multiple rounds to generate multiple training data 240, and the generated training data 240 can be processed in batches. Figure 6 A block diagram of a process 600 for managing the computing system 110 according to an exemplary implementation of the present disclosure is schematically shown. As Figure 6 shown, the training data 240 generated each time can be stored in the training data set 610, and when the number of training data in the training data set 610 reaches a predetermined number, the training process as shown by the arrow 630 is started.

[0037] According to an exemplary implementation of the present disclosure, in order to accelerate the training process, a filter 620 can be used to process the training data set 610. Specifically, the filtering operation can be performed according to the differences between the respective training data in the training data set 610 and the historical training data. It will be understood that the allocation model 210 is a model that has been preliminarily trained, so the allocation model 210 has already accumulated knowledge related to the historical training data used in the preliminary training. When performing subsequent training, it is more desirable to use training data different from the historical training data to obtain new knowledge in other aspects. Therefore, the filter 620 can be used to filter out the training data similar to the historical training data from the training data set 610.

[0038] In the following, refer to Figure 7 to describe more details about the filtering process. Figure 7A block diagram schematically showing a process 700 for filtering a training data set 610 according to an exemplary implementation of the present disclosure. As Figure 7 shown, the dots represent historical training data and the circles represent new training data in the training data set 610. Various training data can be classified (e.g., based on spatial distance), and the similarity between individual training data can be determined according to the obtained clusters. In Figure 7 , a large amount of historical training data is classified into a cluster 710, and a large amount of new training data is classified into a cluster 720.

[0039] According to an exemplary implementation of the present disclosure, if the difference between the training data and the historical training data is lower than a predetermined threshold, it can be considered that the allocation situation represented by the training data has been covered by the historical training data. As Figure 7 shown, the training data 730 is classified into the cluster 710. Since the allocation model 210 currently already includes the knowledge involved in the training data 730, it is not necessary to use the training data 730 to retrain the allocation model 210. In other words, the training data 730 can be deleted from the training data set 610.

[0040] According to an exemplary implementation of the present disclosure, if it is determined that the difference between the training data and the historical training data exceeds a predetermined threshold, the training data can be retained. In Figure 7 , the training data in the cluster 720 and the training data 732 and 734 have a difference from the historical training data that exceeds the predetermined threshold, so these training data can be retained. It will be understood that due to labeling errors and / or other mistakes, some training data in the training data set 610 may be abnormal. These abnormal data do not perform training in a direction that helps improve the performance of the computing system 110, so these abnormal data need to be deleted. Further filtering can be performed on the retained training data. For example, abnormal data can be removed based on the similarity between the retained training data.

[0041] In Figure 7Among them, a large amount of training data is classified into cluster 720, which indicates that these training data have similarities and can reflect the distribution not covered by historical training data. Therefore, the training data in cluster 720 can be retained. That is to say, multiple training data that meet the following conditions can be retained: there are differences between the multiple training data and the historical training data, and there are similarities between the multiple training data. According to an exemplary implementation of the present disclosure, since the training data 732 and 734 outside cluster 720 do not have similarities, these training data can be removed. Alternatively and / or additionally, the training data 732 and 734 can be further submitted to the technical expert 270 for manual confirmation of whether these training data should be removed. By using the exemplary implementation of the present disclosure, the training data that is not helpful for improving the performance of the computing system 110 can be deleted from the training dataset 610. In this way, the training process can be accelerated and the training efficiency can be improved.

[0042] According to an exemplary implementation of the present disclosure, in order to ensure the integrity of the training dataset 610, a part of the training data similar to the historical training data can be retained. Assuming that it is determined that the training dataset 610 includes 1000 training data similar to the historical training data, a predetermined proportion (for example, 50% or other values) of the training data can be deleted from the training dataset 610. By using the exemplary implementation of the present disclosure, on the one hand, the number of training data can be reduced to improve the training efficiency, and on the other hand, the accuracy of the allocation model 210 can be improved based on strengthening the allocation knowledge related to the historical training data.

[0043] According to an exemplary implementation of the present disclosure, the method 300 described above can be periodically executed until the training process meets a predetermined convergence condition. According to an exemplary implementation of the present disclosure, the method 300 can be re-executed when the number of computing units in the computing system 110 changes. Alternatively and / or additionally, the method 300 can be re-executed when the requirements of the administrator of the computing system 110 change. By using the exemplary implementation of the present disclosure, instead of retraining a new allocation model, the training efficiency can be improved based on active learning and expert knowledge from technical experts.

[0044] The process for determining the training dataset 610 and updating the allocation model 210 based on the training dataset 610 has been described above. Further, the updated allocation model 210 can be used to predict new allocation actions for allocating a newly received set of operations to multiple computing units in the computing system 110. According to an exemplary implementation of the present disclosure, another set of operations to be executed on multiple computing units can be received. Another allocation action for allocating the another set of operations to the multiple computing units can be determined based on the another set of operations, the current state of the multiple computing units, and the updated allocation model.

[0045] According to an exemplary implementation of the present disclosure, the allocation model is a model updated using the training dataset 610. The allocation model can include the latest expert knowledge and can cover a more comprehensive allocation scenario. At this time, the determined allocation action can enable the computing system 110 to operate in a direction that is more conducive to improving performance. The determined allocation action can be executed in the computing system 110 so as to make full use of the available resources in the multiple computing units in the computing system 110 in an optimized manner.

[0046] In the above, it has been referred to Figures 2 to 7 Examples of the method according to the present disclosure have been described in detail above. Implementations of the corresponding apparatus will be described below. According to an exemplary implementation of the present disclosure, there is provided an apparatus for managing a computing system. The apparatus includes: an obtaining module configured to obtain a set of operations to be executed on multiple computing units in the computing system; a determining module configured to determine an allocation action for allocating the set of operations to the multiple computing units and a reward for the allocation action based on the set of operations, the state of the multiple computing units, and an allocation model, the allocation model describing the association relationship between the set of operations, the state of the multiple computing units, the allocation action for allocating the set of operations to the multiple computing units, and the reward for the allocation action; a receiving module configured to receive an adjustment for the reward in response to determining that the matching degree between the reward for the allocation action and the performance metric of the computing system after executing the allocation action meets a predetermined condition; and a generating module configured to generate training data for updating the allocation model based on the adjustment. According to an exemplary implementation of the present disclosure, the apparatus further includes a module for performing other steps in the method 300 described above.

[0047] Figure 8A block diagram of a device 800 for managing data patterns according to an exemplary implementation of the present disclosure is schematically shown. As shown, the device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0048] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0049] Each of the processes and processes described above, such as the method 300, can be executed by the processing unit 801. For example, in some implementations, the method 300 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some implementations, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU 801, one or more steps of the method 300 described above can be executed. Alternatively, in other implementations, the CPU 801 can also be configured in any other appropriate manner to implement the above process / method.

[0050] According to an exemplary implementation of the present disclosure, an electronic device is provided, including: at least one processor; a volatile memory; and a memory coupled to the at least one processor, the memory having instructions stored therein, the instructions, when executed by the at least one processor, causing the device to execute a method for managing a computer system. The method includes: obtaining a set of operations to be executed on a plurality of computing units in a computing system; determining, based on the set of operations, the states of the plurality of computing units, and an allocation model, an allocation action for allocating the set of operations to the plurality of computing units and a reward for the allocation action, the allocation model describing the association relationship between the set of operations, the states of the plurality of computing units, the allocation action for allocating the set of operations to the plurality of computing units, and the reward for the allocation action; in response to determining that the matching degree between the reward of the allocation action and the performance metric of the computing system after executing the allocation action meets a predetermined condition, receiving an adjustment for the reward; and generating training data for updating the allocation model based on the adjustment.

[0051] According to an exemplary implementation manner of the present disclosure, the allocation model is generated based on expert knowledge for allocating a set of operations to a plurality of computing units.

[0052] According to an exemplary implementation manner of the present disclosure, the predetermined condition includes: the direction of the reward is opposite to the change direction of the performance metric.

[0053] According to an exemplary implementation manner of the present disclosure, the performance metric includes at least any one of the following: the waiting time of an operation in a set of operations; and the cumulative workload of a computing unit in a plurality of computing units.

[0054] According to an exemplary implementation manner of the present disclosure, receiving the adjustment includes: receiving an adjustment from a technical expert managing the computing system, and wherein the adjustment further includes an adjustment to the action space of the allocation model.

[0055] According to an exemplary implementation manner of the present disclosure, the method further includes: in response to determining that the matching degree does not meet the predetermined condition, generating training data for updating the allocation model based on the reward.

[0056] According to an exemplary implementation manner of the present disclosure, the method further includes at least any one of the following: in response to determining that the difference between the training data and the historical training data for training the allocation model exceeds a predetermined threshold, retaining the training data; and in response to determining that the difference does not exceed the predetermined threshold, deleting the training data.

[0057] According to an exemplary implementation manner of the present disclosure, the allocation model is implemented based on reinforcement learning, and the computing unit includes a graphics processing unit in the computing system.

[0058] According to an exemplary implementation of the present disclosure, the method further includes: obtaining another set of operations to be executed on a plurality of computing units; determining another allocation action for allocating the another set of operations to the plurality of computing units based on the another set of operations, the states of the plurality of computing units, and the updated allocation model; and performing the another allocation action in the computing system.

[0059] According to an exemplary implementation of the present disclosure, the method further includes: updating the allocation model by using training data.

[0060] According to an exemplary implementation of the present disclosure, there is provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions for performing the method according to the present disclosure.

[0061] According to an exemplary implementation of the present disclosure, there is provided a computer-readable medium. Machine-executable instructions are stored on the computer-readable medium, and when the machine-executable instructions are executed by at least one processor, the at least one processor is caused to implement the method according to the present disclosure.

[0062] The present disclosure may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0063] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0064] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0065] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some implementations, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0066] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0067] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0068] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0069] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0070] The foregoing has described various implementations of the present disclosure. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skill in the art in the field to understand the implementations disclosed herein.

Claims

1. A method for managing a computing system, comprising: obtaining a set of operations to be executed on a plurality of computing units in the computing system; determining, based on the set of operations, the states of the plurality of computing units, and an allocation model, an allocation action for allocating the set of operations to the plurality of computing units and a reward for the allocation action, the allocation model being implemented based on reinforcement learning and describing the association relationship between a set of operations, the states of a plurality of computing units, an allocation action for allocating the set of operations to the plurality of computing units, and a reward for the allocation action; filtering the reward; responsive to determining that the match between the reward of the allocation action and a performance metric of the computing system after executing the allocation action satisfies a predetermined condition, receiving an adjustment for the reward from a source external to the allocation model; and generating training data for updating the allocation model based on the adjustment.

2. The method according to claim 1, wherein the allocation model is generated based on expert knowledge for allocating a set of operations to a plurality of computing units.

3. The method according to claim 1, wherein the predetermined condition comprises: the direction of the reward is opposite to the change direction of the performance metric.

4. The method according to claim 3, wherein the performance metric includes at least any one of the following: the waiting time of an operation in the set of operations; and the cumulative workload of a computing unit in the plurality of computing units.

5. The method according to claim 1, wherein receiving the adjustment comprises: receiving the adjustment from a technical expert managing the computing system, and wherein the adjustment further includes an adjustment to the action space of the allocation model.

6. The method according to claim 1, further comprising: responsive to determining that the match does not satisfy the predetermined condition, generating training data for updating the allocation model based on the reward.

7. The method according to claim 1, further comprising at least any one of the following: responsive to determining that the difference between the training data and historical training data for training the allocation model exceeds a predetermined threshold, retaining the training data; and responsive to determining that the difference does not exceed the predetermined threshold, deleting the training data.

8. The method according to claim 7, further comprising: updating the allocation model using the training data.

9. The method according to claim 8, further comprising: obtaining another set of operations to be executed on the plurality of computing units; determining another allocation action for allocating the another set of operations to the plurality of computing units based on the another set of operations, the states of the plurality of computing units, and the updated allocation model; and executing the another allocation action in the computing system.

10. The method according to claim 1, wherein the computing unit includes a graphics processing unit in the computing system.

11. An electronic device, comprising: at least one processor; volatile memory; and A memory coupled to the at least one processor, the memory having instructions stored therein, which when executed by the at least one processor cause the device to perform a method for managing a computing system, the method comprising: Obtaining a set of operations to be executed on a plurality of computing units in the computing system; Based on the set of operations, the states of the plurality of computing units, and an allocation model, determining an allocation action for allocating the set of operations to the plurality of computing units and a reward for the allocation action, the allocation model being implemented based on reinforcement learning and describing the association relationship between a set of operations, the states of a plurality of computing units, an allocation action for allocating the set of operations to the plurality of computing units, and the reward for the allocation action; Filtering the reward; In response to determining that the match between the reward of the allocation action and the performance metric of the computing system after executing the allocation action satisfies a predetermined condition, receiving an adjustment for the reward from a source external to the allocation model; and Generating training data for updating the allocation model based on the adjustment.

12. The apparatus according to claim 11, wherein the allocation model is generated based on expert knowledge for allocating a set of operations to a plurality of computing units.

13. The apparatus according to claim 11, wherein the predetermined condition comprises: The direction of the reward is opposite to the change direction of the performance metric.

14. The apparatus according to claim 13, wherein the performance metric comprises at least any one of the following: The waiting time of the operations in the set of operations; and The cumulative workload of the computing units in the plurality of computing units.

15. The apparatus according to claim 11, wherein receiving the adjustment comprises: Receiving the adjustment from a technical expert managing the computing system, and wherein the adjustment further comprises an adjustment to the action space of the allocation model.

16. The apparatus according to claim 11, wherein the method further comprises: In response to determining that the match does not satisfy the predetermined condition, generating training data for updating the allocation model based on the reward.

17. The apparatus according to claim 11, wherein the method further comprises at least any one of the following: In response to determining that the difference between the training data and the historical training data for training the allocation model exceeds a predetermined threshold, retaining the training data; and In response to determining that the difference does not exceed the predetermined threshold, deleting the training data.

18. The apparatus according to claim 17, wherein the computing unit comprises a graphics processing unit in the computing system, and the method further comprises: Updating the allocation model using the training data.

19. The apparatus according to claim 18, wherein the method further comprises: Obtaining another set of operations to be executed on the plurality of computing units; Based on the another set of operations, the states of the plurality of computing units, and the updated allocation model, determining another allocation action for allocating the another set of operations to the plurality of computing units; and Perform the other allocation action in the computing system.

20. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions for performing the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Computing Resources Workload Scheduling

    US20170109205A1

  • Determining an allocation of computing resources for a job

    US20200117508A1

  • Systems and Methods for Simulating a Complex Reinforcement Learning Environment

    US20200250575A1