Methods, apparatus and computer program products for processing computational tasks
By acquiring the status and parameter information of computing resources, and selecting appropriate computing resource groups to process neural network tasks, the problems of idle computing resources and unbalanced load in the resource pool are solved, thereby improving the efficiency and resource utilization of computing tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EMC IP HLDG CO LLC
- Filing Date
- 2018-04-20
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to effectively utilize multiple computing resources within a resource pool, leading to idle computing resources and unbalanced workloads, especially when handling complex neural network computing tasks.
By acquiring the status information of multiple computing resources, the layer configuration information associated with the neural network model is determined, parameter data is obtained, and a suitable group of computing resources is selected from multiple computing resources based on the status information and parameter data to handle the computing task.
This enables more efficient use of computing resources in the resource pool, improves the processing efficiency and resource utilization of computing tasks, and reduces idle computing resources and load imbalance.
Smart Images

Figure CN110389824B_ABST
Abstract
Description
Technical Field
[0001] Implementations of this disclosure generally relate to computing systems including dedicated computing resources, and more specifically, to methods, apparatus, and computer program products for processing computing tasks. Background Technology
[0002] Applications on the client side can be designed to utilize computing resources, such as processing and storage resources, to perform various processing or analytical tasks. As the demands and complexity of tasks such as machine learning, deep learning, and data mining increase, a large and / or variable amount of computing resources is required to support the operation of these applications. This can be achieved through machines or systems with multiple dedicated computing resources, where applications can be scheduled to run on one or more dedicated computing resources of that machine or system. For example, cloud-based computing systems have been developed that include machines with one or more dedicated computing resources. Different clients can lease computing resources (e.g., dedicated computing resources) from this system as needed to run their respective applications.
[0003] With the development of computer technology, the types of computing resources have become increasingly diverse, and are no longer limited to traditional computing resources such as central processing units. For example, the computing power of graphics processing units (GPUs) is becoming increasingly powerful. Due to their unique properties, GPUs are particularly suitable for performing computational tasks related to deep learning, high-performance computing, and machine learning. However, for ordinary client devices and conventional cloud computing devices, the performance of their graphics processing units is usually limited, lacking high-performance processing capabilities. Therefore, how to utilize (e.g., remotely) the computing power of the graphics processing units of other devices to handle computational tasks has become a research focus.
[0004] However, some current technical solutions cannot fully and effectively utilize the processing power of remote computing resources (e.g., computing resources in a computing resource pool), and may result in idle computing resources and / or unbalanced workloads within the resource pool. Therefore, there is a need to provide a technical solution that can utilize multiple computing resources in a resource pool to process computing tasks in a simple and effective manner. Summary of the Invention
[0005] Implementations of this disclosure provide methods, apparatus, and corresponding computer program products for processing computational tasks.
[0006] According to a first aspect of this disclosure, a method for processing a computational task is provided. The method includes: acquiring state information of a plurality of computational resources; in response to receiving a computational task based on a neural network model, determining configuration information of a plurality of layers associated with the neural network model; acquiring parameter data associated with at least a portion of the plurality of layers based on the configuration information; and selecting a set of computational resources from the plurality of computational resources for processing the computational task based on the state information and the parameter data.
[0007] According to a second aspect of this disclosure, an apparatus for processing a computational task is provided. The apparatus includes: at least one processor; volatile memory; and a memory coupled to the at least one processor, the memory having instructions stored therein, the instructions causing the apparatus to perform an action when executed by the at least one processor. The action includes: acquiring state information of a plurality of computational resources; determining configuration information of a plurality of layers associated with the neural network model in response to receiving a computational task based on a neural network model; acquiring parameter data associated with at least a portion of the plurality of layers based on the configuration information; and selecting a set of computational resources from the plurality of computational resources for processing the computational task based on the state information and the parameter data.
[0008] According to a third aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform the method according to the first aspect.
[0009] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description
[0010] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.
[0011] Figure 1 A block diagram of an exemplary computing system suitable for implementing the present disclosure is shown schematically;
[0012] Figure 2 A block diagram illustrating a process for processing computational tasks based on a neural network model, according to a technical solution, is shown schematically.
[0013] Figure 3 A block diagram illustrating a processing computation task according to an exemplary implementation of the present disclosure is shown schematically.
[0014] Figure 4 A flowchart illustrating an exemplary implementation of a method for processing computational tasks according to this disclosure is shown schematically;
[0015] Figure 5 A block diagram illustrating an exemplary implementation of this disclosure for acquiring parameter data associated with a neural network is shown schematically.
[0016] Figure 6 The diagram illustrates a block diagram of selecting computational resources for parameters of a sorted layer according to an exemplary implementation of the present disclosure.
[0017] Figure 7 The diagram illustrates a block diagram of selecting computational resources for parameters of a sorted layer according to an exemplary implementation of the present disclosure.
[0018] Figure 8 A block diagram of a device for processing computing tasks according to an exemplary implementation of the present disclosure is shown schematically; and
[0019] Figure 9 A block diagram of an exemplary implementation of a device for processing computing tasks is shown schematically according to the present disclosure. Detailed Implementation
[0020] Preferred implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While preferred implementations of this disclosure are shown in the drawings, it should be understood that this disclosure may be implemented in various forms and should not be limited to the implementations set forth herein. Rather, these implementations are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0021] The term “comprising” and its variations, as used herein, signify open inclusion, i.e., “including but not limited to.” Unless otherwise stated, the term “or” means “and / or.” The term “based on” means “at least partially based on.” The terms “one example implementation” and “one implementation” mean “at least one example implementation.” The term “another implementation” means “at least one additional implementation.” The terms “first,” “second,” etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0022] As described above, dedicated computing resources can be provided locally on the client or by a remote machine or system. In some examples, cloud-based computing systems can be deployed, which include multiple machines with one or more dedicated computing resources. These dedicated computing resources can be used by different clients as needed to schedule the appropriate applications to run on the available dedicated computing resources.
[0023] Figure 1 A schematic diagram of an example computing system 100 in which the implementation of this disclosure may be implemented is shown. Multiple servers for application execution are deployed in the computing system 100, including server 110-1, server 110-2, server 110-3, ..., server 110-U (hereinafter collectively or individually referred to as server 110, where U is a natural number greater than 1). The computing system 100 also includes dedicated computing resources 160-1, dedicated computing resources 160-2, dedicated computing resources 160-3, ..., dedicated computing resources 160-V (hereinafter collectively or individually referred to as dedicated computing resources 160, where V is a natural number greater than 1). Each server 110 may have one or more dedicated computing resources 160.
[0024] exist Figure 1 In the example, server 110-1 has dedicated computing resource 160-1, server 110-2 has dedicated computing resource 160-2, and server 110-U has dedicated computing resource 160-V. It will be understood that this does not limit each server to having only one computing resource, but rather a server can have one or more computing resources. Therefore, the values of U and V can be unequal.
[0025] In the context of this disclosure, examples of dedicated computing resources 160 may include, but are not limited to, graphics dedicated computing resources (GPUs), field-programmable gate arrays (FPGAs), etc. For ease of discussion, some implementations will be described using GPUs as an example of dedicated computing resources. In addition to dedicated computing resources 160, server 110 may also include one or more general-purpose processing units (not shown), such as a central processing unit (CPU).
[0026] Figure 1 Multiple clients 120-1, 120-2, ..., 120-P (hereinafter collectively or individually referred to as client 120, where P is a natural number greater than 1) are also shown, each with an application 150-1, 150-2, ..., 150-Q (hereinafter collectively or individually referred to as application 150, where Q is a natural number greater than 1) to run. Application 150 can be any application that can run on the machine, and this application can be designed to perform corresponding data processing or analysis tasks. As an example, application 150 can perform data processing or analysis tasks related to neural networks. It will be understood that there is no restriction on each client having only one application, but rather that a client can have one or more applications. Therefore, the values of P and Q can be unequal here.
[0027] To enable fast and efficient operation of these applications and / or to conserve local computing resources, client 120 may request dedicated computing resources 160 on server 110 to run these applications 150. In such an implementation, client 120 may connect to one or more servers 110 via interconnect network 130 and delegate application 150 to one or more dedicated computing resources 160 on server 110 for execution. Depending on the interfaces supported by client 120, server 110, and / or dedicated computing resources 160, interconnect network 130 may support different types of wired or wireless connections based on various network transport technologies such as Remote Direct Memory Access (RDMA) and Transmission Control Protocol (TCP).
[0028] It should be understood that Figure 1 The devices and / or arrangements shown are merely examples. In other examples, the computing system 100 may include any suitable number of servers 110 and clients 120. Each server 110 may be equipped with any suitable number of dedicated computing resources 160, and each client 120 may have multiple applications 150 to run. Furthermore, although shown separately, the scheduler 140 may be implemented in practice by other devices independent of the servers 110, or may be implemented partially or wholly on one or more servers 110.
[0029] For clarity and conciseness, this disclosure will primarily use GPU cores as an example to describe the sample implementation in detail. As is known, a GPU, as a dedicated processor, derives its powerful computing capabilities from its large number of cores and high-bandwidth memory. In GPU hardware architecture, a GPU typically has a large number of GPU cores, such as 5120 or close to 10,000 cores. A GPU core, as a dedicated computing resource, is the most basic processing unit and is also known as a stream processor (SP). Instructions and tasks are ultimately processed on the GPU core. Multiple GPU cores execute instructions simultaneously, thus enabling parallel computing on the GPU. Multiple SPs, along with other resources such as registers and shared memory, can form a stream multiprocessor (SM).
[0030] However, it should be understood that the GPU is merely an exemplary dedicated computing resource and is not intended to limit the scope of this disclosure. The spirit and principles described herein can be applied to other dedicated computing resources, such as computing resources in accelerators like field-programmable gate arrays (FPGAs), whether currently known or to be developed in the future, and are not limited to GPU cores.
[0031] With the development of cloud computing, technical solutions based on cloud architecture for processing computing tasks have been proposed. For example, application 120 in client 150 can request computing resources 160 in server 110. It should be noted that due to the complexity of computing tasks, computing tasks typically require the invocation of multiple computing resources 160. In the following sections, a computing task based on a neural network model will be used as a specific example to describe in more detail the implementation of this disclosure. Figure 2 A block diagram 200 schematically illustrates a process for handling a computational task 210 based on a neural network model, according to a technical solution. For example... Figure 2 As shown, computational task 210 can be a computational task based on a neural network model, which may involve multiple layers, such as layer 1, layer 2, ..., layer N, as indicated by reference numerals 212, 214, ..., 216. It will be understood that each layer from layer 1 to layer N will involve a large number of parameters defining the neural network model, such as gradients, weights, biases, etc. The amount of data for the parameters involved will vary significantly for different layers; for example, the number of parameters can vary from tens to millions or even more. Therefore, how to handle computational task 210 in a balanced manner with multiple computational resources (e.g., computational resources 160-1 to 160-V) becomes a challenge.
[0032] It will be understood that existing technologies, such as parameter server technology, are already available for processing computational tasks based on neural network models. However, these existing technologies cannot effectively utilize the computational performance of multiple computing resources in a resource pool. Based on the shortcomings of the prior art, this disclosure proposes a method for processing computational tasks.
[0033] Figure 3 A block diagram 300 illustrating a processing computation task according to an exemplary implementation of this disclosure is shown schematically. Figure 3 As shown, the status information 330 of multiple computing resources 160-1 to 160-V in resource pool 320 can be obtained. Upon receiving a computing task based on neural network model 210, multiple layers (e.g., layers (e.g., ...)) associated with neural network model 210 can be determined. Figure 2 The configuration information of layers 1, 2, ..., N (shown by reference numerals 212, 214, ..., 216) is shown in the attached figures. Based on the configuration information, parameter data 310 associated with at least a portion of the layers can be obtained. Subsequently, based on the state information 330 and the parameter data 310, a set of computing resources can be selected from multiple computing resources 160 to process the computational task based on the neural network model 320.
[0034] Figure 4A flowchart illustrating an exemplary implementation of a method 400 for processing a computing task according to the present disclosure is shown. At block 410, status information 330 of a plurality of computing resources 160 is obtained. This status information 330 may relate to various aspects of the computing resources. According to an exemplary implementation of the present disclosure, the resource information of the plurality of computing resources includes at least any one of the following metrics: processing power information of the plurality of computing resources, memory resource information, and bandwidth resource information.
[0035] According to an exemplary implementation of this disclosure, memory resource information may include, for example, the size of the available storage space for the GPU. This metric may be measured in absolute values: for example, one GPU may have 8GB of storage space, while another GPU may have 5GB of storage space. When measured in absolute values, to facilitate comparison of the memory resource information of various computing resources, a normalization rule can be set to unify the memory resource information of multiple computing resources to the same standard. For example, assuming that the largest storage capacity among the multiple computing resources includes 10GB, the memory resource information of that computing resource can be set to 1, and the memory resource information of other computing resources can be calculated based on the proportion of storage capacity. For example, the memory resource information of a computing resource including 5GB of storage space can be expressed as 5GB / 10GB = 0.5. Alternatively or additionally, this metric may also be measured in relative values. For example, assuming that a GPU includes 8GB of storage space, and 7GB of it is already used, then the memory resource information of that GPU can be expressed as (8GB-7GB) / 8GB = 0.125.
[0036] According to an exemplary implementation of this disclosure, processing power information may include, for example, a GPU processing power metric, which may be measured in absolute values such as processing frequency, number of processor cores, etc. When measured in absolute values, to facilitate comparison of the processing power information of various computing resources, normalization rules can be set to unify the processing power information of multiple computing resources to the same standard. For example, the processing power information of each computing resource may be determined based on the ratio of its processing frequency to its highest processing frequency. Alternatively or additionally, the metric may also be measured in relative values. For example, assuming the GPU's theoretical processing power is 1, and 50% of its processing power is currently occupied by other computing tasks, the available processing power may be represented as 0.5.
[0037] According to an exemplary implementation of this disclosure, bandwidth resource information can, for example, represent the communication bandwidth of a GPU, and this metric can be measured, for example, in absolute terms. When measured in absolute terms, to facilitate comparison of the bandwidth resource information of various computing resources, a normalization rule can be set to unify the bandwidth resource information of multiple computing resources to the same standard. For example, the communication bandwidth information of each computing resource can be determined based on the ratio of the communication bandwidth of each computing resource to the maximum communication bandwidth. Alternatively or additionally, this metric can also be measured in relative terms. For example, assuming the theoretical bandwidth of the GPU is 4GB / s, and 2GB / s of communication bandwidth is currently being used by other computing tasks, then the bandwidth resource information can be represented as 2 / 4 = 0.5.
[0038] According to an exemplary implementation of this disclosure, for a given computing resource among multiple computing resources, the importance of a corresponding indicator among multiple metrics for the given computing resource can be determined based on the computing task. For example, if the computing task is found to involve a high computational load, a higher importance can be assigned to processing capacity information. If the computing task is found to involve a large amount of data, a higher importance can be assigned to memory resource information. If the computing task is found to involve a large amount of communication, a higher importance can be assigned to bandwidth resources. Subsequently, state information for the given computing resource can be determined based on the importance of the corresponding indicators and the corresponding indicators. For example, the state information can be determined based on the following Formula 1:
[0039] Status(i)
[0040] =Weight processing capacity *ProcessingCapacity
[0041] +Weight memory capacity *MemoryCapacity
[0042] +Weight band width *BandWidth (Formula 1)
[0043] In Formula 1, Status(i) represents the status information of the i-th computing resource in resource pool 320, ProcessingCapacity represents the processing capacity information, and Weight processing capacity The importance of processing power information is indicated by MemoryCapacity, which represents memory resource information. Weight memory capacity The importance of memory resource information is indicated by BandWidth, and the importance of bandwidth resource information is indicated by Weight. band width This indicates the importance of bandwidth resource information.
[0044] According to an exemplary implementation of this disclosure, the state information 330 can be represented in a vector manner, wherein each dimension of the vector represents the state information of the corresponding computing resource.
[0045] At box 420, it is determined whether a computational task based on neural network model 210 has been received. If the determination result is "yes," then at box 430, configuration information of multiple layers associated with neural network model 210 can be determined. According to an exemplary implementation of this disclosure, multiple layers and associated configuration information can be determined based on the definition of neural network model 210. At box 440, parameter data 310 associated with at least a portion of the multiple layers can be obtained based on the configuration information. See below. Figure 5 Describe more details.
[0046] See also Figure 4 At box 450, based on status information 330 and parameter data 310, a set of computing resources 160 is selected from multiple computing resources 160 to process the computing task. Simply put, if parameter data 310 indicates that the computing task involves a high workload, then computing resources in better condition can be selected from the multiple computing resources 160 based on status information 330. If parameter data 310 indicates that the computing task involves only a low workload, then computing resources in average or even poor condition can be selected from the multiple computing resources 160 based on status information 330. In this way, the multiple computing resources 160 in resource pool 320 can be utilized more effectively.
[0047] Figure 5 A block diagram 500 illustrating an exemplary implementation of this disclosure for acquiring parameter data associated with a neural network is shown. Figure 5 As shown, reference numeral 510 schematically illustrates configuration information for a neural network model 210 according to an exemplary implementation. This configuration information 510 defines the multiple layers included in the neural network model 210 and the parameters involved in each layer. By parsing this configuration information 510, parameter data 310 regarding the neural network model 210 can be obtained.
[0048] like Figure 5As shown, parameter data 310 is a specific example of parameter data according to an exemplary implementation of this disclosure. As shown in parameter data 310, the neural network model 210 may include multiple layers, and the field "Param-size" in each row defines the number of parameters associated with each layer. As shown in the first row of parameter data 310, a layer may include 23,232 parameters; as shown in the second row of parameter data 310, a layer may include 64 parameters; and so on. It will be understood that the method of obtaining parameter data 310 is not limited in the context of this disclosure. Rather, those skilled in the art can obtain parameter data 310 according to various technical solutions that have been developed or will be developed in the future.
[0049] According to an exemplary implementation of this disclosure, computational resources for processing the parameters associated with each layer can be selected separately. It will be understood that this is not limited to using the implementation of this disclosure to process multiple layers, but rather that only the implementation of this disclosure can be used to process at least a portion of the multiple layers. For other layers, other methods can be used to select computational resources for processing the parameters associated with other layers. In this implementation, by processing each layer one by one, computational resources appropriate to the number of parameters of each layer involved in the neural network model 210 can be progressively allocated.
[0050] The operation process for a layer will be described in detail below. According to an exemplary implementation of this disclosure, for a first layer in at least a subset of layers, a first number of parameters associated with the first layer can be determined based on parameter data. For example, with Figure 5 Taking the layer indicated by row 520 in the parameter data 310 as an example, the number of parameters associated with that layer is 23232. For another example, using... Figure 5 Taking the layer indicated by row 522 in the parameter data 310 as an example, the number of parameters associated with this layer is 37,748,736.
[0051] Subsequently, based on the status information 330 of multiple computing resources 160, a first computing resource matching the first quantity can be selected from the multiple computing resources 160 to process the parameters associated with the first layer. Analysis of the layer shown in row 520 of the parameter data 310 reveals that the number of parameters involved in this layer is 23232, which is relatively small within the parameter data 310. Therefore, computing resources with a moderate or even poor state can be selected based on the status information 330 of the multiple computing resources 160. For example, for the layer represented by row 522 in the parameter data 310, the number of parameters involved in this layer is 37,748,736, which is a relatively large number. Therefore, computing resources with a better state can be selected based on the status information 330 of the multiple computing resources 160.
[0052] According to an exemplary implementation of this disclosure, the number of parameters involved in each layer can be counted first, so as to prioritize the allocation of computing resources to layers with a larger number of parameters. For example, based on... Figure 5 The parameter data 310 shown can determine the corresponding number of parameters associated with at least a subset of layers. In this example, by extracting the values of the Param-size field in parameter data 310, the number of parameters associated with each layer can be represented as: [23232, 64, 307200, 192, 663552, 384, 1327104, 384, 884736, 256, 37748736, 4096, 16777216, 4096, 4100096, 1001]. Subsequently, at least a subset of layers are sorted based on the corresponding number, and the first layer is selected based on the sorting. In this implementation, the sorting can be done in descending order, as will be seen in the appendix below. Figure 6 Describe more details.
[0053] Figure 6 A block diagram 600 schematically illustrates the selection of computing resources for parameters of a sorted layer according to an exemplary implementation of the present disclosure. According to an exemplary implementation of the present disclosure, resource information of each of a plurality of computing resources can be monitored separately, and then the status information of each resource can be determined based on the resource information of each computing resource, thereby forming a... Figure 6 The status information 620 for multiple computing resources is shown. For example... Figure 6 As shown, the state information 620 is stored in a vector format, where the i-th dimension stores the state information of the i-th computing resource determined based on the method described above. In this example, only the state information for multiple computing resources 1, 2, 3, etc., as shown using the method described above is illustrated. It will be understood that although in Figure 6In the example, the state information is normalized to the numerical range of [0, 100]. In other examples, the state information can also be normalized to other numerical ranges.
[0054] like Figure 6 As shown, box 610 illustrates the sorted number of parameters associated with each layer. Figure 6 As shown, for layer 612, which involves 37,748,736 parameters, since this number is large, the best-performing computing resource 1 can be selected from the multiple computing resources 160 in the resource pool 210.
[0055] According to an exemplary implementation of this disclosure, a first resource allocation required for processing parameters associated with the first layer can be determined, and the status information 330 can be updated based on the first resource allocation. In this way, it can be ensured that the status information 330 is updated in a timely manner, thereby ensuring that the status information 330 accurately reflects the latest status of the plurality of computing resources 160 in the resource pool 210. See below for further details. Figure 7 Describe more details about the update.
[0056] Figure 7 A block diagram 700 schematically illustrates the selection of computational resources for parameters of a sorted layer according to an exemplary implementation of this disclosure. Figure 7 As shown, layer 612 involves 37,748,736 parameters. Assuming that processing layer 612 requires 50% of the resource allocation in computing resource 1, the computing resource 1-related status information in status information 720 can be updated based on this resource allocation. For example, the status information related to computing resource 1 can be updated to: 100 - 100 * 50% = 50. The updated status information is as follows... Figure 7 As shown in circle 722.
[0057] According to an exemplary implementation of this disclosure, other layers can be processed in a similar manner. For example, for a second layer among at least a subset of layers, a second number of parameters associated with the second layer can be determined based on parameter data. Then, based on updated state information, a second computing resource matching the second number is selected from a plurality of computing resources to process the parameters associated with the second layer. See also... Figure 7 You can select computing resources for the parameters of layer 710, which is ranked second. For example... Figure 7As shown, computing resource 2 can be selected to process the parameters involved in layer 710. The number of parameters involved in layer 710 is 16,777,216. Assuming that processing layer 710 requires 60% of the resource allocation in computing resource 1, the status information related to computing resource 2 in status information 720 can be updated based on this resource allocation. For example, the status information related to computing resource 2 can be updated to: 90 - 100 * 60% = 30. Subsequently, the status information related to computing resource 2 represented by circle 724 can be updated to 30 (…). Figure 7 The image only shows the status information "90" before the update, and does not show the status information "30" after the update.
[0058] According to an exemplary implementation of this disclosure, the number of computing resources required to process at least a subset of layers can be determined, and appropriate computing resources can be selected from the resource pool 320 according to the method described above. When the number of selected computing resources reaches the required quantity, further selection from the resource pool 320 is not necessary. Instead, when selecting computing resources for subsequent layers, computing resources for processing parameters associated with other layers in the at least subset of layers are selected from the already selected computing resources. In this way, it can be ensured that the total number of selected computing resources matches the computing task.
[0059] For example, suppose a computational task is determined to require four computing resources to execute. For example... Figure 7 Box 610 in the diagram shows the sorted layers. First, the four best-performing computing resources from resource pool 320 are selected to process the parameters associated with layers ranked 1 to 4. Then, for layers ranked 5th and below, the corresponding computing resources are again selected from the previously chosen four.
[0060] As mentioned above Figures 2 to 7 Examples of methods according to this disclosure are described in detail below, see also: Figure 8 Provide a detailed description of the implementation of the corresponding device. Figure 8A block diagram of an exemplary implementation of a device 800 for processing a computing task according to the present disclosure is shown schematically. The device 800 includes: a state acquisition module 810 configured to acquire state information of a plurality of computing resources; a configuration determination module 820 configured to determine configuration information of a plurality of layers associated with a neural network model in response to receiving a computing task based on a neural network model; a parameter acquisition module 830 configured to acquire parameter data associated with at least a portion of the plurality of layers based on the configuration information; and a selection module 840 configured to select a set of computing resources from the plurality of computing resources for processing the computing task based on the state information and the parameter data. The device 800 may be configured to perform various steps of the methods described above, which will not be repeated here.
[0061] Figure 9 A block diagram of an exemplary implementation of a device for processing computational tasks according to the present disclosure is shown schematically. As shown, device 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. Various programs and data required for the operation of device 900 may also be stored in RAM 903. CPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0062] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0063] The various processes and procedures described above, such as method 400, can be executed by processing unit 901. For example, in some implementations, method 400 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some implementations, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU 901, one or more steps of method 400 described above can be performed. Alternatively, in other implementations, CPU 901 can also be configured in any other suitable manner to implement the above-described processes / methods.
[0064] According to an exemplary implementation of this disclosure, an apparatus for processing a computational task is provided, comprising: at least one processor; volatile memory; and a memory coupled to the at least one processor, the memory having instructions stored therein, the instructions causing the apparatus to perform an action when executed by the at least one processor. The action includes: acquiring state information of a plurality of computational resources; determining configuration information of a plurality of layers associated with the neural network model in response to receiving a computational task based on a neural network model; acquiring parameter data associated with at least a portion of the plurality of layers based on the configuration information; and selecting a set of computational resources from the plurality of computational resources for processing the computational task based on the state information and the parameter data.
[0065] According to an exemplary implementation of this disclosure, for a first layer in at least a subset of layers, a first number of parameters associated with the first layer is determined based on parameter data; and a first computing resource matching the first number is selected from a plurality of computing resources based on state information for processing the parameters associated with the first layer.
[0066] According to an exemplary implementation of this disclosure, based on parameter data, a corresponding number of parameters associated with at least a portion of the layers is determined; the at least a portion of the layers are sorted based on the corresponding number; and a first layer is selected based on the sorting.
[0067] According to an exemplary implementation of this disclosure, a first resource allocation required for processing parameters associated with the first layer is determined; and state information is updated based on the first resource allocation.
[0068] According to an exemplary implementation of this disclosure, for a second layer in at least a subset of layers, a second number of parameters associated with the second layer is determined based on parameter data; and a second computing resource matching the second number is selected from a plurality of computing resources based on updated state information for processing the parameters associated with the second layer.
[0069] According to an exemplary implementation of this disclosure, a number of computing resources required to process at least a portion of the layers is determined; and in response to determining that the number of selected computing resources has reached a certain amount, computing resources for processing parameters associated with a third layer in at least a portion of the layers are selected from the selected computing resources.
[0070] According to an exemplary implementation of this disclosure, resource information of multiple computing resources is monitored; and status information of the multiple computing resources is determined based on the resource information.
[0071] According to an exemplary implementation of this disclosure, the resource information of multiple computing resources includes at least one of the following metrics: processing power information of the multiple computing resources, memory resource information, and bandwidth resource information.
[0072] According to an exemplary implementation of this disclosure, for a given computing resource among a plurality of computing resources, the importance of a corresponding indicator among a plurality of indicators for the given computing resource is determined based on the computing task; and based on the importance of the corresponding indicator and the corresponding indicator, the state information for the given computing resource is determined.
[0073] According to one exemplary implementation of this disclosure, the multiple computing resources are multiple graphics processing units.
[0074] According to an exemplary implementation of this disclosure, a computer program product is provided. The computer program product is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform the method described in this disclosure.
[0075] According to an exemplary implementation of this disclosure, a computer-readable medium is provided. The computer-readable medium stores machine-executable instructions that, when executed by at least one processor, cause the at least one processor to implement the method according to this disclosure.
[0076] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.
[0077] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0078] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0079] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some implementations, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is customized by utilizing the status information of the computer-readable program instructions to execute the computer-readable program instructions, thereby implementing various aspects of this disclosure.
[0080] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0081] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0082] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0084] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. A method for processing computational tasks, comprising: Obtain status information for multiple computing resources; In response to receiving a computational task based on a neural network model, configuration information of multiple layers associated with the neural network model is determined; Based on the configuration information, obtain parameter data associated with one or more of the plurality of layers; Based on the status information and the parameter data, a set of computing resources is selected from the plurality of computing resources to process the computing task; as well as The computing task is processed according to the selected set of computing resources; Obtaining the status information of the plurality of computing resources further includes, for a given computing resource among the plurality of computing resources: Based on the computing task, determine the importance of the corresponding indicators among multiple indicators of resource information for the given computing resources. as well as Based on the importance of the corresponding indicators and the corresponding indicators, obtain the status information of the given computing resource; and Selecting a set of computing resources based on the state information includes, for the first layer of one or more of the plurality of layers: Based on the parameter data, determine a first number of parameters associated with the first layer; as well as Based on the status information, a first computing resource matching the first quantity is selected from the plurality of computing resources to process the parameters associated with the first layer.
2. The method according to claim 1, wherein selecting a set of computing resources based on the state information includes: Based on the parameter data, determine the corresponding number of parameters associated with one or more of the plurality of layers; Sort one or more of the plurality of layers based on the corresponding quantity; as well as The first layer is selected based on the sorting.
3. The method according to claim 1, further comprising: Determine the first resource allocation required for processing the parameters associated with the first layer; as well as Based on the first resource allocation, update the status information.
4. The method of claim 3, further comprising: For the second layer of one or more of the plurality of layers, Based on the parameter data, determine a second number of parameters associated with the second layer; as well as Based on the updated state information, a second computing resource matching the second quantity is selected from the plurality of computing resources to process the parameters associated with the second layer.
5. The method of claim 4, further comprising: Determine the amount of computing resources required to process one or more of the plurality of layers; as well as In response to determining that the number of the selected set of computing resources has reached the specified number, computing resources are selected from the selected set of computing resources for processing parameters associated with a third layer of one or more of the plurality of layers.
6. The method according to claim 1, wherein obtaining the status information of the plurality of computing resources further comprises: Monitor the resource information of the multiple computing resources; as well as The status information of the plurality of computing resources is obtained based on the resource information.
7. The method of claim 6, wherein the resource information of the plurality of computing resources includes at least one of the following indicators: The processing capacity information, memory resource information, and bandwidth resource information of the multiple computing resources.
8. The method of claim 1, wherein the plurality of computing resources are a plurality of graphics processing units.
9. A device for processing computational tasks, comprising: At least one processor; as well as A memory coupled to the at least one processor, the memory having instructions stored therein, the instructions being responsive to execution thereon by the at least one processor, causing the device to perform an action, the action including: Obtain status information for multiple computing resources; In response to receiving a computational task based on a neural network model, configuration information of multiple layers associated with the neural network model is determined; Based on the configuration information, obtain parameter data associated with one or more of the plurality of layers; Based on the status information and the parameter data, a set of computing resources is selected from the plurality of computing resources to process the computing task; and The computing task is processed according to the selected set of computing resources; Obtaining the status information of the plurality of computing resources further includes, for a given computing resource among the plurality of computing resources: Based on the computational task, determine the importance of corresponding indicators among multiple metrics of resource information for the given computational resource; and Based on the importance of the corresponding indicators and the corresponding indicators themselves, obtain the status information of the given computing resource; and Selecting a set of computing resources based on the state information includes, for the first layer of one or more of the plurality of layers: Based on the parameter data, determine a first number of parameters associated with the first layer; and Based on the status information, a first computing resource matching the first quantity is selected from the plurality of computing resources to process the parameters associated with the first layer.
10. The device of claim 9, wherein selecting a set of computing resources based on the status information includes: Based on the parameter data, determine the corresponding number of parameters associated with one or more of the plurality of layers; Sort one or more of the plurality of layers based on the corresponding quantity; as well as The first layer is selected based on the sorting.
11. The device according to claim 9, wherein the action further comprises: Determine the first resource allocation required for processing the parameters associated with the first layer; as well as Based on the first resource allocation, update the status information.
12. The apparatus of claim 11, wherein the acts further comprise: For the second layer of one or more of the plurality of layers, Based on the parameter data, determine a second number of parameters associated with the second layer; as well as Based on the updated state information, a second computing resource matching the second quantity is selected from the plurality of computing resources to process the parameters associated with the second layer.
13. The device according to claim 12, wherein the action further comprises: Determine the amount of computing resources required to process one or more of the plurality of layers; as well as In response to determining that the number of the selected set of computing resources has reached the specified number, computing resources are selected from the selected set of computing resources for processing parameters associated with a third layer of one or more of the plurality of layers.
14. The device according to claim 9, wherein obtaining the status information of the plurality of computing resources further comprises: Monitor the resource information of the multiple computing resources; as well as The status information of the plurality of computing resources is obtained based on the resource information.
15. The device of claim 14, wherein the resource information of the plurality of computing resources includes at least one of the following indicators: The processing capacity information, memory resource information, and bandwidth resource information of the multiple computing resources.
16. The device of claim 9, wherein the plurality of computing resources are a plurality of graphics processing units.
17. A non-transitory computer-readable medium storing machine-executable instructions, which, when executed, cause a machine to implement a method, the method comprising: Obtain status information for multiple computing resources; In response to receiving a computational task based on a neural network model, configuration information of multiple layers associated with the neural network model is determined; Based on the configuration information, obtain parameter data associated with one or more of the plurality of layers; Based on the status information and the parameter data, a set of computing resources is selected from the plurality of computing resources to process the computing task; as well as The computing task is processed according to the selected set of computing resources; Obtaining the status information of the plurality of computing resources further includes, for a given computing resource among the plurality of computing resources: Based on the computing task, determine the importance of the corresponding indicators among multiple indicators of resource information for the given computing resources. as well as Based on the importance of the corresponding indicators and the corresponding indicators, obtain the status information of the given computing resource; and Selecting a set of computing resources based on the state information includes, for the first layer of one or more of the plurality of layers: Based on the parameter data, determine a first number of parameters associated with the first layer; as well as Based on the status information, a first computing resource matching the first quantity is selected from the plurality of computing resources to process the parameters associated with the first layer.
18. The non-transient computer-readable medium of claim 17, wherein the method further comprises: Determine the first resource allocation required for processing the parameters associated with the first layer; as well as Based on the first resource allocation, update the status information.
19. The non-transitory computer-readable medium of claim 18, further comprising: For the second layer of one or more of the plurality of layers, Based on the parameter data, determine a second number of parameters associated with the second layer; as well as Based on the updated state information, a second computing resource matching the second quantity is selected from the plurality of computing resources to process the parameters associated with the second layer.
20. The non-transient computer-readable medium of claim 19, further comprising: Determine the amount of computing resources required to process one or more of the plurality of layers; as well as In response to determining that the number of the selected set of computing resources has reached the specified number, computing resources are selected from the selected set of computing resources for processing parameters associated with a third layer of one or more of the plurality of layers.