Task allocation method and device, electronic equipment and storage medium
By obtaining the running status information of the neural network model, dynamically determine the task type and allocate it to the multi-core processor, the problem of low utilization caused by improper task allocation in asymmetric multi-core processors is solved, and higher processor efficiency is achieved.
Patent Information
- Application Number
- CN202510315346.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
In asymmetric multi-core processors, cores are prone to idle when performing computing tasks and communication tasks, resulting in reduced processor utilization.
By obtaining the running status information during the training process of neural network model, determine the task type of the task in the current stage, and assign tasks to the multi-core processor according to the task type. Specifically, when the forward phase is determined to be a computing intensive task, the task is split and assigned to a processing core with different computing capabilities; when the reverse phase is determined to be a communication intensive task, the communication task is assigned to a processing core with low computing capabilities.
By dynamically adjusting the task allocation strategy, the idle time of processing cores is reduced, the utilization rate of multi-core processors is improved, and the overall performance is improved.
Smart Images

Figure CN120179401A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a task allocation method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] In an asymmetric multi-core processor, two or more cores with different scales can be provided on the same chip. Generally, these cores are used to execute computing tasks or communication tasks, and at any moment, each core executes a fixed task.
[0003] In related technologies, tasks executed by multiple cores of a processor can be specified in advance. For example, large cores are only used for computing tasks, and small cores are only used for communication tasks. Alternatively, both large and small cores are used for executing computing tasks, and small cores are additionally used for communication tasks. In this way, although parallel execution of computing tasks and communication tasks can be achieved, there will still be a situation where large cores or small cores are idle, resulting in reduced utilization of the processor. Summary of the Invention
[0004] The present disclosure proposes a technical solution for task allocation.
[0005] According to an aspect of the present disclosure, there is provided a task allocation method, including: obtaining operation state information during the training process of a neural network model when a working mode category is an automatic switching mode; determining a task type of a current-stage task of the neural network model according to the operation state information; and allocating the current-stage task to a multi-core processor according to the task type of the current-stage task of the neural network model.
[0006] In a possible implementation manner, the operation state information includes first information for indicating that the neural network model is in a forward stage during the training process, and second information for indicating that the neural network model is in a backward stage during the training process. The obtaining operation state information during the training process of the neural network model includes: obtaining the first information in response to detecting that the neural network model calls a data loading function; or obtaining the second information in response to detecting that the neural network model calls a loss function.
[0007] In a possible implementation manner, the task type includes a compute-intensive task and a communication-intensive task. The determining a task type of a current-stage task of the neural network model according to the operation state information includes: determining the task type of the current-stage task of the neural network model as a compute-intensive task according to the first information; or determining the task type of the current-stage task of the neural network model as a communication-intensive task according to the second information.
[0008] In a possible implementation, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. Allocating the current-stage task to the multi-core processor according to the task type of the current-stage task of the neural network model includes: when the task type of the current-stage task of the neural network model is a compute-intensive task, splitting the operator to be processed in the current-stage task into a first split operator and a second split operator, where the computational amount of the first split operator is greater than that of the second split operator; allocating the first split operator to the first processing core and allocating the second split operator to the second processing core.
[0009] In a possible implementation, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. Allocating the current-stage task to the multi-core processor according to the task type of the current-stage task of the neural network model includes: when the task type of the current-stage task of the neural network model is a communication-intensive task, allocating the operator to be processed in the current-stage task to the first processing core and allocating the communication task in the current-stage task to the second processing core.
[0010] In a possible implementation, the method further includes: when the task type of the current-stage task of the neural network model is a compute-intensive task, in response to detecting that the neural network model calls a communication function, allocating the operator to be processed in the current-stage task to the first processing core and allocating the communication task in the current-stage task to the second processing core; or, when the task type of the current-stage task of the neural network model is a compute-intensive task, in response to detecting that the user adjusts the task type of the current-stage task from the compute-intensive task to a communication-intensive task, allocating the operator to be processed in the current-stage task to the first processing core and allocating the communication task in the current-stage task to the second processing core.
[0011] In a possible implementation, the method further includes: when the obtained working mode category is the fixed mode, determining the task type of the current-stage task of the neural network model according to the preset first configuration information; or, when the obtained working mode category is the user control mode, determining the task type of the current-stage task of the neural network model according to the second configuration information input by the user.
[0012] According to one aspect of the present disclosure, a task allocation device is provided, including: an acquisition module, configured to acquire the running state information during the training process of a neural network model when the working mode category is the automatic switching mode; a determination module, configured to determine the task type of the current stage task of the neural network model according to the running state information; and an allocation module, configured to allocate the current stage task to a multi-core processor according to the task type of the current stage task of the neural network model.
[0013] In a possible implementation manner, the running state information includes first information for indicating that the neural network model is in the forward stage during the training process, and second information for indicating that the neural network model is in the backward stage during the training process. The acquisition module is configured to: in response to detecting that the neural network model calls a data loading function, acquire the first information; or, in response to detecting that the neural network model calls a loss function, acquire the second information.
[0014] In a possible implementation manner, the task type includes a computation-intensive task and a communication-intensive task. The determination module is configured to: according to the first information, determine the task type of the current stage task of the neural network model as a computation-intensive task; or, according to the second information, determine the task type of the current stage task of the neural network model as a communication-intensive task.
[0015] In a possible implementation manner, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. The allocation module is configured to: when the task type of the current stage task of the neural network model is a computation-intensive task, split the operator to be processed in the current stage task into a first split operator and a second split operator, where the computation amount of the first split operator is greater than that of the second split operator; allocate the first split operator to the first processing core and allocate the second split operator to the second processing core.
[0016] In a possible implementation manner, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. The allocation module is configured to: when the task type of the current stage task of the neural network model is a communication-intensive task, allocate the operator to be processed in the current stage task to the first processing core and allocate the communication task of the current stage task to the second processing core.
[0017] In a possible implementation, the allocation module is further configured to: when the task type of the current stage task of the neural network model is a compute-intensive task, in response to detecting that the neural network model calls a communication function, allocate the operators to be processed in the current stage task to the first processing core, and allocate the communication tasks in the current stage task to the second processing core; or, when the task type of the current stage task of the neural network model is a compute-intensive task, in response to detecting that the user adjusts the task type of the current stage task from the compute-intensive task to a communication-intensive task, allocate the operators to be processed in the current stage task to the first processing core, and allocate the communication tasks in the current stage task to the second processing core.
[0018] In a possible implementation, the determination module is further configured to: when the obtained working mode category is the fixed mode, determine the task type of the current stage task of the neural network model according to the preset first configuration information; or, when the obtained working mode category is the user control mode, determine the task type of the current stage task of the neural network model according to the second configuration information input by the user.
[0019] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0020] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0021] In the embodiment of the present disclosure, when the working mode category is the automatic switching mode, the running state information during the training process of the neural network model is obtained; according to the running state information, the task type of the current stage task of the neural network model is determined; according to the task type of the current stage task of the neural network model, the current stage task is allocated to a multi-core processor. In this way, the task allocation method in the embodiment of the present disclosure can automatically determine the task type of the current stage task of the neural network model as a type matching the running state information, such as a communication-intensive task or a compute-intensive task, during the training process of the neural network model, so as to allocate the current stage task to the multi-core processor according to the allocation method for the communication-intensive task or the allocation method for the compute-intensive task, reduce the idle time of each processing core in the multi-core processor, and improve the utilization rate of the multi-core processor.
[0022] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings
[0023] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0024] Figure 1 A flowchart showing a task allocation method according to an embodiment of the present disclosure.
[0025] Figure 2 A schematic diagram showing a task allocation method according to an embodiment of the present disclosure.
[0026] Figure 3 A schematic diagram showing obtaining operating state information according to an embodiment of the present disclosure.
[0027] Figure 4 A schematic diagram showing a state machine of a controller according to an embodiment of the present disclosure.
[0028] Figure 5 A block diagram showing a task allocation device according to an embodiment of the present disclosure.
[0029] Figure 6 A block diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0030] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0031] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0032] The term "and / or" herein merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0033] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0034] In the related art, for a multi-core processor that packages two or more computing cores (Cores) in the same integrated circuit, if a method is adopted where a large processing core is only used for computing tasks and a small processing core is only used for communication tasks, in actual applications, there will be a situation where the small processing core is idle for a long time during the forward stage of the neural network model, resulting in a reduction in the utilization rate of the multi-core processor.
[0035] Alternatively, in the related art, if both the large processing core and the small processing core are used for computing tasks and the small processing core is additionally responsible for communication tasks, then during the backward stage of the neural network model, the small processing core simultaneously undertakes computing tasks and communication tasks, and an optimal performance allocation ratio cannot be achieved, resulting in the large processing core completing the computation in advance and waiting for synchronization, which also causes a reduction in the utilization rate of the multi-core processor.
[0036] In view of this, the task allocation method of the embodiments of the present disclosure can automatically switch between a computationally intensive scenario and a communication-intensive scenario during the training process of the neural network model according to the running state information of the neural network model, and automatically determine the task type of the current stage task of the neural network model as a type matching the running state information, such as a communication-intensive task or a computationally intensive task, so as to allocate the current stage task to the multi-core processor according to the allocation method of the communication-intensive task or the allocation method of the computationally intensive task, reduce the idle time of each processing core in the multi-core processor, and improve the utilization rate of the multi-core processor.
[0037] Figure 1 The flowchart showing the task allocation method according to the embodiments of the present disclosure is as Figure 1 shown, and the task allocation method includes:
[0038] In step S11, when the working mode category is the automatic switching mode, obtain the running state information during the training process of the neural network model;
[0039] In step S12, according to the running state information, determine the task type of the current stage task of the neural network model, and the task type includes a computationally intensive task and a communication-intensive task;
[0040] In step S13, according to the task type of the current stage task of the neural network model, allocate the current stage task to the multi-core processor.
[0041] In a possible implementation, the task allocation method can be applied to electronic devices such as terminal devices or servers. The electronic device may include a main control processor and a multi-core processor serving as a coprocessor. The main control processor can call computer-readable instructions stored in the memory to implement the task allocation method of the embodiments of the present disclosure and allocate tasks to the multi-core processor. The electronic device can be a User Equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a Personal Digital Assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.
[0042] Among them, the main control processor may include, but is not limited to: Central Processing Unit (CPU), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Tensor Processing Unit (TPU), Field Programmable Gate Array (FPGA), etc. The multi-core processor may include, but is not limited to, a Graphics Processing Unit (GPU) with multiple cores, a General-Purpose Computing on Graphics Processing Units (GPGPU), etc. The embodiments of the present disclosure do not limit the types of the main control processor and the multi-core processor.
[0043] In a possible implementation, the task allocation method can be executed by a controller. The controller can be software or program code running in the main control processor and can be implemented through a hardware description language, assembly language, high-level language (such as C, C++), or script language; the controller can also be a logic circuit embedded in the main control processor. The embodiments of the present disclosure do not limit the form of the controller.
[0044] In a possible implementation, the task allocation method can be applied to a distributed training scenario where multiple multi-core processors (such as multi-core GPUs, GPUs with multiple stream processor units) and multiple main control processors (such as CPUs) can communicate with each other. The task allocation method of the embodiments of the present disclosure can effectively allocate tasks to the processing cores of the multi-core processor (such as GPU), thereby improving the speed and efficiency of distributed training.
[0045] In a possible implementation, the controller may have multiple working modes, such as including an automatic switching mode, a fixed mode, and a user control mode.
[0046] Optionally, if the working mode category of the controller is set to the automatic switching mode, the controller will automatically determine whether the task type of the current stage task of the neural network model is a compute-intensive task or a communication-intensive task according to preset rules, conditions, or algorithms, so that when the controller allocates the current stage task to the multi-core processor, it can automatically switch between the allocation method according to the compute-intensive task and the allocation method according to the communication-intensive task.
[0047] Optionally, if the working mode category of the controller is set to the fixed mode, the controller will directly determine whether the task type of the current stage task of the neural network model is a compute-intensive task or a communication-intensive task according to the preset and unchanging settings. For example, assuming that the fixed mode is set to the compute-intensive task scenario, even if the task type of the current stage task of the neural network model is a communication-intensive task, the controller will allocate the current stage task to the multi-core processor according to the allocation method of the compute-intensive task. Another example, assuming that the fixed mode is set to the communication-intensive task scenario, even if the task type of the current stage task of the neural network model is a compute-intensive task, the controller will allocate the current stage task to the multi-core processor according to the allocation method of the communication-intensive task.
[0048] Optionally, if the working mode category of the controller is set to the user control mode, the task type of the current stage task of the neural network model, whether it is a compute-intensive task or a communication-intensive task, can be determined according to the user's instructions through the user control interface. For example, assuming that the user specifies that the task type of the current stage task of the neural network model is a compute-intensive task, even if the task type of the current stage task of the neural network model is a communication-intensive task, the controller will allocate the current stage task to the multi-core processor according to the allocation method of the compute-intensive task. Another example, assuming that the user specifies that the task type of the current stage task of the neural network model is a communication-intensive task, even if the task type of the current stage task of the neural network model is a compute-intensive task, the controller will allocate the current stage task to the multi-core processor according to the allocation method of the communication-intensive task.
[0049] In a possible implementation, in step S11, when the working mode category of the controller is obtained as the automatic switching mode, the running state information during the training process of the neural network model can be obtained; the running state information is used to indicate that the neural network model is in the forward stage or the backward stage during the training process. The forward stage is the stage where the neural network model loads the training data and determines the output result according to the training data. The backward stage is the stage where the neural network model determines the loss result according to the loss function.
[0050] Exemplarily, a neural network model (Neural Networks, NN), such as a convolutional neural network (Convolutional Neural Networks, CNN), a de-convolutional neural network (De-convolutional Neural Networks, DN), a deep neural network (Deep Neural Networks, DNN), a recurrent neural network (Recurrent Neural Network, RNN), a backbone neural network (Backbone Neural Network), etc. The embodiments of the present disclosure do not limit the type of the neural network model.
[0051] Exemplarily, the neural network model may include a forward stage and a backward stage during the training process. In the forward stage of the neural network model, the training data can be sent into the neural network model, and the input training data can be transformed into an output result through the neural network model. This output result is the prediction of the neural network model for the training data. In the backward stage of the neural network model, the neural network model can call the loss function to calculate the loss result (such as error) between the output result and the true label, and then use this loss result to update the weight parameters of the neural network model to make the output result of the neural network model closer to the true label. The training process of the neural network model may include multiple training cycles, and each training cycle may be composed of a forward stage and a backward stage. Through iterative training of multiple cycles, the neural network model can gradually learn how to extract useful features from the input training data and generate more accurate prediction results.
[0052] In a possible implementation, when the running state information is obtained in step S11, in step S12, the task type of the current stage task of the neural network model can be determined according to the running state information. The task type includes a compute-intensive task and a communication-intensive task; among them, a compute-intensive task is a task that includes a large amount of mathematical calculations, logical operations, or data processing and requires a large amount of computing resources to complete. A communication-intensive task is a task that frequently performs data communication or data transmission.
[0053] Through the analysis of the training task of the neural network model, the forward stage in the neural network model involves computing tasks, and the backward stage in the neural network model involves computing tasks and a large number of communication tasks. Therefore, if the running status information indicates that the neural network model is in the forward stage of the training process, the task type of the current stage task of the neural network model can be determined as a compute-intensive task; if the running status information indicates that the neural network model is in the backward stage of the training process, the task type of the current stage task of the neural network model can be determined as a communication-intensive task.
[0054] In a possible implementation, after determining the task type of the current stage task of the neural network model in step S12, in step S13, the current stage task can be allocated to the multi-core processor according to the task type of the current stage task of the neural network model.
[0055] For example, assume that the multi-core processor has M (M≥2) processing cores, and the computing capabilities of the M processing cores are different. The computing capability of the Kth (1≤K≤M) processing core is the smallest among the M processing cores. Among them, the smaller the scale of the processing core, the smaller its computing throughput and the worse its corresponding computing capability; the larger the scale of the processing core, the larger its computing throughput and the better its corresponding computing capability. Here, the computing capability of the processing core is the ability to execute a certain number of operations (such as floating-point operations) per second, and the scale of the processing core represents the number of electronic components or logic gates integrated in the processor core.
[0056] If the task type of the current stage task of the neural network model is a compute-intensive task, the current stage task can be split into M parts according to the computing capability ratio among the M processing cores and sent to the M processing cores of the multi-core processor respectively. Among them, the Kth processing core with the smallest computing capability is used to process the computing task.
[0057] If the task type of the current stage task of the neural network model is a communication-intensive task, the communication tasks in the current stage task can be sent to the Kth processing core with the smallest computing capability, and the Kth processing core with the smallest computing capability is used to process the communication tasks; and, the other tasks in the current stage task except the communication tasks are sent to the other processing cores with greater computing capabilities of the multi-core processor for computing.
[0058] In this way, in the communication-intensive stage, the small processing core can be fully or mostly used for communication tasks; in the compute-intensive stage, the small processing core can be fully or mostly used for computing tasks. Thus, the task load of the small processing core is overall optimized, which is beneficial to improving the performance of the multi-core processor.
[0059] Through steps S11 to S13, when the obtained working mode category is the automatic switching mode, the running state information during the training process of the neural network model can be obtained, and according to the task type of the current stage task of the neural network model determined by the running state information, the current stage task can be allocated to the multi-core processor. In this way, according to the characteristics in the neural network model training scenario (such as including the forward stage and the backward stage indicated by the running state information during the training process), the allocation mechanism of the current stage task can be automatically changed, and according to the estimated communication situation (such as a communication-intensive scenario may occur in the backward stage), it can be decided in advance whether to allocate the computing task to a specific core (such as the processing core with the minimum computing power) to improve the overall utilization rate, reduce the latency, and improve the performance.
[0060] Taking the controller as the execution entity as an example below, the task allocation method of the embodiments of the present disclosure will be exemplarily described. The controller can be software or program code running in the main control processor, or a logic circuit embedded in the main control processor. The embodiments of the present disclosure do not limit the form of the controller.
[0061] Figure 2 A schematic diagram showing the task allocation method according to the embodiments of the present disclosure is as Figure 2 shown. The running state information of the neural network model can be specified by the user, and the controller allocates the current stage task to the multi-core processor in the way specified by the user; or, during the training process of the neural network model, the controller can automatically modify the running state information of the neural network model based on the analysis of the data loading function and the loss function, thereby changing the allocation method of the controller to allocate the current stage task to the multi-core processor. Among them, other information may include the state information of the multi-core processor, environmental variables, configuration parameters, etc. The embodiments of the present disclosure do not limit this.
[0062] In step S11, when the working mode category of the controller is the automatic switching mode, the running state information during the training process of the neural network model can be obtained. The running state information includes the first information for indicating that the neural network model is in the forward stage during the training process, and the second information for indicating that the neural network model is in the backward stage during the training process.
[0063] In a possible implementation manner, step S11 may include: in response to detecting that the neural network model calls the data loading function, obtaining the first information; or, in response to detecting that the neural network model calls the loss function, obtaining the second information.
[0064] In the example, the data loading function is a function that reads training data from various data sources (such as files, databases, etc.) into the neural network model. For example, the data loading function can be used to read files, parse data, handle missing values, convert data types, normalize or standardize data, split the dataset, etc. The embodiments of the present disclosure do not limit the specific content of the data loading function.
[0065] In the example, the loss function is a function that maps the output result predicted by the neural network model to a non - negative real number to represent its loss. For example, the loss function can be used to quantify the inconsistency between the prediction result of the neural network model and the actual label, and optimize the parameters of the neural network model by minimizing the loss function, thereby improving the prediction performance of the neural network model. The embodiments of the present disclosure do not limit the specific content of the loss function.
[0066] Figure 3 A schematic diagram showing the acquisition of the running state information according to an embodiment of the present disclosure is as Figure 3 As shown, considering that the forward stage in the neural network model involves computational tasks, while the backward stage involves a large number of communication tasks. And, during the training process of the neural network model, the data loading function is at the beginning of the forward stage of the neural network model, and the loss function is in the middle of the forward stage and the backward stage of the neural network model. Therefore, by adding logic to modify the controller state to the data loading function and the loss function, the automatic switching of the running state information of the neural network model can be achieved.
[0067] As Figure 3 shown, the neural network model will repeatedly execute the forward stage and the backward stage during the training process until the training is completed. The forward stage and the backward stage can be distinguished by the calls of the data loading function and the loss function, which serve as the demarcation points for distinguishing the forward stage and the backward stage. Adding logic to change the state to the two sites as Figure 3 shown (see the data loading function and the loss function in Figure 3 ) can achieve the automatic switching of the running state information of the neural network model.
[0068] For example, in the data loading function and the loss function, an interface for adjusting the state of the controller can be inserted to enable the controller to automatically obtain relevant information during the operation of the neural network model. This state adjustment interface can be a software interface and an interrupt interface, which are used to detect whether the data loading function or the loss function is called in the current stage of the neural network model. If the controller detects that the neural network model calls the data loading function, the running state information will change to indicate that the neural network model is in the forward stage during the training process; if the controller detects that the neural network model calls the loss function, the running state information will change to indicate that the neural network model is in the backward stage during the training process. Through the state adjustment interface of the controller, the running state information of the neural network model during the training process can be obtained.
[0069] By setting the running state information of the neural network model, the running stage of the neural network model can be automatically identified by detecting the call conditions of the data loading function and the loss function, which supports the automatic identification of the running state information of various neural network models and is beneficial to automatically selecting an appropriate allocation method when allocating the current stage tasks to multi-core processors subsequently.
[0070] When the running state information is obtained in step S11 (for example, the first information indicating that the neural network model is in the forward stage during the training process, or the second information indicating that the neural network model is in the backward stage during the training process), the task type of the current stage task of the neural network model can be determined according to the running state information in step S12, and the task type includes computationally intensive tasks and communication-intensive tasks.
[0071] In a possible implementation manner, step S12 may include: determining the task type of the current stage task of the neural network model as a computationally intensive task according to the first information; or determining the task type of the current stage task of the neural network model as a communication-intensive task according to the second information.
[0072] In the forward stage, the neural network model receives the input training data and performs forward propagation through its various network layers (such as convolutional layers, fully connected layers, etc.) to calculate the predicted output result. It can be seen that the forward stage indicated by the first information is a computationally intensive stage, involving a large number of matrix operations and calculations of activation functions, and the execution task in the forward stage of the neural network model can be determined as a computationally intensive task.
[0073] In the backward phase, the neural network model calculates the loss result between the output result determined in the forward phase and the expected result through the loss function, and updates the weights of the neural network model according to the loss result to reduce the loss result. For example, in a distributed training scenario, multiple multi-core processors (such as multi-core GPUs, GPUs with multiple stream processor units), and multiple master processors (such as CPUs) can communicate with each other. The backward phase indicated by the second information includes a large number of communication tasks, which will communicate frequently between multiple multi-core processors (such as multi-core GPUs) and between multi-core processors and master processors (such as CPUs). The proportion of communication-intensive tasks will increase significantly. The execution tasks in the backward phase of the neural network model can be determined as communication-intensive tasks.
[0074] It should be understood that whether it is a GPU with multiple stream processor units or a CPU with several processing cores, in the task splitting of the embodiments of the present disclosure, no additional communication should occur between stream processors or between processing cores (the communication such as its own data synchronization is not within the scope of discussion of this solution). Its communication occurs in the scenario where multiple GPUs and multiple CPUs are interconnected, that is, the distributed communication scenario.
[0075] Determining the task types of the forward phase and the backward phase in the training process of the neural network model as compute-intensive tasks and communication-intensive tasks respectively is beneficial to optimizing resource allocation and improving training efficiency.
[0076] If the task type of the current phase task of the neural network model is determined in step S12, in step S13, the current phase task can be allocated to the multi-core processor according to the task type of the current phase task of the neural network model.
[0077] As Figure 2 shown, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. When the controller determines that the neural network model is in a compute-intensive scenario, the splitting task of the operator to be processed is normally performed. Otherwise, if the neural network model is in a communication-intensive scenario, the communication task is preferentially allocated to the second processing core as a small core, and the operator to be processed is not decomposed.
[0078] Among them, the operator to be processed can be a convolution operator, a deconvolution operator, an activation operator, a fully connected operator, a pooling operator, a batch normalization operator, etc., or a combination of various operators. The embodiments of the present disclosure do not specifically limit the operator to be processed.
[0079] In a possible implementation, step S13 may include: when the task type of the current stage task of the neural network model is a compute-intensive task, splitting the operator to be processed in the current stage task into a first split operator and a second split operator, where the computational amount of the first split operator is greater than that of the second split operator; allocating the first split operator to the first processing core and the second split operator to the second processing core.
[0080] Exemplarily, assume that the ratio of the computing power of the first processing core to that of the second processing core in a multi-core processor is S1:S2, that is, the ratio of the number of operations executed by the first processing core and the second processing core per unit time is S1:S2, where S1 > S2. For a compute-intensive task, the controller may split the operator to be processed in the current stage task into a first split operator and a second split operator according to the ratio S1:S2 of the computing power of the first processing core to that of the second processing core. The computational amount of the first split operator is S1 / (S1 + S2) of the computational amount of the operator to be processed, and the computational amount of the second split operator is S2 / (S1 + S2) of the computational amount of the operator to be processed. The first split operator with a computational amount of S1 / (S1 + S2) of the computational amount of the operator to be processed may be allocated to the first processing core via the main computation stream, and the second split operator with a computational amount of S2 / (S1 + S2) of the computational amount of the operator to be processed may be allocated to the second processing core via the secondary computation stream.
[0081] In this way, the processing cores can be allocated according to the computational amount of the split operators, which can make more reasonable use of the computing resources of the multi-core processor. For the first split operator with a larger computational amount, it is allocated to the first processing core with higher performance. For the second split operator with a smaller computational amount, it can be allocated to the second processing core with slightly lower performance but still meeting the requirements. Allocating each split operator to different processing cores according to the size of its computational amount can achieve a more balanced load distribution, thereby avoiding resource waste and improving the processing efficiency of the multi-core processor.
[0082] In a possible implementation, step S13 may include: when the task type of the current stage task of the neural network model is a communication-intensive task, allocating the operator to be processed in the current stage task to the first processing core and allocating the communication task in the current stage task to the second processing core.
[0083] The communication-intensive task may include the operator to be processed (i.e., the computing task) and the communication task. The controller may allocate the operator to be processed to the first processing core via the main computation stream, while the communication task is allocated to the second processing core via the secondary computation stream. The computing and communication processes can be separated, and then specific optimizations can be carried out for the computing task and communication, improving the overall execution efficiency.
[0084] It can be seen that the controller can adjust whether to split the operator to be processed and the timing of allocating the operator to be processed or the split operators after splitting to each processing core in the multi-processor according to the task type of the current stage task, so as to avoid resource waste and improve the processing efficiency of the multi-core processor.
[0085] In order to comprehensively analyze the operation state information of the neural network model and reduce the abnormal decomposition in some special scenarios, an automatic state machine can be designed for the controller to realize automatic state switching. Among them, by adding a state machine to the controller, the controller can comprehensively process the existing operation state information, accept the automatically changed operation state information from the neural network model, process the mode changes specified by the user, and minimize the occurrence of abnormal allocation in some special cases, so as to support its own functions.
[0086] Figure 4 A schematic diagram of the state machine of the controller according to an embodiment of the present disclosure is shown, as Figure 4 shown. In the initial state, an automatic switching mode or a fixed mode can be selected according to the configured mode. For example, Figure 3 the other information in may include control parameters for indicating the working mode category of the configured controller. The controller can, according to the indication of the control parameters, set the working mode category to the fixed mode when the control parameter is the first identifier; and set the working mode category to the automatic switching mode when the control parameter is the second identifier.
[0087] Optionally, when the obtained working mode category is the fixed mode, determine the task type of the current stage task of the neural network model according to the preset first configuration information, where the task type includes a computation-intensive task and a communication-intensive task; and allocate the current stage task to the multi-core processor according to the task type of the current stage task of the neural network model.
[0088] For example, if the configuration parameter in the first configuration information is the first value, regardless of whether the task type of the current stage task of the neural network model is a computation-intensive task or a communication-intensive task, the task type of the current stage task of the neural network model can be determined as a computation-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the computation-intensive task. Among them, the allocation method of the computation-intensive task can be called the data splitting mode, and this allocation method will split the operator to be processed in the current stage task and allocate each split operator after splitting to the corresponding processing core in the multi-core processor. The small-scale processing cores (corresponding to the processing cores with low computing power) in the multi-core processor will also receive the corresponding split operators.
[0089] For another example, if the configuration parameter in the first configuration information is the second value, regardless of whether the task type of the current stage task of the neural network model is a compute-intensive task or a communication-intensive task, the task type of the current stage task of the neural network model can be determined as a communication-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the communication-intensive task. Among them, the allocation method of the communication-intensive task can be called the communication priority mode. This allocation method will not split the operators to be processed in the current stage task, but will allocate the communication tasks in the current stage task to the small-scale processing cores (corresponding to the processing cores with low computing power) in the multi-core processor, and allocate the operators to be processed in the current stage task to the large-scale processing cores (corresponding to the processing cores with high computing power) in the multi-core processor.
[0090] Optionally, in the case where the obtained working mode category is the automatic switching mode, since the training of the neural network model first enters the forward stage, the task type of the initial stage task of the neural network model can be preferentially determined as a compute-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the compute-intensive task (that is, the data splitting mode).
[0091] Then, during the training process of the neural network model, the running stage of the neural network model, whether it is the forward stage or the backward stage, can be automatically identified by detecting the call conditions of the data loading function and the loss function. If it is identified that the running stage of the neural network model is the forward stage, the task type of the current stage task of the neural network model can be determined as a compute-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the compute-intensive task (that is, the data splitting mode); if it is identified that the running stage of the neural network model is the backward stage, the task type of the current stage task of the neural network model can be determined as a communication-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the communication-intensive task (that is, the communication priority mode).
[0092] In the scenario of the automatic switching mode, the controller can automatically switch between the communication priority mode and the data splitting mode according to the stage change of the neural network model.
[0093] In a possible implementation manner, if the working mode category of the controller is the automatic switching mode, when the task type of the current stage task of the neural network model is a compute-intensive task, in response to detecting that the neural network model calls the communication function, the operators to be processed in the current stage task are allocated to the first processing core, and the communication tasks in the current stage task are allocated to the second processing core.
[0094] Alternatively, when the task type of the current-stage task of the neural network model is a compute-intensive task, in response to detecting that the user adjusts the task type of the current-stage task from the compute-intensive task to a communication-intensive task, allocate the operators to be processed in the current-stage task to the first processing core, and allocate the communication tasks in the current-stage task to the second processing core.
[0095] As Figure 4 shown, in the scenario of the automatic switching mode, if a special communication scenario occurs, such as the communication-computation hybrid mode in large model training, the communication function will be called in the data splitting mode, and the way for the controller to allocate the current-stage task to the multi-core processor can be changed from the data splitting mode to the communication priority mode.
[0096] In addition, since the forward stage in the distributed training scenario of large language models also belongs to communication-intensive tasks, at this time, according to the configuration information or environment variables input by the user, the task type of the current-stage task is automatically adjusted to a communication-intensive task, and the controller can automatically correct the preset compute-intensive scenario and switch to the communication-intensive scenario to further optimize the performance of the multi-core processor.
[0097] In this way, the applicable scope of the task allocation method of the embodiments of the present disclosure can be extended, and the resource utilization rate of the multi-core processor can be further improved.
[0098] Optionally, when the obtained working mode category is the user control mode, determine the task type of the current-stage task of the neural network model according to the second configuration information input by the user. Among them, the user control mode has the highest priority. Whether in the automatic switching mode or in the fixed model, the specified mode of user control can be entered through the user control interface.
[0099] For example, if the configuration parameter in the second configuration information input by the user is the third value, regardless of the task allocation method determined by the automatic switching mode and the fixed model, whether the task type of the current-stage task of the neural network model is a compute-intensive task or a communication-intensive task, the task type of the current-stage task of the neural network model can be determined as a compute-intensive task, and the current-stage task can be allocated to the multi-core processor according to the allocation method of the compute-intensive task (that is, the data splitting mode).
[0100] For another example, if the configuration parameter in the second configuration information input by the user is the fourth value, regardless of the task allocation method determined by the automatic switching mode and the fixed model, whether the task type of the current stage task of the neural network model is a compute-intensive task or a communication-intensive task, the task type of the current stage task of the neural network model can be determined as a communication-intensive task, and the current stage task can be allocated to the multi-core processor according to the allocation method of the communication-intensive task (i.e., the communication priority mode).
[0101] In this way, the allocation method of the controller for allocating the current stage task to the multi-core processor can be adjusted through the user interface or automatic recognition, which is more flexible and efficient.
[0102] According to the task allocation method of the embodiments of the present disclosure, when the obtained working mode category is the automatic switching mode, the running state information during the training process of the neural network model can be obtained, and the current stage task of the neural network model can be allocated to the multi-core processor according to the task type of the current stage task of the neural network model determined by the running state information. In this way, according to the characteristics in the neural network model training scenario (such as the forward stage and the backward stage indicated by the running state information during the training process), the allocation mechanism of the current stage task can be automatically changed, and according to the estimated communication situation (such as the communication-intensive scenario that will occur in the backward stage), it can be determined in advance whether to allocate the computing task to a specific core (such as the processing core with the minimum computing power) to improve the overall utilization rate, reduce the delay, and improve the performance.
[0103] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0104] In addition, the present disclosure also provides a task allocation device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any task allocation method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated further.
[0105] Figure 5 The block diagram showing the task allocation device according to the embodiments of the present disclosure is as Figure 5 shown, and the device includes:
[0106] An obtaining module 51, configured to obtain the running state information during the training process of the neural network model when the working mode category is the automatic switching mode;
[0107] A determination module 52, configured to determine a task type of the current-stage task of the neural network model according to the operation status information;
[0108] An allocation module 53, configured to allocate the current-stage task to a multi-core processor according to the task type of the current-stage task of the neural network model.
[0109] In a possible implementation manner, the operation status information includes first information for indicating that the neural network model is in a forward stage during a training process, and second information for indicating that the neural network model is in a backward stage during the training process. The acquisition module 51 is configured to: in response to detecting that the neural network model calls a data loading function, acquire the first information; or, in response to detecting that the neural network model calls a loss function, acquire the second information.
[0110] In a possible implementation manner, the task type includes a computation-intensive task and a communication-intensive task. The determination module 52 is configured to: according to the first information, determine the task type of the current-stage task of the neural network model as a computation-intensive task; or, according to the second information, determine the task type of the current-stage task of the neural network model as a communication-intensive task.
[0111] In a possible implementation manner, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. The allocation module 53 is configured to: when the task type of the current-stage task of the neural network model is a computation-intensive task, split the operator to be processed in the current-stage task into a first split operator and a second split operator, where the amount of computation of the first split operator is greater than that of the second split operator; allocate the first split operator to the first processing core, and allocate the second split operator to the second processing core.
[0112] In a possible implementation manner, the multi-core processor includes at least a first processing core and a second processing core, and the computing power of the first processing core is greater than that of the second processing core. The allocation module 53 is configured to: when the task type of the current-stage task of the neural network model is a communication-intensive task, allocate the operator to be processed in the current-stage task to the first processing core, and allocate the communication task in the current-stage task to the second processing core.
[0113] In a possible implementation, the allocation module 53 is further configured to: when the task type of the current stage task of the neural network model is a compute-intensive task, in response to detecting that the neural network model calls a communication function, allocate the operators to be processed in the current stage task to the first processing core, and allocate the communication tasks in the current stage task to the second processing core; or, when the task type of the current stage task of the neural network model is a compute-intensive task, in response to detecting that the user adjusts the task type of the current stage task from the compute-intensive task to a communication-intensive task, allocate the operators to be processed in the current stage task to the first processing core, and allocate the communication tasks in the current stage task to the second processing core.
[0114] In a possible implementation, the determination module 52 is further configured to: when the obtained working mode category is the fixed mode, determine the task type of the current stage task of the neural network model according to the preset first configuration information; or, when the obtained working mode category is the user control mode, determine the task type of the current stage task of the neural network model according to the second configuration information input by the user.
[0115] This method has a specific technical association with the internal structure of the computer system and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing the data storage amount, reducing the data transmission amount, and increasing the hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in line with the natural law.
[0116] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.
[0117] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0118] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.
[0119] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of the electronic device, the processor in the electronic device executes the above methods.
[0120] The electronic device may be provided as a terminal, a server, or a device in other forms.
[0121] Figure 6 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 6 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0122] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Microsoft Server Operating System (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.
[0123] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0124] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0125] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0126] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0127] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0128] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0129] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0130] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0131] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0132] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), and so on.
[0133] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0134] Those skilled in the art can understand that in the above methods of the specific implementation manners, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0135] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is considered consent to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0136] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A task allocation method, characterized in that: The method comprises: When the working mode category is the automatic switching mode, the operating status information during the training process of the neural network model is obtained; Determine the task type of the current stage task of the neural network model according to the running status information; According to the task type of the current stage task of the neural network model, the current stage task is allocated to the multi-core processor.
2. The method according to claim 1, characterized in that The running status information includes first information for indicating that the neural network model is in a forward phase during training, and second information for indicating that the neural network model is in a reverse phase during training. The obtaining of the running status information during the training of the neural network model includes: In response to detecting that the neural network model calls a data loading function, acquiring the first information; Alternatively, in response to detecting that the neural network model calls a loss function, the second information is obtained.
3. The method according to claim 2, characterized in that The task types include computationally intensive tasks and communication intensive tasks, and determining the task type of the task of the current stage of the neural network model according to the running status information includes: According to the first information, determining the task type of the task of the current stage of the neural network model as a computationally intensive task; Alternatively, based on the second information, the task type of the task at the current stage of the neural network model is determined to be a communication intensive task.
4. The method according to any one of claims 1 to 3, characterized in that The multi-core processor includes at least a first processing core and a second processing core, wherein the computing capability of the first processing core is greater than the computing capability of the second processing core. The allocating the current stage task to the multi-core processor according to the task type of the current stage task of the neural network model includes: In the case where the task type of the task at the current stage of the neural network model is a computationally intensive task, splitting the to-be-processed operator of the task at the current stage into a first split operator and a second split operator, the computation amount of the first split operator being greater than the computation amount of the second split operator; The first split operator is assigned to the first processing core, and the second split operator is assigned to the second processing core.
5. The method according to any one of claims 1 to 3, characterized in that: The multi-core processor includes at least a first processing core and a second processing core, wherein the computing capability of the first processing core is greater than the computing capability of the second processing core. The allocating the current stage task to the multi-core processor according to the task type of the current stage task of the neural network model includes: When the task type of the current stage task of the neural network model is a communication intensive task, the to-be-processed operators of the current stage task are allocated to the first processing core, and the communication tasks of the current stage task are allocated to the second processing core.
6. The method according to claim 5, characterized in that The method further includes: when the task type of the task of the current stage of the neural network model is a computationally intensive task, in response to detecting that the neural network model calls a communication function, allocating the to-be-processed operator of the task of the current stage to the first processing core, and allocating the communication task in the task of the current stage to the second processing core; Alternatively, in a case where the task type of the current stage task of the neural network model is a computationally intensive task, in response to detecting that the user adjusts the task type of the current stage task from the computationally intensive task to a communication intensive task, the to-be-processed operators of the current stage task are allocated to the first processing core, and the communication tasks in the current stage task are allocated to the second processing core.
7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In the case where the acquired working mode category is a fixed mode, determining the task type of the task of the current stage of the neural network model according to the preset first configuration information; Alternatively, when the acquired working mode category is the user control mode, the task type of the current stage task of the neural network model is determined according to the second configuration information input by the user.
8. A task allocation device, characterized in that: include: An acquisition module, used to acquire the operation status information during the training process of the neural network model when the working mode category is the automatic switching mode; A determination module, used to determine the task type of the task at the current stage of the neural network model according to the running status information; The allocation module is used to allocate the current stage tasks to the multi-core processor according to the task type of the current stage tasks of the neural network model.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.