Task scheduling method of multi-core processor, storage medium, electronic device and computer program product

By scheduling layer fusion convolution tasks on multi-core processors through the Dynamic Memory Inter-Layer Allocation (DMIA) mechanism, the problems of low computing performance and memory access efficiency of multi-core processors are solved, and more efficient data processing and storage are achieved.

CN121597352APending Publication Date: 2026-03-03ZTE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411169309.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Multi-core processors have many problems in terms of computing performance and memory access efficiency, including inter-core system overhead, data contention and load imbalance, which affect the computing performance and memory access efficiency of neural network models.

Method used

The Dynamic Memory Interlayer Allocation (DMIA) mechanism is adopted. By determining the task information and memory flow order of the layer fusion convolution task, multiple cores are scheduled to execute the layer fusion convolution task, thereby realizing parallel processing and memory flow of multi-core processors and optimizing data storage and access.

Benefits of technology

It improves the processing efficiency of multi-core processors, reduces memory access conflicts, and enhances computing performance and memory access efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597352A_ABST
    Figure CN121597352A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task scheduling method of a multi-core processor, a storage medium, an electronic device and a computer program product. The method comprises the following steps: determining task information of a layer fusion convolution task and a memory circulation sequence of a plurality of cores in the multi-core processor; and scheduling the multiple core execution layer fusion convolution tasks according to the task information and the memory circulation sequence. According to the embodiment of the invention, a multi-core processor processing layer can be utilized to fuse a convolution task and perform parallel processing on multiple cores, so that the processing efficiency is improved, memory circulation can be performed among the multiple cores, the data storage and access efficiency is improved, the memory access conflict is reduced, and the memory access efficiency is improved. Therefore, the technical problem that the computing performance and the memory access efficiency of a multi-core processor are affected in the related art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to a task scheduling method for a multi-core processor, a storage medium, an electronic device, and a computer program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, neural network models have demonstrated outstanding performance in fields such as image recognition, natural language processing, and autonomous driving. However, these models typically require enormous computing resources and data throughput capabilities. To address this issue, engineers have developed specialized hardware accelerators, such as Neural Processing Units (NPUs), which are optimized for the computational characteristics of neural networks in order to achieve higher processing efficiency and lower energy consumption.

[0003] The development of hardware accelerators has evolved from single-core to multi-core. While single-core accelerators offered good performance in the early stages, their performance improvement gradually became limited as model complexity increased. Multi-core accelerators, through parallel processing techniques, can execute computational tasks simultaneously on multiple cores, thereby significantly improving processing power.

[0004] Many data scheduling models specifically for convolutional neural network processors have been proposed in related technologies. Single-core data flow includes layer fusion scheduling, skewed layer fusion scheduling, and row stationary scheduling. Multi-core data flow typically stacks the number of cores on top of single-core data flow, using parallel processing techniques to improve system performance. However, multiple cores also present numerous problems such as inter-core system overhead, data contention and conflicts, and load imbalance, which directly affect the overall system's computational performance and memory access efficiency. Therefore, deploying neural network models on multi-core processors introduces greater complexity to data flow design, and the data flow pattern significantly impacts the processor's computational and memory access performance.

[0005] In conclusion, there is still no good solution in the relevant technologies to address the above problems. Summary of the Invention

[0006] This application provides a task scheduling method, storage medium, electronic device, and computer program product for a multi-core processor, to at least solve the technical problem that the computing performance and memory access efficiency of multi-core processors are affected in the related art.

[0007] According to one embodiment of this application, a task scheduling method for a multi-core processor is provided. The method includes: determining task information of a layer fusion convolution task and the memory flow order of multiple cores in a multi-core processor; and scheduling the multiple cores to execute the layer fusion convolution task according to the task information and the memory flow order.

[0008] According to yet another embodiment of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein the computer program is executed by a processor to perform the steps in any of the above method embodiments.

[0009] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0010] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0011] This application presents a task scheduling method for multi-core processors, which can utilize the multi-core processor to process fused convolutional tasks in parallel across multiple cores, thereby improving processing efficiency. Furthermore, by transferring memory between multiple cores, it can improve data storage and access efficiency and reduce memory access conflicts. This can solve the technical problem in related technologies where the computing performance and memory access efficiency of multi-core processors are affected. Attached Figure Description

[0012] Figure 1 This is a hardware structure block diagram of a computer terminal for a task scheduling method for a multi-core processor according to an embodiment of this application.

[0013] Figure 2 This is a flowchart of a task scheduling method for a multi-core processor according to an embodiment of this application;

[0014] Figure 3 This is a schematic diagram of a sub-layer fusion convolution task in one embodiment of this application;

[0015] Figure 4 This is a schematic diagram of layer-by-layer fusion convolution scheduling and layer-by-layer convolution scheduling in one embodiment of this application;

[0016] Figure 5 This is a schematic diagram of data flow in a multi-core system according to an embodiment of this application;

[0017] Figure 6 This is a timing diagram of a multi-core layer fusion convolution task with dynamic memory inter-layer allocation in one embodiment of this application;

[0018] Figure 7 This is a timing diagram of a multi-core layer fusion convolution task with non-dynamic memory allocation between layers in related technologies. Detailed Implementation

[0019] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0021] This application provides a task scheduling method for multi-core processors, which can schedule layer fusion convolution tasks in multi-core processors. This scheme is based on data flow between multiple cores and is also known as Dynamic Memory Inter-layer Allocation (DMIA).

[0022] The methods and embodiments provided in this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for a task scheduling method for a multi-core processor according to an embodiment of this application, as shown below. Figure 1 As shown, a hardware board may include one or more ( Figure 1 Only one is shown in the diagram. A processor 12 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 14 for storing data are also shown. The computer terminal may further include a transmission device 16 for communication functions and an input / output device 18. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0023] The memory 14 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the task scheduling method of the multi-core processor in this embodiment. The processor 12 executes various functional applications and the task scheduling method of the multi-core processor by running the computer program stored in the memory 14, thus implementing the above-described method. The memory 14 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 14 may further include memory remotely located relative to the processor 12, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0024] The transmission device 16 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a telecommunications provider. In one example, the transmission device 16 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 16 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0025] This embodiment provides a task scheduling method for a multi-core processor running on the aforementioned computer terminal. Figure 2 This is a flowchart of a task scheduling method for a multi-core processor according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0026] Step S202: Determine the task information of the layer fusion convolution task and the memory flow order of multiple cores in the multi-core processor;

[0027] Step S204: Schedule the multiple cores to execute the layer fusion convolution task according to the task information and the memory flow order.

[0028] In this embodiment, the layer fusion convolution task is used to fuse feature map data from multiple convolutional layers into feature map data from a single convolutional layer. The fusion method involves dividing an input feature map into multiple tiles, performing convolution calculations on each tile individually, and then synthesizing the fused results of each tile across multiple convolutional layers into a single output feature map. Therefore, the task information for the layer fusion convolution task includes, but is not limited to, the number of convolutional layers, the processing order of the convolutional layers, the feature map size, the tile size, and the number of tiles. For example, since the layer fusion convolution task in this application is executed by a multi-core processor, the task information may also include information about the multi-core processor executing the task.

[0029] Through the above steps, multi-core processors can be used to process layer fusion convolution tasks, with multiple cores processing in parallel, thus improving processing efficiency. Furthermore, memory transfer between multiple cores improves data storage and access efficiency and reduces memory access conflicts, thereby solving the technical problem of multi-core processors affecting computing performance and memory access efficiency in related technologies.

[0030] In this embodiment, the memory flow order is the flow order of intermediate feature data of the layer fusion convolution task among multiple cores. Among them, between two adjacent cores, the intermediate feature data output by the previous core is the input data of the next core, and the intermediate feature data output by the last core among multiple cores is the input data of the first core.

[0031] In some embodiments, the task information includes the number of convolutional layers, the number of slices, and the convolutional processing order of the multiple convolutional layers. The layer fusion convolution task is used to perform convolutional processing layer by layer on the feature data of slices with the same number in multiple convolutional layers according to the convolutional processing order. Each convolutional layer is divided into the same number of slices, and the slices in each convolutional layer are numbered sequentially. Further, the number of slices in each convolutional layer is the number of slices specified in the task information, the number of layers to be convolved in each slice is the number of convolutional layers specified in the task information, and the order in which each slice is convolved is the convolutional processing order specified in the task information.

[0032] In this embodiment, the convolution processing order is used to describe the data dependencies between multiple convolutional layers. Except for the first convolutional layer, the calculations of other convolutional layers all use the calculation results of the previous convolutional layer as input data. Therefore, there are data dependencies between different convolutional layers. For example, in a layer fusion convolution task, if the convolution processing order is L1-L2-L3, then each slice must complete the convolution operation of L1 before starting the convolution operation of L2. This also means that the result of the convolution operation of each slice in L1 is the input data for the convolution operation of that slice in L2, that is, the data of L2 depends on the data of L1 (slice numbers are the same).

[0033] In this embodiment, according to the convolution processing order, the input data of the first convolutional layer is the initial input feature map of the entire fusion convolution task. The initial input feature map needs to be read from external memory, and only one slice of the first convolutional layer is read at a time. In two adjacent convolutional layers, the output data of the previous convolutional layer is the input data of the next convolutional layer, and the output data of the last convolutional layer among multiple convolutional layers is the target output feature map of the layer fusion convolution task.

[0034] In the embodiments of this application, the processing unit of the layer fusion convolution task is a stack, which is a series of consecutive convolutional layers that are fused together for computation. A convolutional network can be divided into multiple stacks.

[0035] In some embodiments, step S204 may include the following steps:

[0036] Step S2042: Divide the layer fusion convolution task into multiple subtasks according to the number of convolutional layers and the number of slices. Each subtask is used to instruct one core of the multi-core processor to perform convolution processing on the feature data of one slice in a convolutional layer.

[0037] Step S2044: Determine the timing of multiple subtasks on multiple cores based on the convolution processing order and memory flow order;

[0038] Step S2046: According to the task time sequence, multiple cores are scheduled to execute multiple sub-tasks to complete the layer fusion convolution task.

[0039] In some embodiments, step S2044, which determines the task timing of multiple subtasks on multiple cores based on the convolution processing order and memory flow order, may include: determining the core number and time period number of multiple subtasks based on preset task scheduling rules, convolution processing order, and memory flow order; and determining the task timing of multiple subtasks on multiple cores based on the core number and time period number of multiple subtasks.

[0040] In some embodiments, the preset task scheduling rules include:

[0041] Task scheduling rule 1: For multiple subtasks with the same slice number in different convolutional layers, the multiple subtasks are scheduled to execute convolution processing sequentially in multiple consecutive time periods according to the convolution processing order, and the core for executing the subtask is selected from multiple cores according to the memory flow order.

[0042] Task scheduling rule 2: For multiple subtasks with different slice numbers in the same convolutional layer, the multiple subtasks are scheduled to be executed sequentially by the same core in multiple consecutive time periods based on the number of cores in the multi-core processor.

[0043] In this embodiment, subtasks can be marked and grouped according to convolutional layer number and slice number. For example, multiple subtasks with the same slice number can be grouped together and scheduled within the group according to task scheduling rule 1. Alternatively, multiple subtasks with the same convolutional layer number can be grouped together and scheduled within the group according to task scheduling rule 2.

[0044] In an exemplary embodiment, the specific implementation of task scheduling rule 1 may include: assuming the slice number is T1, the subtasks of T1 in multiple convolutional layers can be sequentially scheduled to execute convolution processing in multiple consecutive time periods according to the convolution processing order. Further, if the number of convolutional layers is 5, the convolution processing order is L1-L2-L3-L4-L5, and the time period number of L1T1 is P1, then the time period numbers of the subtasks in the 5 convolutional layers corresponding to T1 are P1, P2, P3, P4, and P5, respectively. Other slices are similar to T1 and will not be described in detail here.

[0045] In one exemplary embodiment, the specific implementation of task scheduling rule 1 may further include: assuming the slice number is T1, the number of convolutional layers is 5, the convolution processing order is L1-L2-L3-L4-L5, the memory flow order is a loop of C1-C2-C3, and the core number of L1T1 is C1, then the core numbers of the subtasks of the 5 convolutional layers corresponding to T1 are C1, C2, C3, C1, and C2, respectively. Other slices are similar to T1 and will not be described in detail here.

[0046] In an exemplary embodiment, the specific implementation of task scheduling rule 2 may include: assuming the number of cores is C, then C consecutive subtasks corresponding to C consecutive slices in the same convolutional layer are sequentially executed by the same core for C consecutive time periods, where C is an integer greater than 1. It should be noted that when the number of remaining slices is less than C, the number of subtasks in the same convolutional layer processed by the same core is equal to the number of remaining slices. For example, assuming the convolutional layer is numbered L1 and the number of cores is 3, if the core number of L1T1 is C1, then the core numbers of L1T2 and L1T3 are also C1. If the time period number of L1T1 is P1, then the time period numbers of L1T2 and L1T3 are P2 and P3, respectively. Other slices in this convolutional layer are similar to T1-T3, scheduled in groups of three based on the number of cores (3). Other convolutional layers are similar to L1 and will not be described further here.

[0047] In this embodiment, the convolution operation is performed on the input feature data and weight data. Since multiple slices in the same convolutional layer have the same weight, based on the task scheduling rule 2 mentioned above, multiple slices of the same convolutional layer can be continuously calculated in the same core. Each core only needs to read the weight data once to continuously process the convolution operation of multiple slices, thereby reducing the latency of reading the weight data and improving the overall processing efficiency.

[0048] In an exemplary embodiment, the specific implementation of task scheduling rule 2 may further include: if, in the t-th time period, the c-th core processes the subtask corresponding to the x-th slice in the last convolutional layer of the convolutional processing sequence, then in the (t+1)-th time period, the (c+1)-th core processes the subtask corresponding to the (x+C)-th slice in the first convolutional layer of the convolutional processing sequence, where t, c, and x are positive integers, C is the number of cores, T is the number of slices, and C and T are integers greater than 1. It should be noted that the memory flow order is a cyclic relationship; if the c-th core is the last core in the memory flow order, then the (c+1)-th core is the first core in the memory flow order.

[0049] In this embodiment, based on the task scheduling rules 1 and 2 above, multiple cores process layer fusion convolution calculations in parallel in a pipeline manner in space, so that the convolution operation timing of multiple cores can meet the following conditions: different cores process different slices on different convolutional layers within the same time period; the same core continuously calculates different slices on the same convolutional layer in multiple consecutive time periods; the intermediate feature data of the same slice on different convolutional layers flows within multiple cores until the corresponding slice of the last convolutional layer is calculated before being output to the outside of the core.

[0050] In some embodiments, before determining the task timing of multiple subtasks on multiple cores based on the convolution processing order and memory transfer order in step S2044, the method further includes the following steps:

[0051] S2043-2, Number the plurality of subtasks according to the convolutional layer and the slice corresponding to each subtask;

[0052] S2043-4, multiple subtasks with the same slice number are identified as a single slice task to obtain multiple slice tasks, wherein the number of slice tasks is equal to the number of slices, and the multiple subtasks in each slice task are processed sequentially according to the convolution processing order.

[0053] In some embodiments, step S2044, determining the task timing of multiple subtasks on multiple cores based on the convolution processing order and memory transfer order, may include:

[0054] Step S2044-2: Determine the number of cores in the multi-core processor;

[0055] Step S2044-4: Determine the core number and time period number of the first subtask in each slice task according to the memory flow order, number of cores, number of slices and number of convolutional layers. The first subtask is the subtask that performs convolution processing first in each slice task, determined according to the convolution processing order.

[0056] Step S2044-6: Determine the core number and time period number of each subtask other than the first subtask in each slice task according to the convolution processing order and memory flow order.

[0057] Step S2044-8: Determine the task sequence of multiple subtasks on multiple cores based on the core number and time cycle number of multiple subtasks.

[0058] In some embodiments, step S2044-4 may include: determining the numbering order of multiple cores in a multi-core processor according to the memory flow order; dividing multiple slices into multiple slice groups according to the number of cores and the number of slices, wherein the multiple slice groups are numbered sequentially, the number of slices in the last slice group is less than or equal to the number of cores, and the number of slices in the other slice groups is equal to the number of cores; determining the core number of the first subtask in at least one slice task corresponding to each slice group according to the number of cores and the number of convolutional layers, wherein the core number of the first subtask in at least one slice task corresponding to each slice group is the same; determining the time period number of the first subtask in at least one slice task corresponding to each slice group according to the number of cores and the number of convolutional layers, wherein the time period number of the first subtask in at least one slice task corresponding to each slice group is consecutive.

[0059] In this embodiment, if the number of cores is C, each core will process C slices from the same convolutional layer sequentially over C consecutive time periods, unless the number of remaining slices is less than C. Correspondingly, only the last slice group may have fewer than C slices; all other slice groups include C slices, and each slice corresponds to one slice task.

[0060] In some embodiments, determining the core number of the first subtask in at least one slice task corresponding to each slice group based on the number of cores and the number of convolutional layers in step S2044-4 includes: determining the core number C of at least one first subtask corresponding to the i-th slice group according to the following method. i :

[0061] C i= ((i-1)×L)%C+1;

[0062] Among them, C i Let i be the core number of at least one first subtask corresponding to the i-th slice group, where i is a positive integer, L is the number of convolutional layers, and C is the number of cores.

[0063] In an exemplary embodiment, if T=6, L=5, C=3, and i=1, then the core number of the three first subtasks in the first slice group (corresponding to the first to third slices in the first convolutional layer), namely L1T1, L1T2, and L1T3, is 1; and if i=2, then the core number of the three first subtasks in the second slice group (corresponding to the fourth to sixth slices in the first convolutional layer), namely L1T4, L1T5, and L1T6, is 3.

[0064] In some embodiments, step S2044-4, determining the time period number of the first subtask in at least one slice task corresponding to each slice group based on the number of cores and the number of convolutional layers, includes:

[0065] The first time period number of the i-th slice group is determined as follows: T i = (i-1)×L+1; where, T i Let i be the first time period number, where the first time period number is the time period number of the first subtask corresponding to the first slice in each slice group, i is a positive integer, and L is the number of convolutional layers;

[0066] By incrementing one by one in the i-th first time period number, we obtain the other time period numbers of the i-th slice group, where the other time period numbers are the time period numbers of the first subtask corresponding to the other slices in each slice group.

[0067] In an exemplary embodiment, if T = 6, L = 5, C = 3, and i = 1, then the time period number of the first subtask of the first slice in the first slice group (corresponding to the first slice in the first convolutional layer), i.e., L1T1, is 1; the time period number of the first subtask of the second slice in the first slice group (corresponding to the second slice in the first convolutional layer), i.e., L1T2, is 2; and the time period number of the first subtask of the third slice in the first slice group (corresponding to the third slice in the first convolutional layer), i.e., L1T3, is 3. Let i = 2, then the time period number of the first subtask of the first slice in the second slice group (corresponding to the fourth slice in the first convolutional layer), i.e., L1T4, is 6; the time period number of the first subtask of the second slice in the second slice group (corresponding to the fifth slice in the first convolutional layer), i.e., L1T5, is 7; and the time period number of the first subtask of the third slice in the second slice group (corresponding to the sixth slice in the first convolutional layer), i.e., L1T6, is 8.

[0068] In some embodiments, steps S2044-6 may include: determining the order of multiple subtasks in each slice task according to the convolution processing order; determining the numbering order of multiple cores in the multi-core processor according to the memory flow order; in each slice task, traversing from the second subtask, if the core number of the previous subtask is less than the maximum core number, determining the core number of the current subtask to be the core number of the previous subtask plus one, and if the core number of the previous subtask is equal to the maximum core number, determining the core number of the current subtask to be the core number of the first core among the multiple cores; in each slice task, traversing from the second subtask, determining the time period number of the current subtask to be the time period number of the previous subtask plus one.

[0069] In an exemplary embodiment, if T = 6, L = 5, C = 3, and the core number of each first subtask L1T1, L1T2, L1T3 in the first slice group has been determined to be 1 through S2044-4, then the core number of each second subtask L2T1, L2T2, L2T3 in the first slice group can be determined to be 2, the core number of each third subtask L3T1, L3T2, L3T3 in the first slice group can be determined to be 3, the core number of each fourth subtask L4T1, L4T2, L4T3 in the first slice group can be determined to be 1, and the core number of each fifth subtask L5T1, L5T2, L5T3 in the first slice group can be determined to be 2.

[0070] In an exemplary embodiment, if T = 6, L = 5, C = 3, and the time period numbers of the first subtasks L1T1, L1T2, and L1T3 in the first slice group have been determined to be 1, 2, and 3 respectively through S2044-4, then the time period numbers of the second subtasks L2T1, L2T2, and L2T3 in the first slice group can be determined to be 2, 3, and 4 respectively, the time period numbers of the third subtasks L3T1, L3T2, and L3T3 in the first slice group can be determined to be 3, 4, and 5 respectively, the time period numbers of the fourth subtasks L4T1, L4T2, and L4T3 in the first slice group can be determined to be 4, 5, and 6 respectively, and the time period numbers of the fifth subtasks L5T1, L5T2, and L5T3 in the first slice group can be determined to be 5, 6, and 7 respectively.

[0071] In some embodiments, step S2046 may include: scheduling the core corresponding to the core number to execute multiple subtasks sequentially according to the order of the time period numbers, wherein each subtask is used to perform convolution processing on feature data on a slice in a convolutional layer according to the task type of the subtask.

[0072] In this embodiment, the task sequence includes multiple sub-task sequences corresponding to multiple cores. Each sub-task sequence corresponds to a core number, and each sub-task sequence includes multiple sub-tasks with consecutive time period numbers.

[0073] In some embodiments, the task type of a subtask includes at least one of the following:

[0074] The first task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights;

[0075] The second task type is used to instruct the current core to read the input feature data and weights from external memory and perform convolution operations based on the input feature data and weights;

[0076] The third task type is used to instruct the current core to read weights from external memory and perform convolution operations based on the intermediate feature data and weights output by the previous core.

[0077] The fourth task type is used to instruct the current core to read input feature data from external memory and perform convolution operations based on the input feature data and pre-read weights;

[0078] The fifth task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights, and write the operation results to external memory;

[0079] The sixth task type is used to instruct the current core to read weights from external memory, perform convolution operations based on the intermediate feature data and weights output by the previous core, and write the results to external memory.

[0080] In some embodiments, multiple cores of a multi-core processor can each use an independent memory space, and multiple cores can also share memory with each other; this application does not impose any restrictions on this. The multi-core memory sharing in this application is based on the task scheduling principle described above, which calculates slices with the same number on different cores according to the convolution processing order.

[0081] In some embodiments, when two adjacent cores share memory in the memory flow sequence, the current core reads intermediate feature data from the first memory shared with the previous core, and / or writes the convolution operation result of the current core as intermediate feature data into the second memory shared with the next core.

[0082] In this embodiment, two adjacent cores can share memory to save the time of intermediate feature data transfer between the two memory, thereby optimizing the core data flow process and improving the efficiency of each core in a multi-core processor in reading and writing intermediate feature data.

[0083] In some embodiments, given the number of slices T, the number of convolutional layers L, and the number of cores in a multi-core processor C, the theoretical number of cycles for the layer fusion convolution task can be determined based on the following:

[0084]

[0085] Among them, P DMIA This refers to the theoretical number of cycles for the layer fusion convolution task under the Dynamic Memory Inter-layer Allocation (DMIA) mechanism, i.e., after executing the task scheduling method for the multi-core processor described in this application. h represents the number of slice groups. % indicates the modulo operation. This indicates rounding up to the nearest integer.

[0086] In this embodiment, firstly, if "one core processes C slices on the same convolutional layer for C consecutive cycles" is considered as a single process A, then a multi-core processor needs a total of h such processes A to process T slices on L convolutional layers. These h processes A are sequentially arranged on C cores, with the core having the most processes A containing a total of h processes A. Process A, therefore The calculation considers the total runtime of the core with the most cores. Then, `h%C` calculates which core the last time cycle terminates on. If it terminates on Core_n, the idle time of the initial n-1 time cycles needs to be added, so `(h%C-1)` needs to be added. Finally, `T` is not necessarily divisible by `C`, meaning the last process A is not necessarily complete. For cases where `T` is not divisible by `C`, the remainder needs to be subtracted. This part of the time was calculated.

[0087] For example, if L=5, T=5, C=3, then h=10, meaning there are 10 processes A. However, the last group (the last 5 processes A) is incomplete; each process A has only 2 time periods (instead of 3). The last group of processes A needs to be subtracted by 1 as a correction, i.e., substituting L=5, T=5, C=3 into the equation. We get -1.

[0088] In this embodiment, the dynamic memory inter-layer allocation mechanism has a significant advantage in data access compared to other multi-core data flow technologies in related technologies. For example, if each core in a multi-core processor still uses layer fusion scheduling but does not perform inter-core data flow scheduling, multiple cores will read or write data at the same time, resulting in port conflicts and greatly reducing performance.

[0089] Through the embodiments of this application, intermediate feature data can be retained on-chip for data transfer between multiple cores; simultaneously, multiple cores will not concurrently access feature data memory externally or concurrently access weight data memory in the global buffer, greatly reducing the data access latency of the multi-core system to the global buffer and external DRAM. Furthermore, in each data transfer cycle, only one core will perform data read or write operations with the external DRAM, further reducing the data width requirement of the bus.

[0090] Figure 3 This is a schematic diagram of a multi-level fusion convolution task in one embodiment of this application, as shown below. Figure 3 As shown, the processing unit for the layer fusion convolution task is a stack containing two convolutional layers, each of which is divided into the same number of slices.

[0091] In this embodiment, L1 and L2 are convolutional layer numbers, and T1 to T7 are slice numbers. However, this application does not limit the number of convolutional layers and slices for the layer fusion convolution task.

[0092] In one exemplary embodiment, each stack is divided into sections starting from the last layer, according to a certain size. The size of the section needs to take into account the core capabilities. This application does not limit the specific method of section division. Adjacent sections may or may not overlap, as described in the subtasks of this application.

[0093] In this embodiment, each convolutional layer is divided into multiple tiles. Each layer is divided from bottom to top according to the convolution processing order, and each layer has the same number of tiles. Each tile can be processed independently by a core. The tiles are relatively small, and the output data can be stored in the core's local buffer without accessing off-chip memory, such as dynamic random access memory (DRAM).

[0094] Figure 4 This is a schematic diagram of layer-by-layer fusion convolution scheduling and layer-by-layer convolution scheduling in one embodiment of this application, as shown below. Figure 4 As shown, the left side represents the layer-by-layer convolution scheduling in related technologies, while the right side represents the layer fusion convolution scheduling in the embodiments of this application.

[0095] The storage system in this application is a multi-level storage system, including: external memory, global buffer (GB), and local buffer (LB). The external memory is located outside the chip and can be DRAM, while the global buffer and local buffer are located inside the chip. The local buffer is the cache of a single core, and the global buffer is the cache of the entire multi-core processor.

[0096] In layer-by-layer (LBL) convolution scheduling, the input feature map I and convolution weights W of L1 are sequentially transferred from off-chip DRAM to the computing array for convolution computation, resulting in a complete L1 output feature map O. This data is then sequentially returned to off-chip storage. At this point, the L1 feature map data is theoretically used up and can be discarded before L2 convolution computation begins. Layer-by-layer convolution only begins computation after the previous layer has been fully completed. For neural networks with large feature maps, such as Fast Super-Resolution Convolutional Neural Networks (FSRCNN), a large on-chip buffer is required to buffer a complete intermediate feature map. Alternatively, feature map data may need to be sequentially transferred from off-chip memory to on-chip global buffers and on-chip local buffers. In this case, the large amount of data, including feature maps, can put pressure on the storage system and may become a technical bottleneck. In terms of memory bandwidth and power consumption, transferring feature map data to and from external DRAM is expensive.

[0097] In layer-fusion convolution scheduling, a feature map is first divided into multiple tiles, and convolution computation is performed on a tile-by-tile basis. This means that only a portion of the feature map is computed at a time, rather than the entire feature map, and the data is transferred between memory levels. This scheduling method reduces the size of data transferred in a single batch, thus enabling the use of smaller, more efficient memory levels to transfer this data. For example, tiles of intermediate feature map data can be transferred on a local buffer instead of on DRAM or a global buffer, which is more efficient, energy-intensive, and has higher latency than a local buffer.

[0098] In layer-fusion convolution, adjacent slices can overlap, both horizontally and vertically. In single-core layer-fusion processing in related technologies, overlapping data can be recalculated or cached in an on-chip buffer for inclusion when computing adjacent slices. In this embodiment, the multi-core processor considers inter-core scheduling on a slice-by-slice basis. Each core needs to receive the data of a complete slice, so the overlapping data can be recalculated without the need for additional buffers.

[0099] Figure 5 This is a schematic diagram of data flow in a multi-core system according to one embodiment of this application, as shown below. Figure 5 As shown, a multi-core system includes a system-on-chip (SoC) and off-chip memory.

[0100] In this embodiment, the off-chip memory can be DRAM, used to store the input feature data of the entire convolutional network.

[0101] In this embodiment, the system-on-a-chip includes a global buffer (GB) and multiple cores. Each core typically has a computing array and several cooperating local buffers (LB). The global buffer is a global buffer shared by all cores and is typically used to store weight data.

[0102] In this embodiment, the operation performed on each core is the computation of a slice in the layer fusion process. The computation array is used to perform convolution operations on the data in the cache. The local cache can store the weight data and input data required for the convolution operation, and can also store the output data of the convolution operation. For example, the weight data can be stored in the W LB, the input data can be stored in the I LB, and the output data can be stored in the O LB.

[0103] In this embodiment, the multi-core data stream is divided into two parts: intra-core data stream and inter-core data stream, which are described below:

[0104] Intra-core data flow: ILB stores input data of one slice size, WLB stores weight data, and OLB stores the corresponding output data. The input data and weight data perform convolution calculations on one slice on the computation array. In this embodiment, data flows sequentially between multiple cores. For example, the memory flow order of intra-core data flow between multiple cores can be core 1 → core 2 → core 3 → core 4 → core 1, and so on.

[0105] Inter-core data flow: DRAM is used to store the complete input feature maps and the final output feature maps of the entire convolutional network. Intermediate feature map data flows on-chip and is not stored or accessed through DRAM. Therefore, only the reading of all slices from the first convolutional layer and the writing back of all slices from the last convolutional layer involve access to off-chip DRAM and the bus (BUS). All weight data for the entire convolutional network is stored in the on-chip global cache. Before a core performs convolutional computation on a slice, the weight data of the corresponding convolutional layer is loaded from the global cache into the kernel's WLB. Feature data flows between cores in a pipelined manner; the output feature data of the previous core is passed from the OLB to the ILB of the next core as input feature data. During on-chip flow, only slices from the first and last convolutional layers involve access to off-chip memory.

[0106] In one exemplary embodiment, ILB and OLB can constitute a ping-pong cache to improve throughput. The ping-pong cache works based on a mechanism of alternating data writing and reading. When data is written to memory A, the computing array can read data from memory B and process it. When memory A is full, writing switches to memory B, while the processing unit begins reading data from memory A. This alternating pattern enables continuous data writing and reading, improving overall data processing efficiency.

[0107] In another exemplary embodiment, the 0 LB of the previous core and the I LB of the next core can be the same memory shared by the two cores. That is, the computing array of core 1 can write the output data of core 1 to the common memory 1, and the computing array of core 2 can read the input data from core 2. Sharing memory between cores can further improve data read and write efficiency and reduce the latency of data flow between cores.

[0108] In this embodiment, multiple cores of the multi-core processor process layer fusion convolution computations in parallel in a pipelined manner in space. The multi-core scheduling principle for the layer fusion convolution task is as follows:

[0109] Principle 1: Within the same cycle, different cores process different slices on different convolutional layers;

[0110] Principle 2: The same core is used to compute different slices of the same convolutional layer over multiple consecutive cycles;

[0111] Principle 3: Data from the same slice across different convolutional layers flows through multiple cores (within the chip) until the data for the corresponding slice of the last convolutional layer is calculated and output. Multiple cores can share memory.

[0112] In this embodiment, a cycle refers to the time it takes for a core to process one slice on a convolutional layer, which can be represented by P.

[0113] In this embodiment, let the number of cores be C, the number of slices be T, and the number of convolutional layers be L. Then, for a layer fusion convolution task with any number of cores, slices, and convolutional layers, multi-core scheduling can be performed through the following actions:

[0114] Action 1, Algorithm Startup: When P=1, Core1 processes L1T1 (using LxTy to refer to the calculation of the y-th slice on the x-th convolutional layer), while the other cores are idle.

[0115] Action 2, Configure slices with the same slice number in different convolutional layers: Process the transfer of the same slice between multiple cores in different layers: The transfer order is Core1, Core2, ..., Core C, Core1, until the current slice calculates the output of the last convolutional layer;

[0116] For example, if when P=1, Core1 starts processing L1T1, then when P=2, L2T1 flows to Core2, when P=3, L3T1 flows to Core3, and so on, flowing between multiple cores as described above, until slice T1 calculates the last convolutional layer and outputs it outside the slice;

[0117] Action 3, configuring slices according to a pattern, is also a boundary case of slice flow between cores: when slice x flows between cores for processing, until the last convolutional layer is calculated and output, the next P and the next core start processing slice x+C, and then continue to follow the above flow order until slice x+C of the last convolutional layer is calculated; then slice x+2C is processed, and so on.

[0118] Action 4: Within the same kernel, sequentially process C consecutive slices from the same convolutional layer over C consecutive cycles (P). Because the same kernel prioritizes processing different slices from the same convolutional layer in time, the next slice can be loaded while the previous slice is being computed, thus reducing latency.

[0119] For example, if Core1 starts processing L1T1 when P=1, then when P=2, Core1 processes the next slice on the same convolutional layer, i.e., L1T2, and so on, until L1TC is calculated when P=C. Furthermore, for Core x, if Core x starts calculating LxT1 when P=x, then Core x calculates LxT2 when P=x+1, and so on, until LxTC is calculated when P=x+C-1.

[0120] Furthermore, after the aforementioned C consecutive cycles, the slice flow sequence described in Action 2 can continue to be followed to calculate the slices that flowed from the previous P and the previous Core to the current core.

[0121] The task scheduling method for multi-core processors in this application will be described in detail below based on the specific number of cores, convolutional layers, and slices.

[0122] Figure 6 This is a timing diagram of a multi-core layer fusion convolution task with dynamic memory allocation between layers in one embodiment of this application, as shown below. Figure 6 As shown, there are 3 cores, numbered Core1 to Core3, 5 convolutional layers, numbered L1 to L5, and 6 slices, numbered T1 to T6.

[0123] In this embodiment, dynamic memory layer allocation refers to the flow of data streams (such as intermediate feature data of convolution operations) between multiple cores.

[0124] In this embodiment, the layer fusion convolution task processes a 5-layer convolutional network, with each layer divided into 6 slices. Therefore, the layer fusion convolution task can be divided into 30 subtasks, each of which can be represented by the corresponding convolutional layer number and slice number as L1T1 to L5T6.

[0125] In this embodiment, determining the task timing can be divided into the following steps:

[0126] Step S601: Determine the starting point (corresponding to action 1). When P=1, start Core1 to process L1T1.

[0127] Step S602: According to the principle of the same slice flowing between cores (corresponding to action 2), the calculation of the 5 convolutional layers T1L1 to T1L5 is allocated.

[0128] Step S603: According to the principle of handling the transition boundary situation (corresponding to action 3), allocate the calculation of all convolutional layers Tile_(1+C), Tile_(1+2C), ..., that is, allocate the calculation of all convolutional layers T4L1 to T4L5 of T4.

[0129] Step S604: According to the processing order of the slices in the time dimension of the same core (corresponding to action 4), allocate the computation of the first convolutional layer of the remaining slices, such as allocating T2L1, T3L1, T5L1, and T6L1.

[0130] Step S605: According to the principle of flow between the same slices in the core (corresponding to action 2), allocate the computation of all convolutional layers of the remaining slices.

[0131] Step S606: Check that the computation of all convolutional layers in all slices is distributed across multiple time periods on 3 cores, and that no core should be idle during any intermediate time period except for the task start-up and task end-up phases.

[0132] Through the above steps S601 to S606, it can be determined that... Figure 6 The task sequence shown is such that subtasks with the same gray level are assigned in the same step, while subtasks with different gray levels are assigned in different steps.

[0133] In some embodiments, the task type of a subtask includes at least one of the following:

[0134] The first task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights;

[0135] The second task type is used to instruct the current core to read the input feature data and weights from external memory and perform convolution operations based on the input feature data and weights;

[0136] The third task type is used to instruct the current core to read weights from external memory and perform convolution operations based on the intermediate feature data and weights output by the previous core.

[0137] The fourth task type is used to instruct the current core to read input feature data from external memory and perform convolution operations based on the input feature data and pre-read weights;

[0138] The fifth task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights, and write the operation results to external memory;

[0139] The sixth task type is used to instruct the current core to read weights from external memory, perform convolution operations based on the intermediate feature data and weights output by the previous core, and write the results to external memory.

[0140] In this embodiment, the above-mentioned task types can be identified in the following ways:

[0141] The first task type is unmarked. For example, when Core2 processes L2T2 and L2T3, the input data used for the convolution operation is obtained from the previous core, and the weight data used for the convolution operation has also been pre-read when Core2 processes L2T1. That is, the data is cached in the local buffer of this core, and there is no need to interact with external memory or global buffer.

[0142] The second task type is identified as Id I+W. For example, in Core1's processing of L1T1, the core needs to read the input feature data and weight data of the first convolutional layer from external memory. The input feature data can be obtained from off-chip memory, and the weight data can be read from the global buffer.

[0143] The third task type is identified as IdW. For example, Core2's processing of L2T1.

[0144] The fourth task type is identified by Id I. For example, Core1's processing of L1T2. The core needs to read the input feature data from external memory and perform convolution operations based on the input feature data and pre-read weights.

[0145] The fifth task type is identified as sd O. For example, in Core2's processing of L5T2 and L5T3, the input data used for the convolution operation is obtained from the previous core, and the weight data used for the convolution operation has also been pre-read when Core2 processes L5T1. Since the core processes a slice of the last convolutional layer, the core will input the convolution operation result into off-chip memory.

[0146] The sixth task type is identified as Id W sd O. For example, when Core2 processes L5T1, the core needs to read weight data from external memory. Since it is processing a slice of the last convolutional layer, the core will input the convolution operation result into off-chip memory.

[0147] In this embodiment, only the data of L1 (the first convolutional layer) needs to be read from the off-chip DRAM bus into the corresponding on-chip buffer, i.e., L1 performs the Id I operation; and only the data of L5 (the last convolutional layer) needs to be written back from the on-chip on-chip buffer to the DRAM bus, i.e., L5 performs the sd O operation.

[0148] In this embodiment, the theoretical number of time cycles for the layer fusion convolution task can be determined as follows:

[0149]

[0150] Given T=6, L=5, and C=3, the theoretical number of time cycles P for the layer fusion convolution task is... DMIA =12, and Figure 6 The number of time periods in the data matches.

[0151] In this embodiment, it is assumed that the time required for one core to process one slice is t. P After the slicing is completed, the area of ​​a tile is A. Tile ic j Oc represents the input channel size of the j-th layer of the convolutional neural network. j Let represent the output channel size of the j-th layer of the convolutional neural network, and BW represent the bandwidth of the off-chip DRAM bus. Given that off-chip data read / write is single-port, the total computational latency t under Dynamic Memory Inter-Layer Allocation (DMIA) can be calculated as follows: DMIA :

[0152]

[0153] Among them, P DMIA This represents the theoretical number of time cycles for a layer fusion convolution task under a dynamic memory inter-layer allocation mechanism (this value is 12 in the example where T / C / L = 6 / 3 / 5).

[0154] In this embodiment, under the DMIA mechanism, since the off-chip data access of different cores is not performed simultaneously, it can be concurrent with the data flow between cores. Therefore, only the data reading delay of the first convolutional layer and the data write-back delay of the last convolutional layer need to be considered. The calculated delay data amount is the data amount of the first slice plus the data amount of the last slice, and this part of the delay only involves one core.

[0155] Through the embodiments of this application, dynamic memory allocation for multi-core systems can be realized. Combined with pipelined data scheduling of multi-core and layer fusion, there will be no off-chip write-back of feature data between the fused convolutional network layers. All intermediate feature data flows between the cores on the chip, reducing off-chip memory access latency and speeding up processing.

[0156] In this embodiment, no more than one core will simultaneously perform read or write operations on the bus within the same time period. This means that the read and write ports of the off-chip DRAM bus only need to meet the data read and write requirements of a single core, which greatly reduces the bus pressure on off-chip data access. Furthermore, in this embodiment, except for the off-chip read operation of Core 1 at P1 time and the off-chip write-back operation of Core 1 at P12 time, other off-chip read and write operations can be processed in parallel with the corresponding inter-core data transfer operations, further reducing system time overhead.

[0157] Therefore, the task scheduling method for multi-core processors in this application embodiment can achieve dynamic memory allocation, which can not only process layer fusion operations in parallel, but also effectively reduce data access to the off-chip DRAM bus and access to weighted data on the global buffer, thereby solving the technical problem that the computing performance and memory access efficiency of multi-core processors are affected in related technologies.

[0158] Figure 7 This is a timing diagram illustrating a multi-core layer fusion convolution task with non-dynamic memory allocation between layers in related technologies, such as... Figure 7 As shown, the number of cores is 3, the number of convolutional layers is 5, and the number of slices is 6.

[0159] In this embodiment, non-dynamic memory layer allocation means that the data flow (such as intermediate feature data of convolution operations) does not flow between multiple cores, but rather a single core processes multiple convolutional layers of the same slice sequentially. Multiple cores process different slices of the same convolutional layer simultaneously.

[0160] like Figure 7 As shown, a single core performs convolution processing on the same slice of different convolutional layers in adjacent time periods, requiring the reading of the corresponding convolutional layer's weight data from the global buffer each time. Furthermore, all three cores simultaneously perform read operations (Id I) or write operations (sd O) to the external memory / bus. Due to the bus's limitation of only being able to execute one read or write operation at a time, this results in severe access latency for the entire multi-core system.

[0161] In this embodiment, the theoretical number of time cycles for the layer fusion convolution task without inter-kernel data exchange can be expressed as:

[0162]

[0163] Where T represents the slice data, C represents the number of cores, and L represents the number of convolutional layers. Since C=3, L=5, and T=6, P can be calculated. N-DMIA =10, and Figure 7 The number of time periods matches.

[0164] In this embodiment, it is assumed that the time required for one core to process one slice is t. P After the slicing is completed, the area of ​​a tile is A. Tile ic j Oc represents the input channel size of the j-th layer of the convolutional neural network. j Let represent the output channel size of the j-th layer of the convolutional neural network, and BW represent the bandwidth of the off-chip DRAM bus. Given that off-chip data read / write is single-port, the total computational latency t under non-dynamic memory inter-layer allocation (N-DMIA) can be calculated as follows: N-DMIA :

[0165]

[0166] Among them, P N-DMIA This represents the theoretical number of time cycles for a layer fusion convolution task under a non-dynamic memory inter-layer allocation mechanism (this value is 10 in the example where T / C / L = 6 / 3 / 5).

[0167] In this embodiment, under the N-DMIA mechanism, all off-chip reads and writes of the cores occur simultaneously. Therefore, the computation array is actually blocked during this period. The total time for reading data from the first convolutional layer and writing back data from the last convolutional layer is [missing information]. Each time, all cores read or write simultaneously, so this time is within the DMIA mechanism. Furthermore, under the N-DMIA mechanism, the data in the intermediate layer also needs to be imported and exported through DRAM. Considering that the time cycle P is not significantly different between the two schemes, the layer fusion convolution task scheduling of the N-DMIA mechanism will result in a higher latency than that of the DMIA mechanism due to its severe memory access bottleneck. Therefore, the DMIA proposed in this application is an effective solution to the memory access bottleneck in multi-core systems, which can improve data storage and access efficiency, reduce memory access conflicts, and thus solve the technical problem of the impact on the computing performance and memory access efficiency of multi-core processors in related technologies.

[0168] In other embodiments of this application, the DMIA mechanism described above is at the forefront of hardware chip design architecture and at a more fundamental, principle-oriented level in the application-principle vertical dimension, thus having a very wide range of applications. For example, the DMIA mechanism can be applied to the following aspects:

[0169] 1. Dataflow deployment for multi-core convolutional neural network accelerators: used to extend single-core layer fusion dataflow to multi-core and efficiently optimize memory bandwidth.

[0170] 2. Applications in the field of neural network processor front-end design: used for architecture exploration, parameter exploration, and data flow solution exploration in the early stages of hardware design for multi-core NPU systems.

[0171] 3. A new neural network algorithm model is deployed and applied to a dedicated accelerator: data flow deployment and efficient compiler design for the accelerator.

[0172] 4. Architecture and dataflow design of digital circuits for large-scale general-purpose computing platforms: solutions for large-scale general-purpose computing.

[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0174] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to perform the steps in any of the above method embodiments.

[0175] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0176] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0177] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0178] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods described in various embodiments of this application.

[0179] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0180] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0181] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A task scheduling method for a multi-core processor, characterized in that, The method includes: Determine the task information for layer fusion convolution tasks and the memory flow order of multiple cores in a multi-core processor; Based on the task information and the memory flow order, the multiple cores are scheduled to execute the layer fusion convolution task.

2. The method according to claim 1, characterized in that, The task information includes the number of convolutional layers, the number of slices, and the convolution processing order of multiple convolutional layers. The layer fusion convolution task is used to perform convolution processing on the feature data on the slices with the same number in the multiple convolutional layers layer by layer according to the convolution processing order. Each convolutional layer is divided into the same number of slices, and the multiple slices in each convolutional layer are numbered sequentially.

3. The method according to claim 2, characterized in that, The step of scheduling the multiple cores to execute the layer fusion convolution task according to the task information and the memory flow order includes: The layer fusion convolution task is divided into multiple subtasks based on the number of convolutional layers and the number of slices, wherein each subtask is used to instruct one core of the multi-core processor to perform convolution processing on the feature data of one slice in a convolutional layer; The timing sequence of the multiple subtasks on the multiple cores is determined based on the convolution processing order and the memory flow order. According to the task timing, the multiple cores execute the multiple sub-tasks to complete the layer fusion convolution task.

4. The method according to claim 3, characterized in that, The timing of the multiple subtasks on the multiple cores is determined based on the convolution processing order and the memory flow order, including: Based on the preset task scheduling rules, the convolution processing order, and the memory flow order, the core number and time period number of the multiple subtasks are determined. The task sequence of the multiple subtasks in the multiple cores is determined based on the core number and the time period number of the multiple subtasks.

5. The method according to claim 4, characterized in that, The task scheduling rules include: For multiple subtasks with the same slice number in different convolutional layers, the multiple subtasks are scheduled to execute convolution processing sequentially in multiple consecutive time periods according to the convolution processing order, and the core for executing the subtask is selected from the multiple cores according to the memory flow order. For multiple subtasks with different slice numbers in the same convolutional layer, the multiple subtasks are scheduled to be executed sequentially by the same core in multiple consecutive time periods according to the number of cores of the multi-core processor.

6. The method according to claim 3, characterized in that, Before determining the task timing of the multiple subtasks on the multiple cores based on the convolution processing order and the memory flow order, the method further includes: The multiple subtasks are numbered according to the convolutional layer and the slice corresponding to each subtask; Multiple subtasks with the same slice number are identified as a single slice task, resulting in multiple slice tasks. The number of slice tasks is equal to the number of slices. The multiple subtasks in each slice task are processed sequentially according to the convolution processing order.

7. The method according to claim 6, characterized in that, The timing of the multiple subtasks on the multiple cores is determined based on the convolution processing order and the memory flow order, including: Determine the number of cores in the multi-core processor; The core number and time period number of the first subtask in each slice task are determined according to the memory flow order, the number of cores, the number of slices, and the number of convolutional layers. The first subtask is the subtask that is first to undergo convolution processing in each slice task, as determined according to the convolution processing order. The core number and time period number of each subtask other than the first subtask in each slice task are determined according to the convolution processing order and the memory flow order. The task sequence of the multiple subtasks in the multiple cores is determined based on the core number and the time period number of the multiple subtasks.

8. The method according to claim 7, characterized in that, The core number and time period number of the first subtask in each slice task are determined based on the memory flow order, the number of cores, the number of slices, and the number of convolutional layers, including: The numbering order of the multiple cores in the multi-core processor is determined according to the memory flow order; The multiple slices are divided into multiple slice groups according to the number of cores and the number of slices. The multiple slice groups are numbered sequentially. The number of slices in the last slice group is less than or equal to the number of cores. The number of slices in the other slice groups is equal to the number of cores. The core number of the first subtask in at least one slice task corresponding to each slice group is determined according to the number of cores and the number of convolutional layers, wherein the core number of the first subtask in at least one slice task corresponding to each slice group is the same; The time period number of the first subtask in at least one slice task corresponding to each slice group is determined according to the number of cores and the number of convolutional layers, wherein the time period numbers of the first subtask in at least one slice task corresponding to each slice group are consecutive.

9. The method according to claim 7, characterized in that, The core number and time period number of each subtask other than the first subtask in each slice task are determined according to the convolution processing order and the memory flow order, including: The order of the multiple subtasks in each slice task is determined according to the convolution processing order; The numbering order of the multiple cores in the multi-core processor is determined according to the memory flow order; In each slice task, starting from the second subtask, if the core number of the previous subtask is less than the maximum core number, the core number of the current subtask is determined to be the core number of the previous subtask plus one; if the core number of the previous subtask is equal to the maximum core number, the core number of the current subtask is determined to be the core number of the first core among the plurality of cores. In each slice task, starting from the second subtask, the time period number of the current subtask is determined to be the time period number of the previous subtask plus one.

10. The method according to claim 3, characterized in that, in, The task timing sequence includes multiple sub-task timing sequences corresponding to the multiple cores. Each sub-task timing sequence corresponds to a core number, and each sub-task timing sequence includes multiple sub-tasks with consecutive time period numbers. The multiple cores are scheduled to execute the multiple sub-tasks according to the task timing sequence to complete the layer fusion convolution task, including: The core corresponding to the core number is scheduled to execute multiple subtasks sequentially according to the order of the time period number, wherein each subtask is used to perform convolution processing on feature data on a slice in a convolutional layer according to the task type of the subtask.

11. The method according to claim 10, characterized in that, The task type of the subtask includes at least one of the following: The first task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights; The second task type is used to instruct the current core to read input feature data and weights from external memory, and to perform convolution operations based on the input feature data and the weights; The third task type is used to instruct the current core to read the weights from the external memory and perform convolution operations based on the intermediate feature data output by the previous core and the weights. The fourth task type is used to instruct the current core to read the input feature data from the external memory and perform convolution operations based on the input feature data and pre-read weights; The fifth task type is used to instruct the current core to perform convolution operations based on the intermediate feature data output by the previous core and the pre-read weights, and to write the operation result into the external memory; The sixth task type is used to instruct the current core to read the weights from the external memory, perform convolution operations based on the intermediate feature data output by the previous core and the weights, and write the operation result into the external memory.

12. The method according to claim 11, characterized in that, In the case where two adjacent cores share memory in the memory flow sequence, the current core reads the intermediate feature data from the first memory shared with the previous core, and / or writes the convolution operation result of the current core as the intermediate feature data into the second memory shared with the next core.

13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is executed by a processor to perform the method described in any one of claims 1 to 12.

14. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method as described in any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 12.