A data processing apparatus and a data processing method
By introducing a synchronous scheduler and a direct memory access controller into the chip, the problems of limited functionality and low memory access efficiency in existing multi-core chips are solved, enabling more efficient data processing and flexible task allocation, thereby improving the chip's computing power and processing capabilities.
Patent Information
- Application Number
- CN202080096325.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-04
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2040-03-04
AI Technical Summary
Existing chips suffer from limited functionality, lack of flexibility, and low storage access efficiency when processing multi-core data, making it difficult to efficiently utilize the computing power of multiple processing cores.
The design employs a synchronous scheduler and a direct memory access controller. The synchronous scheduler instructs the direct memory access controller to read programs from external storage units and send them to the processing cores. This avoids delays caused by the processing cores reading data autonomously and allows each processing core to execute different programs, thus achieving flexible task allocation.
It improves the chip's computing power, enhances the efficiency and flexibility of the processing cores, and can fully utilize the computing power of each processing core to adapt to the data processing needs of different fields.
Smart Images

Figure CN115151892B_ABST
Abstract
Description
BACKGROUND
[0001] With the development of science and technology, human society is rapidly entering an intelligent era. An important feature of the intelligent era is that people obtain more and more types of data, obtain more and more amounts of data, and require higher and higher speed of processing data.
[0002] A chip is the cornerstone of data processing, which fundamentally determines the ability of people to process data. From the application field, there are mainly two routes of chips: one is the general chip route, such as central processing unit (CPU) and the like, which can provide great flexibility, but the effective computing power is relatively low when processing algorithms in specific fields; the other is the special chip route, such as tensor processing unit (TPU) and the like, which can exert higher effective computing power in certain specific fields, but their processing capacity is poor or even unable to process when facing flexible and versatile general fields.
[0003] Because of the variety and huge amount of data in the intelligent era, the chip is required to have both extremely high flexibility to process different fields and rapidly changing algorithms and extremely strong processing capacity to quickly process extremely large and rapidly growing data amounts. SUMMARY
[0004] (I) Invention purposes
[0005] The purpose of the present application is to provide a data processing device and a data processing method, the data processing device is provided with a synchronous scheduler and a direct memory access controller, the synchronous scheduler instructs the direct memory access controller to read the programs corresponding to each processing core from the external storage unit and sends them to the corresponding processing core. On the one hand, the data processing device provided by the embodiment of the present application does not need the processing core to take data from the external storage unit, avoiding the delay caused by multiple processing cores reading data, improving the computing power of the chip. On the other hand, the data processing device provided by the embodiment of the present application, the programs executed by each processing core can be the same or different, and the synchronous scheduler can respond to the update of multiple processing core programs, which can flexibly allocate tasks and fully exert the computing power of each processing core, further improving the computing power of the device.
[0006] (II) Technical solutions
[0007] To solve the above problems, the first aspect of the present application provides a data processing device, which comprises at least two processing cores, a synchronous scheduler for generating and sending a configuration signal in response to a program update signal of each processing core connected to the synchronous scheduler, and a direct memory access controller for reading and sending a program corresponding to each processing core from an external storage unit based on the configuration signal.
[0008] The data processing device provided by the embodiments of the present application is provided with a synchronous scheduler and a direct memory access controller, the synchronous scheduler instructs the direct memory access controller to read a program from an external storage unit and send it to a processing core corresponding to the program, on the one hand, the data processing device provided by the embodiments of the present application does not need the processing core to read data from the external storage unit, avoids the delay caused by reading data by multiple processing cores, and improves the computing power of the chip, on the other hand, the data processing device provided by the embodiments of the present application can flexibly allocate tasks and fully utilize the computing power of each processing core to further improve the computing power of the chip, because the programs executed by each processing core can be the same or different and the synchronous scheduler responds to the update of the programs of multiple processing cores.
[0009] Further, the synchronous scheduler is further configured to send a synchronous running signal to each processing core connected to the synchronous scheduler in response to the program update signal, the synchronous running signal is sent to each processing core by the synchronous scheduler after receiving a predetermined number of program update signals, and the synchronous running signal is used to instruct each processing core to start executing the respective program at the same time.
[0010] Further, the program comprises multiple program segments.
[0011] Further, the processing core is configured to send the program update signal to the synchronous scheduler after executing each program segment.
[0012] Further, the direct access controller is configured to read the program segment corresponding to each processing core from the external storage unit and send the program segment to the corresponding processing core, and the direct access controller is further configured to send a configuration completion signal to the synchronous scheduler after sending the program segment to the corresponding processing core.
[0013] Further, the synchronous scheduler is further configured to send a synchronous running signal to each processing core connected to the synchronous scheduler in response to the program update signal, the synchronous running signal is sent to each processing core by the synchronous scheduler after receiving a predetermined number of program update signals, and the synchronous running signal is used to instruct each processing core to start executing the respective program at the same time.
[0014] Further, the program comprises operation instructions and program updating instructions; the processing core is configured to execute the program updating instructions based on the operation instructions being completed, and send the program updating signal based on the program updating instructions being completed.
[0015] Further, the synchronization scheduler comprises a counter; the counter is configured to record the number of received program updating signals; the synchronization scheduler is further configured to prepare to send the synchronization running signal in response to the number of program updating signals being equal to the predetermined number.
[0016] Further, the synchronization scheduler is further configured to prepare to send the synchronization running signal in response to the number of program updating signals being equal to the predetermined number, comprising: the synchronization scheduler is further configured to send a configuration signal in response to the number of program updating signals being equal to the predetermined number, and send the synchronization running signal after receiving a configuration completion signal sent by the direct access controller.
[0017] Further, the synchronization scheduler receives the predetermined number of program updating signals, comprising: the synchronization scheduler receives the predetermined number of program updating signals sent by all the processing cores connected to the synchronization scheduler; or the synchronization scheduler receives the predetermined number of program updating signals, comprising: the synchronization scheduler receives the predetermined number of program updating signals sent by each of the processing cores connected to the synchronization scheduler.
[0018] Further, the synchronization scheduler is configured to generate and send a configuration signal in response to the program updating signal of each processing core, comprising: the synchronization scheduler is configured to send the configuration signal after receiving the program updating signal sent by each processing core connected to the synchronization scheduler.
[0019] Further, the at least two processing cores comprise a first processing core and a second processing core; the programs executed by the first processing core and the second processing core are different, and the calculation result of the program executed by the first processing core is the input of the program executed by the second processing core.
[0020] According to a second aspect of the present application, a chip is provided, comprising one or more data processing apparatuses provided by the first aspect.
[0021] According to a third aspect of the present application, a card board is provided, comprising one or more chips provided by the second aspect.
[0022] According to a fourth aspect of the present application, an electronic device is provided, comprising one or more card boards provided by the third aspect.
[0023] According to a fifth aspect of the present application, there is provided a data processing method, comprising: a processing core executing a program; a synchronization scheduler responding to a program update signal of each processing core connected to the synchronization scheduler, generating and sending a configuration signal; and a direct memory access controller reading a program corresponding to each processing core from an external storage unit based on the configuration signal and sending the program to the corresponding processing core.
[0024] According to a sixth aspect of the present application, there is provided a computer storage medium, having stored thereon a computer program, which, when executed by a processor, implements the data processing method of the fifth aspect.
[0025] According to a seventh aspect of the present application, there is provided an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the data processing method of the fifth aspect when executing the program.
[0026] According to an eighth aspect of the present application, there is provided a computer program product, comprising computer instructions, which, when executed by a computing device, enable the computing device to perform the data processing method of the fifth aspect.
[0027] (III) Beneficial Effects
[0028] The above technical solutions of the present application have the following beneficial technical effects:
[0029] The data processing device provided by the embodiments of the present application is provided with a synchronization scheduler and a direct memory access controller, the synchronization scheduler instructs the direct memory access controller to read a program from an external storage unit and send the program to a processing core corresponding to the program, on the one hand, the data processing device provided by the embodiments of the present application does not need to take data from the external storage unit by the processing core, avoiding the delay caused by reading data by multiple processing cores, improving the computing power of the chip, on the other hand, the programs executed by each processing core of the data processing device provided by the embodiments of the present application can be the same or different, and the synchronization scheduler can respond to the update of the programs of multiple processing cores, so that the computing power of each processing core can be fully utilized, and the computing power of the chip is further improved. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is a structural schematic diagram of a chip provided by the prior art;
[0031] Figure 2 is a structural schematic diagram of a chip provided by another prior art;
[0032] Figure 3 is a structural schematic diagram of a data processing device according to the present application;
[0033] Figure 4 is a structural schematic diagram of another data processing device provided according to the present application;
[0034] Figure 5 is a structural schematic diagram of a neural network provided according to the present application;
[0035] Figure 6 is a schematic diagram of operation of a data processing device applied to a neural network according to the present application;
[0036] Figure 7 is a schematic diagram of scheduling of processing cores by a synchronization scheduler according to the present application;
[0037] Figure 8 is a flowchart of a data processing method provided according to the present application. DETAILED DESCRIPTION
[0038] To make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. In addition, in the following description, the description of well-known structures and techniques is omitted to avoid unnecessary confusion of the concept of the present application.
[0039] In neural network computation, chips with multi-core or many-core architecture are often used. The processing cores in a general multi-core or many-core architecture chip have certain independent data processing capabilities and also have relatively large memory spaces, which are generally used to store programs, data and weights of the chip itself.
[0040] How to make the numerous cores efficiently exert computing power is the key to determining the performance of the entire chip. The exertion of computing power of each core depends on multiple factors, such as scheduling and distribution of tasks, architecture of the chip, structure of the core, circuit of the core, etc. Among them, the scheduling and distribution of tasks is a very important factor. If the scheduling and distribution of tasks are reasonable, the effective computing power of each core can be fully exerted, otherwise the effective computing power of each core is low.
[0041] Figure 1 is a structural schematic diagram of a chip provided by the prior art.
[0042] As shown in Figure 1 , the chip includes a scheduler and a plurality of processing cores C1 to Cn, in Figure 1In the shown chip, the scheduler receives instructions sent from outside the chip, for example, the scheduler receives instructions sent from an instruction source outside the chip, and then transmits the instructions to the processing cores according to a preset strategy (for example, according to a preset order), and the processing cores execute the same instructions but process different data. For example, the instruction is to process a+b, but a or b of the two processing cores can be different values, so the data processed by the two processing cores is different data.
[0043] For Figure 1 In the shown chip architecture, each processing core can be a relatively simple structure, for example, a single instruction multiple data (SIMD) structure or a single instruction multiple threads (SIMT) structure.
[0044] Generally, this approach has the following disadvantages:
[0045] The scheduler can only passively receive instructions from the outside and then distribute them to the processing cores. Whether it is a SIMD structure or a SIMT structure, each processing core can only execute the same instructions, resulting in a single function of the chip and a lack of flexibility.
[0046] Figure 2 Another prior art provides a structure diagram of a chip.
[0047] As Figure 2 shown, the chip includes a plurality of processing cores C1 to Cn and a memory unit memory. In Figure 2 In the shown chip, each core can independently read instructions from the memory (for example, DDR) and perform operations. Generally, each core has a complete control circuit, a register group, and other circuits. This structure is common in multi-core CPUs or ASICs.
[0048] Generally, this approach has the following disadvantages:
[0049] (1) Each processing core has high autonomy and can independently execute instructions. However, due to the high autonomy of the processing cores, it is difficult for multiple processing cores to efficiently complete a complete task by cooperating with each other.
[0050] (2) The circuit control in the chip is relatively complex. Each core is almost a complete CPU. If we want to use the efficient cooperation between the processing cores to complete a complete task, the design difficulty of the circuit is high, and the power consumption and area are large.
[0051] (3) Multiple processing cores may frequently access the instruction storage area, causing a decrease in storage access efficiency, which in turn affects the performance of the chip.
[0052] To solve the above problems, the technical scheme of the present application is proposed.
[0053] The chip provided by an embodiment of the present application will be described in detail below. In the description of the present application, it should be noted that the terms "first", "second", "third", and "fourth" are only used for descriptive purposes and should not be understood as indicating or implying relative importance. In addition, the technical features involved in the different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0054] Figure 3 is a structural schematic diagram of a data processing device according to a first embodiment of the present application.
[0055] As Figure 3 shown, the data processing device comprises at least two processing cores, a synchronizer and scheduler (S_S) and a direct memory access controller (DMAC).
[0056] The S_S is connected to the at least two processing cores, which can be all the processing cores in the data processing device, for example, processing core C1 to processing core Cn, and the DMAC is connected to the at least two processing cores, the S_S and an external memory unit Memory.
[0057] The S_S is configured to generate and send a configuration signal in response to a program update signal from each processing core connected to the S_S.
[0058] The DAMC is configured to read and send a program corresponding to each processing core from the external Memory to the corresponding processing core based on the configuration signal.
[0059] The data processing device provided by the embodiment of the present application is provided with a synchronizer and scheduler and a direct memory access controller. The synchronizer and scheduler instructs the direct memory access controller to read a program from an external memory unit and send it to a processing core corresponding to the program. On the one hand, the data processing device provided by the embodiment of the present application does not need to take data from the external memory unit, avoiding the delay caused by multiple processing cores reading data, and improving the computing power of the chip. On the other hand, the programs executed by each processing core of the data processing device provided by the embodiment of the present application can be the same or different, and the synchronizer and scheduler can respond to the update of the programs of multiple processing cores, which can flexibly allocate tasks and fully utilize the computing power of each processing core, further improving the computing power of the chip.
[0060] In one embodiment, the S_S is further configured to send a synchronization running signal to the processing cores connected to the S_S in response to the program update signal, the synchronization running signal being sent to the processing cores after the S_S receives a predetermined number of the program update signals, and the synchronization running signal being used to instruct the processing cores to start executing the respective programs at the same time.
[0061] Further, the S_S is further configured to send the configuration signal first in response to the number of the program update signals being equal to the predetermined number, and send the synchronization running signal after receiving a configuration completion signal sent by the direct access controller.
[0062] In one embodiment, the program includes a plurality of program segments.
[0063] The processing cores are configured to send the program update signal to the S_S after executing each of the program segments.
[0064] Preferably, the number of the program segments executed by each of the processing cores is the same.
[0065] Optionally, the S_S includes a first counter configured to record the number of the program update signals received.
[0066] The S_S is further configured to prepare to send the synchronization running signal in response to the number of the program update signals recorded by the first counter being equal to the predetermined number.
[0067] In the embodiment of the application, the number of the first counters can be one or more.
[0068] Optionally, when the number of the first counters is one, the predetermined number can be the sum of the program update signals sent by the processing cores and received by the S_S, i.e., the predetermined number is the product of the number of the processing cores connected to the S_S and the number of the program segments.
[0069] Optionally, the preset number recorded by the first counter can be one or more.
[0070] For example, the data processing device includes two processing cores, and each of the processing cores executes four program segments.
[0071] When the preset number recorded by the first counter is set to one, the preset number recorded by the first counter is set to eight, i.e., when the first counter records eight program update signals, the S_S sends the configuration signal to the DMAC, and after the S_S receives the configuration completion signal returned by the DMAC, the S_S sends the synchronization signal to the processing cores, and at this time, the first counter is cleared and starts counting again.
[0072] When the preset number of the first counter is multiple values, the preset number of the counter is 8, 16, 24, and the like. The preset number is the cumulative number of the program update signals received by the synchronization controller. When the cumulative number of the program update signals received by the first counter is 8, the S_S sends a configuration signal to the DMAC, and when the configuration completion signal returned by the DMAC is received, the S_S sends a synchronization signal to each core. At this time, the counter is not cleared, and the counting continues. When the cumulative number of the program update signals received by the first counter reaches 16, the S_S sends a configuration signal to the DMAC, and when the configuration completion signal returned by the DMAC is received, the S_S sends a synchronization signal to each core. The first counter continues to count, and when the next preset number is reached, the S_S repeats the above steps.
[0073] Optionally, when the number of the first counter is multiple, the preset number can be the number of program segments.
[0074] At this time, for example, each processing core connected to the S_S corresponds to a first counter, which is used to record the number of program update signals received from the processing core connected to the S_S. After the S_S receives the predetermined number of program update signals sent by each processing core, that is, each first counter receives the predetermined number of program update signals, a configuration signal is sent to the DMAC. After the configuration completion signal sent by the DMAC is received, the S_S sends a synchronization signal to each core.
[0075] The above method of setting the counter is only an exemplary description. Any software and hardware structure and the like that can realize the configuration of the S_S to the DMAC and the synchronization operation of the S_S to each processing core under certain conditions can be used in the embodiments of the present application, and no more description is made herein.
[0076] In a preferred embodiment, the S_S, in response to the program update signals of each of the processing cores, generates and sends a configuration signal, and includes: the S_S, after receiving the program update signals sent by each of the processing cores connected to the S_S, sends the configuration signal. For example, the number of the processing cores connected to the S_S is 5, and then the S_S sends the configuration signal after receiving the program update signals sent by the 5 processing cores.
[0077] In a preferred embodiment, the direct access controller is used to read a program segment from an external memory and send the program segment to the processing core corresponding to the program segment. The external memory is, for example, a storage unit on the host.
[0078] The direct access controller is also used to send a configuration completion signal to the S_S after sending the program segment to the processing core corresponding to the program segment.
[0079] In one embodiment, the S_S is further configured to send a synchronization running signal to the processing cores connected to the S_S in response to the program update signal, including:
[0080] The S_S is further configured to send a synchronization running signal to each of the processing cores connected to the S_S in response to the program update signal and the configuration completion signal.
[0081] Specifically, the S_S is configured to send a configuration signal to the DMAC in response to the number of program update signals recorded by the counter being equal to a predetermined number, and the DMAC sends a configuration completion signal to the S_S after sending the program or program segment indicated by the configuration signal to the corresponding processing core, and the S_S sends a synchronization running signal to each of the processing cores connected to the S_S after receiving the configuration completion signal.
[0082] Here, the program indicated by the configuration signal can be the same as or different from the program just executed by the processing core.
[0083] In one embodiment, the program executed by the processing core includes an operation instruction and a program update instruction, and the processing core is configured to execute the program update instruction after completing the operation instruction, and generate and send a program update signal based on the program update instruction.
[0084] When the program executed by the processing core includes a plurality of program segments, each program segment includes an operation instruction and a program update instruction.
[0085] In one embodiment, the processing core is provided with a storage module PRAM, which is configured to store and receive the program sent by the DMAC, and the program executed by the processing core is read from the PRAM of the processing core itself. In this embodiment, each processing core reads the instructions contained in the program from the PRAM provided by itself, without the need to read from an external memory, so that a complex cache circuit can be avoided, and compared with the prior art, the need to read data from the memory is eliminated, the delay is reduced, and the execution efficiency of the instructions is greatly improved.
[0086] Preferably, the storage space of the PRAM is greater than or equal to 16 KB.
[0087] Optionally, the programs stored by any two processing cores are the same or different.
[0088] In a preferred embodiment, the data processing device includes at least two processing cores, including a first processing core and a second processing core. The first and second processing cores execute different programs, and the calculation result of the program executed by the first processing core can serve as the input to the program executed by the second processing core. Using the calculation result of the program executed by the first processing core as the input to the second processing core enables the chip provided in this embodiment to be used for neural network computation. Furthermore, by using S_S to allow each processing core to simultaneously run its stored program, the processing cores can exchange data in an orderly manner and cooperate efficiently to complete a complete task.
[0089] The data processing device provided in this invention allocates programs to each processing core through S_S and DMAC. The programs executed by each processing core and the exchange of data are set before the program runs. The top-level microcontroller unit (MCU) of the chip or the host of the system only needs to configure the counter of S_S to implement the predetermined strategy. Moreover, the MCU or the host of the system can change the programs executed by each core, as well as the allocation and scheduling of programs, by changing the configuration of the counter of S_S and the programs stored in the memory. This facilitates the modification of the allocation and scheduling of tasks of each processing core in the chip and can efficiently utilize the computing power of the processing core.
[0090] According to another aspect of the present invention, a chip is provided, comprising one or more of the data processing means provided in the above aspects.
[0091] For example, when the chip includes multiple data processing devices, the chip may include multiple S_S, each S_S being connected to multiple processing cores and a direct access controller.
[0092] According to another embodiment of the present invention, a cardboard is provided, including one or more chips provided in the above embodiments.
[0093] According to another embodiment of the present invention, an electronic device is provided, including one or more of the cards provided in the above embodiments.
[0094] Figure 4 This is a schematic diagram of the data processing device provided by the present invention.
[0095] like Figure 4 As shown, the device includes a first processing core C1, a second processing core C2, S_S, and a DMAC. Each processing core is equipped with PRAM.
[0096] Figure 5 This is a schematic diagram of the neural network structure provided by the present invention.
[0097] This embodiment uses a two-layer neural network as an example, such as... Figure 5As shown, the neural network is a 2-layer structure, the control program of each layer of the neural network is 128KB, the calculation amount of each layer is the same, the whole network can be assigned to two processing core pipelines according to the equal task allocation strategy for calculation, that is, C1 and C2 each take charge of the calculation of one layer of the network, and each runs the program of the corresponding layer. For example, C1 calculates the first layer network Layer1, and C2 calculates the second layer network Layer2. The input data is sent to C1, C1 performs the first layer processing on the input data, and sends the first layer processing result to C2, C2 takes the first layer processing result as input and performs the second layer processing to obtain the final result and then outputs, that is, the data flows through Layer1 and Layer2 in turn to realize the operation of the whole neural network, and finally the output is obtained.
[0098] Figure 6 The chip provided by the application is applied to the operation of a neural network.
[0099] As shown in Figure 6 , Input1-1 represents the input of the whole neural network, and also represents Input1-1 as the input of the first layer neural network layer1 at the starting point of the t1 time period, Input2-1 represents the calculation result of layer1 in the t1 time period, and also represents the input of the second layer neural network layer2 at the starting point of the t2 time period, output1 represents the calculation result of layer2 after the t2 time period, and the calculation result is also the output result of the neural network, when C1 processes the pipeline of layer1, the input is the input of the neural network data, the output is the input of C2, and the output of C2 is the final output result.
[0100] The control program of each layer is set to 128KB, since the PRAM of each processing core only has 32KB, the control program of each layer needs to be scheduled according to a certain strategy to update the program of each core. For example, the program of each core can be transmitted to the corresponding core in four program segments, and each time 32KB is transmitted.
[0101] The calculation process of the processing core in the chip when running the neural network is as follows:
[0102] At the initial time, C1 receives the first program segment of the first program at time t1, executes the first program segment, and sends a program update instruction PU_S1 to S_S after executing the first program segment. Since the output of C1 is the input of C2, S_S can be configured to send a configuration signal to DMAC each time it receives the program update signal sent by C1 until C1 executes the last program segment of the first program, at which time S_S sends a configuration signal to DMAC. DMAC reads the first segment of the second program executed by C1 and the first segment of the first program executed by C2 from the external memory and sends them to C1 and C2. When the sending is completed, DMAC sends a configuration completion signal to S_S, and S_S sends a synchronous execution signal Sync to C1 and C2, instructing C1 and C2 to simultaneously start executing the received program segments.
[0103] When C1 or C2 finishes executing the operation instruction of the program segment, a program update instruction PUpdate is executed, which sends a program update signal PU_s to S_S, indicating that the program in PRAM needs to be updated. After receiving the signal, S_S determines whether it has received the PU_s sent by all the processing cores connected to S_S. If S_S has not received the PU_s sent by all the processing cores connected to S_S, it will be in a waiting state until it receives the PU_s sent by all the processing cores connected to S_S. If S_S receives the PU_s sent by all the processing cores connected to S_S, it sends a configuration signal to DMAC, which reads the program segment indicated by the configuration signal from the external memory and sends it to the processing core corresponding to the program segment to update the PRAM of each core.
[0104] After receiving the program update signal sent by all the processing cores connected to S_S for the fourth time, it indicates that all the program segments of each core and each layer have been executed, and S_S first sends a configuration signal to DMAC. DMAC reads the first program segment of the program indicated by the configuration signal from the external memory and sends it to the corresponding processing core. When the sending is completed, DMAC sends a configuration completion signal.
[0105] After receiving the configuration completion signal sent by S_S, DMAC generates and sends a synchronous execution signal Sync, indicating that the programs of each core need to be executed simultaneously. After receiving the synchronous execution signal, each core can start exchanging data, i.e., C1 sends the calculation result to C2.
[0106] It should be noted that the program indicated by the configuration signal can be the same as or different from the program just executed.
[0107] It also needs to be explained that at the initial time, C1 can also receive the first program segment of the first program at t1, C2 also receives the first program segment of the first program at t1, C1 and C2 execute the first program segment, and the input of the first program segment of the first program executed by C2 can be set as a preset value, so that C1 and C2 execute the first program at the same time.
[0108] Figure 7 is a schematic diagram of scheduling of the processing core by the synchronization scheduler provided by the application.
[0109] As shown in Figure 7 C1 and C2 receive the Sync signal sent by S_S, and simultaneously start running the program stored in PRAM from the beginning. After C1 executes the operation instruction in the first program segment, it executes the last instruction in the first program segment, i.e., the update instruction PUpdate, indicating that the program segment has been executed, PUpdate generates the program update signal PU_s1 and sends the update signal to S_S, and then C1 starts waiting.
[0110] S_S receives PU_s1 and finds that it has not received the update signal sent by each processing core connected with S_S, i.e., PU_s2 has not been received, and continues to wait.
[0111] After C2 executes the operation instruction in the first program segment, it executes the last instruction in the first program segment, i.e., the update instruction PUpdate, indicating that the program segment has been executed, PUpdate generates the program update signal PU_s2 and sends it to S_S, and then C2 starts waiting.
[0112] S_S receives PU_s2, and the first counter finds that it has received the program update signal sent by each processing core connected with S_S, and configures DMAC to start the program update of C1 and C2.
[0113] DMAC reads the new program segment of C1 and C2 from Memory and sends it to the PRAM of C1 and C2, respectively, until the new program segment is updated.
[0114] When the first counter of S_S records that it has received 8 program update signals, it indicates that the two processing cores have completed four program segments respectively, i.e., the entire program of the two processing cores has been executed, S_S resets the first counter to start recording the update number from the beginning; first configures DMAC to start the program update of C1 and C2, i.e., reloads the first program segment of the next program, and when DMAC sends the program segment, it sends a configuration completion signal to S_S, S_S generates a synchronization running signal Sync and sends it to each core, indicating that each core starts working at the same time, at this time, each core can transmit data to each other.
[0115] It should be noted that the resetting of the first counter, the working order of configuring the DMAC and sending the synchronization running signal can not be in sequence.
[0116] The above are described by taking the 2-layer neural network as an example and taking the calculation result of C1 as the input data of C2. Of course, the data processing apparatus provided by the embodiment of the application can be used for any neural network, and C1 and C2 can also have no data association.
[0117] Figure 8 is a flowchart of the data processing method provided by the application.
[0118] As shown in Figure 8 , the data processing method comprises the following steps.
[0119] In step S101, the processing core executes a program.
[0120] In step S102, the synchronization scheduler generates and sends a configuration signal in response to a program update signal of each processing core connected to the synchronization scheduler.
[0121] In step S103, the direct memory access controller reads the program corresponding to each processing core from an external storage unit based on the configuration signal and sends the program to the corresponding processing core.
[0122] In one embodiment, the synchronization scheduler also sends a synchronization running signal to each processing core connected to the S_S in response to the program update signal, and the synchronization running signal is used to indicate that each processing core connected to the synchronization scheduler starts to execute the respective program at the same time.
[0123] Specifically, the S_S also sends a synchronization running signal to each processing core connected to the S_S in response to the program update signal, including: when the S_S receives a predetermined number of program update signals, the S_S sends a synchronization running signal to each processing core connected to the S_S.
[0124] In one embodiment, the program executed by the processing core comprises a plurality of program segments.
[0125] The processing core sends the program update signal to the synchronization scheduler after executing each program segment.
[0126] Preferably, the number of program segments executed by each processing core is the same.
[0127] Preferably, the first counter of the S_S records the number of received program update signals.
[0128] The S_S also prepares to send the synchronization running signal in response to the number of program update signals recorded by the first counter being equal to the predetermined number.
[0129] In the embodiment of the present application, the number of the first counters can be one or more.
[0130] Optionally, when the number of the first counters is one, the predetermined number can be the total number of the program update signals sent by each processing core connected to the S_S, i.e., the predetermined number is the product of the number of the processing cores connected to the S_S and the number of the program segments.
[0131] Optionally, the preset number recorded by the first counter can be one or more. For example, the data processing device includes two processing cores, and each processing core executes four program segments.
[0132] When the preset number recorded by the first counter is one, the preset number recorded by the first counter is set to eight, i.e., when the first counter receives eight program update signals, the S_S sends a configuration signal to the DMAC, and after receiving a configuration completion signal returned by the DMAC, the S_S sends a synchronization signal to each processing core. At this time, the first counter is cleared and starts counting again.
[0133] When the preset number recorded by the first counter is more than one, the preset number recorded by the first counter can be eight, 16, 24, etc. These preset numbers are all the cumulative number of the program update signals received by the synchronization controller. When the first counter receives eight program update signals, the S_S sends a configuration signal to the DMAC, and after receiving a configuration completion signal returned by the DMAC, the S_S sends a synchronization signal to each processing core. At this time, the first counter is not cleared and continues counting. When the first counter receives 16 program update signals, the S_S sends a configuration signal to the DMAC, and after receiving a configuration completion signal returned by the DMAC, the S_S sends a synchronization signal to each processing core. The first counter continues counting, and when the next preset number is reached, the S_S repeats the above steps.
[0134] Optionally, when the number of the first counters is more than one, the predetermined number can be the number of the program segments.
[0135] At this time, for example, each processing core connected to the S_S corresponds to a first counter, which is used to record the number of the program update signals sent by the processing core connected to the S_S. After the S_S receives a predetermined number of program update signals from each processing core, i.e., each first counter receives a preset number of program update signals, a configuration signal is sent to the DMAC. After receiving a configuration completion signal sent by the DMAC, the S_S sends a synchronization signal to each processing core.
[0136] Optionally, the S_S generates and sends the configuration signal in response to the program update signal of each processing core, including that the S_S sends the configuration signal after receiving the program update signal sent by each processing core connected to the S_S. For example, the number of processing cores connected to the synchronous scheduler is 5, and the synchronous scheduler sends the configuration signal after receiving the program update signal sent by the 5 processing cores.
[0137] Optionally, the DMAC also reads the program segment corresponding to each processing core from the external memory and sends the program segment to the corresponding processing core. The DMAC also sends the configuration completion signal to the S_S after sending the program segment to the corresponding processing core.
[0138] In one embodiment, the S_S also sends the synchronization running signal to each processing core connected to the S_S in response to the program update signal, including that the S_S is also configured to send the synchronization running signal to each processing core connected to the S_S in response to the program update signal and the configuration completion signal.
[0139] Specifically, the S_S sends the configuration signal to the DMAC in response to the number of program update signals recorded by the counter being equal to the predetermined number, and the DMAC sends the configuration completion signal to the S_S after sending the program or program segment indicated by the configuration signal to the corresponding processing core. The S_S sends the synchronization running signal to each processing core connected to the S_S after receiving the configuration completion signal.
[0140] In one embodiment, the program includes an operation instruction and a program update instruction; the processing core executes the program update instruction after the operation instruction is completed, and generates and sends the program update signal based on the completion of the program update instruction. The completion of the operation instruction refers to when or after the operation instruction is completed.
[0141] In one embodiment, the at least two processing cores include a first processing core and a second processing core; the programs executed by the first processing core and the second processing core are different, and the calculation result of the program executed by the first processing core is the input of the program executed by the second processing core.
[0142] According to another embodiment of the present application, a computer storage medium is provided, and the computer storage medium stores a computer program. The program is executed by a processor to implement the data processing method provided by the above-mentioned embodiments.
[0143] According to another embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to implement the data processing method provided by the above-mentioned embodiments.
[0144] According to another aspect of the present application, a computer program product is provided, which comprises computer instructions, and when the computer instructions are executed by a computing device, the computing device can perform the data processing method provided by the above-mentioned embodiments.
[0145] In the data processing method provided by the embodiments of the present application, the direct memory access controller is instructed by the synchronization scheduler to read a program from an external storage unit and send the program to the processing core corresponding to the program. On the one hand, the processing core does not need to fetch data from the external storage unit, avoiding the delay caused by multiple processing cores reading data, improving the computing power of the chip. On the other hand, the programs executed by the processing cores can be the same or different, and the synchronization scheduler can respond to the update of the programs of multiple processing cores, so that the task allocation can be flexible, the computing power of each processing core can be fully utilized, and the computing power of the chip can be further improved.
[0146] It should be understood that the above specific embodiments of the present application are only used for illustrative or explanatory purposes of the principles of the present application, and do not constitute a limitation of the present application. Therefore, any modification, equivalent replacement, improvement, etc. made without departing from the spirit and scope of the present application shall be included in the protection scope of the present application. In addition, the appended claims of the present application are intended to cover all variations and modifications falling within the scope and boundary of the appended claims, or the equivalent forms of such scope and boundary.
Claims
1. A data processing apparatus, characterized by, The data processing apparatus comprises: at least two processing cores; a synchronization scheduler configured to generate and send a configuration signal in response to a program update signal from each of the processing cores connected to the synchronization scheduler; a direct memory access controller configured to read and send a program corresponding to each of the processing cores from an external storage unit based on the configuration signal; the synchronization scheduler is further configured to send a synchronization run signal to each of the processing cores connected to the synchronization scheduler in response to the program update signal, the synchronization run signal being sent to each of the processing cores after the synchronization scheduler receives a predetermined number of the program update signals, the synchronization run signal being used to instruct each of the processing cores to start executing the respective program at the same time.
2. The data processing apparatus according to claim 1, wherein the program comprises a plurality of program segments.
3. The data processing apparatus according to claim 2, wherein each of the processing cores is configured to send the program update signal to the synchronization scheduler after executing each of the program segments.
4. The data processing apparatus according to claim 2, wherein the direct memory access controller is configured to read and send each of the program segments corresponding to each of the processing cores from the external storage unit to the respective processing core; the direct memory access controller is further configured to send a configuration completion signal to the synchronization scheduler after sending each of the program segments to the respective processing core.
5. The data processing apparatus according to claim 4, characterized in that, the synchronization scheduler is further configured to send the synchronization run signal to each of the processing cores connected to the synchronization scheduler in response to the program update signal and the configuration completion signal. the synchronization scheduler comprises a counter configured to record the number of received program update signals; 6. The data processing apparatus according to any one of claims 1 to 5, characterized in that, the synchronization scheduler is further configured to prepare to send the synchronization run signal in response to the number of received program update signals being equal to the predetermined number.
7. The data processing apparatus according to any one of claims 1-5, wherein each of the processing cores executes the same number of program segments; the predetermined number is the number of program segments, or the predetermined number is the product of the number of processing cores connected to the synchronization scheduler and the number of program segments. the at least two processing cores comprise a first processing core and a second processing core; 8. The data processing apparatus according to any one of claims 1 to 5, characterized by, the first processing core and the second processing core execute different programs, and the calculation result of the program executed by the first processing core is an input of the program executed by the second processing core. The data processing apparatus comprises:
9. A data processing method, characterized by, a processing core executing a program; a synchronization scheduler configured to generate and send a configuration signal in response to a program update signal from each of the processing cores connected to the synchronization scheduler; a direct memory access controller configured to read and send the program corresponding to each of the processing cores from an external storage unit based on the configuration signal to the respective processing core; The synchronization scheduler sends a synchronization running signal to each processing core connected to the synchronization scheduler in response to the program update signal, the synchronization running signal being sent to each processing core after the synchronization scheduler receives a predetermined number of the program update signals, and the synchronization running signal being used to instruct each processing core to start executing the respective program at the same time.
Citation Information
Patent Citations
Prefetching instruction blocks
CN108027766A