Artificial intelligence acceleration method and device, chip, electronic device, storage medium
By arranging the controller and data scheduling unit vertically and establishing a propagation link, the problem of irregular layout of artificial intelligence accelerators was solved, achieving the effect of reasonable layout and efficient data flow.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-03-02
- Publication Date
- 2026-06-19
AI Technical Summary
The irregular layout of existing artificial intelligence accelerators leads to difficulties in layout and wiring, and low area utilization.
The controller, k-level data scheduling unit, and first direct memory access (DMA) interface are arranged vertically to establish a propagation link to realize the reverse and forward propagation of data scheduling commands, and data scheduling operations are executed in each level of data scheduling unit.
This approach achieves a reasonable layout for the AI accelerator, reduces wiring complexity, improves chip area utilization, and ensures efficient data flow.
Smart Images

Figure CN116738130B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an artificial intelligence acceleration method and device, chip, electronic device, and storage medium. Background Technology
[0002] Currently, the architecture of artificial intelligence accelerators mainly includes convolutional processing arrays, data scheduling units, and controllers.
[0003] For AI accelerators, such as Figure 1 As shown, in order to achieve efficient data flow, for example, for the operation unit E of the convolution processing array, it can receive the operation result of the operation unit C while generating the corresponding operation result based on the image data provided by the second-level data scheduling unit D. Therefore, the controller must be designed in the lower right corner and connected to the first-level data scheduling unit B. This results in an irregular layout of the entire artificial intelligence accelerator. As a result, the layout and wiring of the chip back end of the deployment of the artificial intelligence accelerator is more difficult and the area utilization is low. Summary of the Invention
[0004] This application provides an artificial intelligence acceleration method, device, chip, electronic device, and storage medium.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides an artificial intelligence accelerator, the layout rules of which include: a controller, a k-level data scheduling unit and a first direct memory access (DMA) interface, where k is a natural number greater than 1;
[0007] The controller and the k-level data scheduling unit, from the k-level data scheduling unit to the first-level data scheduling unit, are arranged vertically in sequence.
[0008] The first DMA interface is deployed on the k-th level data scheduling unit, and the first DMA interface is connected to the controller for receiving data scheduling commands sent by the controller;
[0009] A first propagation link is established between the k-level data scheduling units to propagate the data scheduling command from the k-level data scheduling unit in reverse to the first-level data scheduling unit, and then propagate the data scheduling command from the first-level data scheduling unit in forward to the k-level data scheduling unit.
[0010] Each data scheduling unit of the k-level data scheduling unit is used to execute data scheduling operations according to the corresponding sub-commands in the data scheduling command when the data scheduling command is propagated forward into the unit.
[0011] The aforementioned artificial intelligence accelerator also includes: a convolutional processing array;
[0012] The convolution processing array includes k rows of operation units, with each row of operation units arranged side by side with a first-level data scheduling unit in the k-level data scheduling unit;
[0013] Each data scheduling unit of the k-level data scheduling unit is connected to an adjacent operation unit in a row of parallel operation units. Specifically, it is used to read the corresponding image data according to the corresponding sub-command in the data scheduling command and propagate the corresponding image data to the connected operation unit. The image data corresponding to each data scheduling unit is the data of the feature image on one input channel.
[0014] In each of the k rows of operation units, the operation unit connected to the data scheduling unit is used to propagate the obtained image data horizontally to each operation unit in the same row in sequence, so that the operation units in the same row can obtain the same image data.
[0015] The aforementioned artificial intelligence accelerator also includes: an accumulator;
[0016] In the k-row operation units, the first row operation unit is arranged alongside the first-level data scheduling unit, and the k-row operation unit is arranged alongside the k-th level data scheduling unit.
[0017] The accumulator is located side-by-side with the controller and adjacent to the k-th row of arithmetic units; each arithmetic unit of the k-th row of arithmetic units is connected to the accumulator.
[0018] Each arithmetic unit in the first row is used to perform multiplication and accumulation operations on the obtained image data and output the operation result to the arithmetic unit in the same column of the next row;
[0019] In the k-row operation units, each operation unit of each row operation unit between the first row operation unit and the k-th row operation unit is used to perform multiplication and accumulation operations on the obtained image data, and then accumulate the operation result with the operation result output by the operation unit in the same column of the previous row, and output it to the operation unit in the same column of the next row.
[0020] Each operation unit in the k-th row is used to perform multiplication and accumulation operations on the obtained image data, and then add the operation result to the operation result output by the operation unit in the same column of the previous row, and output the result to the accumulator.
[0021] The aforementioned artificial intelligence accelerator also includes: k memories;
[0022] In the k memories, each memory is parallel to and connected to a first-level data scheduling unit in the k-level data scheduling unit;
[0023] Each level of the k-level data scheduling unit is specifically used to read the corresponding image data from the connected memory.
[0024] In the aforementioned artificial intelligence accelerator, the first DMA interface is also used to receive initialization information;
[0025] The first propagation link is also used to propagate the initialization information back from the k-th level data scheduling unit to the first level data scheduling unit;
[0026] Each data scheduling unit of the k-level data scheduling unit is further configured to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
[0027] The aforementioned AI accelerator also includes: a second DMA interface;
[0028] The second DMA interface is deployed on the k-th level data scheduling unit and is used to receive initialization information;
[0029] A second propagation link is established between the k-level data scheduling units. The second propagation link is used to propagate the initialization information from the k-level data scheduling unit back to the first-level data scheduling unit.
[0030] Each data scheduling unit of the k-level data scheduling unit is further configured to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
[0031] This application provides an artificial intelligence acceleration method applied to an artificial intelligence accelerator. The artificial intelligence accelerator layout includes: a controller, k-level data scheduling units, and a first direct memory access (DMA) interface, where k is a natural number greater than 1. The controller and the k-level data scheduling units, from the k-th level to the first level, are arranged vertically in sequence. The first DMA interface is deployed on the k-th level data scheduling unit and connected to the controller. A first propagation link is established between the k-level data scheduling units. The method includes:
[0032] Using the first DMA interface, receive the data scheduling command sent by the controller;
[0033] Using the first propagation link, the data scheduling command is propagated backward from the k-th level data scheduling unit to the first level data scheduling unit, and then the data scheduling command is propagated forward from the first level data scheduling unit to the k-th level data scheduling unit.
[0034] Using each level of the k-level data scheduling unit, when the data scheduling command is propagated forward into the unit, the data scheduling operation is executed according to the corresponding sub-command in the data scheduling command.
[0035] This application provides a chip that includes the aforementioned artificial intelligence accelerator.
[0036] This application provides an electronic device, including: an artificial intelligence accelerator, a memory for storing computer programs capable of running on the artificial intelligence accelerator, and a communication bus;
[0037] The communication bus is used to realize the communication connection between the artificial intelligence accelerator and the memory;
[0038] The artificial intelligence accelerator is used to execute the computer program stored in the memory to implement the above-described artificial intelligence acceleration method.
[0039] This application provides a computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the above-described artificial intelligence acceleration method.
[0040] This application provides an artificial intelligence acceleration method and device, chip, electronic device, and storage medium, including an AI accelerator layout rule. The rule includes: a controller, k-level data scheduling units, and a first direct memory access (DMA) interface, where k is a natural number greater than 1. The controller and the k-level data scheduling units, from the k-th level to the first level, are arranged vertically. The first DMA interface is deployed on the k-th level data scheduling unit and connected to the controller to receive data scheduling commands sent by the controller. A first propagation link is established between the k-level data scheduling units to propagate data scheduling commands backward from the k-th level data scheduling unit to the first level data scheduling unit, and then forward from the first level data scheduling unit to the k-th level data scheduling unit. Each level data scheduling unit within the k-level data scheduling unit executes a data scheduling operation according to the corresponding sub-command in the data scheduling command when the data scheduling command propagates forward into the unit. The AI accelerator provided in this application not only deploys the controller on one side of the k-level data scheduling unit, resulting in a reasonable and regular layout, but also configures a specific data flow direction to ensure efficient data flow. Attached Figure Description
[0041] Figure 1 A schematic diagram of the structure of an artificial intelligence accelerator provided by existing technology;
[0042] Figure 2 This is a schematic diagram of the structure of an artificial intelligence accelerator provided in an embodiment of this application;
[0043] Figure 3 A schematic diagram illustrating the correspondence of a convolutional processing array provided in an embodiment of this application;
[0044] Figure 4 A flowchart illustrating an artificial intelligence acceleration method provided in an embodiment of this application;
[0045] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0047] The technical solutions of this application and how they solve the aforementioned technical problems will be described in detail below through embodiments and in conjunction with the accompanying drawings. The embodiments below can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0048] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0049] This application provides an artificial intelligence accelerator. Figure 2 This is a schematic diagram of the structure of an artificial intelligence accelerator provided in an embodiment of this application. Figure 2 As shown in the embodiments of this application, the layout rules of the artificial intelligence accelerator 1 include: a controller 10, a k-level data scheduling unit 11, and a first direct memory access (DMA) interface 12, where k is a natural number greater than 1.
[0050] The controller 10 and the k-level data scheduling unit 11, from the k-level data scheduling unit to the first-level data scheduling unit, are arranged vertically in sequence.
[0051] The first DMA interface 12 is deployed on the k-th level data scheduling unit. The first DMA interface 12 is connected to the controller 10 and is used to receive data scheduling commands sent by the controller 10.
[0052] A first propagation link 13 is established between k-level data scheduling units 11, which is used to propagate data scheduling commands from the k-level data scheduling unit in reverse to the first-level data scheduling unit, and then propagate data scheduling commands from the first-level data scheduling unit in forward to the k-level data scheduling unit.
[0053] Each level of the k-level data scheduling unit 11 is used to execute data scheduling operations according to the corresponding sub-commands in the data scheduling command when the data scheduling command is propagated forward into the unit.
[0054] It should be noted that, in the embodiments of this application, the k-level data scheduling unit 11, wherein the first-level data scheduling unit is the first data scheduling unit used to execute the corresponding sub-command in the data scheduling command, and the k-level data scheduling unit is the last data scheduling unit used to execute the corresponding sub-command in the data scheduling command. The specific number of k-level data scheduling units 11 can be set according to actual needs and application scenarios, and is not limited in the embodiments of this application.
[0055] It should be noted that, in the embodiments of this application, each level of data scheduling unit can not only be used for data storage and retrieval, but also has vector calculation function to realize vector calculation of data.
[0056] It is understood that, in the embodiments of this application, as Figure 2 As shown, the controller 10, the k-th level data scheduling unit, the (k-1)-th level data scheduling unit, ... the first level data scheduling unit are arranged vertically in sequence. The k-th level data scheduling unit is the closest to the controller 10, which is the last level data scheduling unit, and the one farthest from the controller 10 is the first level data scheduling unit. This makes the layout of the artificial intelligence accelerator 1 reasonable and regular, reduces the difficulty of the back-end wiring of the chip for deploying the artificial intelligence accelerator 1, and also improves the utilization rate of the chip area.
[0057] It should be noted that, in the embodiments of this application, the transmission and reception of data scheduling commands between the k-level data scheduling unit 11 and the controller 10 is achieved through the first DMA interface 12 deployed on the k-level data scheduling unit. The first DMA interface 12 is connected to the controller 10, thereby receiving the data scheduling commands sent by the controller 10. Then, through the first propagation link 13 established between the k-level data scheduling units 11, the commands are propagated to each level of data scheduling unit, so that each level of data scheduling unit can execute data scheduling operations according to the corresponding sub-commands in the data scheduling commands.
[0058] It is understandable that for the k-level data scheduling unit 11, the data scheduling operation needs to be executed sequentially by each level of data scheduling unit according to the corresponding sub-command in the data scheduling command, in order from the first-level data scheduling unit to the k-level data scheduling unit, to ensure efficient data flow. However, in the embodiments of this application, in order to achieve a regular layout of the artificial intelligence accelerator 1, such as... Figure 2 As shown, the controller 10 is deployed above the k-th level data scheduling unit. The data scheduling command is received by the k-th level data scheduling unit through the first DMA interface 12. Therefore, the data scheduling command needs to be propagated backward from the k-th level data scheduling unit to the first level data scheduling unit through the first propagation link 13. Then, the data scheduling command is propagated forward from the first level data scheduling unit to the k-th level data scheduling unit. In this way, each level data scheduling unit can perform data scheduling operations according to the corresponding sub-command in the data scheduling command when the data scheduling command is propagated forward into the unit.
[0059] Specifically, in the embodiments of this application, such as Figure 2 As shown, the artificial intelligence accelerator 1 also includes: a convolution processing array 14;
[0060] The convolution processing array 14 includes k rows of operation units, with each row of operation units arranged side by side with a first-level data scheduling unit in the k-level data scheduling unit 11;
[0061] Each level of the k-level data scheduling unit 11 is connected to an adjacent operation unit in a row of parallel operation units. Specifically, it is used to read the corresponding image data according to the corresponding sub-command in the data scheduling command and to propagate the corresponding image data to the connected operation unit. The image data corresponding to each data scheduling unit is the data of the feature image on one input channel.
[0062] In each row of k operation units, the operation unit connected to the data scheduling unit is used to propagate the obtained image data horizontally to each operation unit in the same row in sequence, so that the operation units in the same row can obtain the same image data.
[0063] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 further includes a convolution processing array 14, which is composed of computing units, each of which is capable of performing multiplication and accumulation operations, i.e., multiply-accumulate operations.
[0064] Figure 3 This is a schematic diagram illustrating the correspondence of a convolutional processing array provided in an embodiment of this application. For example... Figure 3As shown in the embodiments of this application, each operation unit in the convolution processing array 14 can perform multiplication and accumulation operations. Each column of operation units is preloaded with the weight data of a convolution kernel. Different operation units in the same column correspond to the weight data of the same convolution kernel for different input channels. Different operation units in each row correspond to the weight data of different convolution kernels for the same input channel.
[0065] It is understood that, in the embodiments of this application, for Figure 3 The correspondence shown in the convolution processing array 14 is that, in fact, in the convolution processing array 14, the image data that each row of operation units needs to process is the data of the feature image in one input channel. Therefore, after obtaining the image data of the feature image in one input channel provided by the data scheduling unit, the operation units connected to the data scheduling unit in each row can propagate horizontally in sequence, so that the operation units in the same row can all obtain the same image data for processing.
[0066] Specifically, in the embodiments of this application, such as Figure 2 As shown, the artificial intelligence accelerator 1 also includes: an accumulator 15;
[0067] In the k-row arithmetic units, the first row arithmetic unit is located alongside the first-level data scheduling unit, and the k-row arithmetic unit is located alongside the k-th level data scheduling unit.
[0068] Accumulator 15 is placed side by side with controller 10 and adjacent to the k-th row of arithmetic units; each arithmetic unit of the k-th row of arithmetic units is connected to accumulator 15.
[0069] Each operation unit in the first row is used to perform multiplication and accumulation operations on the obtained image data, and output the operation result to the operation unit in the same column of the next row;
[0070] In the k-row arithmetic units, each arithmetic unit in each row between the first row arithmetic unit and the k-th row arithmetic unit is used to perform multiplication and accumulation operations on the obtained image data, and then accumulate the operation result with the operation result output by the arithmetic unit in the same column of the previous row, and output it to the arithmetic unit in the same column of the next row.
[0071] Each operation unit in the k-th row is used to perform multiplication and accumulation operations on the obtained image data, and then add the result of the operation to the result of the operation unit in the same column of the previous row, and output it to the accumulator 15.
[0072] It should be noted that in the embodiments of this application, the accumulator 15 can store the operation result output by the k-th row operation unit. Of course, it can also be accumulated according to the actual situation. For example, if the number of input channels of the feature image exceeds k, that is, the k-th row operation unit cannot meet the processing of all input channel image data at one time, the k-th row operation unit can be used repeatedly for processing, and finally accumulated in the accumulator 15. This application embodiment does not limit this.
[0073] It is understood that in the embodiments of this application, the number of rows in each of the k-row operation units is actually the same as the level of the data scheduling units arranged side by side. That is, the first row of operation units is arranged side by side with the first-level data scheduling unit, the second row of operation units is arranged side by side with the second-level data scheduling unit, and so on, until the k-th level data scheduling unit is arranged side by side with the k-th row of operation units. Each operation unit in the first row of operation units only needs to perform multiplication and accumulation operations on the obtained image data and output it to the operation unit in the same column of the next row. Each operation unit in each row of operation units between the first row of operation units and the k-th row of operation units not only needs to perform multiplication and accumulation operations on the obtained image data, but also needs to accumulate the result of the operation with the result output by the operation unit in the same column of the previous row, and then output it to the operation unit in the same column of the next row. The k-th row of operation units also needs to accumulate the result, the difference being that the result is output to the accumulator 15.
[0074] Specifically, in the embodiments of this application, such as Figure 2 As shown, the artificial intelligence accelerator 1 also includes: k memory units 16;
[0075] In the k memory units 16, each memory unit is parallel to and connected to the first-level data scheduling unit in the k-level data scheduling unit 11.
[0076] Each level of the k-level data scheduling unit 11 is specifically used to read the corresponding image data from the connected memory.
[0077] It should be noted that, in the embodiments of this application, the format of the data scheduling command is as shown in Table 1 below:
[0078]
[0079] Here, SRAM1 to SRAMk each represent a memory, and readen represents read enable. For example, if a convolution command needs to be executed, the command identifier is set to convolution command. The feature image has 8 input channels. Therefore, SRAM1_readen to SRAM8_readen need to be enabled respectively, and then issued through controller 10, indicating that the data of the feature image in input channels 1 to 8 needs to be read from the SRAM and sent to the convolution processing array 14 for convolution operation. This command first reaches the first-level data scheduling unit. After receiving the command, the first-level data scheduling unit parses the command in the first clock cycle and finds that SRAM1_readen is enabled, requiring the reading of the feature image data in input channel 1 from SRAM1. Therefore, the image data in SRAM1 passes through the first-level data scheduling unit to the rightmost operation unit in the first row of operation units, where a single-channel multiplication and accumulation operation is performed. Simultaneously, the command is transmitted from the first-level data scheduling unit to the second... In the second clock cycle, the data scheduling unit at the first level sends the accumulated result of the rightmost operation unit in the first row of operation units to the operation unit in the same column of the second row of operation units, i.e., the rightmost operation unit in the second row of operation units. At the same time, the second-level data scheduling unit parses the command and finds that SRAM2_readen is valid. Therefore, it retrieves the image data of the feature image on input channel 2 from SRAM2 and sends it to the rightmost operation unit in the second row of operation units. At this time, since the operation result of the rightmost operation unit in the first row of operation units has already been sent to the rightmost operation unit in the second row of operation units, the two operation results are accumulated to obtain the convolution result of the feature image of the two input channels, which continues to propagate to the upper layer. This process continues in the same way. During the propagation of the command, it can be ensured that each operation unit simultaneously receives the image data obtained from SRAM and the operation result propagated from the operation unit in the same column of the previous row, and continues to perform the accumulation operation.
[0080] It should be noted that, in the embodiments of this application, the above description refers to the vertical data propagation process. Horizontal data propagation also exists, such as... Figure 2 As shown, each arithmetic unit performs a multiplication-accumulation operation and propagates the result forward while simultaneously propagating the image data to the left. Similarly, the timing of the arithmetic unit reading image data from the right is the same as the timing of obtaining the result from the arithmetic unit in the same column of the previous row. All arithmetic units propagate and process data in a pulsating manner, with the data pulsating uniformly once per clock cycle.
[0081] Specifically, in the embodiments of this application, the first DMA interface 12 is also used to receive initialization information;
[0082] The first propagation link 13 is also used to propagate the initialization information from the k-th level data scheduling unit back to the first level data scheduling unit;
[0083] Each data scheduling unit of the k-level data scheduling unit 11 is also used to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
[0084] It should be noted that in the embodiments of this application, the first DMA interface 12 and the first propagation link 13 can be time-division multiplexed. That is, the first DMA interface 12 and the first propagation link 13 can be used to initialize k memories 16 first, and then the first DMA interface 12 and the first propagation link 13 can be used to propagate data scheduling commands.
[0085] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 may further include: a second DMA interface (not shown in the figure);
[0086] The second DMA interface is deployed on the k-th level data scheduling unit and is used to receive initialization information;
[0087] A second propagation link (not shown in the figure) is established between the k-level data scheduling units 11. The second propagation link is used to propagate the initialization information from the k-level data scheduling unit back to the first-level data scheduling unit.
[0088] Each data scheduling unit of the k-level data scheduling unit 11 is also used to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
[0089] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 may also deploy a second DMA interface and establish a second propagation link between the k-level data scheduling units 11. In this way, the initialization of the k memories 16 and the propagation of data scheduling commands are implemented using different DMA interfaces and propagation links, and are independent of each other.
[0090] This application provides an artificial intelligence accelerator with a layout rule, including: a controller, k-level data scheduling units, and a first direct memory access (DMA) interface, where k is a natural number greater than 1; the controller and the k-level data scheduling units, from the k-th level to the first level, are arranged vertically; the first DMA interface is deployed on the k-th level data scheduling unit and connected to the controller for receiving data scheduling commands sent by the controller; a first propagation link is established between the k-level data scheduling units for propagating data scheduling commands backward from the k-th level data scheduling unit to the first level data scheduling unit, and then propagating data scheduling commands forward from the first level data scheduling unit to the k-th level data scheduling unit; each level data scheduling unit of the k-level data scheduling unit is used to execute data scheduling operations according to the corresponding sub-commands in the data scheduling command when the data scheduling command is propagated forward into the unit. The artificial intelligence accelerator provided in this application not only deploys the controller on one side of the k-level data scheduling unit, making the layout reasonable and regular, but also configures a specific data flow direction to ensure efficient data flow.
[0091] This application provides an artificial intelligence acceleration method, applied to the aforementioned artificial intelligence accelerator 1. Figure 4 This is a flowchart illustrating an artificial intelligence acceleration method provided in an embodiment of this application. Figure 4 As shown, the artificial intelligence acceleration method mainly includes the following steps:
[0092] S201. Receive the data scheduling command sent by the controller using the first DMA interface.
[0093] In the embodiments of this application, see Figure 2 The artificial intelligence accelerator 1 uses the first DMA interface 12 to receive data scheduling commands sent by the controller 10.
[0094] It is understood that, in the embodiments of this application, as Figure 2 As shown, the first DMA interface 12 can be connected to the controller 10, thereby receiving data scheduling commands sent by the controller 10 to provide to the k-level data scheduling unit 11 in the artificial intelligence accelerator 1.
[0095] S202. Using the first propagation link, the data scheduling command is propagated backward from the k-th level data scheduling unit to the first level data scheduling unit, and then the data scheduling command is propagated forward from the first level data scheduling unit to the k-th level data scheduling unit.
[0096] In the embodiments of this application, see Figure 2A first propagation link 13 is established between the k-level data scheduling units 11. After the artificial intelligence accelerator 1 receives the data scheduling command sent by the controller 10 using the first DMA interface 12, it can use the first propagation link 13 to propagate the data scheduling command from the k-level data scheduling unit in reverse to the first-level data scheduling unit, and then propagate the data scheduling command from the first-level data scheduling unit in forward to the k-level data scheduling unit.
[0097] S203. Utilizing each level of the k-level data scheduling unit, when the data scheduling command is propagated forward into the unit, execute the data scheduling operation according to the corresponding sub-command in the data scheduling command.
[0098] In the embodiments of this application, see Figure 2 The artificial intelligence accelerator 1 can utilize each level of the k-level data scheduling unit 11 to execute data scheduling operations according to the corresponding sub-commands in the data scheduling command when the data scheduling command is propagated forward into the unit.
[0099] It should be noted that, in the embodiments of this application, see... Figure 2 The AI accelerator 1 also includes a convolution processing array 14 composed of k rows of operation units, an accumulator 15, and k memories 16. Specifically, the AI accelerator 1 can utilize each level of the k-level data scheduling unit 11 to read the corresponding image data according to the corresponding sub-command in the data scheduling command, and propagate the corresponding image data to the parallel and connected operation units. Then, using the operation units in each row of the k-level operation units connected to the data scheduling unit, the obtained image data is propagated horizontally to each operation unit in the same row in sequence, so that the operation units in the same row can obtain the same image data.
[0100] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 specifically utilizes each operation unit of the first row of operation units to perform multiplication and accumulation operations on the obtained image data, and outputs the operation result to the operation unit in the same column of the next row; utilizes each operation unit of each row of operation units between the first row of operation units and the kth row of operation units to perform multiplication and accumulation operations on the obtained image data, and adds the operation result to the operation result output by the operation unit in the same column of the previous row, and outputs it to the operation unit in the same column of the next row; utilizes each operation unit of the kth row of operation units to perform multiplication and accumulation operations on the obtained image data, and adds the operation result to the operation result output by the operation unit in the same column of the previous row, and outputs it to the accumulator 15.
[0101] It should be noted that, in the embodiments of this application, see... Figure 2The AI accelerator 1 also includes k memories 16, each of which is parallel to and connected to a first-level data scheduling unit in the k-level data scheduling unit 11. The AI accelerator 1 can use each level of the k-level data scheduling unit 11 to read the corresponding image data from the connected memories.
[0102] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 can also use the first DMA interface 12 to receive initialization information, and then use the first propagation link 13 to propagate the initialization information from the k-th level data scheduling unit to the first level data scheduling unit. In addition, each level data scheduling unit of the k-th level data scheduling unit 11 can use the corresponding sub-information in the initialization information to initialize the connected memory when the initialization information is propagated back to the unit.
[0103] It should be noted that, in the embodiments of this application, the artificial intelligence accelerator 1 may further include: a second DMA interface; the second DMA interface is deployed on the k-th level data scheduling unit, and the artificial intelligence accelerator 1 may also use the second DMA interface to receive initialization information. A second propagation link is established between the k-th level data scheduling units 11, and the artificial intelligence accelerator 1 may also use the second propagation link to propagate the initialization information from the k-th level data scheduling unit back to the first level data scheduling unit. In this way, each level data scheduling unit of the k-th level data scheduling unit 11 can be used to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
[0104] This application also provides a chip including the aforementioned artificial intelligence accelerator 1, thereby supporting the execution of corresponding artificial intelligence acceleration methods.
[0105] This application also provides an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device includes: the aforementioned artificial intelligence accelerator 1, a memory 2 for storing computer programs capable of running on the artificial intelligence accelerator 1, and a communication bus 3;
[0106] The communication bus 3 is used to realize the communication connection between the artificial intelligence accelerator 1 and the memory 2;
[0107] The artificial intelligence accelerator 1 is used to execute the computer program stored in the memory 2 to implement the above-mentioned artificial intelligence acceleration method.
[0108] This application provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the aforementioned artificial intelligence acceleration method. The computer-readable storage medium can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or it can be a device including one or any combination of the above-mentioned memories, such as a mobile phone, computer, tablet device, or personal digital assistant.
[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0110] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the schematic and / or block diagrams can be implemented by computer program instructions, and combinations of blocks in the schematic and / or block diagrams can be implemented. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the schematic and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An artificial intelligence accelerator, comprising: The layout rules for the artificial intelligence accelerator include: a controller, a k-level data scheduling unit, and a first direct memory access (DMA) interface, where k is a natural number greater than 1. The controller and the k-level data scheduling unit, from the k-level data scheduling unit to the first-level data scheduling unit, are arranged vertically in sequence. The first DMA interface is deployed on the k-th level data scheduling unit, and the first DMA interface is connected to the controller for receiving data scheduling commands sent by the controller; A first propagation link is established between the k-level data scheduling units to propagate the data scheduling command from the k-level data scheduling unit in reverse to the first-level data scheduling unit, and then propagate the data scheduling command from the first-level data scheduling unit in forward to the k-level data scheduling unit. Each data scheduling unit of the k-level data scheduling unit is used to execute data scheduling operations according to the corresponding sub-commands in the data scheduling command when the data scheduling command is propagated forward into the unit.
2. The artificial intelligence accelerator of claim 1, wherein, Also includes: Convolutional processing array; The convolution processing array includes k rows of operation units, with each row of operation units arranged side by side with a first-level data scheduling unit in the k-level data scheduling unit; Each data scheduling unit of the k-level data scheduling unit is connected to an adjacent operation unit in a row of parallel operation units. Specifically, it is used to read the corresponding image data according to the corresponding sub-command in the data scheduling command and propagate the corresponding image data to the connected operation unit. The image data corresponding to each data scheduling unit is the data of the feature image on one input channel. In each of the k rows of operation units, the operation unit connected to the data scheduling unit is used to propagate the obtained image data horizontally to each operation unit in the same row in sequence, so that the operation units in the same row can obtain the same image data.
3. The artificial intelligence accelerator of claim 2, wherein, Also includes: accumulator; In the k-row operation units, the first row operation unit is arranged alongside the first-level data scheduling unit, and the k-row operation unit is arranged alongside the k-th level data scheduling unit. The accumulator is located side-by-side with the controller and adjacent to the k-th row of arithmetic units; each arithmetic unit of the k-th row of arithmetic units is connected to the accumulator. Each arithmetic unit in the first row is used to perform multiplication and accumulation operations on the obtained image data and output the operation result to the arithmetic unit in the same column of the next row; In the k-row operation units, each operation unit of each row operation unit between the first row operation unit and the k-th row operation unit is used to perform multiplication and accumulation operations on the obtained image data, and then accumulate the operation result with the operation result output by the operation unit in the same column of the previous row, and output it to the operation unit in the same column of the next row. Each operation unit in the k-th row is used to perform multiplication and accumulation operations on the obtained image data, and then add the operation result to the operation result output by the operation unit in the same column of the previous row, and output the result to the accumulator.
4. The artificial intelligence accelerator of claim 2, wherein, Also includes: k memory units; In the k memories, each memory is parallel to and connected to a first-level data scheduling unit in the k-level data scheduling unit; Each level of the k-level data scheduling unit is specifically used to read the corresponding image data from the connected memory.
5. The artificial intelligence accelerator according to claim 4, characterized in that, The first DMA interface is also used to receive initialization information; The first propagation link is also used to propagate the initialization information back from the k-th level data scheduling unit to the first level data scheduling unit; Each data scheduling unit of the k-level data scheduling unit is further configured to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
6. The artificial intelligence accelerator of claim 1, wherein, Also includes: Second DMA interface; The second DMA interface is deployed on the k-th level data scheduling unit and is used to receive initialization information; A second propagation link is established between the k-level data scheduling units. The second propagation link is used to propagate the initialization information from the k-level data scheduling unit back to the first-level data scheduling unit. Each data scheduling unit of the k-level data scheduling unit is further configured to initialize the connected memory using the corresponding sub-information in the initialization information when the initialization information is propagated back to the unit.
7. An artificial intelligence acceleration method, characterized by, The method is applied to an artificial intelligence accelerator, wherein the accelerator layout rules include: a controller, a k-level data scheduling unit, and a first direct memory access (DMA) interface, where k is a natural number greater than 1; the controller and the k-level data scheduling units, from the k-th level to the first level, are arranged vertically in sequence; the first DMA interface is deployed on the k-th level data scheduling unit and connected to the controller; a first propagation link is established between the k-level data scheduling units; and the method includes: Using the first DMA interface, receive the data scheduling command sent by the controller; Using the first propagation link, the data scheduling command is propagated backward from the k-th level data scheduling unit to the first level data scheduling unit, and then the data scheduling command is propagated forward from the first level data scheduling unit to the k-th level data scheduling unit. Using each level of the k-level data scheduling unit, when the data scheduling command is propagated forward into the unit, the data scheduling operation is executed according to the corresponding sub-command in the data scheduling command.
8. A chip, characterized by Including the artificial intelligence accelerator as described in any one of claims 1-6.
9. An electronic device, comprising: include: Artificial intelligence accelerator, memory for storing computer programs capable of running on the artificial intelligence accelerator, and communication bus; The communication bus is used to realize the communication connection between the artificial intelligence accelerator and the memory; The artificial intelligence accelerator is used to execute the computer program stored in the memory to implement the artificial intelligence acceleration method of claim 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the artificial intelligence acceleration method as claimed in claim 7.