An AI accelerator, processor, chip, board and electronic device

By designing a control module, on-chip storage module, instruction queue module, and data transport module in the AI ​​accelerator, and stitching image data into regional data blocks, the problem of bandwidth waste in existing technologies is solved, and the performance of the AI ​​accelerator is improved.

CN120950453BActive Publication Date: 2026-02-03BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511468839.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-03
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

When processing image data, existing AI accelerators have insufficient on-chip storage modules to store an entire image, leading to repeated reading of image data and wasting bandwidth.

Method used

Design an AI accelerator, including a control module, an on-chip storage module, a transport instruction queue module, a data transport module, and a computing array. By stitching together regional data blocks of image data, it avoids the import and export of duplicate data and reduces bandwidth consumption.

Benefits of technology

By avoiding the import and export of duplicate data, bandwidth waste is reduced and the performance of the AI ​​accelerator is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950453B_ABST
    Figure CN120950453B_ABST
Patent Text Reader

Abstract

The application provides an AI accelerator, a processor, a chip, a board card and electronic equipment, the AI accelerator comprises a control module, an on-chip storage module, a data carrying instruction queue module, a data carrying module and a calculation array, wherein: the on-chip storage module is used for caching region data, data blocks, intermediate data and output results of image data; the data carrying instruction queue module is used for controlling the data carrying module to splice the data blocks; the data carrying module is used for splicing the data blocks corresponding to the region data based on the region data of the image data and the last spliced data block corresponding to the region data; the calculation array is used for performing calculation in a neural network based on the data blocks, obtaining the intermediate data and the output results; and the control module is used for controlling the operation of the calculation array. The AI accelerator, the processor, the chip, the board card and the electronic equipment provided in the application embodiment reduce the waste of bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to an AI accelerator, processor, chip, board, and electronic device. Background Technology

[0002] An AI (Artificial Intelligence) accelerator is a hardware device specifically designed for efficiently processing artificial intelligence tasks. Its core objective is to accelerate the massive computational operations in AI algorithms, thereby improving the speed and efficiency of AI applications.

[0003] In existing technologies, such as Figure 1 As shown, the AI ​​accelerator includes an on-chip storage module, a data transport module, and a computing array. The on-chip storage is used to cache external data and internal computational data. The computing array is used to process calculations such as convolution and vector operations in neural networks. The data transport module is used to organize the data before and after computation, ensuring that the data conforms to the data structure requirements of the neural network model. Images that need to be processed by the AI ​​accelerator are first stored in the on-chip storage module. Typically, the on-chip storage module is insufficient to store an entire image, requiring repeated readings of image data. Because of data overlap in these repeated readings, bandwidth is wasted. Therefore, how to propose an AI accelerator that can reduce bandwidth waste has become a crucial issue that urgently needs to be addressed in this field. Summary of the Invention

[0004] In view of the problems in the prior art, embodiments of this application provide an AI accelerator, processor, chip, board and electronic device, which can at least partially solve the problems existing in the prior art.

[0005] In a first aspect, this application proposes an AI accelerator, comprising a control module, an on-chip storage module, a instruction transfer queue module, a data transfer module, and a computing array, wherein:

[0006] The on-chip storage module is used to cache the region data, data blocks, intermediate data and output results of the image data.

[0007] The transport instruction queue module is used to control the data transport module to splice data blocks;

[0008] The data transfer module is used to stitch together the data blocks corresponding to the region data based on the region data of the image data and the previously stitched data block corresponding to the region data.

[0009] The computing array is used to perform calculations in the neural network based on data blocks to obtain the intermediate data and the output results.

[0010] The control module is used to control the operation of the computing array.

[0011] Furthermore, the on-chip storage module includes multiple storage blocks, with data blocks for splicing and data blocks for computation in the neural network stored in different storage blocks.

[0012] Furthermore, the transport instruction queue module is specifically used for:

[0013] After determining that the operating status of the data transport module and the data status of the data block corresponding to the current region data meet the splicing rules, a transport instruction is issued to the data transport module to splice the data block corresponding to the current region data; wherein, the transport instruction carries first data guidance information and second data guidance information of the data block corresponding to the current region data; the current region data refers to the region data of the latest cached image data of the on-chip storage module.

[0014] Furthermore, the data transfer module is specifically used for:

[0015] According to the first data guidance information, the current region data is obtained as the first part of the data block corresponding to the current region data, and according to the second data guidance information, the second part of the data block corresponding to the current region data is obtained from the previous spliced ​​data block corresponding to the current region data.

[0016] The first part of the data and the second part of the data are concatenated.

[0017] Furthermore, the splicing rules include the data transport module being in an idle state and the data status of the previous spliced ​​data block corresponding to the current region data being unaccessed.

[0018] Furthermore, the data transport module is also used to organize the data format of data blocks and intermediate data.

[0019] Furthermore, the image data includes multiple region data, and there is no overlap between the region data; the first region data corresponds to a complete data block, and the i-th region data is the data after removing duplicate data from the i-th data block and removing duplicate data from the (i-1)-th data block; the (i-1)-th data block is obtained based on the (i-1)-th region data; where i is a positive integer greater than or equal to 2, and i is less than or equal to the number of region data in the image data.

[0020] Secondly, this application proposes an artificial intelligence image signal processor, including at least one AI accelerator as described in any of the above embodiments.

[0021] Thirdly, this application proposes a chip including at least one artificial intelligence image signal processor as described in the above embodiments.

[0022] Fourthly, this application proposes a board that includes at least one chip as described in the above embodiments.

[0023] Fifthly, this application proposes an electronic device including at least one board as described in the above embodiments.

[0024] The AI ​​accelerator, processor, chip, board, and electronic device provided in this application include a control module, an on-chip storage module, a transfer instruction queue module, a data transfer module, and a computing array. The on-chip storage module is used to cache region data, data blocks, intermediate data, and output results of image data. The transfer instruction queue module is used to control the data transfer module to perform data block splicing. The data transfer module is used to splice data blocks corresponding to the region data based on the region data of the image data and the previously spliced ​​data block corresponding to the region data. The computing array is used to perform calculations in the neural network based on the data blocks to obtain the intermediate data and the output results. The control module is used to control the operation of the computing array. Since the import and export of duplicate data can be avoided when processing image data, the waste of bandwidth is reduced. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0026] Figure 1 This is a schematic diagram of the structure of an AI accelerator in existing technology.

[0027] Figure 2 This is a schematic diagram of the data processing flow of an AI accelerator provided in an embodiment of the present invention.

[0028] Figure 3A This is a schematic diagram of the processing of the first block of data provided in an embodiment of the present invention.

[0029] Figure 3B This is a schematic diagram of the processing of the second block of data provided in an embodiment of the present invention.

[0030] Figure 4 This is a schematic diagram of the structure of an AI accelerator provided in an embodiment of the present invention.

[0031] Figure 5 This is a schematic diagram of data processing for an AI accelerator provided in an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with relevant laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant means, and the acquisition, storage, use, and processing of user information have been authorized and agreed upon by the customer.

[0033] To facilitate understanding of the technical solution provided in this application, the relevant content of the technical solution in this application will be explained below.

[0034] Factors limiting the performance of AI accelerators include: (1) chip size; and (2) data bandwidth of external DDR memory and AI accelerator internal scratch pad memory (SPM).

[0035] Chip size is positively correlated with cost and manufacturing process. Tape-out cost largely determines the chip's scale and performance. Larger investments and more advanced processes result in better performance for AI accelerators. Providing sufficiently large on-chip cache and efficient bus data transfer capabilities allows AI accelerators to enjoy ample bandwidth for higher performance. Furthermore, leveraging data structures and the computational characteristics of deep learning networks to fully utilize data reuse and storage methods can reduce bandwidth consumption and improve AI accelerator performance.

[0036] This invention proposes a high-performance, low-bandwidth-consumption solution aimed at improving AI accelerator performance by reducing bandwidth consumption.

[0037] The process of AI accelerators processing image data, such as Figure 2 As shown, the AI ​​accelerator first reads image data from external memory, then reads weight data. Based on the image data and weight data, it performs calculations in neural networks such as convolution, activation, and pooling to obtain feature map results, and finally obtains the output result. From the moment the image data is read in, it is stored in the AI ​​accelerator's on-chip memory. Due to cost considerations, the on-chip memory of the AI ​​accelerator is usually insufficient to store the image data of an entire image. Therefore, each time image data is processed, the on-chip memory stores a portion of the entire image data for calculation.

[0038] like Figure 3A As shown, the image data within the blue dashed box, as the first piece of data segmented from the entire image data, enters the AI ​​accelerator for computation, yielding feature map results ① from neural network calculations such as convolution and pooling. Based on the characteristics of neural networks, feature map result ① continues to undergo convolution, pooling, and other neural network calculations to obtain feature value results ②, and so on, until the final image data processing result is obtained. Because convolution and pooling operations reduce the image height, starting from feature map result ①, the feature image height decreases with each round of calculation, resulting in the final result R0, which is then output. The receptive field of result R0 is defined as the initial image data within the blue dashed box, i.e. Figure 3A The image data is in the blue dashed box on the far left.

[0039] like Figure 3B As shown, the image data within the red dashed box represents the second segment of the entire image data. This second segment is fed into the AI ​​accelerator for computation, yielding feature map results ① from neural network calculations such as convolution and pooling. Based on the characteristics of neural networks, feature map result ① continues to undergo convolution and pooling calculations to obtain feature value results ②, and so on, until the final image data processing result R1 is calculated. Then, result R1 is output. The receptive field of result R1 is the initial image data within the red dashed box, i.e. Figure 3B The image data is shown in the leftmost red dashed box.

[0040] Comparing the first block of data (the image data within the blue dashed box) and the second block of data (the image data within the red dashed box) segmented from the entire image data reveals a significant overlap, with numerous duplicate data points. These duplicate data points are imported into the AI ​​accelerator twice: once when the first block enters the accelerator for computation, and again when the second block enters the accelerator for computation. Furthermore, the processed result R0 of the first block includes the processing results corresponding to the duplicate data, and the processed result R1 of the second block also includes the processing results corresponding to the duplicate data, indicating that the processing results for the duplicate data are exported twice.

[0041] Similarly, when the AI ​​accelerator processes the third, fourth, and fifth data sets, it will generate duplicate data, resulting in repeated data imports and exports. Furthermore, with a large number of convolutional layers, the overlapping receptive fields of the final result will gradually accumulate, leading to numerous instances of duplicate data imports and exports.

[0042] The aforementioned image data processing methods waste bandwidth between the DDR memory and the AI ​​accelerator's internal memory due to the import and export of duplicate data. Therefore, this application proposes an AI accelerator that avoids the import and export of duplicate data when processing image data, thereby reducing bandwidth waste.

[0043] Figure 4 This is a schematic diagram of the structure of an AI accelerator provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the AI ​​accelerator provided in this embodiment includes a control module 1, an on-chip storage module 2, a command transfer queue module 3, a data transfer module 4, and a computing array 5, wherein:

[0044] On-chip storage module 2 is used to cache region data, data blocks, intermediate data and output results of image data;

[0045] The data transfer instruction queue module 3 is used to control the data transfer module 4 to splice data blocks;

[0046] The data transfer module 4 is used to stitch together the data blocks corresponding to the region data based on the region data of the image data and the previously stitched data block corresponding to the region data.

[0047] The computation array 5 is used to perform computations in the neural network based on data blocks to obtain the intermediate data and the output results;

[0048] The control module 1 is used to control the operation of the computing array.

[0049] Specifically, the image data is divided into multiple regions. The first region corresponds to a complete data block, and the first data block can be obtained based on the first region. The second region is the data after removing duplicate data from the second data block. The third region is the data after removing duplicate data from the third data block. The fourth region is the data after removing duplicate data from the fourth data block. And so on, the nth region is the data after removing duplicate data from the nth data block.

[0050] The on-chip storage module 2 loads and caches the first region data of the image data from memory. The data transfer module 4 converts the first region data into a data format suitable for neural network processing, forming the first data block, which is then stored in the on-chip storage module 2. Under the control of the control module 1, the computing array 5 reads the first data block from the on-chip storage module 2 and performs calculations in the neural network until the final result corresponding to the first data block is obtained, and then stores the final result corresponding to the first data block in the on-chip storage module 2. When performing calculations in the neural network on the first data block, the calculations are performed along the neural network, with multiple rounds of calculations. After each round of calculation, the feature map result obtained becomes smaller and smaller in height. The calculations in the neural network include, but are not limited to, operations such as convolution, pooling, and activation, which are set according to actual needs and are not limited in this embodiment. The number of rounds of calculation is set according to actual needs and is not limited in this embodiment.

[0051] On-chip storage module 2 loads the second region data from DDR. This second region data is not a complete data block; it requires partial data from the first data block, which is then concatenated with the second region data to form the second data block. The data transfer instruction queue module 3 controls the data transfer module 4 to perform the concatenation of the second data block. According to the control instructions from the data transfer instruction queue module 3, the data transfer module 4 retrieves the second region data from the on-chip storage module 2 and the data required for concatenating the second data block from the first data block. It then concatenates the data and converts it into a data format suitable for neural network processing, forming the second data block, which is stored in the on-chip storage module 2. Under the control of the control module 1, the computing array 5 reads the second data block from the on-chip storage module 2 and performs calculations in the neural network until the final result corresponding to the second data block is obtained. The final result corresponding to the second data block is then stored in the on-chip storage module 2. During the neural network calculations on the second data block, the calculations are performed along the neural network, undergoing multiple rounds. After each round of calculation, the resulting feature map becomes progressively smaller in height. The first data block is the previous concatenated data block corresponding to the second region data.

[0052] On-chip storage module 2 loads the third region data from DDR. This third region data is not a complete data block; it requires partial data from the second data block, which is then concatenated with the third region data to form the third data block. The data transfer instruction queue module 3 controls the data transfer module 4 to perform the concatenation of the third data block. According to the control instructions from the data transfer instruction queue module 3, the data transfer module 4 retrieves the third region data from the on-chip storage module 2 and the data required for concatenating the third data block from the second data block. It then concatenates the data and converts it into a data format suitable for neural network processing, forming the third data block, which is stored in the on-chip storage module 2. Under the control of the control module 1, the computing array 5 reads the third data block from the on-chip storage module 2 and performs calculations in the neural network until the final result corresponding to the third data block is obtained. This final result is then stored in the on-chip storage module 2. During the neural network calculations on the third data block, the calculations are performed along the neural network, undergoing multiple rounds. After each round of calculation, the resulting feature map becomes progressively smaller in height. The second data block is the previous concatenated data block corresponding to the third region data.

[0053] On-chip storage module 2 loads the fourth region data from DDR. This fourth region data is not a complete data block; it requires partial data from the third data block, which is then concatenated with the fourth region data to form the fourth data block. The data transfer instruction queue module 3 controls the data transfer module 4 to concatenate the fourth data block. According to the control instructions from the data transfer instruction queue module 3, the data transfer module 4 retrieves the fourth region data from on-chip storage module 2 and the data required for concatenating the fourth data block from the third data block. It then concatenates the data and converts it into a data format suitable for neural network processing, forming the fourth data block, which is stored in on-chip storage module 2. Under the control of control module 1, computing array 5 reads the fourth data block from on-chip storage module 2 and performs calculations in the neural network until the final result corresponding to the fourth data block is obtained. This final result is then stored in on-chip storage module 2. During the neural network calculations on the fourth data block, the calculations are performed along the neural network, undergoing multiple rounds. After each round of calculation, the resulting feature map becomes progressively smaller in height. The third data block corresponds to the previous concatenated data block corresponding to the fourth region data.

[0054] Following this pattern, the final results corresponding to the fifth data block, the sixth data block, the seventh data block, the eighth data block, ..., the (n-1)th data block, and the nth data block can be obtained. Under the control of the control module 1, the computing array 5 reads the final results corresponding to the first data block, the second data block, the third data block, the fourth data block, ..., the (n-1)th data block, and the nth data block from the on-chip storage module 2, and combines these n final results to obtain the output result of the image data processed by the neural network.

[0055] The feature map results obtained after each round of calculation for each data block, and the final result corresponding to each data block, are used as intermediate data and can be cached in on-chip storage module 2.

[0056] The AI ​​accelerator provided in this application includes a control module, an on-chip storage module, a transfer instruction queue module, a data transfer module, and a computing array. The on-chip storage module is used to cache region data, data blocks, intermediate data, and output results of image data. The transfer instruction queue module is used to control the data transfer module to perform data block splicing. The data transfer module is used to splice data blocks corresponding to the region data based on the region data of the image data and the previously spliced ​​data block corresponding to the region data. The computing array is used to perform calculations in the neural network based on the data blocks to obtain the intermediate data and the output results. The control module is used to control the operation of the computing array. Since the import and export of duplicate data can be avoided when processing image data, the waste of bandwidth is reduced.

[0057] Based on the above embodiments, the on-chip storage module 2 further includes multiple storage blocks, with data blocks for splicing and data blocks for computation in the neural network stored in different storage blocks.

[0058] Specifically, since the on-chip storage module 2 is divided into multiple storage blocks, the data blocks used for splicing and the data blocks used for computation in the neural network can be stored in different storage blocks. That is, the data blocks that the data transport module 4 needs to read when splicing data blocks and the data blocks that the computation array 5 needs to read when performing computation in the neural network are stored in different storage partitions. Since the splicing of data blocks and the computation of data blocks are read from different storage partitions, they can be performed simultaneously, which can avoid data reading conflicts between data splicing and data computation and avoid performance degradation due to data splicing.

[0059] Based on the above embodiments, the transport instruction queue module 3 is further specifically used for:

[0060] After determining that the operating status of the data transfer module 4 and the data status of the previously completed data block corresponding to the current area data meet the splicing rules, a transfer instruction is issued to the data transfer module 4 to splice the data block corresponding to the current area data; the transfer instruction carries the first data guidance information and the second data guidance information of the data block corresponding to the current area data; the current area data refers to the area data of the latest cached image data of the on-chip storage module.

[0061] Specifically, if the data transfer instruction queue module 3 determines that the running status of the data transfer module 4 and the data status of the previously spliced ​​data block corresponding to the current region data satisfy the splicing rules, then it sends a transfer instruction to the data transfer module 4, causing the data transfer module 4 to splice the data block corresponding to the current region data based on the current region data and the previously spliced ​​data block corresponding to the current region data. The transfer instruction carries first and second data guidance information for the data block corresponding to the current region data. The first data guidance information instructs the data transfer module 4 to acquire the current region data, and the second data guidance information instructs the data transfer module 4 to acquire a portion of the data from the previously spliced ​​data block corresponding to the current region data. The splicing rules are preset and can be set according to actual needs; this embodiment does not limit them. The correspondence between the first guidance information and the current region data is predetermined, and the correspondence between the second guidance information and a portion of the data from the previously spliced ​​data block corresponding to the current region data is predetermined. The portion of the data from the previously spliced ​​data block corresponding to the current region data, i.e., the duplicate data between the data block corresponding to the current region data and the previously spliced ​​data block corresponding to the current region data, is predetermined.

[0062] The data transfer instruction queue module 3 can obtain the operating status of the data transfer module 4, as well as the data status of the last completed data block corresponding to the current region data. The operating status of the data transfer module 4 is divided into occupied and idle. Occupied indicates that the data transfer module 4 is working and cannot perform data block splicing for the current region data; idle indicates that the data transfer module 4 is not working and can perform data block splicing for the current region data. The data block status can be divided into accessed and unaccessed. Accessed indicates that the data block is being read; unaccessed indicates that the data block is idle and has not been read. The current region data refers to the region data of the latest cached image data in the on-chip storage module 2.

[0063] For example, the splicing rules include the data transport module 4 being in an idle state and the data status of the previous spliced ​​data block corresponding to the current area being unaccessed.

[0064] Based on the above embodiments, the data transfer module 4 is further specifically used for:

[0065] According to the first data guidance information, the current region data is obtained as the first part of the data block corresponding to the current region data, and according to the second data guidance information, the second part of the data block corresponding to the current region data is obtained from the previous spliced ​​data block corresponding to the current region data.

[0066] The first part of the data and the second part of the data are concatenated.

[0067] Specifically, the data transfer module 4 can obtain the current region data as the first part of the data block corresponding to the current region data based on the first data guidance information, and obtain the second part of the data block corresponding to the current region data from the previous completed data block corresponding to the current region data based on the second data guidance information. Then, it concatenates the first and second parts of the data to obtain the data block corresponding to the current region data. The first data guidance information can be a data identifier, uniquely corresponding to the current region data. The second guidance identifier can be a data identifier, corresponding to a portion of the data in the previous completed data block corresponding to the current region data.

[0068] For example, the second part of the data block corresponding to the current region data is appended to the first part of the data block corresponding to the current region data to form the data block corresponding to the current region data. Based on the above embodiments, the data transport module 4 is further used to organize the data format of the data blocks and intermediate data.

[0069] Specifically, the data transfer module 4 can convert the data format of the data block into a data format suitable for neural network processing, such as converting the data block into the HWC4 data format, and then storing it in the on-chip storage module 2.

[0070] The data transfer module 4 can organize the data format of intermediate data, such as converting intermediate data into HWC32 format and then storing it in the on-chip storage module 2.

[0071] Based on the above embodiments, the splicing rules further include that the data transport module is in an idle state and the data status of the previous spliced ​​data block corresponding to the current region data is unaccessed.

[0072] Specifically, when the data transport module 4 is idle and the data status of the previously assembled data block corresponding to the current area data is unaccessed, the transport instruction queue module 3 can issue a transport instruction to the data transport module 4 to assemble the data block corresponding to the current area data.

[0073] Based on the above embodiments, the image data further includes multiple region data, and there is no duplicate data between the region data; the first region data corresponds to a complete data block, and the i-th region data is the data after removing the duplicate data with the (i-1)-th data block from the i-th data block; the (i-1)-th data block is obtained based on the (i-1)-th region data; where i is a positive integer greater than or equal to 2, and i is less than or equal to the number of region data included in the image data.

[0074] Specifically, the image data is divided into multiple regions, with no overlap between regions and no overlap between adjacent regions. The first region corresponds to a complete data block, and the first data block can be obtained based on the first region. The second region is the data after removing duplicate data from the second data block; the third region is the data after removing duplicate data from the third data block; the fourth region is the data after removing duplicate data from the fourth data block; and so on, with the nth region being the data after removing duplicate data from the (n-1)th data block. In other words, the i-th region is the data after removing duplicate data from the (i-1)-th data block, where i is a positive integer greater than or equal to 2 and less than or equal to n, where n represents the number of regions included in the image data.

[0075] This application provides an Artificial Intelligence Image Signal Processor (AI-ISP), which includes at least one AI accelerator as described in any of the above embodiments.

[0076] This application provides a chip that includes at least one artificial intelligence image signal processor as described in the above embodiments.

[0077] This application provides a board card, which includes at least one chip as described in the above embodiments.

[0078] This application provides an electronic device, which includes the chip or board described in the above embodiments. The electronic device includes, but is not limited to, user equipment (UE), mobile devices, user terminals, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, and other devices with AI application requirements.

[0079] The following specific embodiment illustrates the data processing process of the AI ​​accelerator provided in this invention, such as... Figure 5 As shown.

[0080] Image data X is divided into m regions: region 1, region 2, region 3, ..., region m-1, region m. Region 1 corresponds to a complete data block, and there is no overlap between adjacent regions.

[0081] On-chip storage module 2 loads region data 1, which corresponds to a complete data block without the need for data concatenation. Data transfer module 4 formats region data 1, converting it into a format suitable for neural network processing, specifically HWC4 format data, which is then stored as the first data block (region data 1) in storage block C0. Under the control of control module 1, computing array 5 reads the first data block from storage block C0 and performs multi-round neural network calculations, obtaining feature map results ①, ②, and ③ corresponding to region data 1.

[0082] On-chip storage module 2 loads region data 2, which is a part of a data block and requires data concatenation. After determining that the operating status of data transfer module 4 and the data status of the data block corresponding to region data 2 meet the concatenation rules, transfer instruction queue module 3 issues a transfer instruction to data transfer module 4. The transfer instruction carries first and second data guidance information for the data block corresponding to region data 2. Based on the first data guidance information, data transfer module 4 obtains region data 2 as the first part of the data block corresponding to region data 2, and based on the second data guidance information, obtains the second part of the data block corresponding to region data 2 from the previously concatenated data block (i.e., the first data block). Then, it concatenates the first and second parts of the data block corresponding to region data 2 to obtain the second data block. The second data block is then stored in storage block C1. Under the control of the control module 1, the computing array 5 reads the second data block from the storage block C1 and performs calculations in the multi-round neural network to obtain the feature map result ①, feature map result ② and feature map result ③ corresponding to the region data 2.

[0083] On-chip storage module 2 loads region data 3, which is a part of a data block and requires data concatenation. After determining that the operating status of data transfer module 4 and the data status of the data block corresponding to region data 3 meet the concatenation rules, transfer instruction queue module 3 issues a transfer instruction to data transfer module 4. The transfer instruction carries first and second data guidance information for the data block corresponding to region data 3. Based on the first data guidance information, data transfer module 4 obtains region data 3 as the first part of the data block corresponding to region data 2, and based on the second data guidance information, obtains the second part of the data block corresponding to region data 2 from the previously concatenated data block (i.e., the second data block). Then, it concatenates the first and second parts of the data block corresponding to region data 3 to obtain the third data block. The second data block is then stored in storage block C1. Under the control of the control module 1, the computing array 5 reads the second data block from the storage block C1 and performs calculations in the multi-round neural network to obtain the feature map result ①, feature map result ② and feature map result ③ corresponding to the region data 3.

[0084] By analogy, we can obtain feature map results ①, ②, and ③ corresponding to m regions of data. Then, the final results corresponding to the m regions of data (i.e., feature map result ③) are combined to form the final result of image data X after processing by the neural network.

[0085] Understandably, due to the limited storage space of on-chip storage module 2, the AI ​​accelerator cannot process data from m regions simultaneously, and can therefore process data according to... Figure 5 The data in the central region is processed sequentially from top to bottom. The AI ​​accelerator controls the input speed of the region data in image data X while ensuring no data loss. The computational data blocks required for neural networks, along with related intermediate data, can be stored simultaneously in on-chip storage module 2. Completed data blocks can be overwritten by subsequent data blocks. The number of data blocks that the AI ​​accelerator can process simultaneously depends on the size of on-chip storage module 2 and the amount of intermediate data generated by the neural network computations.

[0086] In AI-ISP network computing, feature map sizes can reach the megabyte level. If images with larger height and width cannot be imported at once, the image computation overlap will increase many times over. Furthermore, in vision chips with AI-ISP capabilities, DDR bandwidth resources are extremely precious, needing to be supplied simultaneously to the ISP, video processing unit (VPU), and image sensor. Therefore, in AI-ISP application scenarios, this invention provides an AI accelerator, which can significantly reduce the bandwidth consumption of the AI ​​processor, thereby improving the performance of the AI-ISP. With limited area and bandwidth resources, the AI ​​processor achieves higher performance with less bandwidth.

[0087] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0088] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0089] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0090] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An AI accelerator, characterized in that, It includes a control module, an on-chip storage module, a transfer instruction queue module, a data transfer module, and a computing array, among which: The on-chip storage module is used to cache the region data, data blocks, intermediate data and output results of the image data. The transport instruction queue module is used to control the data transport module to splice data blocks; The data transfer module is used to stitch together the data blocks corresponding to the region data based on the region data of the image data and the previously stitched data block corresponding to the region data. The computing array is used to perform calculations in the neural network based on data blocks to obtain the intermediate data and the output results. The control module is used to control the operation of the computing array; Specifically, the transport instruction queue module is used for: After determining that the operating status of the data transport module and the data status of the data block corresponding to the current region data meet the splicing rules, a transport instruction is issued to the data transport module to splice the data block corresponding to the current region data; wherein, the transport instruction carries first data guidance information and second data guidance information of the data block corresponding to the current region data; the current region data refers to the region data of the latest cached image data of the on-chip storage module; Specifically, the data transfer module is used for: According to the first data guidance information, the current region data is obtained as the first part of the data block corresponding to the current region data, and according to the second data guidance information, the second part of the data block corresponding to the current region data is obtained from the previous spliced ​​data block corresponding to the current region data. The first part of the data and the second part of the data are concatenated.

2. The AI ​​accelerator according to claim 1, characterized in that, The on-chip storage module includes multiple storage blocks, with data blocks for splicing and data blocks for computation in neural networks stored in different storage blocks.

3. The AI ​​accelerator according to claim 1, characterized in that, The splicing rules include the data transport module being in an idle state and the data status of the previous spliced ​​data block corresponding to the current region being unaccessed.

4. The AI ​​accelerator according to claim 1, characterized in that, The data transfer module is also used to organize the data format of data blocks and intermediate data.

5. The AI ​​accelerator according to any one of claims 1 to 4, characterized in that, The image data includes multiple region data, and there is no overlap between the region data; the first region data corresponds to a complete data block, and the i-th region data is the data after removing duplicate data from the i-th data block and removing duplicate data from the (i-1)-th data block; the (i-1)-th data block is obtained based on the (i-1)-th region data; where i is a positive integer greater than or equal to 2, and i is less than or equal to the number of region data in the image data.

6. An artificial intelligence image signal processor, characterized in that, Includes at least one AI accelerator as described in any one of claims 1 to 5.

7. A chip, characterized in that, It includes at least one artificial intelligence image signal processor as described in claim 6.

8. A circuit board, characterized in that, It includes at least one chip as described in claim 7.

9. An electronic device, characterized in that, It includes at least one board as described in claim 8.

Citation Information

Patent Citations

  • Data transfer method, related product and computer storage medium

    CN109992542A

  • Configurable universal convolutional neural network accelerator

    WO2020258528A1