Data access method for chip and chip

By using the address rearrangement module in the chip to reorder the data in blocks to different memory banks, the bank conflict problem caused by reading multiple rows of data in the same clock cycle in the chip is solved, and data reading efficiency is improved.

CN119938549AActive Publication Date: 2025-05-06MOORE THREADS TECH CO LTD

Patent Information

Application Number
CN202411997631.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the chip, reading multiple rows of data in the same memory bank from the on-chip memory within the same clock cycle will lead to memory access conflicts, affecting data reading and processing efficiency.

Method used

By introducing a first address rearrangement module into the chip, the write addresses of the data units are rearranged, and the data within the same group is chunked and rearranged into different memory banks in the on-chip memory, thereby reducing bank conflicts.

Benefits of technology

It realizes data units that read multiple data blocks from on-chip memory within the same clock cycle, improves data reading efficiency and reduces the occurrence of bank conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938549A_ABST
    Figure CN119938549A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data access method for a chip and the chip, and the method comprises the steps: responding to a data loading instruction, and reading target data to be written into an on-chip memory from an off-chip memory; the target data comprises at least one group of data blocks, and each data block comprises at least one data unit; rearranging the write-in address of each data unit by using a first address rearrangement module to obtain a rearranged address of each data unit; wherein the data blocks in the same group correspond to different memory banks in the on-chip memory after being rearranged, and the rearrangement addresses of the data units in the data blocks are located in the memory banks corresponding to the data blocks; and writing the data unit into the on-chip memory according to the rearrangement address of the data unit. Thus, the data reading efficiency can be improved, meanwhile, bank conflicts can be reduced, the complexity of software processing logic in the chip application process can be reduced, the number of instructions can be reduced, and the universality of chip application can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular to a data access method for a chip and a chip. Background Art

[0002] In order to improve data processing efficiency in a chip, the data to be processed is usually loaded from an off-chip memory to an on-chip memory, and the required data can be quickly read from the on-chip memory during data processing. However, in the process of data processing, the chip in the related art may read multiple rows of data in the same memory bank from the on-chip memory in the same clock cycle, which will cause a memory bank access conflict (i.e., bank conflict), thereby affecting the data reading efficiency and subsequent data processing efficiency. Summary of the invention

[0003] In view of this, the embodiments of the present disclosure at least provide a data access method for a chip and a chip.

[0004] The technical solution of the embodiment of the present disclosure is implemented as follows:

[0005] The present disclosure provides a data access method for a chip, the method comprising:

[0006] In response to a data loading instruction, target data to be written into the on-chip memory is read from the off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit;

[0007] Using a first address rearrangement module, the write address of each data unit is rearranged to obtain a rearranged address of each data unit; wherein each data block in the same group corresponds to a different storage body in the on-chip memory after rearrangement, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block;

[0008] The data unit is written into the on-chip memory according to the rearranged address of the data unit.

[0009] The present disclosure provides a chip, the chip comprising:

[0010] A data storage engine and an on-chip memory, wherein the data storage engine comprises a first data reading module, a first address rearrangement module and a first data writing module; wherein:

[0011] The first data reading module is used to: in response to a data loading instruction, read the target data to be written into the on-chip memory from the off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit;

[0012] The first address rearrangement module is used to: rearrange the write address of each data unit to obtain the rearranged address of each data unit; wherein each data block in the same group corresponds to different storage bodies in the on-chip memory after rearrangement, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block;

[0013] The first data writing module is used to write the data unit into the on-chip memory according to the rearranged address of the data unit.

[0014] In the disclosed embodiment, in response to a data loading instruction, target data to be written to an on-chip memory is read from an off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit; the write address of each data unit is rearranged using a first address rearrangement module to obtain a rearranged address of each data unit; wherein, after rearrangement, each data block in the same group corresponds to a different storage body in the on-chip memory, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block; and the data unit is written to the on-chip memory according to the rearranged address of the data unit. In this way, on the one hand, the data blocks in the same group can be rearranged to different storage bodies in the on-chip memory for storage, so that in the same clock cycle, data units in multiple data blocks in the same group can be read from the on-chip memory, thereby increasing data reading efficiency and reducing the occurrence of bank conflicts; on the other hand, by utilizing the first address rearrangement module to rearrange the write address, the address rearrangement logic can be hardware-based on the chip, without the need to add additional address rearrangement logic at the software level, thereby reducing the complexity of the software processing logic during the chip application process, and reducing the number of instructions, thereby reducing the development burden of software developers, which is conducive to improving the breadth of chip applications.

[0015] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used to illustrate the technical solutions of the present disclosure together with the specification.

[0017] Figure 1A A schematic diagram of an implementation flow of a data access method for a chip provided in an embodiment of the present disclosure;

[0018] Figure 1B A schematic diagram of the composition structure of an on-chip memory provided by an embodiment of the present disclosure;

[0019] Figure 2 A schematic diagram of an implementation flow of a data access method for a chip provided in an embodiment of the present disclosure;

[0020] Figure 3A A schematic diagram of the composition structure of a chip provided in an embodiment of the present disclosure;

[0021] Figure 3B A schematic diagram of the composition structure of a chip provided in an embodiment of the present disclosure Figure 2 ;

[0022] Figure 4A Schematic diagram 3 of the composition structure of a chip provided in an embodiment of the present disclosure;

[0023] Figure 4B A schematic diagram of storing the A matrix in the local memory in the related art;

[0024] Figure 4C A schematic diagram of the arrangement of data blocks in a local memory after address rearrangement in a matrix operation scenario of a chip provided by an embodiment of the present disclosure;

[0025] Figure 4D A schematic diagram of arrangement of data blocks in a local memory after address rearrangement when the data type of matrix elements in a chip is FP16 provided in an embodiment of the present disclosure;

[0026] Figure 4E A schematic diagram of arrangement of data blocks in a local memory after address rearrangement when the data type of a matrix element in a chip is TF16 provided in an embodiment of the present disclosure;

[0027] Figure 4F Schematic diagram 1 of the currently available row size of a local memory in a chip provided by an embodiment of the present disclosure;

[0028] Figure 4G A schematic diagram of the currently available row size of a local memory in a chip provided by an embodiment of the present disclosure Figure 2 ;

[0029] Figure 4H A schematic diagram of arrangement of data blocks in a local memory after address re-arrangement without introducing an address interval step in a convolution operation scenario provided by a chip according to an embodiment of the present disclosure;

[0030] Fig. 4I A schematic diagram of the arrangement of data blocks in a local memory after a chip introduces an address interval step to perform address re-arrangement in a convolution operation scenario in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.

[0032] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure belongs. The terms used herein are only for the purpose of describing the present disclosure and are not intended to limit the present disclosure.

[0035] In order to better understand the solution of the embodiment of the present disclosure, the data access method of the chip in the related art is first described below.

[0036] In a chip in the related art, if multiple rows of data need to be read from the same storage bank in an on-chip memory within one clock cycle, a bank conflict will occur.

[0037] In some related technologies, in the process of loading the data to be processed from the off-chip memory to the on-chip memory, only the minimum size of data supported by the chip in a single clock cycle is read each time, so that the data to be read in a single clock cycle is stored in the same row of storage units in the on-chip memory. In this way, in the process of the chip reading data from the on-chip memory for processing, only the same row of storage units in the on-chip memory need to be accessed in a single clock cycle, so that no bank conflict will occur. However, this solution will reduce the data reuse rate and the memory access bandwidth for reading data during data processing due to the small amount of data loaded from the off-chip memory each time, and will increase the number of accesses to the off-chip memory, affecting the overall efficiency of data processing.

[0038] In view of this, an embodiment of the present disclosure provides a data access method for a chip. The method can be executed by a chip. In some embodiments, the chip can be any suitable processor chip. For example, the processor chip can include but is not limited to at least one of a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a data processing unit (DPU), etc. Figure 1A A schematic diagram of an implementation flow of a data access method for a chip provided in an embodiment of the present disclosure is shown in FIG. Figure 1A As shown, the method may include the following steps S101 to S103:

[0039] Step S101, in response to a data loading instruction, read target data to be written into an on-chip memory from an off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit.

[0040] Here, the off-chip memory may include but is not limited to at least one of the main memory, global memory, etc. The on-chip memory may include but is not limited to at least one of the local memory in the chip, the register file, etc.

[0041] The data loading instruction is used to instruct to read the target data from the off-chip memory and write the read target data to the on-chip memory. The target data can be any suitable data to be processed by the chip, and the embodiments of the present disclosure are not limited to this. In some embodiments, the data loading instruction can be sent to the chip by the main processor. According to the data loading instruction, the storage address of each data unit in the target data in the off-chip memory and the write address corresponding to each data unit in the on-chip memory can be determined.

[0042] It can be understood that the data volume of the target data, the number of groups of data blocks in the target data, the number of data blocks included in each group of data blocks, and the number of data units included in each data block can all be determined based on actual data processing requirements.

[0043] In some implementations, the amount of target data may be determined based on the available storage capacity of the on-chip memory. For example, the amount of target data may be equal to the available storage capacity of the on-chip memory, or may be less than the available storage capacity of the on-chip memory.

[0044] In some embodiments, a group of data blocks may be data that the chip needs to read from the on-chip memory in a single clock cycle for processing. For example, in the case where the chip is used to perform a matrix product operation A×B, the target data may include matrix data of the matrix A / B, a group of target data blocks may include a matrix slice in the matrix A / B, a data unit may include an element in the matrix A / B, and the chip needs to read a matrix slice of the matrix A / B from the on-chip memory in a single clock cycle to perform the matrix operation. For another example, in the case where the chip is used to perform a convolution operation, the target data may include image data, a group of data blocks may include an image slice in the image data, a data unit may include pixel data in the image slice, and the chip needs to read an image slice in the image data from the on-chip memory in a single clock cycle to perform the matrix operation.

[0045] Step S102, using the first address rearrangement module, rearrange the write address of each of the data units to obtain the rearranged address of each of the data units; wherein, after rearrangement, each of the data blocks in the same group corresponds to different storage bodies in the on-chip memory, and the rearranged address of each of the data units in the data block is located in the storage body corresponding to the data block.

[0046] Here, the chip may include a first address reordering module. When implemented, the first address reordering module may use any appropriate address reordering logic to reorder the write addresses of each data unit so that the reordered addresses of the data units in different data blocks in the same group fall into different storage bodies in the on-chip memory respectively, and the embodiments of the present disclosure are not limited. For example, the first address reordering module may use at least one of a reordering logic based on a bitwise exclusive OR (XOR) operation, a reordering logic based on a sequence offset, etc. to reorder the write addresses of each data unit.

[0047] In some implementations, the address offset corresponding to each write address may be predetermined, and the first address rearrangement module may offset the write address of each data unit according to the corresponding address offset to obtain the rearrangement address of each data unit.

[0048] In some implementations, the correspondence between each write address and the corresponding rearrangement address may be predetermined, and the first address rearrangement module may query the correspondence according to the write address of each data unit to obtain the rearrangement address of each data unit.

[0049] Step S103: writing the data unit into the on-chip memory according to the rearranged address of the data unit.

[0050] In the disclosed embodiment, in response to a data loading instruction, target data to be written to an on-chip memory is read from an off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit; the write address of each data unit is rearranged using a first address rearrangement module to obtain a rearranged address of each data unit; wherein, after rearrangement, each data block in the same group corresponds to a different storage body in the on-chip memory, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block; and the data unit is written to the on-chip memory according to the rearranged address of the data unit. In this way, on the one hand, the data blocks in the same group can be rearranged to different storage bodies in the on-chip memory for storage, so that in the same clock cycle, data units in multiple data blocks in the same group can be read from the on-chip memory, thereby increasing data reading efficiency and reducing the occurrence of bank conflicts; on the other hand, by utilizing the first address rearrangement module to rearrange the write address, the address rearrangement logic can be hardware-based on the chip, without the need to add additional address rearrangement logic at the software level, thereby reducing the complexity of the software processing logic during the chip application process, and reducing the number of instructions, thereby reducing the development burden of software developers, which is conducive to improving the breadth of chip applications.

[0051] In some embodiments, the above step S102 may include the following step S111:

[0052] Step S111 : using a first address rearrangement module to convert the write address of the data unit into a rearrangement address of the data unit based on an address conversion parameter.

[0053] Here, the address conversion parameters may include but are not limited to at least one of the row size of the on-chip memory, the block size of the data block, the address interval step between two adjacent data blocks in the same group of data blocks, and the like.

[0054] It is understandable that different address conversion parameters can be set in different application scenarios, so that the first address rearrangement module can convert the write address of the data unit according to the address conversion parameters and adopt appropriate address conversion logic. In this way, the flexibility of address rearrangement can be improved to better adapt to different application scenarios and improve the versatility of the first address rearrangement module.

[0055] In some embodiments, Figure 1BAs shown, the on-chip memory 100 includes a plurality of storage blocks 110 arranged in an array, the storage block 110 includes at least one storage unit 111 with a continuous address, at least one storage unit 111 in the storage block 110 is used to store at least one data unit in the data block, and each storage block 110 in the same column belongs to the same storage body (e.g., bank0, bank1, ...). The address conversion parameter includes the row size of a single row of storage blocks in the on-chip memory.

[0056] The above-mentioned step S111, based on the address conversion parameter, converting the write address of the data unit into the rearrangement address of the data unit, may include the following steps S121 to S123:

[0057] Step S121 : determining a first row identifier and a first intra-row address offset corresponding to the write address based on the write address of the data unit and the row size.

[0058] Here, the row size of a single row storage block in the on-chip memory is the current available row size of the on-chip memory. For example, the current available row size of the on-chip memory may include but is not limited to 128 bytes or 256 bytes.

[0059] In some embodiments, the currently available row size of the on-chip memory may be determined based on the maximum access bandwidth provided by the on-chip memory. For example, in the case where the on-chip memory is a local memory, the computing engine in the chip may directly read data from the local memory for calculation, and the maximum access bandwidth provided by the local memory for the computing engine is 256 Bytes, then the row size of the local memory may be 256 Bytes. For another example, in the case where the on-chip memory is a local memory, the chip may read from the local memory to the register stack through the local memory access pipeline, and the computing engine may read data from the register stack for calculation, and the maximum access bandwidth provided by the local memory for the local memory access pipeline is 128 Bytes, then the currently available row size of the on-chip memory may be 128 Bytes.

[0060] In some implementations, the write address of the data unit may be divided by the row size to obtain a first row identifier corresponding to the write address; and the write address modulo the row size may be taken to obtain a first intra-row address offset corresponding to the write address.

[0061] For example, the write address of the data unit is lms_addr, and the row size of the single row storage block in the on-chip memory is line_size, then the first row identifier line_id corresponding to the write address is lms_addr / line_size, and the first intra-row address offset line_offset1 corresponding to the write address is lms_addr%line_size.

[0062] In some implementations, the write address of the data unit may be divided by an integer multiple of the row size to obtain a first row identifier corresponding to the write address; the write address may be modulo an integer multiple of the row size to obtain a first intra-row address offset corresponding to the write address.

[0063] Step S122: determining a second intra-row address offset based on the first row identifier and the first intra-row address offset.

[0064] The second intra-row address offset refers to the intra-row address offset of the data unit after the address is rearranged. In implementation, the second intra-row address offset corresponding to the data unit can be determined based on the first row identifier and the first intra-row address offset corresponding to the write address of the data unit according to the address rearrangement logic actually adopted, so that the second intra-row address offsets corresponding to the data units in different data blocks in the same group fall into different storage bodies in the on-chip memory respectively.

[0065] In some embodiments, each row identifier may correspond to a different address offset step. By offsetting the address offset within the first row according to the address offset step corresponding to the first row identifier, the address offset within the second row may be obtained, wherein the address offset step may be an integer multiple of the block size of the data block.

[0066] Step S123: determining a reordering address of the data unit based on the first row identifier and the second intra-row address offset.

[0067] In some implementations, based on the first row identifier corresponding to the data unit, the row start address of the row where the data unit is located can be determined, and based on the row start address and the second intra-row address offset, the reordered address of the data unit can be obtained.

[0068] For example, the product line_id×line_size of the first row identifier line_id corresponding to the data unit and the row size line_size of a single row storage block in the on-chip memory can be determined as the row starting address of the row where the data unit is located; when the address offset in the second row corresponding to the data unit is line_offset2, the reordering address of the data unit can be determined as line_id×line_size+line_offset2.

[0069] In some implementations, based on the first row identifier corresponding to the data unit, the row start address of the row where the data unit is located can be determined, and based on the row start address, the reserved size within the row, and the second address offset within the row, the reordering address of the data unit can be obtained. The reserved size within the row can be the size of the address space reserved in a single row storage block for storing a specific flag field or attribute field.

[0070] For example, the product line_id×line_size of the first row identifier line_id corresponding to the data unit and the row size line_size of a single row storage block in the on-chip memory can be determined as the row starting address of the row where the data unit is located; the reserved size line_offset0 within the row where the data unit is located; when the second row address offset corresponding to the data unit is line_offset2, the reordering address of the data unit can be determined as line_id×line_size+line_offset0+line_offset2.

[0071] In the above embodiment, based on the write address of the data unit and the row size of the single row storage block in the on-chip memory, the first row identifier and the first row address offset corresponding to the write address are determined; based on the first row identifier and the first row address offset, the second row address offset is determined; based on the first row identifier and the second row address offset, the data unit reordering address is determined. In this way, the write address of the data unit can be quickly and accurately converted into the reordering address according to the row size of the single row storage block in the on-chip memory, so that the first address reordering module can adapt to the on-chip memory with different row sizes.

[0072] In some embodiments, the address conversion parameters also include the block size of the data block. The above step S122 may include the following steps S131 to S133:

[0073] Step S131 : determining a first block identifier and a first intra-block address offset corresponding to the write address based on the first intra-row address offset and the block size.

[0074] Here, the block size of a data block refers to the amount of data contained in the data block, that is, the sum of the data amounts of each data unit in the data block.

[0075] In some implementations, the data size of the data unit may be determined according to the data type of the data unit, and the block size of the data block may be determined according to the data size and the number of data units included in the data block. For example, if the data type of the data unit is FP16 and the number of data units included in a single data block is 8, the data size of the single data unit is 2 bytes, and the block size of the data block is 16 bytes.

[0076] In some implementations, the first intra-row address offset corresponding to the write address may be divided by the block size to obtain the first block identifier corresponding to the write address; the first intra-row address offset corresponding to the write address may be modulo the block size to obtain the first intra-block address offset corresponding to the write address.

[0077] For example, the first intra-line address offset corresponding to the write address lms_addr is line_offset1, and the block size of the data block is sector_sz, then the first block identifier sector1_in_line corresponding to the write address is line_offset1 / sector_sz, and the first intra-block address offset offset_in_sector corresponding to the write address is line_offset1%sector_sz.

[0078] In some embodiments, the first in-row address offset corresponding to the write address can be divided by an integer multiple of the block size to obtain the first block identifier corresponding to the write address; the first in-row address offset corresponding to the write address can be modulo the integer multiple of the block size to obtain the first in-block address offset corresponding to the write address.

[0079] Step S132: determining a second block identifier based on the first row identifier and the first block identifier.

[0080] The second block identifier refers to the block identifier of the storage block corresponding to the data unit after the address is rearranged. In implementation, the second block identifier corresponding to the data unit can be determined based on the first row identifier and the first block identifier corresponding to the write address of the data unit according to the address rearrangement logic actually adopted, so that the second block identifiers corresponding to the data units in different data blocks in the same group fall into different storage bodies in the on-chip memory respectively.

[0081] In some implementations, each row identifier may correspond to a different block offset, and the first block identifier may be offset according to the block offset corresponding to the first row identifier to obtain the second block identifier.

[0082] In some implementations, each combination of a row identifier and a block identifier may correspond to a different block offset, and the first block identifier may be offset according to the block offset corresponding to the first row identifier and the first block identifier to obtain the second block identifier.

[0083] Step S133: determining the second intra-row address offset based on the second block identifier and the first intra-block address offset.

[0084] In some implementations, based on the second block identifier corresponding to the data unit, the starting address of the data block where the data unit is located can be determined, and based on the starting address of the data block and the address offset within the first block, the second row address offset corresponding to the data unit can be obtained.

[0085] For example, the second block identifier sector2_in_line corresponding to the data unit and the block size sector_sz of the data block can be summed up to obtain the second row address offset line_offset2 corresponding to the data unit.

[0086] In some implementations, based on the second block identifier corresponding to the data unit, the starting address of the data block where the data unit is located can be determined, and based on the starting address of the data block, the reserved size within the block, and the first address offset within the block, the second intra-row address offset corresponding to the data unit can be obtained. The reserved size within the block can be the size of the address space reserved in the storage block for storing a specific flag field or attribute field.

[0087] In the above embodiment, based on the first intra-row address offset and the block size, the first block identifier and the first intra-block address offset corresponding to the write address are determined; based on the first row identifier and the first block identifier, the second block identifier is determined; based on the second block identifier and the first intra-block address offset, the second intra-row address offset is determined. In this way, the second intra-row address offset corresponding to each data unit can be determined quickly and accurately according to the block size of the data block, so that the write address of the data unit is converted into a rearrangement address, so that the first address rearrangement module can adapt to the address rearrangement of data units of different data types.

[0088] In some embodiments, the address conversion parameter also includes the address interval step between two adjacent data blocks in the same group of data blocks. The above step S132 may include the following steps S141 to S143:

[0089] Step S141, determining a row interval step for address rearrangement based on the address interval step and the block size.

[0090] Here, the address interval step length between two adjacent data blocks in the same group of data blocks can be determined according to the distance between two adjacent data blocks that need to be read within a single clock cycle in an actual data processing scenario. For example, the address interval step length can include 256Byte, 128Byte, 64Byte or 32Byte, etc.

[0091] In some embodiments, the address interval step between two adjacent data blocks in the same group of data blocks can be determined according to the current calculation mode of the chip. For example, when the current calculation mode of the chip is matrix multiplication operation, each data block in the same group corresponds to multiple storage blocks in the same storage body, so the address interval step between two adjacent data blocks in the same group of data blocks can be equal to the row size of a single row storage block in the on-chip memory. For another example, when the current calculation mode of the chip is convolution operation, the target data includes each pixel data of the image, and the write address of each pixel data is continuous. Assuming that the data size of each pixel data in the image is 32Byte, each pixel data includes two parts of channel data, high 16Byte and low 16Byte, and the high 16Byte data and the low 16Byte data are respectively used as a data block. Then, when the low 16Byte data in a group of pixels needs to be read in a single clock cycle, the address interval step between two adjacent data blocks in the same group of data blocks can be equal to the data size of each pixel data, that is, 32Byte. In this way, the first address rearrangement module can adapt to different computing modes.

[0092] The row spacing step of address rearrangement refers to the number of rows included in the row range for address rearrangement, that is, the address rearrangement can be performed in a single address rearrangement cycle using the row spacing step as the address rearrangement cycle.

[0093] In some implementations, the address spacing step may be divided by the block size to obtain the row spacing step of the address rearrangement. For example, when the address spacing step is stride and the block size is sector_sz, the row spacing step of the address rearrangement may be stride / sector_sz.

[0094] In some implementations, the address interval step may be divided by an integer multiple of the block size to obtain the row interval step of the address rearrangement.

[0095] Step S142: determining a second row identifier based on the first row identifier and the row interval step.

[0096] Here, the second row identifier represents a row identifier of an address row corresponding to a write address of a data unit within a row interval step.

[0097] In some implementations, the first line identifier may be modulo the line spacing step to obtain the second line identifier. For example, when the first line identifier is line_id and the line spacing step of the address rearrangement is swizzle_group, the second line identifier swizzle_line_id may be line_id%swizzle_group.

[0098] In some implementations, the first row identifier may be modulo an integer multiple of the row spacing step length to obtain the second row identifier.

[0099] Step S143: determine the second block identifier based on the second row identifier and the first block identifier.

[0100] In some implementations, each row identifier within the row interval step may correspond to a different block offset, and the first block identifier may be offset according to the block offset corresponding to the second row identifier to obtain the second block identifier.

[0101] In some implementations, the second row identifier and the first block identifier may be bitwise XORed to obtain the second block identifier.

[0102] In the above embodiment, the row spacing step of the address rearrangement is determined based on the address spacing step and the block size; the second row identifier is determined based on the first row identifier and the row spacing step; and the second block identifier is determined based on the second row identifier and the first block identifier. In this way, the second block identifier corresponding to each data unit can be quickly and accurately determined based on the address spacing step between two adjacent data blocks in the same group of data blocks, so that the write address of the data unit is converted into a rearrangement address, so that the first address rearrangement module can adapt to different data processing scenarios.

[0103] In some embodiments, the above data access method may further include the following step S151:

[0104] Step S151, based on the data loading instruction, determine the address conversion parameters; the address conversion parameters include at least one of the following: the row size of a single row storage block in the on-chip memory, the block size of the data block, and the address interval step between two adjacent data blocks in the same group of data blocks.

[0105] In some implementations, the data loading instruction may directly carry the address conversion parameter, and the chip may directly obtain the address conversion parameter from the data loading instruction.

[0106] In some implementations, the data loading instruction may carry indication information for indicating the address conversion parameters, and the address conversion parameters may be obtained by parsing the indication information in the data loading instruction.

[0107] In the above embodiment, by configuring the address conversion parameters through the data loading instruction, the flexibility of address rearrangement can be further improved, and the software logic in the data processing process can be simplified to reduce the number of instructions.

[0108] The embodiment of the present disclosure provides a data access method for a chip. The method can be executed by the chip. Figure 2 A schematic diagram of an implementation flow of a data access method for a chip provided in an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the method may include the following steps S201 to S207:

[0109] Step S201, in response to a data loading instruction, read target data to be written into an on-chip memory from an off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit.

[0110] Step S202, using the first address rearrangement module, rearrange the write address of each of the data units to obtain the rearranged address of each of the data units; wherein, after rearrangement, each of the data blocks in the same group corresponds to different storage bodies in the on-chip memory, and the rearranged address of each of the data units in the data block is located in the storage body corresponding to the data block.

[0111] Step S203: writing the data unit into the on-chip memory according to the rearranged address of the data unit.

[0112] Here, the above steps S201 to S203 correspond to steps S101 to S103 in the above embodiments respectively, and the implementation of the above steps S101 to S103 may be referred to during implementation.

[0113] Step S204, obtaining a data read instruction, wherein the data read instruction is used to instruct to read a group of target data blocks from the on-chip memory, wherein the target data blocks include at least one target data unit.

[0114] Step S205 , determining a read address of each of the target data units based on the data read instruction.

[0115] Step S206: using a second address rearrangement module, rearrange the read addresses of the target data units to obtain rearranged addresses of the target data units.

[0116] Step S207 , reading each of the target data units in parallel from the on-chip memory according to the rearrangement address of each of the target data units.

[0117] The data read instruction may be sent to the chip by the main processor, or may be generated by a computing engine in the chip. The target data block may be a data block currently to be processed by the chip.

[0118] The second address rearrangement module may be the same address rearrangement module as the first address rearrangement module, or may be another address rearrangement module having the same rearrangement logic as the first address rearrangement module.

[0119] The process of rearranging the read address of the target data unit in the above step S206 corresponds to the process of rearranging the write address of the data unit in the above step S102. When implemented, the read address of the target data unit can be rearranged with reference to the implementation method of the above step S102.

[0120] In the disclosed embodiment, a data read instruction is obtained, the data read instruction is used to instruct to read a group of target data blocks from the on-chip memory, the target data blocks including at least one target data unit; based on the data read instruction, the read address of each target data unit is determined; the read address of each target data unit is rearranged by using the second address rearrangement module to obtain the rearranged address of each target data unit; and each target data unit is read in parallel from the on-chip memory according to the rearranged address of each target data unit. In this way, since each target data block in the same group corresponds to different storage bodies in the on-chip memory after rearrangement, the rearranged address of each target data unit in the target data block is located in the storage body corresponding to the data block, so that in the same clock cycle, the target data units in multiple target data blocks in the same group can be read from the on-chip memory, so that the occurrence of bank conflicts can be reduced while increasing the data reading efficiency.

[0121] In some embodiments, the above method may further include the following step S211:

[0122] Step S211, determining address conversion parameters based on the data read instruction; the address conversion parameters include at least one of the following: the row size of a single row storage block in the on-chip memory, the block size of the data block, and the address interval step between two adjacent data blocks in the same group of data blocks.

[0123] The above step S206 may include the following step S212:

[0124] Step S212: using a second address rearrangement module, based on the address conversion parameter, rearrange the read address of each of the target data units to obtain a rearranged address of each of the target data units.

[0125] In some implementations, the data read instruction may directly carry the address conversion parameter, and the chip may directly obtain the address conversion parameter from the data read instruction.

[0126] In some implementations, the data read instruction may carry indication information for indicating the address conversion parameters, and the address conversion parameters may be obtained by parsing the indication information in the data read instruction.

[0127] In different application scenarios, the second address rearrangement module can convert the read address of the target data unit according to the corresponding address conversion parameters and adopt appropriate address conversion logic. In this way, the flexibility of address rearrangement can be improved to better adapt to different application scenarios and improve the versatility of the second address rearrangement module. In addition, by configuring the address conversion parameters through the data read instruction, the flexibility of address rearrangement can be further improved, and the software logic in the data processing process can be simplified to reduce the number of instructions.

[0128] In some embodiments, the target data includes matrix data, a group of the target data blocks includes a target matrix slice in the matrix data, and the target data unit includes a matrix element in the target matrix slice. The above method further includes the following step S221:

[0129] Step S221: input each of the matrix elements in the read target matrix slice into a computing unit array to perform a matrix product operation.

[0130] The chip may include a computing unit array for performing data operations. The target matrix slice may be a matrix slice that the computing unit array needs to read and perform matrix product operations on in a single clock cycle.

[0131] It should be noted that, in the process of performing the matrix product operation A×B of matrix A and matrix B, the target matrix slice can be matrix A or matrix B, and the embodiment of the present disclosure does not limit this. Among them, both matrix A and matrix B can be read from the off-chip memory and stored in the on-chip memory using the data access method in the embodiment of the present disclosure, and read from the on-chip memory and output to the computing unit array to complete the matrix product operation.

[0132] In some embodiments, the target data includes image data, a group of the target data blocks includes a target image slice in the image data, and the target data block includes channel data of a single pixel in the target image slice; the above method further includes the following step S231:

[0133] Step S231, inputting the read channel data of each pixel in the target image slice into the computing unit array to perform a convolution operation.

[0134] In some implementations, the channel data of each pixel in the target image slice may correspond to a data block respectively.

[0135] The present disclosure provides a chip. Figure 3A A schematic diagram of the composition structure of a chip provided in an embodiment of the present disclosure is shown in FIG. Figure 3A As shown, the chip 300 includes a data storage engine 310 and an on-chip memory 320, and the data storage engine 310 includes a first data reading module 311, a first address rearrangement module 312 and a first data writing module 313; wherein:

[0136] The first data reading module 311 is used to: in response to a data loading instruction, read the target data to be written into the on-chip memory 320 from the off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit;

[0137] The first address rearrangement module 312 is used to rearrange the write addresses of each data unit to obtain the rearranged addresses of each data unit; wherein each data block in the same group corresponds to different storage bodies in the on-chip memory 320 after rearrangement, and the rearranged addresses of each data unit in the data block are located in the storage body corresponding to the data block;

[0138] The first data writing module 313 is used to write the data unit into the on-chip memory 320 according to the rearranged address of the data unit.

[0139] In some embodiments, the data storage engine may include a tensor storage engine (TME) in a GPGPU chip.

[0140] In the disclosed embodiment, by setting a first address reordering module in the data storage engine of the chip, on the one hand, each data block in the same group can be reordered to different storage bodies in the on-chip memory for storage, so that in the same clock cycle, data units in multiple data blocks in the same group can be read from the on-chip memory, thereby increasing data reading efficiency and reducing the occurrence of bank conflicts; on the other hand, by using the first address reordering module to realize the reordering of write addresses, the address reordering logic can be hardware-based in the chip, without adding additional address reordering logic at the software level, thereby reducing the complexity of software processing logic in the chip application process, and reducing the number of instructions, thereby reducing the development burden of software developers, which is conducive to improving the wide application of the chip.

[0141] In some embodiments, the first address rearrangement module is further used to: convert the write address of the data unit into a rearrangement address of the data unit based on an address conversion parameter.

[0142] In some embodiments, the data storage engine further includes:

[0143] A first parameter determination module is used to: determine the address conversion parameters based on the data loading instruction, and send the address conversion parameters to the first address rearrangement module; the address conversion parameters include at least one of the following: the row size of a single row storage block in the on-chip memory, the block size of the data block, and the address interval step between two adjacent data blocks in the same group of data blocks.

[0144] In some embodiments, Figure 3B As shown, the chip 300 further includes a computing engine 330, the computing engine 330 includes an address determination module 331, a second address rearrangement module 332, a second data reading module 333 and a computing unit array 334, the computing unit array 334 includes a plurality of computing units; wherein:

[0145] The address determination module 331 is used to: obtain a data read instruction, wherein the data read instruction is used to instruct to read a group of target data blocks from the on-chip memory 320, wherein the target data blocks include at least one target data unit; and determine a read address of each of the target data units based on the data read instruction;

[0146] A second address rearrangement module 332, configured to rearrange the read addresses of the target data units to obtain rearranged addresses of the target data units;

[0147] The second data reading module 333 is used to read each target data unit in parallel from the on-chip memory 320 according to the rearrangement address of each target data unit, and input each target data unit in the read set of target data blocks into the computing unit array 334.

[0148] In some embodiments, the computing engine further comprises:

[0149] A second parameter determination module is used to: determine an address conversion parameter based on the data read instruction; the address conversion parameter includes at least one of the following: a row size of a single row storage block in the on-chip memory, a block size of the data block, and an address interval step between two adjacent data blocks in the same group of the data blocks;

[0150] The second address rearrangement module is further used to rearrange the read addresses of each of the target data units based on the address conversion parameters to obtain the rearranged addresses of each of the target data units.

[0151] In some embodiments, when the target data includes matrix data, a group of the target data blocks includes a target matrix slice in the matrix data, the target data unit includes a matrix element in the target matrix slice, and the calculation unit array is used to perform a matrix product operation based on each of the matrix elements in the read target matrix slice;

[0152] In the case where the target data includes image data, a group of the target data blocks includes target image slices in the image data, and the target data blocks include channel data of a single pixel in the target image slice; the computing unit array is used to perform convolution operations based on the channel data of each pixel in the read target image slice.

[0153] The following describes the application of the data access method for a chip provided by an embodiment of the present disclosure in a practical scenario.

[0154] In the related art, artificial intelligence (AI) accelerators generally use ATOMIC_C×ATOMIC_K computing unit configuration. Each time an image is read, only one pixel data needs to be read. Each pixel data contains ATOMIC_C channel data. The storage unit is generally designed specifically, and there is generally no bank conflict. GPGPU needs to support general computing, and in order to accelerate matrix multiplication and convolution operations, tensor cores are added. If the data source of the tensor core is a register, it is necessary to reasonably map the data and the register to avoid register access conflicts. Although the tensor core can read data from the register, it is also necessary to move the data from the local memory to the register, and the arrangement of matrix data in the global memory is linearly continuous. If the data of multiple rows of storage units in the same bank is read at the same time, a bank conflict will occur. One solution is to only read the minimum size of data supported by the tensor core when moving data from the global memory to the local memory, so that the data will be continuously stored in the same row of storage units in the local memory, which can avoid bank conflicts, but it will lead to reduced data reuse and increased access to the global memory.

[0155] The present disclosure provides a chip, such as Figure 4AAs shown, the chip 400 includes a tensor storage engine 410, a local memory 420 and a tensor computing engine 430. A hardware address reordering module (i.e., a first address reordering module) 411 is provided in the tensor storage engine 410. The first address reordering module 411 can realize the address reordering of data from the global memory to the local memory 420. At the same time, a second address reordering module 431 which is the same as the first address reordering module 411 is provided in the tensor computing engine 430. The tensor storage engine 410 reads the target data from the global memory, and before writing the target data into the local memory 420, the first address reordering module 411 is used to perform address reordering transformation on the original write address of each data unit in the target data. The transformation uses the row in the local memory 420 as the boundary, and reorders the write address of each data unit in the same row, so that the write address of the data unit in the data block corresponding to different rows in the same bank will be staggered and fall in different banks after transformation. When the tensor computing engine 430 reads data from the local memory 420, the same clock cycle needs to read multiple rows of data blocks in the same bank. At this time, the second address rearrangement module 431, which is the same as the first address rearrangement module 411, can be used to rearrange addresses to reduce bank conflicts. In this way, the chip provided by the embodiment of the present disclosure can reduce bank conflicts in the process of data memory access, and can improve the memory access bandwidth requirements of data in the process of matrix calculation and convolution calculation by the tensor computing engine 430, reducing the burden on software personnel.

[0156] In some embodiments, the module adopts a rearrangement logic based on an XOR operation, and the user configures the current available line size line_size of the local memory, the block size sector_sz of the data block, and the address interval step length stride between two adjacent data blocks in the same group of data blocks in the data loading instruction and / or the data reading instruction according to the usage scenario, wherein the line size line_size represents the amount of data that can be stored in a single row in the local memory, the block size sector_sz represents the amount of data with continuous addresses that needs to be read from a single row of the local memory in a single clock cycle, and the address interval step length stride represents the address interval between two adjacent data blocks in the same group of data blocks read in a single clock cycle. In this way, the first address rearrangement module and the second address rearrangement module can support the configurability of address conversion parameters, thereby being able to meet a variety of application scenarios and improve the flexibility of the chip.

[0157] In some embodiments, see Figure 4A, the tensor computing engine 430 may include a tile_M×tile_N dot multiplication matrix 432 (corresponding to the computing unit array in the aforementioned embodiment), and each dot multiplication unit DP in the dot multiplication matrix 432 may perform tile_K multiplication and addition operations, such as A0×B0+A1×B1+…. Wherein, tile_M, tile_N and tile_K are all positive integers. The dot multiplication matrix 432 may be used to perform general matrix dot multiplication operations and / or convolution operations on the data read from the local memory 420.

[0158] In some embodiments, when the computing mode of the chip is a matrix dot multiplication operation, the tensor storage engine can read the matrix data from the global memory and write it to the local memory. For example, if the product A×B between matrix A and matrix B is to be calculated, the tensor storage engine can read the matrix data of matrix A / B from the global memory and write it to the local memory, and the memory access control unit in the tensor computing engine can read the matrix data of matrix A / B to be subjected to the matrix product operation from the local memory. Among them, in each clock cycle, the memory access control unit in the tensor computing engine needs to read the matrix data of the elements of tile_M×tile_K of matrix A from the local memory (corresponding to the target matrix slicing in the aforementioned embodiment). Figure 4B FIG. 1 is a schematic diagram of storing the A matrix in the local memory in the related art, such as Figure 4B As shown, the target matrix slice includes 8×8 elements A0_0 to A7_7 of matrix A, and the target matrix slice includes 8 rows of elements from Line0 to Line7. Each row of elements (corresponding to the data blocks in the aforementioned embodiment) is in the same storage body bank0. Therefore, bank conflicts will occur in the process of reading the multiple rows of elements, resulting in a decrease in the memory access bandwidth of the local memory, affecting the computing efficiency of the computing engine.

[0159] In the disclosed embodiment, the amount of data of the matrix A that needs to be read in each row of the tensor computing engine in each clock cycle can be taken as a data block. Taking the data structure in which tile_K is 8 and the matrix elements are FP16 as an example, the block size of the data block is 16 bytes, and the size of each row in the local memory is 256 bytes. Then, each row of the local memory has 16 data blocks, and the addresses of each matrix element in each data block are continuous. The matrix data in each row of the local memory is address-rearranged at the granularity of the data block, and the rearranged address of each matrix element in each data block can be obtained. Figure 4CA schematic diagram of the arrangement of data blocks in a local memory after address rearrangement in a matrix operation scenario of a chip provided in an embodiment of the present disclosure, wherein the block size of the data block is 16 bytes, the size of each row in the local memory is 256 bytes, and each row in the local memory includes 16 data blocks, namely data blocks 0, 1, 2, ... 15. In the original data row before address rearrangement, the data blocks 0, 1, 2, ... 15 in each row are arranged in sequence, so that the data blocks with the same sequence number in each row (corresponding to a group of data blocks in the aforementioned embodiment) are located in the same column in the local memory, that is, they fall in the same bank. In this way, when reading the data blocks in the same group (such as data blocks 0, or data blocks 1, etc.) in a single clock cycle, there will be a bank conflict problem. In the rows s_line0 to s_line15 after the address is rearranged, the data blocks with the same sequence number in each row are located in different columns in the local memory, that is, they fall in different banks. Therefore, when the data blocks in the same group (for example, data blocks 0 or data blocks 1, etc.) are read in a single clock cycle, there will be no bank conflict problem.

[0160] For matrix dot multiplication operations, the data arrangement of the input matrix can be row-major, that is, the matrix is ​​stored row by row in the memory, so that each data unit in a single data block can correspond to a row of matrix elements in the target matrix slice; the data arrangement of the input matrix can also be column-major, that is, the matrix is ​​stored column by column in the memory, so that each data unit in a single data block can correspond to a column of matrix elements in the target matrix slice. It can be understood that for the matrix dot multiplication operation of matrix A and matrix B, the matrix data of matrix A can be stored in row-major order, the matrix data of matrix B can be stored in column-major order, and the address rearrangement method of each data block in the matrix data of matrix B can be similar to the address rearrangement method of each data block in the matrix data of matrix A mentioned above.

[0161] In some embodiments, the block size of the data block may correspond to the data type of each matrix element in the matrix data. For example, if the matrix element is of INT8 type, the tensor computing engine needs to read tile_MByte of matrix data per row in each clock cycle, that is, the block size of the data block is tile_M; if the matrix element is of FP16 type, the tensor computing engine needs to read tile_M×2Byte of matrix data per row in each clock cycle, that is, the block size of the data block is tile_M×2; if the matrix element is of TF16 type, the tensor computing engine needs to read tile_M×4Byte of matrix data per row in each clock cycle, that is, the block size of the data block is tile_M×4.

[0162] Figure 4D A schematic diagram of the arrangement of each data block in a local memory after address rearrangement when the data type of a matrix element in a chip provided by an embodiment of the present disclosure is FP16, such as Figure 4D As shown, taking tile_M as 16 as an example, when the data type of the matrix element is FP16, the block size of the data block is 32Byte, the size of each row in the local memory is 256Byte, and each row in the local memory includes 8 data blocks, namely data blocks 0, 1, 2, ...7.

[0163] Figure 4E A schematic diagram of the arrangement of each data block in a local memory after address rearrangement when the data type of a matrix element in a chip is TF16 provided in an embodiment of the present disclosure, such as Figure 4E As shown, taking tile_M as 16 as an example, when the data type of the matrix element is TF16, the block size of the data block is 64Byte, the size of each row in the local memory is 256Byte, and each row in the local memory includes 4 data blocks, namely data blocks 0, 1, 2, and 3.

[0164] In some embodiments, in addition to reading data directly from the local memory for calculation, the tensor computing engine can also read data from the register file for calculation. In the case of reading data from the register file for calculation, the input data of the matrix operation / convolution operation can be read from the local memory to the register file through the local memory access (Local Memory Access, LMA) pipeline. Since the data access bandwidth provided by the local memory to the tensor computing engine and the LMA pipeline may be different, when the tensor computing engine directly reads data from the local memory for calculation, the current available line size line_size of the local memory can be the first size, and when the tensor computing engine reads data from the register file for calculation, the current available line size line_size of the local memory can be the second size. Figure 4F A schematic diagram of the currently available row size of a local memory in a chip provided by an embodiment of the present disclosure is shown in FIG. Figure 4F As shown, if the tensor computing engine directly reads data from the local memory for computing, the data access bandwidth provided by the local memory to the tensor computing engine may be 256 Bytes, and the current available line size line_size of the local memory may be 256 Bytes. Figure 4G A schematic diagram of the currently available row size of a local memory in a chip provided by an embodiment of the present disclosure Figure 2 ,like Figure 4GAs shown in FIG. 1 , if the tensor computing engine reads data from the register file for computing, the data access bandwidth provided by the local memory for the LMA pipeline can be 128 Byte, and the current available line size line_size of the local memory can be 128 Byte. In this way, by adding the current available line size of the local memory as a configurable address conversion parameter, the software can configure the appropriate available line size to adapt to different application scenarios.

[0165] In some embodiments, when the computing mode of the chip is a convolution operation, the tensor storage engine can read image data from the global memory and write it to the local memory. For convolution calculations, image data is stored in the local memory in the format of CxNDHW, and is stored in a memory interleaved form according to the channel dimension to ensure the address continuity of pixel data in the channel dimension; where C represents the number of input channels, N represents the batch sample size of the convolution operation, D represents the number of output channels, and H and W represent the height and width of the convolution feature map, respectively. Figure 4H A schematic diagram of the arrangement of each data block in a local memory after address re-arrangement without introducing an address interval step in a convolution operation scenario provided by a chip according to an embodiment of the present disclosure, assuming that the data size of each pixel data in the image is 32 bytes, each pixel data includes two parts of channel data, high 16 bytes and low 16 bytes, the high 16 bytes data and the low 16 bytes data are respectively used as a data block, the memory interleave granularity (Interleave Granularity, IG) used for storing image data in the local memory is 32 bytes, the width W of the convolution feature map corresponding to the image data is 32, for pixel Wi (such as pixel W0, W1, W2...etc.), Wi_lo is the low 16 bytes data of the pixel Wi, and Wi_hi is the high 16 bytes data of the pixel Wi. If the lower 16 bytes of data W0_lo to W15_lo of pixels W0 to W15 need to be read, since W0_lo to W15_lo are located in different columns in the local memory, that is, in different banks, no bank conflict will occur when data blocks W0_lo to W15_lo are read in the same clock cycle.

[0166] Continue to see Figure 4H, if the lower 16-byte data W13_lo to W28_lo of pixels W13 to W28 need to be read, since W13_lo and W28_lo are located in the same column in the local memory, that is, in the same bank, a bank conflict will occur when reading data blocks W13_lo to W28_lo in the same clock cycle. In view of this situation, the address interval step parameter is introduced to characterize the address interval between two adjacent data blocks in the same group of data blocks read in a single clock cycle, and the addresses of each data block are rearranged according to the address interval step. Fig. 4I A schematic diagram of the arrangement of data blocks in a local memory after address rearrangement is performed by introducing an address interval step in a convolution operation scenario in a chip provided by an embodiment of the present disclosure. When the memory interleaving granularity is 32Byte, the address interval step stride between two adjacent data blocks in the same group of data blocks is set to 32Byte, and the block size of the data blocks is 16Byte. It is only necessary to perform address rearrangement within the range of 2 rows (32Byte / 16Byte), so that the computing engine can read any 16 consecutive pixels of single-channel data (high 16Byte or low 16Byte) in a single clock cycle without bank conflict.

[0167] In the disclosed embodiment, by implementing a hardware address reordering module in TCE and TME, the occurrence of bank conflicts is reduced, the complexity of the software calculation of address reordering logic is simplified, and the number of instructions is reduced. By adding three configurable parameters in the instructions, namely, the current available row size line_size of the local memory, the block size sector_sz of the data block, and the address interval step length stride between two adjacent data blocks in the same group of data blocks, the flexibility of address reordering is improved, and the first address reordering module and / or the second address reordering module can be made more universal and can support various calculation modes such as matrix dot multiplication operations and convolution operations.

[0168] It should be pointed out here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above chip embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the chip provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. For technical details not disclosed in the embodiments of the device of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0169] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present disclosure, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution. The execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure are for description only and do not represent the advantages and disadvantages of the embodiments.

[0170] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0171] In the several embodiments provided in the present disclosure, it should be understood that the disclosed chips and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0172] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present disclosure may all be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0173] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0174] The above description is only an implementation mode of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed in the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A data access method for a chip, characterized in that: The method comprises: In response to a data loading instruction, target data to be written into the on-chip memory is read from the off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit; Using a first address rearrangement module, the write address of each data unit is rearranged to obtain a rearranged address of each data unit; wherein each data block in the same group corresponds to a different storage body in the on-chip memory after rearrangement, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block; The data unit is written into the on-chip memory according to the rearranged address of the data unit.

2. The data access method according to claim 1, characterized in that: The using the first address rearrangement module to rearrange the write addresses of the data units to obtain the rearranged addresses of the data units includes: The write address of the data unit is converted into a rearrangement address of the data unit based on an address conversion parameter by using a first address rearrangement module.

3. The data access method according to claim 2, characterized in that: The on-chip memory includes a plurality of storage blocks arranged in an array, the storage block includes at least one storage unit with a continuous address, the at least one storage unit is used to store at least one data unit in the data block, the storage blocks in the same column belong to the same storage body, and the address conversion parameter includes a row size of a single row of the storage blocks; The converting the write address of the data unit into the rearrangement address of the data unit based on the address conversion parameter includes: Based on the write address of the data unit and the row size, determining a first row identifier and a first intra-row address offset corresponding to the write address; Determine a second intra-row address offset based on the first row identifier and the first intra-row address offset; A reordering address of the data unit is determined based on the first row identifier and the second intra-row address offset.

4. The data access method according to claim 3, characterized in that: The address conversion parameters also include the block size of the data blocks; The determining the second intra-row address offset based on the first row identifier and the first intra-row address offset includes: Determine a first block identifier and a first intra-block address offset corresponding to the write address based on the first intra-row address offset and the block size; Determining a second block identifier based on the first row identifier and the first block identifier; The second intra-row address offset is determined based on the second block identifier and the first intra-block address offset.

5. The data access method according to claim 4, characterized in that: The address conversion parameters also include the address interval step between two adjacent data blocks in the same group of data blocks; The determining the second block identifier based on the first row identifier and the first block identifier includes: Determining a row spacing step of address rearrangement based on the address spacing step and the block size; Determine a second row identifier based on the first row identifier and the row interval step size; The second block identifier is determined based on the second row identifier and the first block identifier.

6. The data access method according to claim 2, characterized in that: The method further comprises: Based on the data loading instruction, the address conversion parameters are determined; the address conversion parameters include at least one of the following: the row size of a single row storage block in the on-chip memory, the block size of the data block, and the address interval step between two adjacent data blocks in the same group of data blocks.

7. The data access method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining a data read instruction, wherein the data read instruction is used to instruct to read a group of target data blocks from the on-chip memory, wherein the target data blocks include at least one target data unit; Based on the data read instruction, determining a read address of each of the target data units; Using a second address rearrangement module, rearrange the read addresses of the target data units to obtain rearranged addresses of the target data units; The target data units are read in parallel from the on-chip memory according to the rearranged addresses of the target data units.

8. The data access method according to claim 7, characterized in that: The method further comprises: Based on the data read instruction, determining an address conversion parameter; the address conversion parameter includes at least one of the following: a row size of a single row storage block in the on-chip memory, a block size of the data block, and an address interval step between two adjacent data blocks in the same group of the data blocks; The using the second address rearrangement module to rearrange the read addresses of the target data units to obtain the rearranged addresses of the target data units includes: The second address rearrangement module is used to rearrange the read addresses of the target data units based on the address conversion parameters to obtain rearranged addresses of the target data units.

9. The data access method according to claim 7, characterized in that: The target data includes matrix data, a group of the target data blocks includes a target matrix slice in the matrix data, and the target data unit includes a matrix element in the target matrix slice; The method further comprises: Each of the matrix elements in the read target matrix slice is input into a computing unit array to perform a matrix product operation.

10. The data access method according to claim 7, characterized in that: The target data includes image data, a group of the target data blocks includes a target image slice in the image data, and the target data block includes channel data of a single pixel in the target image slice; the method further includes: The channel data of each pixel in the read target image slice is input into the computing unit array to perform a convolution operation.

11. A chip, characterized in that: include: A data storage engine and an on-chip memory, wherein the data storage engine comprises a first data reading module, a first address rearrangement module and a first data writing module; wherein: The first data reading module is used to: in response to a data loading instruction, read the target data to be written into the on-chip memory from the off-chip memory; the target data includes at least one group of data blocks, and the data block includes at least one data unit; The first address rearrangement module is used to: rearrange the write address of each data unit to obtain the rearranged address of each data unit; wherein each data block in the same group corresponds to different storage bodies in the on-chip memory after rearrangement, and the rearranged address of each data unit in the data block is located in the storage body corresponding to the data block; The first data writing module is used to write the data unit into the on-chip memory according to the rearranged address of the data unit.

12. The chip according to claim 11, characterized in that: The first address rearrangement module is further used to convert the write address of the data unit into the rearrangement address of the data unit based on an address conversion parameter.

13. The chip according to claim 12, characterized in that: The data storage engine also includes: A first parameter determination module is used to: determine the address conversion parameters based on the data loading instruction, and send the address conversion parameters to the first address rearrangement module; the address conversion parameters include at least one of the following: the row size of a single row storage block in the on-chip memory, the block size of the data block, and the address interval step between two adjacent data blocks in the same group of data blocks.

14. The chip according to any one of claims 11 to 13, characterized in that: The chip further includes a computing engine, which includes an address determination module, a second address rearrangement module, a second data reading module and a computing unit array, wherein the computing unit array includes a plurality of computing units; wherein: The address determination module is used to: obtain a data read instruction, the data read instruction is used to instruct to read a group of target data blocks from the on-chip memory, the target data blocks including at least one target data unit; based on the data read instruction, determine the read address of each of the target data units; The second address rearrangement module is used to rearrange the read addresses of the target data units to obtain the rearranged addresses of the target data units; The second data reading module is used to read each target data unit in parallel from the on-chip memory according to the rearrangement address of each target data unit, and input each target data unit in a group of read target data blocks into the computing unit array.

15. The chip according to claim 14, characterized in that: The computing engine also includes: A second parameter determination module is used to: determine an address conversion parameter based on the data read instruction; the address conversion parameter includes at least one of the following: a row size of a single row storage block in the on-chip memory, a block size of the data block, and an address interval step between two adjacent data blocks in the same group of the data blocks; The second address rearrangement module is further used to rearrange the read addresses of each of the target data units based on the address conversion parameters to obtain the rearranged addresses of each of the target data units.

16. The chip according to claim 14, characterized in that: In the case where the target data includes matrix data, a group of the target data blocks includes a target matrix slice in the matrix data, the target data unit includes a matrix element in the target matrix slice, and the calculation unit array is used to perform a matrix product operation based on each of the matrix elements in the read target matrix slice; In the case where the target data includes image data, a group of the target data blocks includes target image slices in the image data, and the target data blocks include channel data of a single pixel in the target image slice; the computing unit array is used to perform convolution operations based on the channel data of each pixel in the read target image slice.

Citation Information

Patent Citations

  • Memory data migration method and device and storage medium

    CN111143241A

  • Data reading method and data reading circuit

    CN112506567A

  • Heterogeneous structure system packet processing method

    CN114995882A

  • Data storage method and device, electronic equipment and storage medium

    CN117742594A

  • Data processing method and corresponding data storage device

    CN117806532A

Cited By

  • AI intelligent calculation platform reasoning acceleration method and system

    CN121860070A

  • Data access and storage method for chip, and chip

    WO2026144973A1