Methods for accessing memory data and storing data in memory
By using on-chip memory buffered data in the artificial intelligence chip, the performance degradation caused by non-aligned access is solved, efficient and secure memory access is achieved, and the overall performance of the artificial intelligence model is improved.
Patent Information
- Application Number
- CN202410341030.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-03-22
AI Technical Summary
The performance degradation and hardware abnormal problems caused by non-aligned access in the existing memory pool access mechanism are especially high when the system overhead is frequently applied for and freed during the training and inference of artificial intelligence models.
By utilizing the on-chip memory buffer data in the artificial intelligence chip, access to any pointer is achieved, avoiding the alignment of pointer offsets during the compilation process, and reducing the copying of data in high-bandwidth memory, and loading and storing memory data in the first and second granularities.
It improves the efficiency and security of memory access, reduces system overhead, and improves the performance of artificial intelligence models.
Smart Images

Figure CN118152130B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence services and chip technology, and more specifically to a method for accessing memory data, a method for storing data in a memory, an electronic device, and a storage medium. Background Art
[0002] During the training and inference of AI models, the frequent allocation and release of large amounts of temporary memory leads to high system overhead. To optimize memory utilization, a memory pool mechanism has been proposed. This mechanism pre-allocates a large memory space as a memory pool, which can store many tensors. Memory blocks are allocated from the memory pool when needed and returned to the pool for recycling after use. This mechanism reduces frequent memory allocation and release operations, lowers system overhead, improves memory utilization, and accelerates memory allocation and deallocation, effectively enhancing overall performance.
[0003] However, in the current memory pool access mechanism, unaligned data access is very likely to occur. This is because most modern processor architectures require that memory access addresses must be aligned with the data to be accessed according to specific boundaries. For example, when accessing an element in a tensor (for example, it may be a 4-byte data), it is often required to obtain the data corresponding to the element from an address that is an integer multiple of the element's size (that is, the address size needs to be an integer multiple of 4 bytes). Failure to do so will result in performance degradation or hardware exceptions. Because the current memory pool allocation mechanism typically dynamically allocates memory blocks, the first address of the accessed memory may not always be aligned with the granularity of the element's corresponding data (for example, not an integer multiple of 4 bytes), which may lead to unaligned access. Unaligned access reduces the efficiency of CPU memory access and affects the training and inference of artificial intelligence models.
[0004] To this end, it is necessary to improve the current memory data access mechanism to ensure that memory access is efficient, safe and reliable. Summary of the Invention
[0005] The embodiments of the present disclosure provide a method for accessing memory data, a method for storing data in a memory, an electronic device, and a storage medium.
[0006] An embodiment of the present disclosure provides a method for accessing memory data, wherein the access granularity of the memory data is a first granularity, and the method includes: determining that a starting address storing the memory data has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity; based on the starting address storing the memory data, loading the memory data into an on-chip memory; and accessing the memory data from the on-chip memory at the first granularity.
[0007] For example, loading the memory data into the on-chip memory includes: loading a first portion of the memory data into the on-chip memory at the second granularity and loading a second portion of the memory data into the on-chip memory at the first granularity.
[0008] For example, the memory data includes M*N bytes of data, M is the first granularity, the first part of the memory data includes the first k1 bytes of data in the memory data and the last k2 bytes of data in the memory data, k1+k2=M, M, N, k1 and k2 are all integers, and the second part of the memory data includes the M*(N-1) bytes of data in the middle of the memory data.
[0009] For example, loading the first portion of the memory data into the on-chip memory at the second granularity includes: loading the first portion of the memory data into the execution unit at the second granularity; and storing the first portion of the memory data from the execution unit to the on-chip memory at the second granularity.
[0010] For example, loading the second part of the memory data into the on-chip memory at the first granularity includes: loading the second part of the memory data into the execution unit at the first granularity; and storing the second part of the memory data from the execution unit to the on-chip memory at the first granularity.
[0011] For example, loading the second portion of the memory data into the on-chip memory at the first granularity includes: loading the second portion of the memory data into the execution unit at the first granularity; storing the second portion of the memory data from the execution unit to a register at the first granularity; loading the second portion of the memory data from the register to the execution unit at the second granularity; and storing the second portion of the memory data from the execution unit to the on-chip memory at the second granularity.
[0012] For example, loading the memory data into the on-chip memory includes: loading the memory data into the on-chip memory using a direct memory access mechanism.
[0013] For example, loading the memory data into the on-chip memory includes: using a dedicated data loading instruction by the execution unit to load the memory data into the on-chip memory.
[0014] An embodiment of the present disclosure provides a method for storing data in a memory, wherein the access granularity of the data is a first granularity, and the method includes: determining that a starting address in the memory where the data is to be stored has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity; based on the starting address where the data is to be stored, storing the data from an execution unit to an on-chip memory at the first granularity; and storing the data from the on-chip memory to a memory space corresponding to the starting address where the data is to be stored.
[0015] For example, storing the data in the memory space corresponding to the first address where the data is to be stored includes: storing the first part of the data in the memory at the second granularity and storing the second part of the data in the memory at the first granularity.
[0016] For example, the data includes M*N bytes of data, M is the first granularity, the first part of the data includes the first k1 bytes of data in the data and the last k2 bytes of data in the data, k1+k2=M, M, N, k1 and k2 are all integers, and the second part of the data includes M*(N-1) bytes of data in the middle of the data.
[0017] For example, storing the first portion of the data in the memory at the second granularity includes: loading the first portion of the data into an execution unit at the second granularity; and storing the first portion of the data from the execution unit to the memory at the second granularity.
[0018] For example, storing the second portion of the data in the memory at the first granularity includes: loading the second portion of the data into the execution unit at the first granularity; and storing the second portion of the data from the execution unit to the memory at the first granularity.
[0019] For example, storing the second portion of the data in the memory at the first granularity includes: loading the second portion of the data into the execution unit at the second granularity; storing the second portion of the data from the execution unit to the register at the second granularity; loading the second portion of the data from the register to the execution unit at the first granularity; and storing the second portion of the data from the execution unit to the memory at the first granularity.
[0020] For example, storing the data in the memory space corresponding to the first address where the data is to be stored includes: storing the data in a memory using a direct memory access mechanism.
[0021] For example, storing the data in the memory space corresponding to the first address where the data is to be stored includes: the execution unit storing the data in the memory using a dedicated data load instruction.
[0022] An embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory, wherein the memory stores a computer executable program, and when the computer executable program is executed by the processor, the above-mentioned method of accessing memory data and / or the method of storing data in the memory is executed.
[0023] An embodiment of the present disclosure provides a device, including: a processor; a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the method for accessing memory data and / or the method for storing data in the memory is implemented.
[0024] An embodiment of the present disclosure provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed by a processor, the method for accessing memory data and / or the method for storing data in the memory are implemented.
[0025] According to another aspect of the present disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable medium and executes the computer instructions, causing the computer device to perform the methods provided in the various aspects or various optional implementations of the various aspects described above.
[0026] Various embodiments of this disclosure utilize on-chip memory (buffers) in AI chips to improve current memory data access mechanisms. This allows the on-chip memory to buffer data when the processor encounters unaligned access, enabling access to arbitrary pointers. This eliminates the need to align pointer offsets during compilation, reducing the limitations of AI chips. Furthermore, by using on-chip memory to buffer data, data copying in High Bandwidth Memory (HBM) is avoided, thereby improving operator performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. The drawings described below are only exemplary embodiments of the present disclosure.
[0028] Figure 1A A schematic diagram showing a processor reading data from the HBM via a pointer aligned with an address in the HBM.
[0029] Figure 1B A schematic diagram showing a processor storing data into the HBM through a pointer aligned with an address in the HBM.
[0030] Figure 1C A schematic diagram showing a processor accessing data through a pointer that is not aligned with an address in the HBM.
[0031] Figure 2 is a flowchart illustrating a method for accessing memory data according to an embodiment of the present disclosure.
[0032] Figure 3 is a schematic diagram illustrating a method for accessing memory data according to an embodiment of the present disclosure.
[0033] Figure 4 is another schematic diagram illustrating a method for accessing memory data according to an embodiment of the present disclosure.
[0034] Figure 5 is another schematic diagram illustrating a method for accessing memory data according to an embodiment of the present disclosure.
[0035] Figure 6 is another schematic diagram illustrating a method for accessing memory data according to an embodiment of the present disclosure.
[0036] Figure 7 is a flowchart illustrating a method for storing data in a memory according to an embodiment of the present disclosure.
[0037] Figure 8 is a schematic diagram illustrating a method for storing data in a memory according to an embodiment of the present disclosure.
[0038] Figure 9 is another schematic diagram illustrating a method for storing data in a memory according to an embodiment of the present disclosure.
[0039] Figure 10 is another schematic diagram illustrating a method for storing data in a memory according to an embodiment of the present disclosure.
[0040] Figure 11 is another schematic diagram illustrating a method for storing data in a memory according to an embodiment of the present disclosure.
[0041] Figure 12 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0042] Figure 13An architectural diagram of a computing device according to an embodiment of the present disclosure is shown.
[0043] Figure 14 A schematic diagram of a computer-readable storage medium according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present disclosure more apparent, the following will describe in detail exemplary embodiments of the present disclosure with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0045] In this specification and the accompanying drawings, substantially the same or similar steps and elements are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance or ranking.
[0046] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.
[0047] Optionally, the various models involved in the embodiments of the present disclosure may be artificial intelligence models, in particular artificial intelligence-based neural network models. Typically, artificial intelligence-based neural network models are implemented as acyclic graphs in which neurons are arranged in different layers. Typically, a neural network model includes an input layer and an output layer, which are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. The network nodes (i.e., neurons) are fully connected to the nodes in the adjacent layers via edges, and there are no edges between the nodes in each layer. The data received at the nodes of the input layer of the neural network are propagated to the nodes of the output layer via any one of the hidden layer, activation layer, pooling layer, convolutional layer, etc. The input and output of the neural network model can take various forms, and the present disclosure is not limited thereto.
[0048] Artificial intelligence models are particularly well-suited for processing video images. A single image or video frame can be considered a three-dimensional tensor. Since videos are composed of multiple frames, they can be considered four-dimensional tensors. Due to the underlying structure and design of AI models, they are well-suited for efficient parallel operations on tensor data. However, in practical applications, the data format to be processed often does not fully match the model's tensor representation. Specifically, when the video codec output format is NV12, the Y (luminance) and UV (chrominance) channels are typically stored in a contiguous block of memory. To save space, the addresses of the UV channels may not be aligned to the element size or pixel size (e.g., 4 bytes), meaning that the memory address may not be an integer multiple of 4 bytes. Accessing these non-aligned addresses on standard CPUs can result in performance degradation or trigger hardware exceptions.
[0049] Currently, industry and academia have proposed AI chips that allow for unaligned data access, but these chips often suffer from poor performance. Most mainstream chips currently do not allow unaligned access, and any unaligned access operation can result in incorrect operator results.
[0050] To this end, this paper utilizes the on-chip memory (buffer) in AI chips to improve the current memory pool access mechanism. This allows the processor to use on-chip memory to buffer data when accessing data in an unaligned manner, enabling access to arbitrary pointers. This eliminates the need to align pointer offsets during compilation, reducing the limitations of AI chips. Furthermore, by using on-chip memory to buffer data, data copying in High Bandwidth Memory (HBM) is avoided, thereby improving operator performance.
[0051] Various embodiments of the present disclosure are described in detail through the following drawings.
[0052] Figure 1A A schematic diagram showing a processor reading data from the HBM via a pointer aligned with an address in the HBM. Figure 1B A schematic diagram showing a processor storing data into the HBM through a pointer aligned with an address in the HBM. Figure 1C A schematic diagram showing a processor accessing data through a pointer that is not aligned with an address in the HBM.
[0053] The execution unit (Execution Unit) is a key component of modern processor microarchitecture, responsible for actually performing various arithmetic and logical operations and memory access operations. Processor cores (including CPU cores and GPU cores) typically contain multiple execution units, each dedicated to executing a specific type of instruction. Major execution units include: Integer Arithmetic Logic Unit (Integer ALU), Floating Point Unit (FPU), Load / Store Unit (Load / Store Unit), Branch Prediction Unit (BPU), SIMD Unit (Single Instruction Multiple Data), and so on.
[0054] like Figure 1A and Figure 1B As shown, the execution unit can read and store data normally and efficiently for memory pointers that are aligned with the element size. The "element size alignment" here means that the memory address in the memory (for example, HBM) must be an integer multiple of the size of a specific data type. For example, for a 4-byte integer, its address must be divisible by 4. For example, if you want to access a 4-byte integer, the aligned access address for a 32-bit system can be 0x00000000, 0x00000004, 0x00000008, etc. When the memory pointer address accessed by the program meets this alignment requirement, the execution unit can directly issue a read and write request to the specified address without special processing. Specifically, as Figure 1A As shown, the execution unit (such as the load unit) can directly issue a load instruction to the HBM to directly obtain data from the HBM. Figure 1B As shown, the execution unit (eg, store unit) can directly issue a store instruction to the HBM to directly store data into the HBM.
[0055] However, in reality, not all memory addresses can be "element-size aligned", and there may be misalignment. Figure 1CAs shown in the figure, when the processor accesses a memory pointer address that does not meet the "element-size alignment" requirement, known as an unaligned access, the execution unit cannot directly and efficiently read the data. For example, if a 4-byte integer is to be accessed, the unaligned access address on a 32-bit system might be 0x00000002, 0x00000007, and so on. In this case, the traditional approach is to allocate a temporary element-size aligned memory pointer and then use Direct Memory Access (DMA) to copy the data from the original unaligned pointer to this aligned memory area. The processor then accesses this aligned memory area to retrieve the complete data. Similarly, when the processor attempts to write data to a memory address that does not meet the "element-size alignment" requirement, the execution unit cannot directly use a "store" instruction to store the data to this "unaligned" address. DMA is also required to attempt to store the data to the "unaligned" address.
[0056] Although through Figure 1C The method in this paper can avoid the performance degradation or even abnormal interruption caused by direct access to non-aligned addresses, but the additional memory copy operation of this solution brings additional performance overhead.
[0057] To this end, an embodiment of the present disclosure proposes a method for accessing memory data, wherein the access granularity of the memory data is a first granularity, the starting address storing the memory data has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity. The method for accessing memory data includes: loading the memory data into an on-chip memory based on the starting address storing the memory data; and accessing the memory data from the on-chip memory at the first granularity.
[0058] To this end, an embodiment of the present disclosure proposes a method for storing data in a memory, wherein the access granularity of the data is a first granularity, and the method includes: determining that the starting address of the data to be stored in the memory has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity; based on the starting address of the data to be stored, storing the data from the execution unit to the on-chip memory at the first granularity; and storing the data from the on-chip memory to the memory space corresponding to the starting address of the data to be stored.
[0059] Embodiments of the present disclosure also provide a device for accessing memory data, a device for storing data in memory, an electronic device, and a storage medium, which are used to implement the method for accessing memory data or the method for storing data in memory in the above embodiments.
[0060] Various embodiments of the present disclosure utilize on-chip memory in AI chips to improve current memory data access mechanisms. This allows the processor to buffer data when unaligned, enabling access to arbitrary pointers without requiring alignment of pointer offsets during compilation, thus reducing the limitations of AI chips. Furthermore, by using on-chip memory to buffer data, direct data copying to High Bandwidth Memory (HBM) is avoided, thereby improving operator performance.
[0061] The following combination Figures 2 to 14 All or part of the embodiments according to the present disclosure will be described in more detail.It should be noted that the same reference numerals in different drawings will be used to refer to the same elements described.
[0062] Figure 2 FIG. 2 is a flow chart illustrating a method 20 for accessing memory data according to an embodiment of the present disclosure.
[0063] For example, Figure 2 As shown, at least one embodiment of the present disclosure provides a method for accessing memory data, wherein the access granularity of the memory data is a first granularity. For example, the memory data access method includes the following steps S210 to S230.
[0064] In step S210 , it is determined that a first address storing the memory data has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity.
[0065] In step S220, the memory data is loaded into an on-chip memory based on the first address of the memory data.
[0066] In step S230 , the memory data is accessed from the on-chip memory at the first granularity.
[0067] For example, the first granularity is the basic access granularity corresponding to the memory data, which is equal to the size of the data type, such as 8 bytes for double-precision floating-point numbers (double), 4 bytes for single-precision floating-point numbers (float), 2 bytes for BF16, and 1 byte for characters (char). For example, assuming a 32-bit system, the execution unit may need to access a tensor that continuously stores N elements, where each element in the tensor is 4 bytes of data, that is, the first granularity is 4 bytes. At this time, the tensor requires a total of 4*N bytes of storage space. Optionally, the storage space is located on HBM.
[0068] In step S210, based on the first address storing the memory data, it can be determined whether an unaligned access occurs. For example, it can be determined whether the first address storing the memory data is divisible by a first granularity. Specifically, if in S210, it is determined that the first address storing the memory data is offset by a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity, then an unaligned access must occur.
[0069] For example, the first element of the storage tensor corresponds to the starting address 0x00000006. This starting address is offset by 2 bytes relative to an integer multiple of 4 bytes. Assuming the second granularity is 1 byte, 2 bytes of data for the second granularity indicate that unaligned access is inevitable.
[0070] In order to avoid the problem of performance degradation or even abnormal interruption caused by unaligned access, in step S220, the memory data is buffered by copying it from the HBM to the on-chip memory to achieve access to any pointer. The on-chip memory can rearrange and align the memory data. Therefore, in step S230, the execution unit can directly access the data rearranged by the on-chip memory without the need for additional memory copying or other overhead. The present disclosure does not require the program to align the offset of the address pointer, reducing the limitations of the artificial intelligence chip.
[0071] At the same time, Figure 1C Compared with the traditional solution in
[15] , the present disclosure utilizes on-chip memory to buffer data and avoids reapplying for storage space and copying data in HBM, thereby improving the performance of the operator.
[0072] Figure 3 2 is a schematic diagram illustrating a method 20 for accessing memory data according to an embodiment of the present disclosure, which illustrates optional details of steps S220 and S230.
[0073] Specifically, in step S220 , the first portion of the memory data may be loaded into the on-chip memory at the second granularity, and the second portion of the memory data may be loaded into the on-chip memory at the first granularity.
[0074] For example, in order to successfully access each element in the tensor, the memory data can be divided into two parts according to the logical view: the first part is non-aligned (see Figure 3Assume that the memory data includes M*N bytes of data, where M is a first granularity, the first part of the memory data includes the first k1 bytes of data in the memory data and the last k2 bytes of data in the memory data, k1+k2=M, where M, N, k1, and k2 are all integers, and the second part of the memory data includes the middle M*(N-1) bytes of data in the memory data.
[0075] For example, assuming the first granularity is 4 bytes and the second granularity is 1 byte, the first unaligned portion is the 2 bytes at the beginning and the 2 bytes at the end of the memory data. The second aligned portion is the portion from the 3rd byte to the 4*N-2th byte of the memory data.
[0076] In the case where the on-chip memory supports non-aligned storage of data (that is, the on-chip memory can store data at different granularities), the loading of the first portion of the memory data into the on-chip memory at the second granularity includes: loading the first portion of the memory data into the execution unit at the second granularity; storing the first portion of the memory data from the execution unit to the on-chip memory at the second granularity. The loading of the second portion of the memory data into the on-chip memory at the first granularity includes: loading the second portion of the memory data into the execution unit at the first granularity; storing the second portion of the memory data from the execution unit to the on-chip memory at the first granularity. Of course, the present disclosure is not limited to this.
[0077] Specifically, the execution unit can use the "load" instruction at the second granularity to first load the 2 bytes of data at the head of the memory data from the HBM to the execution unit. Then use the "store" instruction to store the 2 bytes of data at the head of the memory data from the execution unit to the on-chip memory. Subsequently, the execution unit can use the "load" instruction at the first granularity to load the part of the memory data from the 3rd byte to the 4*N-2th byte from the HBM to the execution unit in sequence at the first granularity, and then similarly use the "store" instruction to store these data in the on-chip memory. Finally, the execution unit can move the 2 bytes of data at the tail of the memory data from the HBM to the on-chip memory in a similar manner at the second granularity.
[0078] During the entire process, the first part of the memory data can be logically regarded as data of the second granularity, and the address pointer corresponding to this part of the data is also of the second granularity. At the same time, the second part of the memory data can be logically regarded as data of the first granularity, and the address pointer corresponding to this part of the data is also of the first granularity. Since moving data from HBM to on-chip memory does not involve any calculation process (that is, the value of the data will not be changed during this process), and the execution unit reads data directly from the on-chip memory at the first granularity, it will not affect subsequent calculations.
[0079] Figure 4 is another schematic diagram illustrating the method 20 for accessing memory data according to an embodiment of the present disclosure, which illustrates optional details of steps S220 and S230.
[0080] In the case that the on-chip memory does not support non-aligned data storage (i.e., the on-chip memory can only store data at the same granularity), the memory data can only be stored in the on-chip memory in sequence at a smaller granularity (i.e., the second granularity). Figure 3 In the manner shown, similarly, the first portion of the memory data is first loaded into the execution unit and then stored into the on-chip memory at the second granularity.
[0081] As for the second part of the memory data, loading the first part of the memory data into the on-chip memory at the second granularity includes: loading the second part of the memory data into the execution unit at the first granularity; storing the second part of the memory data from the execution unit to the register at the first granularity; loading the second part of the memory data from the register to the execution unit at the second granularity; and storing the second part of the memory data from the execution unit to the on-chip memory at the second granularity.
[0082] like Figure 4 As shown, the execution unit first stores the data from the 3rd byte to the 4N-2th byte in the memory data into the register at the first granularity, and then uses the second granularity to move the data in the register to the on-chip memory in sequence. After that, the memory data in the on-chip memory is accessed at the first granularity. Since moving data from the HBM or register to the on-chip memory does not involve any calculation process (that is, the value of the data will not change during this process), and the execution unit reads data directly from the on-chip memory at the first granularity, it will not affect subsequent calculations.
[0083] Figure 5 is another schematic diagram illustrating the method 20 for accessing memory data according to an embodiment of the present disclosure, which illustrates optional details of steps S220 and S230.
[0084] Optionally, step S220 includes: loading the memory data into the on-chip memory using a direct memory access mechanism. Figure 5 As shown, if the hardware supports direct memory access (DMA) copy functionality, it can be used to efficiently copy data from any memory location to an on-chip buffer, further enhancing access flexibility. Specifically, when the processor needs to access data in HBM or other unaligned memory areas, the DMA controller can be used to move the data at the hardware level, avoiding processor core intervention. Compared to copying within the processor core, DMA copying can run in parallel outside the processor core, significantly improving the system's parallel processing capabilities. Furthermore, because DMA operates directly at the hardware level, it avoids context switching and other software overhead associated with the processor core.
[0085] Figure 6 is another schematic diagram illustrating the method 20 for accessing memory data according to an embodiment of the present disclosure, which illustrates optional details of steps S220 and S230.
[0086] Optionally, step S220 includes: the execution unit uses a dedicated data loading instruction to load the memory data into the on-chip memory. Figure 6 As shown, the execution unit can issue a dedicated data load instruction based on the first address of the memory data, specifying the address range of the data to be loaded in the HBM and the target buffer address in the on-chip memory. This directly copies the memory data from the HBM to the on-chip memory. This approach avoids the overhead of multiple data copies between the HBM, cache, and registers, achieving direct data access and improving overall efficiency.
[0087] Figure 7 FIG. 7 is a flow chart illustrating a method 70 for storing data into a memory according to an embodiment of the present disclosure.
[0088] For example, Figure 7 As shown, at least one embodiment of the present disclosure provides a method for storing data in a memory, wherein the access granularity of the data is a first granularity. For example, the method for storing data in a memory includes the following steps S710 to S730.
[0089] In step S710 , it is determined that a first address of the memory in which the data is to be stored has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity.
[0090] In step S720 , based on the first address where the data is to be stored, the data is stored from the execution unit to the on-chip memory at the first granularity.
[0091] In step S730, the data is stored from the on-chip memory to the memory space corresponding to the first address where the data is to be stored.
[0092] For example, the first granularity is the basic access granularity corresponding to the memory data, which is equal to the size of the data type, such as 8 bytes for double-precision floating-point number (double), 4 bytes for single-precision floating-point number (float), 2 bytes for BF16, and 1 byte for character type (char). For example, assuming a 32-bit system, the execution unit may need to store a tensor of N elements continuously in the storage space in the memory, where each element in the tensor is 4 bytes of data, that is, the first granularity is 4 bytes. At this time, the tensor requires a total of 4*N bytes of storage space. Optionally, the storage space is located on HBM.
[0093] In step S710, based on the first address of the data to be stored, it can be determined whether unaligned storage occurs. For example, whether unaligned storage occurs can be determined by determining whether the first address of the data to be stored is divisible by the first granularity. Specifically, if it is determined in S210 that the first address of the data to be stored is offset by a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity, then unaligned storage must occur.
[0094] For example, suppose the first element of the tensor is to be stored in memory with a starting address of 0x00000006. This memory space may be located in HBM. In this case, the starting address is offset by 2 bytes relative to an integer multiple of 4 bytes. Assuming the second granularity is 1 byte, 2 bytes of data for the second granularity constitute the data. Therefore, unaligned storage is inevitable.
[0095] In order to avoid the problem of performance degradation or even abnormal interruption caused by non-aligned storage, in steps S720 to S730, the memory data is buffered by copying the data to the on-chip memory, and then the buffered data is moved to the memory to store the data in the memory space pointed to by any pointer. The on-chip memory can rearrange and align the memory data. Therefore, in step S730, the execution unit can directly access the data rearranged by the on-chip memory without additional memory copying or other overhead. The present disclosure does not require the program to align the offset of the pointer to the memory address, reducing the limitations of the artificial intelligence chip. At the same time, compared with Figure 1C Compared with the traditional solution in
[15] , the present disclosure utilizes on-chip memory to buffer data and avoids reapplying for storage space and copying data in HBM, thereby improving the performance of the operator.
[0096] Figure 8FIG. 7 is a schematic diagram illustrating a method 70 for storing data in a memory according to an embodiment of the present disclosure, which illustrates optional details of steps S720 and S730 .
[0097] Specifically, in step S730, the first portion of the data is stored in the memory at the second granularity and the second portion of the data is stored in the memory at the first granularity.
[0098] For example, in order to successfully store each element in the tensor, the data to be stored can be divided into two parts according to the logical view: the first part is non-aligned (see Figure 3 Assume that the data to be stored includes M*N bytes of data, where M is a first granularity, the first part of the data to be stored includes the first k1 bytes of data of the data to be stored and the last k2 bytes of data of the data to be stored, k1+k2=M, where M, N, k1, and k2 are all integers, and the second part of the data to be stored includes the middle M*(N-1) bytes of data of the data to be stored.
[0099] For example, assuming the first granularity is 4 bytes and the second granularity is 1 byte, the first unaligned portion is the 2 bytes at the beginning and the 2 bytes at the end of the data to be stored. The second aligned portion is the portion of the data to be stored from the 3rd byte to the 4*N-2th byte.
[0100] In the case where the on-chip memory supports unaligned access to data (i.e., data can be retrieved from the on-chip memory at different granularities), storing the first portion of the data into the memory at the second granularity includes: loading the first portion of the data into the execution unit at the second granularity; and storing the first portion of the data from the execution unit into the memory at the second granularity. Storing the second portion of the data into the memory at the first granularity includes: loading the second portion of the data into the execution unit at the first granularity; and storing the second portion of the data from the execution unit into the memory at the first granularity. Of course, the present disclosure is not limited to this.
[0101] Specifically, the execution unit can use the "load" instruction at the second granularity to first load the 2 bytes of data at the head of the data to be stored from the on-chip memory to the execution unit. Then use the "store" instruction to store the 2 bytes of data at the head of the data to be stored from the execution unit to the HBM. Subsequently, the execution unit can use the "load" instruction at the first granularity to load the part of the data to be stored from the 3rd byte to the 4*N-2th byte into the execution unit in sequence at the first granularity, and then similarly use the "store" instruction to store these data into the HBM. Finally, the execution unit can move the 2 bytes of data at the tail of the data to be stored from the on-chip memory to the HBM in a similar manner at the second granularity.
[0102] Throughout the entire process, the first portion of the data to be stored can be logically considered as data of the second granularity, and the address pointer corresponding to this portion of data is also of the second granularity. Simultaneously, the second portion of the data to be stored can be logically considered as data of the first granularity, and the address pointer corresponding to this portion of data is also of the first granularity. Since moving data from on-chip memory to HBM does not involve any computational process (i.e., the data value does not change during this process), it will not affect the subsequent storage of data.
[0103] Figure 9 is another schematic diagram illustrating a method 70 for storing data in a memory according to an embodiment of the present disclosure, which illustrates optional details of steps S720 and S730.
[0104] In the case that the on-chip memory does not support non-aligned access to data (i.e., the execution unit can only obtain data from the on-chip memory at one granularity), the data to be stored can only be stored in the on-chip memory in sequence at a smaller granularity (i.e., the second granularity). Figure 3 In the manner shown, similarly, the first portion of the data to be stored is first loaded into the execution unit and then stored into the HBM at the second granularity.
[0105] For the second part of the data to be stored, loading the first part of the data to be stored into the on-chip memory at the second granularity includes: loading the second part of the data into the execution unit at the second granularity; storing the second part of the data from the execution unit to the register at the second granularity; loading the second part of the data from the register to the execution unit at the first granularity; and storing the second part of the data from the execution unit to the memory at the first granularity.
[0106] like Figure 9As shown, the execution unit first stores the data from the 3rd byte to the 4N-2th byte in the register at the second granularity, and then sequentially moves the data in the register to the HBM at the first granularity. Since moving data from the register to the HBM does not involve any computation (i.e., the data value does not change during this process), it does not affect subsequent data storage.
[0107] Figure 10 is another schematic diagram illustrating a method 70 for storing data in a memory according to an embodiment of the present disclosure, which illustrates optional details of steps S720 and S730.
[0108] Optionally, step S730 includes: storing the data in a memory using a direct memory access mechanism. Figure 10 As shown, if the hardware supports direct memory access (DMA) copy functionality, it can be used to efficiently copy data from the on-chip buffer to the HBM, further enhancing access flexibility. Specifically, when the processor needs to store data in the HBM, the DMA controller can be used to move the data at the hardware level, avoiding processor core intervention. Compared to copying within the processor core, DMA copying can run in parallel outside the processor core, significantly improving the system's parallel processing capabilities. Furthermore, because it operates directly at the hardware level, DMA avoids context switching and other software overhead associated with the processor core.
[0109] Figure 11 is another schematic diagram illustrating a method 70 for storing data in a memory according to an embodiment of the present disclosure, which illustrates optional details of steps S720 and S730.
[0110] Optionally, step S730 includes: the execution unit uses a dedicated data load instruction to store the data into the memory. Figure 11 As shown, the execution unit can issue a dedicated data load instruction based on the starting address of the data to be stored, specifying the data address range in the HBM where the data needs to be stored. This allows data to be directly copied from on-chip memory to the HBM. This approach avoids the overhead of multiple data copies between the HBM, cache, and registers, achieving direct data access and improving overall efficiency.
[0111] According to yet another aspect of the present disclosure, an electronic device is provided for implementing the method 20 or 70 according to the embodiment of the present disclosure. Figure 12 A schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure is shown.
[0112] like Figure 12As shown, the electronic device 2000 may include one or more processors 2010 and one or more memories 2020. The memory 2020 stores computer-readable codes, which, when executed by the one or more processors 2010, may execute the method described above.
[0113] The processor in the embodiments of the present disclosure may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor, and may be an X86 architecture or an ARM architecture.
[0114] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0115] For example, the method or apparatus according to the embodiment of the present disclosure may also be used by means of Figure 13 The architecture of the computing device 3000 shown in FIG. Figure 13 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 13 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 13 One or more components of a computing device are shown.
[0116] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided. Figure 14 A schematic diagram of a storage medium 4000 according to the present disclosure is shown.
[0117] like Figure 14 As shown, the computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0118] The present disclosure also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to the present disclosure.
[0119] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0120] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0121] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for accessing memory data, wherein: The access granularity of the memory data is a first granularity, and the method includes: Determining that a first address storing the memory data has an offset of a second granularity compared to an integer multiple of the first granularity, and the second granularity is smaller than the first granularity; Based on the first address of storing the memory data, loading the memory data into an on-chip memory; and accessing the memory data from the on-chip memory at the first granularity, The loading of the memory data into the on-chip memory comprises: loading a first portion of the memory data into the on-chip memory at the second granularity and loading a second portion of the memory data into the on-chip memory at the first granularity, The memory data includes bytes of data, M is the first granularity, the first part of the memory data includes the first k1 bytes of data in the memory data and the last k2 bytes of data in the memory data, k1+k2=M, M, N, k1 and k2 are all integers, the second part of the memory data includes the middle k1 bytes of data in the memory data bytes of data.
2. The method according to claim 1, wherein The loading the first portion of the memory data into the on-chip memory at the second granularity comprises: loading a first portion of the memory data into an execution unit at the second granularity; A first portion of the memory data is stored from the execution unit to the on-chip memory at the second granularity.
3. The method according to claim 1, wherein The loading the second portion of the memory data into the on-chip memory at the first granularity includes: loading a second portion of the memory data into an execution unit at the first granularity; A second portion of the memory data is stored from the execution unit to the on-chip memory at the first granularity.
4. The method according to claim 1, wherein The loading the second portion of the memory data into the on-chip memory at the first granularity includes: loading a second portion of the memory data into an execution unit at the first granularity; storing a second portion of the memory data from the execution unit to a register at the first granularity; loading a second portion of the memory data from the register to the execution unit at the second granularity; and A second portion of the memory data is stored from the execution unit to the on-chip memory at the second granularity.
5. A method for storing data in a memory, wherein: The data access granularity is a first granularity, and the method includes: Determining that a first address in a memory where the data is to be stored has an offset of a second granularity compared to an integer multiple of the first granularity, where the second granularity is smaller than the first granularity; storing the data from the execution unit to an on-chip memory at the first granularity based on the first address at which the data is to be stored; and Storing the data from the on-chip memory to the memory space corresponding to the first address where the data is to be stored, Wherein, storing the data into the memory space corresponding to the first address where the data is to be stored comprises: storing the first part of the data into the memory at the second granularity and storing the second part of the data into the memory at the first granularity, The data includes bytes of data, M is the first granularity, the first part of the data includes the first k1 bytes of data and the last k2 bytes of data, k1+k2=M, M, N, k1 and k2 are all integers, the second part of the data includes the middle k1 bytes of data bytes of data.
6. The method according to claim 5, wherein: Storing the first portion of the data in the memory at the second granularity includes: loading a first portion of the data into an execution unit at the second granularity; A first portion of the data is stored from the execution unit to the memory at the second granularity.
7. The method according to claim 5, wherein: Storing the second portion of the data in the memory at the first granularity includes: loading a second portion of the data into an execution unit at the first granularity; A second portion of the data is stored from the execution unit to the memory at the first granularity.
8. The method of claim 5, wherein: Storing the second portion of the data in the memory at the first granularity includes: loading a second portion of the data into an execution unit at the second granularity; storing a second portion of the data from the execution unit to a register at the second granularity; loading a second portion of the data from the register to the execution unit at the first granularity; and A second portion of the data is stored from the execution unit to the memory at the first granularity.
9. An electronic device comprising: processor; as well as A memory, wherein a computer executable program is stored in the memory, and when the computer executable program is executed by the processor, the method for accessing memory data as described in any one of claims 1 to 4 and / or the method for storing data in the memory as described in any one of claims 5 to 8 are executed.
10. A computer-readable storage medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the method for accessing memory data according to any one of claims 1 to 4 and / or the method for storing data in a memory according to any one of claims 5 to 8 are implemented.
Citation Information
Patent Citations
Processing method and device for access exception of processor, storage medium and computing equipment
CN113535451A