Memory allocation and data processing method and device based on dual-Die cross storage
By alternately allocating and remapping memory address space for operator tasks in a dual-Die system, and using independent thread groups to bind and load data, the problem of insufficient data transmission speed across Die is solved, and the execution efficiency and performance of operator tasks are improved.
Patent Information
- Application Number
- CN202510855227.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the field of artificial intelligence, when a single Die computing resource cannot meet the operator's parallel computing needs, the cross-Die data transmission speed cannot keep up with the computing speed, resulting in the overall performance of the operator and the failure to perform the strongest hardware performance.
The memory allocation method based on double Die cross storage is adopted. By applying for continuous logical memory address space for the operator task, it is alternately allocated to Die0 and Die1, and remapping it to the physical memory address space. TG0 and TG1 are bound to Die0 and Die1, respectively, and data is loaded and calculated alternately to avoid cross-Die data communication.
Improve data storage efficiency, give full play to hardware performance, reduce data communication between Dies, and improve the overall performance of operator tasks.
Smart Images

Figure CN120371723A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence chip design, and particularly to a memory allocation and data processing method based on dual-Die cross storage. Background Art
[0002] In the field of artificial intelligence, an operator refers to a basic unit for performing specific computing tasks. A chip unit (Die) is an independent computing unit on a physical chip (including hardware resources such as an ALU and a cache). Each operator ultimately needs to be mapped to a specific circuit on the Die for execution.
[0003] In the related art, when the computing resources of a single die cannot meet the parallel computing requirements of operators (such as matrix multiplication, convolution operations, etc.), the operator needs to be split into multiple subtasks and allocated to multiple Dies for collaborative execution. Collaborative execution of multiple Dies usually requires cross-Die data transmission. Limited by the Die-to-Die interconnection bandwidth, the cross-Die data transmission speed cannot keep up with the computing speed, which will slow down the overall performance of the operator and cannot fully utilize the strongest performance of the hardware. Summary of the Invention
[0004] In view of this, this application provides a memory allocation and data processing method and device based on dual-Die cross storage, which can fully utilize the strongest performance of the hardware and improve the overall performance of operator tasks.
[0005] To solve the above technical problems, the technical solution of this application is implemented as follows: In one embodiment, a memory allocation and data processing method based on dual-Die cross storage is provided. The method includes: Apply for a continuous logical memory address space for the operator task, and alternately allocate the continuous logical memory address space to the first chip unit Die0 and the second chip unit Die1 in units of byte blocks with a preset number of bytes; Remap each logical memory address space allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1 respectively, and cross-store the data associated with the operator task into the physical memory address spaces of Die0 and Die1; Apply for thread resources for the operator task and divide them into a first thread group TG0 and a second thread group TG1, and bind TG0 and TG1 to Die0 and Die1 respectively; In the order of alternately allocating the continuous logical memory address space to Die0 and Die1, use TG0 and TG1 to alternately load byte blocks with a preset number of bytes from Die0 and Die1 respectively, calculate the data loaded each time, and store the calculation results.
[0006] In a possible implementation, the alternating allocation includes: Allocating a byte block of the Nth preset number of bytes in the continuous logical memory address space to Die0, and allocating a byte block of the (N + 1)th preset number of bytes in the continuous logical memory address space to Die1, where N is an even number.
[0007] In a possible implementation, the remapping is achieved by deploying an address converter PAGEN module between the system direct memory accessor SDMA and the physical memory address spaces of Die0 and Die1, or by deploying an address converter PAGEN module between the host aperture HA and the physical memory address spaces of Die0 and Die1.
[0008] In a possible implementation, the remapping includes: Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to Die0; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement processing as the physical memory address of Die0 corresponding to the logical memory address. Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to Die1; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement processing as the physical memory address of Die1 corresponding to the logical memory address.
[0009] In a possible implementation, the thread resources applied for the operator task are used to execute multiple subtasks obtained by splitting the operator task, where the subtasks to be executed by the thread resources allocated to TG0 are independent of the subtasks to be executed by the thread resources allocated to TG1.
[0010] In a possible implementation, the thread group binding satisfies: All memory access instructions of TG0 are only executed on Die0, and all memory access instructions of TG1 are only executed on Die1.
[0011] In a possible implementation, the alternating load includes: Using the starting address of the continuous logical memory address space as the base address, and repeatedly performing the following operations until the loading process of all cross-stored data is completed: TG0 performs a loading operation on Die0 to obtain data of a preset number of bytes, and the loading address is: the base address; TG1 performs a loading operation on Die1 to obtain data of a preset number of bytes, and the loading address is: the base address plus the preset number of bytes; Calculate the data loaded by TG0 and TG1, and store the calculation result; Increment the base address by 2 × the preset number of bytes.
[0012] In another embodiment, a memory allocation and data processing device based on dual-Die cross-storage is further provided. The device includes: An allocation unit for applying for a continuous logical memory address space for an operator task, and alternately allocating the continuous logical memory address space to a first chip unit Die0 and a second chip unit Die1 in units of byte blocks of a preset number of bytes; A mapping unit for remapping each logical memory address space allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1 respectively, and cross-storing the data associated with the operator task into the physical memory address spaces of Die0 and Die1; A binding unit for applying for thread resources for the operator task and dividing them into TG0 and TG1, and binding TG0 and TG1 to Die0 and Die1 respectively; A processing unit for alternately loading byte blocks of a preset number of bytes from Die0 and Die1 using TG0 and TG1 respectively in the order of alternately allocating the continuous logical memory address space to Die0 and Die1, calculating the data loaded each time, and storing the calculation result.
[0013] In a possible implementation manner, the allocation unit performs the alternate allocation including: Allocating the Nth byte block of a preset number of bytes in the continuous logical memory address space to Die0, and allocating the (N + 1)th byte block of a preset number of bytes in the continuous logical memory address space to Die1, where N is an even number.
[0014] In a possible implementation manner, the mapping unit realizes the remapping by deploying an address converter PAGEN module between the system direct memory accessor SDMA and the physical memory address spaces of Die0 and Die1, or realizes the remapping by deploying an address converter PAGEN module between the host aperture HA and the physical memory address spaces of Die0 and Die1.
[0015] In a possible implementation, the mapping unit executes the remapping including: Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address assigned to Die0; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement processing as the physical memory address of Die0 corresponding to the logical memory address; Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address assigned to Die1; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement processing as the physical memory address of Die1 corresponding to the logical memory address.
[0016] In a possible implementation, the thread resources applied for the operator task are used to execute multiple subtasks obtained by splitting the operator task, where the subtasks to be executed by the thread resources assigned to TG0 are independent of the subtasks to be executed by the thread resources assigned to TG1.
[0017] In a possible implementation, the thread group binding satisfies: All memory access instructions of TG0 are only executed on Die0, and all memory access instructions of TG1 are only executed on Die1.
[0018] In a possible implementation, the processing unit executes the alternating load including: Using the starting address of the continuous logical memory address space as the base address, and cyclically executing the following operations until the loading process of all cross-stored data is completed: TG0 executes a loading operation on Die0 to obtain data of a preset number of bytes, and the loading address is: the base address; TG1 executes a loading operation on Die1 to obtain data of a preset number of bytes, and the loading address is: the base address plus the preset number of bytes; Calculating the data loaded by TG0 and TG1, and storing the calculation result; Incrementing the base address by 2 × the preset number of bytes.
[0019] In another embodiment, an electronic device is further provided, including: A processor; A memory for storing executable instructions of the processor; Among them, the processor is configured to execute the executable instructions to implement the memory allocation and data processing method based on dual-Die cross storage as described above.
[0020] In another embodiment, a computer-readable storage medium is further provided. When at least one instruction in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device can implement the memory allocation and data processing method based on dual-Die cross storage as described above.
[0021] In another embodiment, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, it implements the memory allocation and data processing method based on dual-Die cross storage as described above.
[0022] As can be seen from the above technical solutions, in the above embodiments, in order to execute the operator task, first apply for a continuous logical memory address space and alternately allocate it to Die0 and Die1 in fixed-length blocks; then remap each logical memory address space allocated to Die0 to different physical memory address spaces of Die0, remap each logical memory address space allocated to Die1 to different physical memory address spaces of Die1, and cross-store the data associated with the operator task into the physical memory address spaces of Die0 and Die1, which can make the data evenly distributed in Die0 and Die1 instead of being concentrated, helping to improve the memory access efficiency of the data in Die0 and Die1, and thus giving full play to the strongest performance of the hardware; then, also apply for thread resources for the operator task and divide them into the first thread group TG0 and the second thread group TG1, and bind the TG0 and the TG1 to the Die0 and the Die1 respectively, so that TG0 can only access the data on Die0 and TG1 can only access the data on Die1, so as to ensure that when alternately loading byte blocks of a preset number of bytes from Die0 and Die1 in the order of alternately allocating the continuous logical memory address space to Die0 and Die1 and processing the loaded data, the data communication between Die0 and Die1 can be reduced, and thus the overall performance of the operator task can be improved. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1Schematic flowchart of the memory allocation and data processing method based on dual-Die cross storage in the embodiments of the present application; Figure 2 Example diagram of deploying the PAGEN module in the embodiments of the present application; Figure 3 Schematic flowchart of alternately allocating a continuous logical memory address space to Die0 and Die1 in the embodiments of the present application; Figure 4 Schematic diagram of the result of alternately allocating a continuous logical memory address space to Die0 and Die1 in the embodiments of the present application; Figure 5 Schematic diagram of the format of the logical memory address in the embodiments of the present application; Figure 6 Example diagram of the distribution of the physical memory address spaces of Die0 and Die1 corresponding to each logical memory address space after remapping in the embodiments of the present application; Figure 7 Schematic flowchart of remapping a logical memory address space of Die0 to a physical memory address space of Die0 in the embodiments of the present invention; Figure 8 Schematic diagram of the remapping of a logical memory address space of a Die to a physical memory address space of the Die in the embodiments of the present application; Figure 9 Schematic flowchart of alternately loading byte blocks of a preset number of bytes from Die0 and Die1 using TG0 and TG1 respectively in the embodiments of the present application; Figure 10 Schematic diagram of the structure of the memory allocation and data processing device based on dual-Die cross storage in the embodiments of the present application; Figure 11 Schematic diagram of the structure of the electronic device in the embodiments of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0026] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe the order or sequence of the objects. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] In the related art, complex operators usually need to be executed across Dies. The cross-die communication results in a longer data transmission path, significantly increasing the communication delay compared to the communication within a single die. And limited by the Die-to-Die interconnection bandwidth, the cross-die data transmission speed does not match the computing speed, thus affecting the execution efficiency of the operator task, reducing the overall performance of the operator task, and also unable to exert the strongest performance of the hardware.
[0028] Based on the above technical problems, the embodiments of the present application disclose a memory allocation and data processing method based on dual-Die cross storage. For an operator task, first apply for a continuous logical memory address space, and alternately allocate the continuous logical memory address space to two chip units: Die0 and Die1 in units of byte blocks with a preset number of bytes. Then, remap the respective logical memory address spaces allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1, and cross-store the data associated with the operator task into the physical memory address spaces of Die0 and Die1. Through remapping, the data is evenly arranged in Die0 and Die1, which helps to improve the memory access efficiency of the data in Die0 and Die1 and can fully exert the strongest performance of the hardware. Then, apply for thread resources for the operator task and divide them into TG0 and TG1, and bind TG0 and TG1 to Die0 and Die1 respectively, so that TG0 can only access the data on Die0 and TG1 can only access the data on Die1. Finally, in the order of alternately allocating the continuous logical memory address space to Die0 and Die1, use TG0 and TG1 to alternately load byte blocks with a preset number of bytes from Die0 and Die1 respectively, calculate the data loaded each time and store the calculation results. During the process of alternately loading data, since TG0 only accesses the data on Die0 and TG1 only accesses the data on Die1, cross-die data communication between Die0 and Die1 is avoided, thus improving the execution efficiency of the operator task and the overall performance of the operator task.
[0029] The following will, in conjunction with the accompanying drawings, elaborate on the specific implementation of the memory allocation and data processing method based on dual-Die cross storage in the embodiments of the present invention.
[0030] Refer to Figure 1 , Figure 1 , which is a schematic flowchart of the memory allocation and data processing method based on dual-Die cross storage in the embodiments of this application. As Figure 1 shown, the method specifically includes the following steps: Step 101: Apply for a continuous logical memory address space for the operator task. Taking byte blocks with a preset number of bytes as units, alternately allocate the continuous logical memory address space to the first chip unit (Die0) and the second chip unit (Die1); In the embodiments of this application, a logical memory address space with a fixed block length can be alternately allocated to Die0 and Die1. The fixed block length can be preset according to requirements. For example, the fixed block length is preset to 512 bytes, that is, the preset number of bytes is 512 bytes.
[0031] In the embodiments of this application, the operation of alternately allocating the continuous logical memory address space to Die0 and Die1 can be implemented by hardware.
[0032] Step 102: Remap the respective logical memory address spaces allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1, and cross-store the data associated with the operator task into the physical memory address spaces of Die0 and Die1; In the embodiments of this application, by remapping the respective logical memory address spaces allocated to Die0 to different physical memory address spaces of Die0, and remapping the respective logical memory address spaces allocated to Die1 to different physical memory address spaces of Die1, and then cross-storing the data associated with the operator task into the physical memory address spaces of Die0 and Die1, the arrangement of the data on Die0 and Die1 is no longer continuous, but scattered and uniform, so that data access will not be concentrated in adjacent physical memory address spaces. This helps the parallel reading of data, and thus helps improve the memory access efficiency of the data in Die0 and Die1, and can fully exert the strongest performance of the hardware.
[0033] In practical applications, the physical memory of the device can be directly operated by the CPU (i.e., the host side), or the physical memory of the device can be accessed through the GPU.
[0034] In a possible implementation manner of the present application, the physical memory of the device is accessed through the GPU. In this case, the remapping of the logical memory address space to the physical memory address space can be achieved by deploying a PAGEN module between the System Direct Memory Access (SDMA) and the physical memory of the device (such as the physical memories of Die0 and Die1). Figure 2 The figure shows the PAGEN module deployed between the SDMA and the physical memory of the device.
[0035] In another possible implementation manner of the present application, the physical memory of the device is accessed through the CPU. In this case, the remapping of the logical memory address space to the physical memory address space can be achieved by deploying a PAGEN module between the Host Aperture (HA) and the physical memory of the device (such as the physical memories of Die0 and Die1). Figure 2 The figure shows the PAGEN module deployed between the SDMA and the physical memory of the device.
[0036] The above PAGEN module is an operating system memory management module for implementing the address conversion from the logical memory address to the physical memory address.
[0037] Step 103: Apply for thread resources for the operator task and divide them into the first thread group (TG0) and the second thread group (TG1), and bind TG0 and TG1 to Die0 and Die1 respectively; In the embodiment of the present application, the operator task can be split into multiple subtasks, and different subtasks are used to complete different functions.
[0038] In the embodiment of the present application, the thread resources applied for the operator task are used to execute the multiple subtasks obtained by splitting the operator task. Among them, the subtasks to be executed by the thread resources divided into TG0 are independent of the subtasks to be executed by the thread resources divided into TG1. The independence of the two subtasks indicates that the two subtasks are executed independently of each other and do not need to communicate with each other during the execution. Therefore, the subtasks to be executed by the thread resources divided into TG0 are independent of the subtasks to be executed by the thread resources divided into TG1, which can avoid the communication between TG0 and TG1 during the execution, that is, there is no need to perform cross-Die data communication between Die0 and Die1, so the execution efficiency of the operator task can be improved and the overall performance of the operator task can be improved.
[0039] In the embodiments of the present application, binding a thread group to a chip unit (Die) can ensure that the thread group can only access the chip unit, that is, all memory access instructions of the thread group are only executed on the chip unit. Therefore, by binding TG0 and TG1 to Die0 and Die1 respectively, it is satisfied that all memory access instructions of TG0 are only executed on Die0, and all memory access instructions of TG1 are only executed on Die1. Through the binding of TG0 to Die0 and TG1 to Die1, cross-Die data communication between Die0 and Die1 can be further avoided, thereby further improving the execution efficiency of the operator task and the overall performance of the operator task.
[0040] Step 104: In the order of alternately allocating consecutive logical memory address spaces to Die0 and Die1, use TG0 and TG1 to alternately load byte blocks of a preset number of bytes from Die0 and Die1 respectively, calculate the data loaded each time, and store the calculation results.
[0041] In the embodiments of the present application, after cross-storing the data associated with the operator task into the physical memory address spaces of Die0 and Die1, and applying for thread resources for the operator task and dividing them into TG0 and TG1 and binding TG0 and TG1 to Die0 and Die1 respectively, TG0 and TG1 can be used to load data from the physical memory address spaces of Die0 and Die1 respectively and perform calculation operations and storage operations, where the alternate loading of data can start from the starting address of the above-mentioned consecutive logical memory address space.
[0042] In some embodiments, in step 101, alternately allocating the consecutive logical memory address spaces to Die0 and Die1, for the specific implementation, refer to Figure 3 .
[0043] Figure 3 FIG. is a schematic flowchart of alternately allocating the consecutive logical memory address spaces to Die0 and Die1 in the embodiments of the present application. The specific steps are as follows: Step 301: Allocate the Nth byte block of a preset number of bytes in the consecutive logical memory address space to Die0; N here is an even number.
[0044] Step 302: Allocate the (N + 1)th byte block of a preset number of bytes in the consecutive logical memory address space to Die1.
[0045] Taking the preset number of bytes as 512 as an example, alternately allocating the consecutive logical memory address spaces to Die0 and Die1, the results are as follows: The 0th, 2nd, 4th, 6th... 512-byte blocks in the consecutive logical memory address space are allocated to Die0.
[0046] The 1st, 3rd, 5th, 7th... 512-byte blocks in the continuous logical memory address space are allocated to Die1.
[0047] Figure 4 The schematic diagram showing the result of alternately allocating the continuous logical memory address space to Die0 and Die1 is as Figure 4 shown. Die0 and Die1 are alternately allocated byte blocks of a plurality of preset byte quantities. Any two byte blocks belonging to Die0 are not adjacent to each other, and any two byte blocks belonging to Die1 are not adjacent to each other.
[0048] In the embodiment of the present application, after the continuous logical memory address space is alternately allocated to Die0 and Die1, a logical memory address can be expressed in the Figure 5 format shown, as Figure 5 shown. This logical memory address includes a section offset (section_offset), a partition identifier (Partion_ID), a secondary cache identifier (L2ID), a chip unit identifier (Die_ID), and a page offset (Byte_offset). Among them, section_offset occupies bits 13 - 38; Partion_ID indicates a resource partition of a Die and occupies bit 12. L2ID indicates a secondary cache in a certain resource partition (indicated by Partion_ID) of a Die and occupies bits 10 - 11. Die_ID indicates a Die and occupies bit 9. Byte_offset occupies bits 0 - 8.
[0049] In the embodiment of the present application, after the continuous logical memory address space is alternately allocated to Die0 and Die1, the logical memory address spaces allocated to Die0 and Die1 can be further remapped to the physical memory address spaces of Die0 and Die1 respectively. Figure 6 The schematic diagram showing the distribution of the physical memory address spaces of Die0 and Die1 corresponding to the logical memory address spaces after remapping is as Figure 6 shown. The continuous logical memory address space is scattered after being remapped to the physical memory address space.
[0050] In some embodiments, step 102 remaps each logical memory address space assigned to Die0 and Die1 to the physical memory address spaces of Die0 and Die1 respectively, including: remapping each logical memory address space assigned to Die0 to different physical memory address spaces of Die0, and remapping each logical memory address space assigned to Die1 to different physical memory address spaces of Die1. Among them, the remapping process from each logical memory address space to the physical memory address space is the same. Taking the remapping of a logical memory address space of Die0 to a physical memory address space of Die0 as an example, the remapping process is described below. For specific implementation, refer to Figure 7 .
[0051] Figure 7 FIG. is a schematic flowchart of remapping a logical memory address space of Die0 to a physical memory address space of Die0 in an embodiment of the present invention. The specific steps are as follows: Step 701: Perform a hash operation on the partition identifier and the secondary cache identifier in the logical memory address assigned to Die0.
[0052] Step 702: Replace the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation results of the partition identifier and the secondary cache identifier respectively.
[0053] Step 703: Use the logical memory address after the above replacement process as the physical memory address of Die0 corresponding to this logical memory address.
[0054] Figure 8 FIG. Figure 8 shows a schematic diagram of remapping a logical memory address space of a Die to a physical memory address space of the Die. As shown in FIG. Figure 8 , only the hash operation is performed on the Partion_ID and L2ID in the logical memory address, so that only the partition identifier (Partion_ID) and the secondary cache identifier (L2ID) are different between the logical memory address and its corresponding physical memory address. Figure 8 As shown in FIG. Figure 8 , only the hash operation is performed on the Partion_ID and L2ID in the logical memory address, so that only the partition identifier (Partion_ID) and the secondary cache identifier (L2ID) are different between the logical memory address and its corresponding physical memory address.
[0055] In some embodiments, in step 104, TG0 and TG1 are used to alternately load byte blocks of a preset number of bytes from Die0 and Die1 respectively. For specific implementation, refer to Figure 9 .
[0056] Figure 9 FIG. is a schematic flowchart of alternately loading byte blocks of a preset number of bytes from Die0 and Die1 using TG0 and TG1 in an embodiment of the present application. The specific steps include the following: Step 901: Use the start address of a continuous logical memory address space as the base address, and loop through the following steps 902 - 905 until all cross - stored data has been loaded and processed.
[0057] Step 902: TG0 performs a load operation on Die0 to obtain data of a preset number of bytes. The load address is: the base address.
[0058] In the embodiments of the present application, the base address is a predefined logical memory address. When performing data loading, it is necessary to first remap this logical memory address to a physical memory address, and then load data of a preset number of bytes starting from this physical memory address. The preset number of bytes can be, for example, 512 bytes.
[0059] Step 903: TG1 performs a load operation on Die1 to obtain data of a preset number of bytes. The load address is: the base address plus the preset number of bytes.
[0060] In the embodiments of the present application, the base address plus the preset number of bytes corresponds to a logical memory address. When performing data loading, this logical memory address can be remapped to a physical memory address first, and then data of a preset number of bytes starting from this physical memory address can be loaded.
[0061] Step 904: Calculate the data loaded by TG0 and TG1, and store the calculation result. Step 905: Increment the base address by 2 × the preset number of bytes (i.e., base address = base address + 2 × the preset number of bytes).
[0062] In the embodiments of the present application, the base address can increase as the number of loop operations for loading cross - stored data increases. Taking the preset number of bytes as 512 bytes as an example, each time a loop operation is performed, the base address is incremented by 1024 bytes (i.e., base address = base address + 1024 bytes).
[0063] In this step, update the base address, and then return to step 902 to perform a new round of data alternating loading until all cross - stored data has been loaded and processed (including the calculation and storage of the loaded data).
[0064] As can be seen from the above embodiments, in the present application, by allocating a continuous logical memory address space for the operator task, taking byte blocks of a preset number of bytes as a unit, the continuous logical memory address space is alternately allocated to two chip units: Die0 and Die1; then, each logical memory address space allocated to Die0 and Die1 is remapped to the physical memory address spaces of Die0 and Die1 respectively, and the data associated with the operator task is cross-stored in the physical memory address spaces of Die0 and Die1, so that the data is evenly arranged in Die0 and Die1, which helps to improve the memory access efficiency of the data in Die0 and Die1 and can give full play to the strongest performance of the hardware; then, thread resources are applied for the operator task and divided into TG0 and TG1, and TG0 and TG1 are respectively bound to Die0 and Die1, so that TG0 can only access the data on Die0 and TG1 can only access the data on Die1; finally, in the process of alternately loading byte blocks of a preset number of bytes from Die0 and Die1 by using TG0 and TG1 respectively in the order of alternately allocating the continuous logical memory address space to Die0 and Die1, calculating the data loaded each time and storing the calculation results, since TG0 only accesses the data on Die0 and TG1 only accesses the data on Die1, cross-Die data communication between Die0 and Die1 is avoided, thus improving the execution efficiency of the operator task and the overall performance of the operator task.
[0065] Any combination of the above all optional technical solutions can form optional embodiments of the present disclosure, which will not be elaborated herein one by one.
[0066] Based on the same inventive concept, an embodiment of the present application also provides a memory allocation and data processing device based on cross-storage of dual Dies. Refer to Figure 10 , Figure 10 is a schematic structural diagram of the memory allocation and data processing device based on cross-storage of dual Dies according to an embodiment of the present application. As Figure 10 shown, the device includes: An allocation unit 1001, configured to apply for a continuous logical memory address space for the operator task, and alternately allocate the continuous logical memory address space to a first chip unit Die0 and a second chip unit Die1 in units of byte blocks of a preset number of bytes; A mapping unit 1002, configured to respectively remap each logical memory address space allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1, and cross-store the data associated with the operator task in the physical memory address spaces of Die0 and Die1; The binding unit 1003 is used to apply for thread resources for the operator task and divide them into the TG0 and the TG1, and bind the TG0 and the TG1 to the Die0 and the Die1 respectively; The processing unit 1004 is used to alternately load byte blocks of a preset number of bytes from the Die0 and the Die1 by using the TG0 and the TG1 respectively according to the order of alternately allocating the continuous logical memory address space to the Die0 and the Die1, calculate the data loaded each time, and store the calculation results.
[0067] In some embodiments, the allocation unit 1001 performing the alternate allocation includes: Allocating the Nth byte block of a preset number of bytes in the continuous logical memory address space to the Die0, and allocating the (N + 1)th byte block of a preset number of bytes in the continuous logical memory address space to the Die1, where N is an even number.
[0068] In some embodiments, the mapping unit 1002 realizes the remapping by deploying an address converter PAGEN module between the system direct memory accessor SDMA and the physical memory address spaces of the Die0 and the Die1, or realizes the remapping by deploying an address converter PAGEN module between the host aperture HA and the physical memory address spaces of the Die0 and the Die1.
[0069] In some embodiments, the mapping unit 1002 performing the remapping includes: Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to the Die0; replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation results of the partition identifier and the secondary cache identifier respectively; using the logical memory address after the above replacement process as the physical memory address of the Die0 corresponding to the logical memory address; Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to the Die1; replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation results of the partition identifier and the secondary cache identifier respectively; using the logical memory address after the above replacement process as the physical memory address of the Die1 corresponding to the logical memory address.
[0070] In some embodiments, the thread resources applied for the operator task are used to execute multiple subtasks obtained by splitting the operator task, where the subtasks to be executed by the thread resources divided into the TG0 are independent of the subtasks to be executed by the thread resources divided into the TG1.
[0071] In some embodiments, the thread group binding satisfies: All memory access instructions of the TG0 are only executed on the Die0, and all memory access instructions of the TG1 are only executed on the Die1.
[0072] In some embodiments, the processing unit 1004 executes the alternating load including: Taking the starting address of the consecutive logical memory address space as the base address, and looping through the following operations until the loading process of all cross-stored data is completed: The TG0 executes a load operation on the Die0 to obtain data of a preset number of bytes, and the load address is: the base address; The TG1 executes a load operation on the Die1 to obtain data of a preset number of bytes, and the load address is: the base address plus the preset number of bytes; Calculating the data loaded by the TG0 and the TG1, and storing the calculation result; Incrementing the base address by 2 × the preset number of bytes.
[0073] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. In some embodiments, the electronic device is a test platform. The electronic device 1100 may vary greatly due to configuration or performance differences, and may include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102. Among them, at least one program code is stored in the memory 1102, and the at least one program code is loaded and executed by the processor 1101 to implement the methods for testing the processor provided in the above various embodiments. Of course, the electronic device 1100 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The electronic device 1100 may also include other components for implementing device functions, which will not be elaborated here.
[0074] In an exemplary embodiment, there is also provided a computer-readable storage medium including at least one instruction, such as a memory including at least one instruction. The at least one instruction can be executed by a processor in a computer device to complete the method for memory allocation and data processing based on dual-Die cross-storage in the above embodiments.
[0075] Optionally, the above computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may include ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.
[0076] In an exemplary embodiment, there is also provided a computer program product including a computer program, which when executed by a processor implements the memory allocation and data processing method based on dual-Die cross storage provided in each of the above embodiments.
[0077] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0078] The flowcharts and block diagrams in the drawings of this application illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments disclosed in this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in the order marked in different drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0079] Those skilled in the art will understand that the features described in the various embodiments and / or claims disclosed in this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, without departing from the spirit and teachings of this application, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, and all such combinations and / or combinations fall within the scope disclosed in this application.
[0080] In this article, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention, and is not used to limit this application. For those skilled in the art, according to the ideas, spirit and principles of the present invention, changes can be made in the specific implementation manners and application scopes, and any modifications, equivalent replacements, improvements, etc. made by them shall be included within the scope protected by this application.
Claims
1. A memory allocation and data processing method based on dual-Die cross storage, characterized in that The method includes: Applying for a continuous logical memory address space for the operator task, and alternately allocating the continuous logical memory address space to the first chip unit Die0 and the second chip unit Die1 in units of byte blocks of a preset number of bytes; Remapping each logical memory address space allocated to Die0 and Die1 to the physical memory address spaces of Die0 and Die1 respectively, and cross-storing the data associated with the operator task into the physical memory address spaces of Die0 and Die1; Applying for thread resources for the operator task and dividing them into a first thread group TG0 and a second thread group TG1, and binding TG0 and TG1 to Die0 and Die1 respectively; In the order of alternately allocating the continuous logical memory address space to Die0 and Die1, using TG0 and TG1 to alternately load byte blocks of a preset number of bytes from Die0 and Die1 respectively, calculating the data loaded each time and storing the calculation results.
2. The method according to claim 1, wherein The alternate allocation includes: Allocating the Nth byte block of the preset number of bytes in the continuous logical memory address space to Die0, and allocating the (N + 1)th byte block of the preset number of bytes in the continuous logical memory address space to Die1, where N is an even number.
3. The method according to claim 1, wherein The remapping is achieved by deploying an address converter PAGEN module between the system direct memory accessor SDMA and the physical memory address spaces of Die0 and Die1; Or, the remapping is achieved by deploying an address converter PAGEN module between the host aperture HA and the physical memory address spaces of Die0 and Die1.
4. The method according to claim 1, characterized in that, The remapping includes: Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to Die0; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement process as the physical memory address of Die0 corresponding to the logical memory address; Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to Die1; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement process as the physical memory address of Die1 corresponding to the logical memory address.
5. The method according to claim 1, characterized in that, The thread resources applied for the operator task are used to execute multiple subtasks obtained by splitting the operator task. Among them, the subtasks to be executed by the thread resources divided into TG0 are independent of the subtasks to be executed by the thread resources divided into TG1.
6. The method according to claim 1, wherein The binding satisfies: All memory access instructions of TG0 are only executed on Die0, and all memory access instructions of TG1 are only executed on Die1.
7. The method according to claim 1, wherein The alternating loading includes: Taking the starting address of the continuous logical memory address space as the base address, and looping through the following operations until all cross-stored data is loaded and processed: The TG0 performs a loading operation on the Die0 to obtain data of a preset number of bytes, and the loading address is: the base address; The TG1 performs a loading operation on the Die1 to obtain data of a preset number of bytes, and the loading address is: the base address plus the preset number of bytes; Calculate the data loaded by the TG0 and the TG1, and store the calculation result; Increment the base address by 2 × the preset number of bytes.
8. A memory allocation and data processing device based on dual-Die cross storage, characterized in that, The device includes: An allocation unit, configured to apply for a continuous logical memory address space for an operator task, and alternately allocate the continuous logical memory address space to a first chip unit Die0 and a second chip unit Die1 in units of byte blocks of a preset number of bytes; A mapping unit, configured to remap each logical memory address space allocated to the Die0 and the Die1 to the physical memory address spaces of the Die0 and the Die1 respectively, and cross-store the data associated with the operator task into the physical memory address spaces of the Die0 and the Die1; A binding unit, configured to apply for thread resources for the operator task and divide them into a first thread group TG0 and a second thread group TG1, and bind the TG0 and the TG1 to the Die0 and the Die1 respectively; A processing unit, configured to alternately load byte blocks of a preset number of bytes from the Die0 and the Die1 using the TG0 and the TG1 in the order of alternately allocating the continuous logical memory address space to the Die0 and the Die1, calculate the data loaded each time, and store the calculation result.
9. The device according to claim 8, characterized in that, The allocation unit, when performing the alternating allocation, includes: Allocating the Nth byte block of the preset number of bytes in the continuous logical memory address space to the Die0, and allocating the (N + 1)th byte block of the preset number of bytes in the continuous logical memory address space to the Die1, where N is an even number.
10. The device according to claim 8, wherein The mapping unit realizes the remapping by deploying an address converter PAGEN module between the system direct memory accessor SDMA and the physical memory address spaces of the Die0 and the Die1, or realizes the remapping by deploying an address converter PAGEN module between the host aperture HA and the physical memory address spaces of the Die0 and the Die1.
11. The device according to claim 8, characterized in that, The mapping unit, when performing the remapping, includes: Performing a hash operation on the partition identifier and the secondary cache identifier in each logical memory address allocated to the Die0; respectively replacing the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation result of the partition identifier and the hash operation result of the secondary cache identifier; using the logical memory address after the above replacement processing as the physical memory address of the Die0 corresponding to the logical memory address. Perform a hash operation on the partition identifier and the secondary cache identifier in each logical memory address assigned to the Die1; replace the original partition identifier and the original secondary cache identifier in the logical memory address with the hash operation results of the partition identifier and the secondary cache identifier respectively; use the logical memory address after the above replacement process as the physical memory address of the Die1 corresponding to the logical memory address.
12. The device according to claim 8, characterized in that, The thread resources applied for the operator task are used to execute multiple subtasks obtained by splitting the operator task. Among them, the subtasks to be executed by the thread resources allocated to the TG0 are independent of the subtasks to be executed by the thread resources allocated to the TG1.
13. The device according to claim 8, wherein The binding satisfies: All memory access instructions of the TG0 are only executed on the Die0, and all memory access instructions of the TG1 are only executed on the Die1.
14. The device according to claim 8, characterized in that, The processing unit, when executing the alternating load, includes: Use the starting address of the continuous logical memory address space as the base address, and loop through the following operations until the loading process of all cross-stored data is completed: The TG0 executes a loading operation on the Die0 to obtain data of a preset number of bytes, and the loading address is: the base address; The TG1 executes a loading operation on the Die1 to obtain data of a preset number of bytes, and the loading address is: the base address plus the preset number of bytes; Calculate the data loaded by the TG0 and the TG1, and store the calculation result; Increment the base address by 2 × the preset number of bytes.
15. An electronic device, characterized in that, Includes: A processor; A memory for storing executable instructions of the processor; Among them, the processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, When at least one instruction in the computer-readable storage medium is executed by the processor of the electronic device, the electronic device can implement the method according to any one of claims 1 to 7.
17. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for dynamically allocating physical address, and electronic device
CN109542798A
Memory device and managing method of memory device
KR1020140133800A
Techniques for efficiently partitioning memory
US10909033B1
Inter-die coherence processing system, method and apparatus, device, and medium
WO2025091909A1
Cited By
Chip architecture capable of selecting cache-free mode
CN122111950A
A chip architecture with optional cacheless mode
CN122111950B