GPGPU (General Purpose Graphics Processing Unit) storage optimization method and device, equipment and storage medium
By building a hybrid storage area of MRAM and SRAM in GPGPU and dynamically adjusting the data storage location, the traditional storage architecture solves the problems of slow reading and writing speed, high energy consumption and easy data loss in GPGPU data storage, and realizes efficient and low-power data storage.
Patent Information
- Application Number
- CN202510226186.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In GPGPU data storage, traditional DRAM has slow reading and writing speed, high energy consumption and easy data loss. Although SRAM read and write speed is fast, it is expensive and has low integration, making it difficult to deploy on a large scale.
By building a storage area including MRAM and SRAM, using SRAM for high access frequency data storage, and using MRAM for low access frequency data storage, and dynamically adjusting the storage location of data between MRAM and SRAM.
It improves data storage efficiency, reduces system power consumption, avoids data loss, and makes full use of the advantages of fast read and write speed of SRAM.
Smart Images

Figure CN120104061A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data access technology, and in particular to a GPGPU storage optimization method, device, equipment and storage medium. Background Art
[0002] In modern high-performance computing scenarios, GPGPU (General-purpose computing on graphics processing units) is increasingly widely used, and its requirements for data storage and processing capabilities are becoming increasingly stringent. Traditional storage architectures have gradually exposed many drawbacks when dealing with GPGPU's large-scale, high-concurrency data reading and writing needs.
[0003] At present, when GPGPU is used for data storage, the commonly used DRAM (Dynamic Random Access Memory) memory has a certain storage capacity, but has relatively slow read and write speeds, high energy consumption, and easy data loss. For example, in the training and inference process of deep learning, frequent data exchange will make the bandwidth of DRAM a performance bottleneck, and its regular refresh operation consumes a lot of electricity; on the other hand, SRAM (Static Random-Access Memory) has fast read and write speeds, but it is expensive and has low integration, making it difficult to deploy on a large scale in GPGPU as the main storage medium. Therefore, how to perform GPGPU data storage with low energy consumption, high data security and fast data read and write speeds has become a technical problem that needs to be solved. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a GPGPU storage optimization method, device, equipment and storage medium, which can improve storage efficiency by using SRAM storage area for data storage, and reduce system power consumption by using MRAM for data storage. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a GPGPU storage optimization method, comprising:
[0006] Obtain the first target task data issued by the GPGPU computing core, and determine the target ratio corresponding to each of the first target task data based on the preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each of the first target task data, the memory access count information is the number of read operations experienced after the memory access address is written into the preset address index entry, the clock count information is the number of clock cycles experienced after the memory access address is written into the preset address index entry, and the target ratio is the ratio of the memory access count information to the clock count information;
[0007] Determine whether the memory access operation corresponding to each of the first target task data is a write operation, and if the memory access operation corresponding to each of the first target task data is a write operation, store each of the first target task data in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area;
[0008] After obtaining a new memory access operation, the storage position of the first target task data in the target storage space is periodically and dynamically adjusted according to the target ratio of the corresponding second target task data and the storage status of the target storage space, so that the second target task data corresponding to the new memory access operation is stored in the corresponding position of the target storage space.
[0009] Optionally, the periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes:
[0010] If the data storage capacity of the SRAM storage area has reached the upper limit, and the data storage capacity of the MRAM storage area has not reached the upper limit, the preset address index entries are sorted in ascending order according to the target ratios, and it is determined whether the first target task data with the smallest target ratio meets the preset hotspot data determination standard within the current preset time period;
[0011] If the first target task data with the smallest target ratio meets the preset hot data determination standard within the current preset time period, a new memory access address corresponding to the new memory access operation is stored in the preset address index entry, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the MRAM storage area; wherein the configurable number of entries of the preset address index entry is the same as the maximum task storage number of the hybrid cache area;
[0012] If the first target task data with the smallest target ratio does not meet the preset hotspot data determination standard within the current preset time period, the first target task data with the smallest target ratio is migrated from the SRAM storage area to the MRAM storage area, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the SRAM storage area.
[0013] Optionally, the periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes:
[0014] If the data storage amounts of the SRAM storage area and the MRAM storage area have both reached upper limits, determining whether a memory access address corresponding to a new memory access operation is in the preset address index entry;
[0015] If the memory access address corresponding to the new memory access operation is in the preset address index entry, determining whether the new memory access operation is a read operation, and if the new memory access operation is a read operation, adjusting the storage position of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation;
[0016] If the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area, the second target task data is deleted from the hybrid cache area, and the storage location corresponding to the second target task data in the preset address index entry is recorded as an invalid cache.
[0017] Optionally, adjusting the storage location of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation includes:
[0018] If the second target task data corresponding to the new memory access operation is stored in the SRAM storage area, and the target ratio corresponding to the second target task data is less than the maximum target ratio corresponding to each of the first target task data in the MRAM storage area, then the storage area of the second target task data is swapped with the storage area of the first target task data corresponding to the maximum target ratio;
[0019] If the second target task data corresponding to the new memory access operation is stored in the MRAM storage area, and the target ratio corresponding to the second target task data is greater than the minimum target ratio corresponding to each of the first target task data in the SRAM storage area, the storage area of the second target task data is swapped with the storage area of the first target task data corresponding to the minimum target ratio.
[0020] Optionally, the periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes:
[0021] If the memory access address corresponding to the new memory access operation is not in the preset address index entry, determine whether the new memory access operation is a read operation; if the new memory access operation is a read operation, determine whether there is an entry in the preset address index entry whose storage location is an invalid cache;
[0022] If there is no entry whose storage location is an invalid cache in the preset address index entry, reading the second target task data corresponding to the new memory access operation from the MRAM main storage area;
[0023] If there is an entry whose storage location is an invalid cache in the preset address index entry, it is determined whether the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area; if the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area, the second target task data corresponding to the new memory access operation is migrated from the MRAM main storage area to the SRAM storage area;
[0024] If the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area.
[0025] Optionally, after storing each target task data in a corresponding position of the target storage space, the method further includes:
[0026] The data to be used is predicted from each target task data according to the data type and data access frequency of each target task data, and the data to be used is loaded from the MRAM main storage area to the hybrid cache area so that the GPGPU computing core processes the data to be used.
[0027] Optionally, the GPGPU storage optimization method further includes: expanding an initial instruction set corresponding to the GPGPU computing core to obtain a corresponding target instruction set, so as to use the target instruction set to control data storage operations corresponding to the MRAM main storage area and the MRAM storage area.
[0028] In a second aspect, the present application provides a GPGPU storage optimization device, comprising:
[0029] A data acquisition module, used for acquiring first target task data issued by a GPGPU computing core, and determining a target ratio corresponding to each first target task data based on a preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each first target task data, the memory access count information is the number of read operations experienced by the memory access address after the preset address index entry is written, the clock count information is the number of clock cycles experienced by the memory access address after the preset address index entry is written, and the target ratio is the ratio of the memory access count information to the clock count information;
[0030] a data storage module, used to determine whether the memory access operation corresponding to each of the first target task data is a write operation, and if the memory access operation corresponding to each of the first target task data is a write operation, then store each of the first target task data in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area;
[0031] A storage position adjustment module is used to periodically and dynamically adjust the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage status of the target storage space after obtaining a new memory access operation, so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space.
[0032] In a third aspect, the present application provides an electronic device, including:
[0033] Memory, used to store computer programs;
[0034] A processor is used to execute the computer program to implement the aforementioned GPGPU storage optimization method.
[0035] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned GPGPU storage optimization method when executed by a processor.
[0036] In the present application, the first target task data issued by the GPGPU computing core is first obtained, and the target ratio corresponding to each of the first target task data is determined based on the preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each of the first target task data, the memory access count information is the number of read operations experienced after the memory access address is written to the preset address index entry, the clock count information is the number of clock cycles experienced after the memory access address is written to the preset address index entry, and the target ratio is the ratio of the memory access count information to the clock count information, then it is determined whether the memory access operation corresponding to each of the first target task data is a write operation, if each If the memory access operation corresponding to the first target task data is a write operation, each first target task data is stored in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area. After obtaining a new memory access operation, the storage position of the first target task data in the target storage space is periodically and dynamically adjusted according to the target ratio of the corresponding second target task data and the storage state of the target storage space, so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space. It can be seen that the present application optimizes the storage structure of GPGPU and constructs a storage area including MRAM and SRAM, so that the MRAM storage area and the SRAM storage area can be used to store data with different access frequencies respectively, which solves the problem that SRAM has low integration and cannot be applied to GPGPU, and fully utilizes the advantages of SRAM's fast reading and writing speed to improve data storage efficiency; by using MRAM and SRAM for data storage, the problem of data loss caused by using DRAM for data storage is avoided; by storing data with low data access frequency in MRAM, frequent exchange of data between the main storage area and the cache area is avoided, thereby reducing energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 A flow chart of a GPGPU storage optimization method disclosed in this application;
[0039] Figure 2 A schematic diagram of a GPGPU storage structure disclosed in this application;
[0040] Figure 3 A schematic diagram of an address index entry structure disclosed in this application;
[0041] Figure 4 A schematic diagram of a specific GPGPU storage optimization method disclosed in this application;
[0042] Figure 5 A flow chart of a data pre-fetching method disclosed in this application;
[0043] Figure 6 A schematic diagram of the structure of a GPGPU storage optimization device disclosed in this application;
[0044] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] At present, when GPGPU performs data storage, there are problems such as relatively slow reading and writing speed, high energy consumption, and easy data loss. To this end, the present application provides a GPGPU storage optimization method, which improves storage efficiency by constructing a storage area including MRAM and SRAM, and uses the SRAM storage area for high-access frequency data storage, and reduces system power consumption by using MRAM for low-access frequency data storage.
[0047] See also Figure 1 As shown, an embodiment of the present invention discloses a GPGPU storage optimization method, comprising:
[0048] Step S11, obtain the first target task data issued by the GPGPU computing core, and determine the target ratio corresponding to each of the first target task data based on the preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each of the first target task data, the memory access count information is the number of read operations experienced after the memory access address is written to the preset address index entry, the clock count information is the number of clock cycles experienced after the memory access address is written to the preset address index entry, and the target ratio is the ratio of the memory access count information to the clock count information.
[0049] This embodiment provides a new GPGPU storage architecture. The specific architecture is as follows: Figure 2 As shown, it is mainly composed of an MRAM (Magnetoresistive Random Access Memory, i.e., a non-volatile magnetic random access memory) main storage area, a hybrid cache layer, an intelligent scheduler, and a GPGPU computing core. The GPGPU computing core is the computing upper layer of the storage architecture of the solution of the present invention, and is responsible for completing specific computing and storage instructions. In order to ensure the upward compatibility of the solution of the present invention with the original GPGPU computing core, the present invention does not change the external interface of the original GPGPU computing core, but achieves performance optimization for the MRAM storage architecture by expanding the GPGPU computing core instruction set and adding GPGPU functional modules.
[0050] The intelligent scheduler is the core control unit of the storage architecture of the present invention. By setting an instruction detection unit and a data detection unit in the GPGPU computing core, maintaining the historical data access track internally, and setting a storage state detection unit in the hybrid cache area and the MRAM main storage area, the data flow detection between the various levels of storage is completed. Through the internal intelligent algorithm, the data allocation and transmission strategy between the MRAM main storage area, the hybrid cache layer and the computing core is dynamically adjusted according to factors such as the GPGPU computing task type, data access frequency and time locality.
[0051] The hybrid cache layer is a transition cache module in the storage architecture of the present invention, located between the MRAM main storage area and the GPGPU computing core, integrating the advantages of SRAM and MRAM. Among them, the SRAM storage area is used to cache hot spot data that is frequently accessed during GPGPU computing, and its high-speed read and write characteristics ensure that the computing core can quickly obtain data. The MRAM storage area is used to store warm spot data that may be accessed in the near future and some data that has relatively low requirements for read and write speeds but needs to be stored for a long time. Its non-volatility and large storage capacity are used to reduce the frequent exchange of data between the main storage area and the cache area, thereby reducing energy consumption. Among them, the hot spot data, warm spot data, and cold spot data in this embodiment are determined by the target ratio corresponding to the task data in the current cycle; the preset address index entries in this embodiment are such as Figure 3 As shown, it includes memory access address information, memory access count information and clock count information. The memory access address information includes the memory access address corresponding to each first target task data. The memory access count information is the number of read operations experienced after the memory access address is written to the preset address index entry. The clock count information is the number of clock cycles experienced after the memory access address is written to the preset address index entry. The above-mentioned target ratio is the ratio of the memory access count information to the clock count information.
[0052] The MRAM main storage area is the core storage module in the storage architecture of the solution of the present invention. It is constructed with multi-module, multi-layer, three-dimensionally stacked MRAM chips, making full use of the storage performance advantages of MRAM to achieve high storage density. At the same time, in the MRAM main storage area of the solution of the present invention, data redundancy backup protection is achieved through RAID5 (a data storage technology) technology to ensure data security. By constructing address index entries, the data access frequency of each task data in the current cycle can be determined based on the total number of calls to the task data and the number of cycles after the task data enters the address index entry, so that the task data can be stored in the corresponding position of the target storage space according to the data access frequency of the task data.
[0053] Step S12, determining whether the memory access operation corresponding to each of the first target task data is a write operation; if the memory access operation corresponding to each of the first target task data is a write operation, storing each of the first target task data in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area.
[0054] The cache management and replacement strategy in this embodiment is centered on the intelligent scheduler and mainly consists of a cache data classification mechanism, a read-write judgment mechanism, and an eviction judgment mechanism. The implementation process is as follows: Figure 4As shown in the figure: the GPGPU computing core sends the instruction stream and data stream information of the computing task, and then the cache data classification mechanism completes the data classification identification and divides the data into hot spot data, warm spot data and cold spot data; then according to the judgment of the hot spot data, warm spot data and cold spot data, the read and write judgment mechanism is entered, and according to the storage status of the MRAM main storage area / hybrid cache area, the data is written / read to the SRAM storage area in the hybrid cache area, the MRAM storage area in the hybrid cache area and the MRAM main storage area. At the same time, if the corresponding cache area is full, it is also necessary to judge the storage location of the evicted data according to the eviction judgment mechanism.
[0055] The above process is completed by the address index entry of configurable depth n maintained in the intelligent scheduler. The content of the entry mainly includes memory access address information, memory access count information, clock count information and storage location information. The configurable depth n is equal to the storable address depth of the SRAM and MRAM storage areas in the hybrid cache area. The memory access address information includes the memory access address sent by the GPGPU core to the intelligent scheduler, and the information maintained by this entry will be continuously refreshed; the memory access count information is the number of read operations experienced after the memory access address is written to the address index entry, and the location information will be cleared with the write operation to the address and the address change; the clock count information is the number of clock cycles experienced after the memory access address is written to the address index entry, and the upper limit is the number of clock cycles set by the system, where the number of clock cycles set by the system is a configurable option and can be dynamically configured according to the type of executed software, and the location information will be cleared with the write operation to the address and the address change; the storage location information is the physical storage location where the address is stored, marked with 00 / 01 / 10 / 11, which respectively represent the SRAM storage area in the hybrid cache area, the MRAM storage area in the hybrid cache area, the MRAM main storage area and the invalidated cache.
[0056] Specifically, when the GPGPU starts running, the address index entry information is empty. At this time, the memory access address is sequentially stored in the memory access address information in the address index entry, and the clock count information experienced by the entry is counted starting from this time as zero. At the same time, since there is no data written to the GPGPU cache area at this time, the data, that is, the first target task data, is marked as hot data, and then the storage location is marked as 00, and the address data is extracted from the MRAM main storage area and stored in the SRAM storage area in the hybrid cache area.
[0057] When the next memory access operation comes, the memory access address will be judged first. When the memory access address is inconsistent with the address information maintained in the existing address index entry, the memory access address will be stored in the direction of increasing the depth of the address index entry, and the subsequent clock cycles will be counted and written into the clock count entry. Since the address index entry has not yet increased to the upper limit of the SRAM capacity in the hybrid cache area at this time, the data is also marked as hot data, the storage location is marked as 00, and the address data is extracted from the MRAM main storage area and stored in the SRAM storage area in the hybrid cache area; when the memory access address is consistent with the address information maintained in the existing address index entry, the memory access address read and write judgment will be performed. If it is a read data operation, the memory access count value will be +1, and the data will be read from the SRAM storage area in the hybrid cache area according to the address; if it is a write data operation, the memory access count and clock count value will be cleared, and the data will be written to the corresponding address of the SRAM storage area in the hybrid cache area and the corresponding address of the MRAM main storage area. By first saving the task data to the SRAM storage area, the access efficiency of data in GPGPU is guaranteed.
[0058] Step S13: After obtaining a new memory access operation, the storage position of the first target task data in the target storage space is periodically and dynamically adjusted according to the target ratio of the corresponding second target task data and the storage status of the target storage space, so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space.
[0059] As new memory access operations continue to arrive, the address index entry depth will increase to the upper limit of the SRAM capacity in the hybrid cache area, indicating that the SRAM storage area in the hybrid cache area of the GPGPU cache system is full of stored data. Therefore, it is necessary to update the address index entry and replace the GPGPU cache system. Accordingly, the above process of periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space may specifically include: if S If the data storage capacity of the RAM storage area has reached the upper limit, and the data storage capacity of the MRAM storage area has not reached the upper limit, the preset address index entries are sorted in order from small to large according to the target ratio, and it is determined whether the first target task data with the smallest target ratio meets the preset hot data judgment standard within the current preset time period; if the first target task data with the smallest target ratio meets the preset hot data judgment standard within the current preset time period, the new memory access address corresponding to the new memory access operation is stored in the preset address index entry, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the MRAM storage area; if the target ratio If the smallest first target task data does not meet the preset hot data judgment standard within the current preset time period, the first target task data with the smallest target ratio is migrated from the SRAM storage area to the MRAM storage area, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the SRAM storage area; that is, the memory access count / clock count ratio in the existing address index entry is arranged from low to high, and it is determined whether its lowest memory access count / clock count ratio meets the hot data standard, that is, the above-mentioned preset hot data judgment standard: if it meets the new memory access address, the new memory access address is stored in the subsequent address index entry, and the subsequent experience is counted. The clock cycle is written into the clock count entry, the storage position is marked as 01, and the address data is extracted from the MRAM main storage area and stored in the MRAM storage area in the hybrid cache area; if the hot data standard is not met, the storage address in the address index entry corresponding to the lowest memory access count / clock count ratio is marked as 01, and the address data is migrated from the SRAM storage area in the hybrid cache area to the MRAM storage area in the hybrid cache area, and the storage position of the new memory access address is marked as 00, and the data corresponding to the new memory access address is extracted from the MRAM main storage area and stored in the reserved position of the SRAM storage area in the hybrid cache area.
[0060] When new memory access operations continue to arrive, and the address index entry depth increases to the upper limit of the total capacity of SRAM and MRAM in the hybrid cache area, it means that the data stored in the SRAM storage area and the MRAM storage area in the hybrid cache area of the GPGPU cache system are full at this time, and then the address index entry is updated and the management strategy of the GPGPU cache system enters the normal operation stage; accordingly, the above-mentioned periodic dynamic adjustment of the storage position of the first target task data in the target storage space according to the target ratio of the second target task data and the storage status of the target storage space, so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space, including: if the data storage capacity of the SRAM storage area and the MRAM storage area has reached the upper limit, then determine whether the memory access address corresponding to the new memory access operation is in the preset address index entry; if the memory access address corresponding to the new memory access operation is in the preset address index entry, then determine whether the new memory access operation is a read operation, and if the new memory access operation is a read operation, then determine whether the new memory access operation is a read operation according to the new memory access operation. The storage position of the first target task data is adjusted according to the target ratio of the second target task data; if the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area, the second target task data is deleted from the hybrid cache area, and the storage position corresponding to the second target task data in the preset address index entry is recorded as an invalid cache; specifically, when the new memory access operation enters the intelligent scheduler, it will first determine whether its address is in the address index entry. If it exists, a read and write judgment of the memory access operation is performed. If it is a read operation, data is first read from the corresponding hybrid cache area according to the storage position in the address index entry, and then the memory access count is updated to calculate its memory access count / clock count ratio; if it is a write operation, the data to be written is first written into the MRAM main memory area, and then the corresponding hybrid cache module is entered according to the storage position 00 / 01 in its entry to clear its data, and then the memory access address, memory access count, and clock count value in the address index entry are cleared, and the storage position is written as 11.
[0061] The process of adjusting the storage location of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation may specifically include: if the second target task data corresponding to the new memory access operation is stored in the SRAM storage area, and the target ratio corresponding to the second target task data is less than the maximum target ratio corresponding to each task data in the MRAM storage area, then exchanging the storage area of the second target task data with the storage area of the first target task data corresponding to the maximum target ratio; if the second target task data corresponding to the new memory access operation is stored in the MRAM storage area, and the target ratio corresponding to the second target task data is greater than the maximum target ratio corresponding to each first target task data in the SRAM storage area, then exchanging the storage area of the second target task data with the storage area of the first target task data corresponding to the maximum target ratio; According to the corresponding minimum target ratio, the storage area of the second target task data is interchanged with the storage area of the first target task data corresponding to the minimum target ratio; that is, if the storage position in the address index entry corresponding to the new memory access operation is 00 and the memory access count / clock count ratio is lower than the maximum memory access count / clock count ratio of the storage position 01 in the address index entry, or the storage position in the address index entry is 01 and the memory access count / clock count ratio is higher than the minimum memory access count / clock count ratio of the storage position 00 in the address index entry, then the values of the two storage positions are replaced and the positions of the two addresses in the SRAM and MRAM storage areas are swapped in the hybrid cache area.
[0062] The above process of periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space may specifically include: if the memory access address corresponding to the new memory access operation is not in the preset address index entry, then determining whether the new memory access operation is a read operation; if the new memory access operation is a read operation, then determining whether there is an entry in the preset address index entry whose storage position is an invalid cache; if there is no entry in the preset address index entry whose storage position is an invalid cache, then reading the second target task data corresponding to the new memory access operation from the MRAM main storage area; if there is an entry in the preset address index entry whose storage position is an invalid cache, then determining whether the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area; if the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area space, the second target task data corresponding to the new memory access operation is migrated from the MRAM main storage area to the SRAM storage area; if the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area; that is, if the address corresponding to the new memory access operation is not in the address index entry, the read and write judgment of the memory access operation is first performed, if it is a write operation, the data is directly written to the MRAM main storage area, if it is a read operation, first determine whether there is an entry position equal to 11 in the storage position in the address index entry, if not, read the data directly from the MRAM main storage area, if so, determine whether all entries with storage positions equal to 00 in the address index entry meet the number of SRAM storage areas, if so, additionally migrate the data from the MRAM main storage area to the MRAM area in the hybrid cache area and write 01 to the storage position of the address index entry, if not, additionally migrate the data from the MRAM main storage area to the SRAM area in the hybrid cache area and write 00 to the storage position of the address index entry.
[0063] It should be noted that in order to ensure the temporal locality of memory access address judgment and the absoluteness of hot spot data, when the clock count of an entry in the address index entry reaches the number of clock cycles set by the system, the solution of the present invention will clear the memory access address, memory access reception, and clock count values in the memory access address entry, write the invalidation mark 11 to the storage location of the entry, and clear the value maintained in the SRAM or MRAM storage area in the hybrid cache area.
[0064] In addition, in order to further adapt the hybrid cache storage architecture in the solution of the present invention, the solution of the present invention further expands the instruction set of the GPGPU computing core and adds specific instructions for MRAM storage operations, that is, the initial instruction set corresponding to the GPGPU computing core is expanded to obtain the corresponding target instruction set, so as to use the target instruction set to control the data storage operations corresponding to the MRAM main storage area and the MRAM storage area. The specific instruction table is as follows:
[0065] Table 1 GPGPU computing core extension instruction set
[0066]
[0067] It can be seen that the present application optimizes the storage structure of GPGPU and constructs a storage area including MRAM and SRAM, so that the MRAM storage area and the SRAM storage area can be used to store data with different access frequencies respectively, which solves the problem that SRAM has low integration and cannot be applied to GPGPU, and fully utilizes the advantages of SRAM's fast reading and writing speed to improve data storage efficiency; by using MRAM and SRAM for data storage, the problem of data loss caused by using DRAM for data storage is avoided; by storing data with low data access frequency in MRAM, frequent exchange of data between the main storage area and the cache area is avoided, thereby reducing energy consumption.
[0068] Based on the above embodiments, the present application describes the overall process of storing data in GPGPU. Next, the present application will describe the process of pre-fetching data in GPGPU. Figure 5 As shown, an embodiment of the present invention discloses a data pre-fetching process, including:
[0069] Step S21 : predicting the data to be used from each target task data according to the data type and data access frequency of each target task data.
[0070] In this embodiment, in order to further improve the storage performance, a data pre-fetching and transmission optimization system is constructed. First, the computing task type is determined by the upper-layer cache data classification module, and then dynamic identification is performed according to the computing task category to determine the overall data and current execution data of the current computing task, and the data that the GPGPU may need in the next few computing cycles is predicted in advance, and these data are pre-fetched from the MRAM main storage area to the hybrid cache area. Specifically, taking the deep learning convolutional neural network training process as an example, during the execution of the GPGPU computing core, the storage architecture of the present invention performs front-end hot spot, warm point, and cold point data reading and writing operations, while the back-end will also predict the data required for the next layer of convolution calculation based on the parameters of the current network layer and the characteristics of the input data, and load it into the cache in advance to reduce the time the computing core waits for data. By predicting the data required by the GPGPU and pre-fetching the data in advance, the data processing efficiency of the GPGPU is improved.
[0071] Step S22: Load the data to be used from the MRAM main storage area to the hybrid cache area so that the GPGPU computing core processes the data to be used.
[0072] In this embodiment, in order to further adapt to the high concurrency characteristics of GPGPU computing tasks, the physical transmission links between the GPGPU computing core and the intelligent scheduler, hybrid cache area, and MRAM main storage area in the solution of the present invention are organized in thread bundles, using thread bundle identification, multi-thread address, and multi-thread data concurrent transmission. At the same time, hot spot, warm spot, and cold spot data judgment, data reading, data writing, and data migration are also performed in thread bundles. At the same time, in the MRAM main storage area, a multi-module multi-layer three-dimensional stacking structure is adopted, in which the layer depth of each module is also matched with the number of threads accommodated by the thread bundle in the transmission link. Multiple independent threads in each thread bundle simultaneously read and write the multi-layer MRAM storage chip of each MRAM module to achieve concurrent data reading and writing. By processing data in parallel, the data processing speed is improved.
[0073] It can be seen that the present application optimizes the storage structure of GPGPU and constructs a storage area including MRAM and SRAM, so that the MRAM storage area and the SRAM storage area can be used to store data with different access frequencies respectively, which solves the problem that SRAM has low integration and cannot be applied to GPGPU, and fully utilizes the advantages of SRAM's fast reading and writing speed to improve data storage efficiency; by using MRAM and SRAM for data storage, the problem of data loss caused by using DRAM for data storage is avoided; by storing data with low data access frequency in MRAM, frequent exchange of data between the main storage area and the cache area is avoided, thereby reducing energy consumption.
[0074] See also Figure 6 As shown, an embodiment of the present invention discloses a GPGPU storage optimization device, comprising:
[0075] A data acquisition module 11 is used to acquire the first target task data issued by the GPGPU computing core, and determine the target ratio corresponding to each of the first target task data based on a preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each of the first target task data, the memory access count information is the number of read operations experienced by the memory access address after the preset address index entry is written, the clock count information is the number of clock cycles experienced by the memory access address after the preset address index entry is written, and the target ratio is the ratio of the memory access count information to the clock count information;
[0076] The data storage module 12 is used to determine whether the memory access operation corresponding to each of the first target task data is a write operation. If the memory access operation corresponding to each of the first target task data is a write operation, then according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data, each of the first target task data is stored in a corresponding position of the target storage space; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area;
[0077] The storage position adjustment module 13 is used to periodically and dynamically adjust the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage status of the target storage space after obtaining a new memory access operation, so as to store the second target task data corresponding to the new memory access operation to the corresponding position of the target storage space.
[0078] It can be seen that the present application optimizes the storage structure of GPGPU and constructs a storage area including MRAM and SRAM, so that the MRAM storage area and the SRAM storage area can be used to store data with different access frequencies respectively, which solves the problem that SRAM has low integration and cannot be applied to GPGPU, and fully utilizes the advantages of SRAM's fast reading and writing speed to improve data storage efficiency; by using MRAM and SRAM for data storage, the problem of data loss caused by using DRAM for data storage is avoided; by storing data with low data access frequency in MRAM, frequent exchange of data between the main storage area and the cache area is avoided, thereby reducing energy consumption.
[0079] In some specific embodiments, the data storage module 12 may specifically include:
[0080] An index entry sorting unit, for sorting the preset address index entries according to the target ratio in ascending order if the data storage capacity of the SRAM storage area has reached the upper limit and the data storage capacity of the MRAM storage area has not reached the upper limit, and determining whether the first target task data with the smallest target ratio meets the preset hotspot data determination standard within the current preset time period;
[0081] A first data migration unit is configured to store a new memory access address corresponding to a new memory access operation into the preset address index entry if the first target task data with the smallest target ratio meets the preset hot data determination standard within the current preset time period, and migrate the second target task data corresponding to the new memory access address from the MRAM main storage area to the MRAM storage area; wherein the configurable number of entries of the preset address index entry is the same as the maximum task storage number of the hybrid cache area;
[0082] The second data migration unit is used to migrate the first target task data with the smallest target ratio from the SRAM storage area to the MRAM storage area if the first target task data with the smallest target ratio does not meet the preset hot spot data judgment standard within the current preset time period, and migrate the second target task data corresponding to the new memory access address from the MRAM main storage area to the SRAM storage area.
[0083] In some specific embodiments, the storage location adjustment module 13 may specifically include:
[0084] A memory access address determination unit, configured to determine whether a memory access address corresponding to a new memory access operation is in the preset address index entry if the data storage amounts of the SRAM storage area and the MRAM storage area have reached upper limits;
[0085] a storage position adjustment submodule, configured to determine whether the new memory access operation is a read operation if the memory access address corresponding to the new memory access operation is in the preset address index entry, and if the new memory access operation is a read operation, adjust the storage position of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation;
[0086] A data deleting unit is used to store the second target task data corresponding to the new memory access operation in the MRAM main storage area if the new memory access operation is a write operation, delete the second target task data from the hybrid cache area, and record the storage location corresponding to the second target task data in the preset address index entry as an invalid cache.
[0087] In some specific embodiments, the storage position adjustment submodule may specifically include:
[0088] A first storage area exchange unit is used to exchange the storage area of the second target task data with the storage area of the first target task data corresponding to the maximum target ratio if the second target task data corresponding to the new memory access operation is stored in the SRAM storage area and the target ratio corresponding to the second target task data is less than the maximum target ratio corresponding to each of the first target task data in the MRAM storage area;
[0089] The second storage area exchange unit is used to exchange the storage area of the second target task data with the storage area of the first target task data corresponding to the minimum target ratio if the second target task data corresponding to the new memory access operation is stored in the MRAM storage area and the target ratio corresponding to the second target task data is greater than the minimum target ratio corresponding to each of the first target task data in the SRAM storage area.
[0090] In some specific embodiments, the storage location adjustment module 13 may specifically include:
[0091] An index entry judgment unit, configured to judge whether the new memory access operation is a read operation if the memory access address corresponding to the new memory access operation is not in the preset address index entry, and if the new memory access operation is a read operation, judge whether there is an entry in the preset address index entry whose storage location is an invalid cache;
[0092] a third data migration unit, configured to determine whether the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area if there is an entry whose storage location is an invalid cache in the preset address index entry, and if the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area, migrate the second target task data corresponding to the new memory access operation from the MRAM main storage area to the SRAM storage area;
[0093] The data storage unit is used to store the second target task data corresponding to the new memory access operation into the MRAM main storage area if the new memory access operation is a write operation.
[0094] In some specific embodiments, the storage location adjustment module 13 further includes:
[0095] A data preloading unit is used to predict the data to be used from each target task data according to the data type and data access frequency of each target task data, and load the data to be used from the MRAM main storage area to the hybrid cache area so that the GPGPU computing core can process the data to be used.
[0096] In some specific embodiments, the GPGPU storage optimization device further includes:
[0097] The instruction set expansion module is used to expand the initial instruction set corresponding to the GPGPU computing core to obtain a corresponding target instruction set, so as to use the target instruction set to control the data storage operations corresponding to the MRAM main storage area and the MRAM storage area.
[0098] Furthermore, the present application also discloses an electronic device. Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0099] Figure 7 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the GPGPU storage optimization method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0100] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0101] In addition, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0102] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the GPGPU storage optimization method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0103] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed GPGPU storage optimization method is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0104] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0105] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0106] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0107] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0108] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A GPGPU storage optimization method, characterized in that: include: Obtain the first target task data issued by the GPGPU computing core, and determine the target ratio corresponding to each of the first target task data based on the preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each of the first target task data, the memory access count information is the number of read operations experienced after the memory access address is written into the preset address index entry, the clock count information is the number of clock cycles experienced after the memory access address is written into the preset address index entry, and the target ratio is the ratio of the memory access count information to the clock count information; Determine whether the memory access operation corresponding to each of the first target task data is a write operation, and if the memory access operation corresponding to each of the first target task data is a write operation, store each of the first target task data in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area; After obtaining a new memory access operation, the storage position of the first target task data in the target storage space is periodically and dynamically adjusted according to the target ratio of the corresponding second target task data and the storage status of the target storage space, so that the second target task data corresponding to the new memory access operation is stored in the corresponding position of the target storage space.
2. The GPGPU storage optimization method according to claim 1, characterized in that: The periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes: If the data storage capacity of the SRAM storage area has reached the upper limit, and the data storage capacity of the MRAM storage area has not reached the upper limit, the preset address index entries are sorted in ascending order according to the target ratios, and it is determined whether the first target task data with the smallest target ratio meets the preset hotspot data determination standard within the current preset time period; If the first target task data with the smallest target ratio meets the preset hot data determination standard within the current preset time period, a new memory access address corresponding to the new memory access operation is stored in the preset address index entry, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the MRAM storage area; wherein the configurable number of entries of the preset address index entry is the same as the maximum task storage number of the hybrid cache area; If the first target task data with the smallest target ratio does not meet the preset hotspot data determination standard within the current preset time period, the first target task data with the smallest target ratio is migrated from the SRAM storage area to the MRAM storage area, and the second target task data corresponding to the new memory access address is migrated from the MRAM main storage area to the SRAM storage area.
3. The GPGPU storage optimization method according to claim 2, characterized in that: The periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes: If the data storage amounts of the SRAM storage area and the MRAM storage area have both reached upper limits, determining whether a memory access address corresponding to a new memory access operation is in the preset address index entry; If the memory access address corresponding to the new memory access operation is in the preset address index entry, determining whether the new memory access operation is a read operation, and if the new memory access operation is a read operation, adjusting the storage position of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation; If the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area, the second target task data is deleted from the hybrid cache area, and the storage location corresponding to the second target task data in the preset address index entry is recorded as an invalid cache.
4. The GPGPU storage optimization method according to claim 3, characterized in that: The step of adjusting the storage location of the first target task data according to the target ratio of the second target task data corresponding to the new memory access operation includes: If the second target task data corresponding to the new memory access operation is stored in the SRAM storage area, and the target ratio corresponding to the second target task data is less than the maximum target ratio corresponding to each of the first target task data in the MRAM storage area, then the storage area of the second target task data is swapped with the storage area of the first target task data corresponding to the maximum target ratio; If the second target task data corresponding to the new memory access operation is stored in the MRAM storage area, and the target ratio corresponding to the second target task data is greater than the minimum target ratio corresponding to each of the first target task data in the SRAM storage area, the storage area of the second target task data is swapped with the storage area of the first target task data corresponding to the minimum target ratio.
5. The GPGPU storage optimization method according to claim 3, characterized in that: The periodically and dynamically adjusting the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage state of the target storage space so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space includes: If the memory access address corresponding to the new memory access operation is not in the preset address index entry, determine whether the new memory access operation is a read operation; if the new memory access operation is a read operation, determine whether there is an entry in the preset address index entry whose storage location is an invalid cache; If there is no entry whose storage location is an invalid cache in the preset address index entry, reading the second target task data corresponding to the new memory access operation from the MRAM main storage area; If there is an entry whose storage location is an invalid cache in the preset address index entry, it is determined whether the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area; if the occupied space of the first target task data in the SRAM storage area is not greater than the total storage space of the SRAM storage area, the second target task data corresponding to the new memory access operation is migrated from the MRAM main storage area to the SRAM storage area; If the new memory access operation is a write operation, the second target task data corresponding to the new memory access operation is stored in the MRAM main storage area.
6. The GPGPU storage optimization method according to claim 1, characterized in that: After storing each target task data in the corresponding position of the target storage space, the method further includes: The data to be used is predicted from each target task data according to the data type and data access frequency of each target task data, and the data to be used is loaded from the MRAM main storage area to the hybrid cache area so that the GPGPU computing core processes the data to be used.
7. The GPGPU storage optimization method according to any one of claims 1 to 6, characterized in that: Also includes: An initial instruction set corresponding to the GPGPU computing core is expanded to obtain a corresponding target instruction set, so as to use the target instruction set to control data storage operations corresponding to the MRAM main storage area and the MRAM storage area.
8. A GPGPU storage optimization device, characterized in that: include: A data acquisition module, used for acquiring first target task data issued by a GPGPU computing core, and determining a target ratio corresponding to each first target task data based on a preset address index entry; wherein the preset address index entry includes memory access address information, memory access count information and clock count information, the memory access address information includes the memory access address corresponding to each first target task data, the memory access count information is the number of read operations experienced by the memory access address after the preset address index entry is written, the clock count information is the number of clock cycles experienced by the memory access address after the preset address index entry is written, and the target ratio is the ratio of the memory access count information to the clock count information; a data storage module, used to determine whether the memory access operation corresponding to each of the first target task data is a write operation, and if the memory access operation corresponding to each of the first target task data is a write operation, then store each of the first target task data in a corresponding position of the target storage space according to the space occupancy of the target storage space and the target ratio corresponding to each of the first target task data; wherein the target storage space includes an MRAM main storage area and a hybrid cache area, and the hybrid cache area includes an SRAM storage area and an MRAM storage area; A storage position adjustment module is used to periodically and dynamically adjust the storage position of the first target task data in the target storage space according to the target ratio of the corresponding second target task data and the storage status of the target storage space after obtaining a new memory access operation, so as to store the second target task data corresponding to the new memory access operation in the corresponding position of the target storage space.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the GPGPU storage optimization method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the GPGPU storage optimization method according to any one of claims 1 to 7.
Citation Information
Patent Citations
GPGPU performance optimization method based on memory access priorities
CN108279981A
Storage device
CN114974340A
Method for determining access times of memory page and computing equipment
CN117033254A
Hybrid cache memory and method for controlling the same
US20210081331A1
Memory management device capable of managing memory address translation table using heterogeneous memories and method of managing memory address thereby
US20210096745A1