Efficient cache architecture construction method and device based on GPGPU (General Purpose Graphics Processing Unit) and medium

By installing the external register function module, allocating the buffer module, determining the data information extraction module and performing mask code configuration in the GPGPU architecture, the problem of insufficient area and performance utilization in the construction of the GPGPU architecture is solved, and efficient cache utilization and use is achieved.

CN119938596APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510010677.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the construction of GPGPU architectures has the problem of insufficient utilization of area and performance.

Method used

By externalizing the preset register function module and using the kernel storage allocation to obtain the to-process buffer module, the data storage bit configuration is performed to obtain the buffer module architecture, the data information extraction module is determined, and the preset mask code is assigned to thread bundles to obtain the mask code configuration.

Benefits of technology

It has achieved the improvement of cache occupation and usage efficiency, and solved the problem of insufficient area and performance utilization in GPGPU architecture construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938596A_ABST
    Figure CN119938596A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient cache architecture construction method and device based on a GPGPU and a medium, and relates to the technical field of cache frameworks. The method comprises the steps that a preset register function module is externally arranged, and a to-be-processed buffer module is obtained through core storage allocation based on the register function module; performing data storage bit configuration on the to-be-processed buffer module to obtain a buffer module architecture; determining a data information extraction module through data information extraction configuration according to the buffer module architecture; and performing thread bundle distribution on a preset mask code to obtain mask code configuration. By means of the method, the technical problem that in the prior art, the utilization degree of the area and performance is insufficient during construction of a GPGPU architecture is solved, and cache occupation, use efficiency and the like are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cache framework technology, and in particular to a method, device and medium for constructing an efficient cache architecture based on GPGPU. Background Art

[0002] With the development of contemporary science and technology, image recognition, artificial intelligence and other fields, people's demand for computing large amounts of data is increasing, and GPGPU is a technical means for processing big data computing. GPGPU is derived from GPU. It uses the high parallel computing capability of GPU to perform non-graphic tasks. Compared with GPU, GPGPU cancels the graphics computing unit and other parts of GPU that are specifically used for graphics processing, and instead focuses on parallel processing of big data.

[0003] In recent years, GPGPU architectures have emerged one after another with the update of GPUs, but the optimization of cache is not particularly rich. The original single-level cache is only gradually replaced by a multi-level cache architecture. The characteristics of the multi-level cache architecture are that the smaller the cache level, the smaller the capacity; the highest level of cache main memory cache speed is the slowest, but due to the small unit area, the capacity can be made very large. The existing technology for the construction of GPGPU architecture has the problem of insufficient utilization of area and performance. Summary of the invention

[0004] The embodiments of the present application provide a method, device and medium for constructing an efficient cache architecture based on GPGPU, which solves the technical problem of insufficient area and performance utilization in the construction of GPGPU architecture in the prior art.

[0005] In a first aspect, an embodiment of the present application provides a method for constructing an efficient cache architecture based on GPGPU, characterized in that the method includes: externalizing a preset register function module, and based on the register function module, obtaining a buffer module to be processed through core storage allocation; configuring data storage bits for the buffer module to be processed to obtain a buffer module architecture; determining a data information extraction module through data information extraction configuration according to the buffer module architecture; and performing thread bundle allocation on a preset mask code to obtain a mask code configuration.

[0006] In one implementation of the present application, based on the register, a buffer module to be processed is obtained through core storage allocation, which specifically includes: configuring the number of cores for the register function module to determine the core storage parameters; based on the core storage parameters, obtaining the number of buffer threads through thread bundle generation; according to the buffer thread bundle, obtaining the buffer module to be processed through buffer space configuration.

[0007] In one implementation of the present application, data storage bits are configured for a buffer module to be processed to obtain a buffer module architecture, specifically including: Head-Buffer bit amplification of the buffer module to be processed to obtain a first data storage architecture; Tail-Buffer bit amplification of the buffer module to be processed to obtain a second data storage architecture; based on the first data storage architecture and the second data storage architecture, the buffer module architecture is obtained by determining the storage data status.

[0008] In one implementation of the present application, based on the first data storage architecture and the second data storage architecture, a buffer module architecture is obtained by storing data status judgment, which specifically includes: configuring the Buffer module storage hierarchy of the first data storage architecture and the second data storage architecture to determine the storage data judgment bit; based on the first data storage architecture, determining the last bit of incoming data, and storing the last bit of incoming data to the Head bit, to obtain a first storage judgment; according to the second data storage architecture, determining the last bit of out-cache data, and storing the last bit of out-cache data to the Tail bit, to obtain a second storage judgment; based on the first storage judgment and the second storage judgment, the buffer module architecture is obtained.

[0009] In one implementation of the present application, according to the buffer module architecture, the data information extraction module is determined through the data information extraction configuration, specifically including: obtaining pipeline key data, and pushing the pipeline key data into the buffer stack to obtain related data to be compared; performing data branch comparison on the related data to be compared to obtain deduplicated data, and based on the deduplicated data, idling the current thread to determine the data information extraction module.

[0010] In one implementation of the present application, data branch comparison is performed on the relevant data to be compared to obtain deduplicated data, specifically including: storing the relevant data to be compared in a Bankstore to determine the comparison portion of data; comparing subsequent data in the Bankstore based on the comparison portion of data to obtain data to be eliminated; and eliminating the data to be eliminated from the thread data to obtain deduplicated data.

[0011] In one implementation of the present application, a preset mask code is assigned to a thread bundle to obtain a mask code configuration, specifically including: obtaining a current thread bundle, and based on the current thread bundle, determining the current thread bundle position through bit width allocation; according to the current thread bundle position, determining the total valid entry information through a sub-valid entry information set; and determining the mask code configuration based on the current thread bundle position, the total valid entry information and the read-write status bit.

[0012] In one implementation of the present application, the total valid entry information is determined according to the current thread warp position through a set of sub-valid entry information, specifically including: based on the current thread warp position, the sub-valid entry information is split to obtain the sub-valid entry information; according to the preset storage requirements, the target information is configured, and the target information and the sub-valid entry information are combined to obtain the total valid entry information.

[0013] In the second aspect, an embodiment of the present application also provides a device for constructing an efficient cache architecture based on GPGPU, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: externalize a preset register function module, and based on the register function module, obtain a buffer module to be processed through core storage allocation; configure data storage bits for the buffer module to be processed to obtain a buffer module architecture; determine the data information extraction module through data information extraction configuration according to the buffer module architecture; and perform thread bundle allocation on the preset mask code to obtain a mask code configuration.

[0014] In the third aspect, the embodiment of the present application also provides a non-volatile computer storage medium constructed based on an efficient cache architecture of GPGPU, storing computer executable instructions, characterized in that the computer executable instructions are set to: externalize a preset register function module, and based on the register function module, obtain a buffer module to be processed through core storage allocation; configure data storage bits for the buffer module to be processed to obtain a buffer module architecture; determine the data information extraction module through data information extraction configuration according to the buffer module architecture; and perform thread bundle allocation on the preset mask code to obtain a mask code configuration.

[0015] The embodiments of the present application provide a method, device and medium for constructing an efficient cache architecture based on GPGPU. By optimizing the working logic of the cache, the technical problem of insufficient area and performance utilization in the construction of the GPGPU architecture in the prior art is solved, and the cache occupancy and usage efficiency are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 A flow chart of a method for constructing an efficient cache architecture based on GPGPU provided in an embodiment of the present application;

[0018] Figure 2 A framework diagram of a method for constructing an efficient cache architecture based on GPGPU provided in an embodiment of the present application;

[0019] Figure 3 A data clearing logic diagram of an efficient cache architecture based on GPGPU provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of Masks of an efficient cache architecture based on GPGPU provided in an embodiment of the present application;

[0021] Figure 5 A schematic diagram of the internal structure of a device for building an efficient cache architecture based on GPGPU provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0023] The embodiments of the present application provide a method, device and medium for constructing an efficient cache architecture based on GPGPU. By optimizing the working logic of the cache, the technical problem of insufficient area and performance utilization in the construction of the GPGPU architecture in the prior art is solved, and the cache occupancy and usage efficiency are improved.

[0024] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0025] Figure 1 A flow chart of a method for constructing an efficient cache architecture based on GPGPU is provided in an embodiment of the present application. Figure 1 As shown, the embodiment of the present application provides a method for constructing an efficient cache architecture based on GPGPU, which specifically includes the following steps:

[0026] Step 101: externalize a preset register function module, and obtain a buffer module to be processed through core storage allocation based on the register function module.

[0027] Specifically, the method includes: configuring the number of cores for the register function module and determining the core storage parameters; obtaining the number of buffer threads through thread bundle generation based on the core storage parameters; and obtaining the buffer module to be processed through buffer space configuration according to the buffer thread bundle.

[0028] The present application realizes the pre-configuration of the buffer module by externalizing the preset register function module and obtaining the buffer module to be processed through core storage allocation based on the register function module.

[0029] In one implementation of the present application, a detailed explanation is given through the following Example 1.

[0030] Example 1: Place the register externally and allocate the corresponding storage space and address to it according to the number of cores.

[0031] The storage space involved in the allocation includes a buffer module in the thread bundle allocation module, which is responsible for buffering threads and forming thread bundles, and forming a Head-Tail buffer space inside the L2 Cache.

[0032] It should be noted that the above technical solution is based on the premise that the register is external.

[0033] Step 102: Perform data storage bit configuration on the buffer module to be processed to obtain a buffer module architecture.

[0034] Specifically, it includes: performing Head-Buffer bit amplification on the buffer module to be processed to obtain a first data storage architecture; performing Tail-Buffer bit amplification on the buffer module to be processed to obtain a second data storage architecture; based on the first data storage architecture and the second data storage architecture, determining the storage data status to obtain a buffer module architecture.

[0035] Based on the first data storage architecture and the second data storage architecture, a buffer module architecture is obtained by determining the storage data status, specifically including: configuring the Buffer module storage hierarchy for the first data storage architecture and the second data storage architecture to determine the storage data determination bit; based on the first data storage architecture, determining the last bit of incoming data, and storing the last bit of incoming data to the Head bit, to obtain a first storage determination; based on the second data storage architecture, determining the last bit of out-cache data, and storing the last bit of out-cache data to the Tail bit, to obtain a second storage determination; based on the first storage determination and the second storage determination, the buffer module architecture is obtained.

[0036] The present application configures the data storage bits of the buffer module to be processed to obtain the buffer module architecture, thereby speeding up the speed at which data enters the main memory, and data that may need to be reused can be directly extracted from the buffer module in a more timely manner, thereby improving the reading and writing efficiency.

[0037] Figure 2 This is a framework diagram of a method for building an efficient cache architecture based on GPGPU, which is explained in detail through the following Example 2 in the embodiments of the present application.

[0038] Example 2: For registers, their own speed is already quite fast, so there is no need to merge with the buffer module. Therefore, the main fusion part is for L1, L2 and main memory. The implementation method is to increase the Head-Buffer bit and Tail-Buffer bit. The capacity of the original register part needs to be increased. With its high-speed reading and writing characteristics, several Buffer modules are allocated to the front end of storage units at different levels. The Head bit is used to store the last data entering the module, and the Tail bit is used to store the last data leaving the cache module and entering the main memory. Through such a logical structure, the speed of data entering the main memory can be accelerated. At the same time, for data that may need to be reused, it can also be directly extracted from the Buffer module in a more timely manner, thereby improving the reading and writing efficiency.

[0039] Since the storage level of L1 is relatively low and is greatly affected by the pipeline, no additional modification is required for it. Therefore, the cache part of L1 (including sharemem) is not modified. In addition to L2, Head-Tail Buffer is added to L3 and main memory.

[0040] Step 103: Determine the data information extraction module according to the buffer module architecture and the data information extraction configuration.

[0041] Specifically, it includes: obtaining pipeline key data and pushing the pipeline key data into the buffer stack to obtain relevant data to be compared; performing data branch comparison on the relevant data to be compared to obtain deduplicated data, and based on the deduplicated data, idling the current thread to determine the data information extraction module.

[0042] Specifically, the method includes: storing the relevant data to be compared into Bankstore to determine the comparison part of the data; comparing the subsequent data in Bankstore according to the comparison part of the data to obtain the data to be eliminated; and eliminating the data to be eliminated from the thread data to obtain the deduplicated data.

[0043] This application determines the data information extraction module based on the buffer module architecture through data information extraction configuration, realizes the elimination of duplicate data in the pipeline thread, and improves the utilization rate of threads and memory space.

[0044] Figure 3 A data clearing logic diagram of an efficient cache architecture based on GPGPU is provided in an embodiment of the present application. In the embodiment of the present application, it is explained in detail through the following Example 3.

[0045] Example 3: Capture and collect relevant information of the data on the pipeline, and cooperate with the buffer module to push the relevant data into the buffer stack composed of Head-Buffer and Tail-Buffer.

[0046] At the same time, the data extraction module is also responsible for a certain degree of duplicate value comparison. It will extract the relevant information carried in the data and store it in the Bankstore in its functional block for data comparison. Subsequently, for the parts carrying the same information in the Bankstore, duplicate elimination operations will be performed (for the same thread bundle, when executing tasks, the data carried are the same. The different results are because each thread encounters a different branch, but there are still many threads that do not encounter branch conditions, and such numbers may account for a considerable proportion in the work of the entire thread bundle and pipeline).

[0047] Therefore, it is very necessary to eliminate duplicate data. After the elimination is completed, the current thread will be idle and a certain amount of memory space will be released.

[0048] After obtaining the data data, the thread position in the thread warp and the valid entry information are collected. When the thread position is the highest among the valid entry information, the data is retained and written (or read from the corresponding cache).

[0049] When the thread position is not the highest in the valid entry information, it is determined whether it is duplicate information. If it is not duplicate information, the changed data is retained (or read from the corresponding cache).

[0050] On the contrary, if it is duplicate information, the data will not be retained, and subsequent threads in the same harness will be detected to remove the related duplicate information.

[0051] It should be noted that the buffer part between these caches is implemented through registers and space allocation and control modules, and the specific implementation method is through the division of address space.

[0052] For example, the buffer inside Core0 is 0000_0000-1FFF_1FFF, and the address space of L3 Cache is 4FFF_FFFF-6FFF_FFFF. The specific division of address space can be achieved through artificial definition.

[0053] The logic of the register and space allocation and control module controlling the Head-Tail Buffer is to control both ends of the buffer. When data enters the buffer, the built-in counter starts counting at the same time. The position of the current data in the buffer can be determined by the number of counters. At the same time, the module also records the relevant information of the cache line position in the mask code of the data.

[0054] When there is data matching, the corresponding data is extracted to the bottom of the stack through the tail end. The data can be directly extracted from the Buffer and returned to the cache that issued the request or the Core on the pipeline.

[0055] Step 104: perform thread warp allocation on the preset mask code to obtain a mask code configuration.

[0056] Specifically, the method includes: obtaining the current thread warp, and determining the current thread warp position through bit width allocation based on the current thread warp; determining the total valid entry information through the sub-valid entry information set according to the current thread warp position; and determining the mask code configuration based on the current thread warp position, the total valid entry information and the read-write status bit.

[0057] According to the current thread warp position, the total valid entry information is determined through the sub-valid entry information set, specifically including: based on the current thread warp position, the sub-valid entry information is split to obtain the sub-valid entry information; according to the preset storage requirements, the target information is configured, and the target information and the sub-valid entry information are combined to obtain the total valid entry information.

[0058] The present application allocates thread bundles to preset mask codes to obtain mask code configuration, thereby achieving the overall area and performance of the GPGPU architecture and improving the cache occupancy status and usage efficiency.

[0059] Figure 4 A Masks schematic diagram of an efficient cache architecture based on GPGPU is provided in an embodiment of the present application. In the embodiment of the present application, it is explained in detail through the following Example 4.

[0060] Example 4: First, in the thread bundle allocation module within the core, assuming that one warp consists of 32 threads, when allocating threads, the required bit width will be allocated. For example, when there are 32 threads, the bit width of 00000-11111 will be allocated to store the position of the thread in the current thread bundle.

[0061] In addition, the operation process of the entire GPGPU joint pipeline and multi-level cache is as follows:

[0062] If the cache line address is found in the cache of the current level, it is considered a hit; if not found, it is a miss, and the read / write request needs to be upgraded to a higher level of storage space to continue searching until a hit is found.

[0063] If the data is not found even at the highest level, a write replacement will generally be performed directly, and the read operation will not fail even at the highest level.

[0064] For the mask encoding method, the original valid entry information is split and corresponds to each target information respectively. Only when the sub-valid entry information is integrated to form the total valid information, the judgment on whether the information is consistent will take effect.

[0065] The combination of the secondary valid entry information + the target information is customized according to human needs. In theory, the more such combinations there are, the less likely they are to be repeated.

[0066] Data stored in a higher cache level will first be placed in the Head-Tail Buffer. After entering the buffer, the cache line addresses will be recorded in the order of entry. If the data needs to be read again during the stack push and pop process (that is, the cache line address matches), it will be directly taken out of the buffer after checking the valid title information without going through the higher-level cache with slower read and write speeds and larger capacity. At the same time, in terms of data clearing, for the case where the total valid entry information is consistent, the read and write status bits are consistent, but the thread positions in the thread bundle are inconsistent, the threads with low position rankings will be cleared and pulled up, leaving only the threads with the highest position ranking to continue working and reading and writing.

[0067] This optimization scheme can maintain the sparseness of threads to the greatest extent. After the redundant threads are released, they can be assigned to other cores through the overall scheduling of the entire GPGPU to form new thread bundles to work, making the most of the high parallelism of GPGPU.

[0068] The interaction logic between the cores and the L3 cache is implemented through a cross switch, which is an arbitration mechanism that can connect the cores with transmission requirements to the L3 cache through timing and logic adjustments. The L3 cache is responsible for the cache outside the entire GPGPU pipeline and is responsible for the storage of interactive data between multiple cores. When there are multiple cores with transmission requirements at the same time, it will ensure the smooth completion of the task of the current core by pulling down the ready signal, and then transmit to other cores after completion.

[0069] L3 is directly connected to the main memory. The main memory is generally used to retain results and send data. The frequency of internal data changes is the lowest, but the capacity is the largest. Each storage has an independent buffer module to ensure the efficiency and stability of data flow.

[0070] Through the present invention, the efficiency of data in GPGPU can be improved, the occupancy of useless threads on the pipeline can be reduced, and the utilization efficiency of GPGPU can be further improved from the perspective of cache and data reading and writing storage optimization.

[0071] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a device for building an efficient cache architecture based on GPGPU, whose structure is as follows Figure 2 shown.

[0072] Figure 5 The schematic diagram of the internal structure of a device for constructing an efficient cache architecture based on GPGPU provided in the embodiment of the present application is as follows. Figure 5 As shown, the device includes:

[0073] at least one processor 501;

[0074] and, a memory 502 communicatively connected to the at least one processor;

[0075] The memory 502 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 501 to enable at least one processor 501 to:

[0076] The preset register function module is externalized, and based on the register function module, a buffer module to be processed is obtained through core storage allocation; data storage bits are configured for the buffer module to be processed to obtain a buffer module architecture; according to the buffer module architecture, a data information extraction module is determined through data information extraction configuration; thread bundles are allocated for the preset mask code to obtain a mask code configuration.

[0077] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium constructed based on a GPGPU efficient cache architecture stores computer executable instructions, wherein the computer executable instructions are set as:

[0078] The preset register function module is externalized, and based on the register function module, a buffer module to be processed is obtained through core storage allocation; data storage bits are configured for the buffer module to be processed to obtain a buffer module architecture; according to the buffer module architecture, a data information extraction module is determined through data information extraction configuration; thread bundles are allocated for the preset mask code to obtain a mask code configuration.

[0079] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0080] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0081] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0082] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0083] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0085] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0086] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0087] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0088] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0089] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for constructing an efficient cache architecture based on GPGPU, characterized in that: The method comprises: The preset register function module is externalized, and based on the register function module, a buffer module to be processed is obtained through core storage allocation; Performing data storage bit configuration on the buffer module to be processed to obtain a buffer module architecture; According to the buffer module architecture, a data information extraction module is determined through data information extraction configuration; The preset mask code is assigned to thread warps to obtain the mask code configuration.

2. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: Based on the register, a buffer module to be processed is obtained through core storage allocation, which specifically includes: Performing core quantity configuration on the register function module and determining core storage parameters; Based on the core storage parameters, the number of buffered threads is obtained by thread warp generation; According to the buffer warp, the buffer module to be processed is obtained through buffer space configuration.

3. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: The data storage bit configuration is performed on the buffer module to be processed to obtain a buffer module architecture, which specifically includes: Performing Head-Buffer bit expansion on the buffer module to be processed to obtain a first data storage architecture; Performing Tail-Buffer bit expansion on the buffer module to be processed to obtain a second data storage architecture; Based on the first data storage architecture and the second data storage architecture, the buffer module architecture is obtained by determining the storage data state.

4. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: Based on the first data storage architecture and the second data storage architecture, the buffer module architecture is obtained by determining the storage data state, which specifically includes: Performing buffer module storage level configuration on the first data storage architecture and the second data storage architecture to determine storage data determination bits; Based on the first data storage architecture, determine the last bit of incoming data, and store the last bit of incoming data in the Head bit to obtain a first storage determination; According to the second data storage architecture, determine the last cache data, and store the last cache data in the Tail position to obtain a second storage determination; The buffer module architecture is obtained based on the first storage determination and the second storage determination.

5. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: According to the buffer module architecture, the data information extraction module is determined through the data information extraction configuration, specifically including: Acquire pipeline key data, and push the pipeline key data into a buffer stack to obtain relevant data to be compared; Data branch comparison is performed on the related data to be compared to obtain deduplicated data, and based on the deduplicated data, the current thread is idle to determine the data information extraction module.

6. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: Performing data branch comparison on the relevant data to be compared to obtain duplicate-free data, specifically including: Storing the relevant data to be compared in Bankstore to determine the comparison part of the data; Comparing the subsequent data in the Bankstore according to the compared partial data to obtain data to be eliminated; The data to be removed is removed from the thread data to obtain the deduplicated data.

7. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: The preset mask code is assigned to thread warps to obtain the mask code configuration, which specifically includes: Obtaining a current thread warp, and determining a current thread warp position based on the current thread warp by bit width allocation; According to the current thread warp position, determine the total valid entry information through the secondary valid entry information set; The mask code configuration is determined based on the current warp position, total valid entry information, and read / write status bits.

8. The method for constructing an efficient cache architecture based on GPGPU according to claim 1, characterized in that: According to the current warp position, the total valid entry information is determined through the secondary valid entry information set, specifically including: Based on the current thread warp position, split the valid entry information to obtain secondary valid entry information; According to the preset storage requirement, the target information is configured, and the target information and the secondary valid entry information are combined to obtain the total valid entry information.

9. A device for constructing an efficient cache architecture based on GPGPU, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: The preset register function module is externalized, and based on the register function module, a buffer module to be processed is obtained through core storage allocation; Performing data storage bit configuration on the buffer module to be processed to obtain a buffer module architecture; According to the buffer module architecture, a data information extraction module is determined through data information extraction configuration; The preset mask code is assigned to thread warps to obtain the mask code configuration.

10. A non-volatile computer storage medium constructed based on a GPGPU efficient cache architecture, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: The preset register function module is externalized, and based on the register function module, a buffer module to be processed is obtained through core storage allocation; Performing data storage bit configuration on the buffer module to be processed to obtain a buffer module architecture; According to the buffer module architecture, a data information extraction module is determined through data information extraction configuration; The preset mask code is assigned to thread warps to obtain the mask code configuration.