Improved OpenCL heterogeneous programming framework memory management method and device

By adopting front-end, middle-end and back-end cache memory management methods in OpenCL, the problem of low memory allocation and access efficiency in the existing technology is solved, and more efficient memory management and computing performance improvement is achieved.

CN120029929APending Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057690.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing OpenCL memory management mechanism is inefficient, resulting in low allocation and access efficiency of global memory, affecting computing performance.

Method used

The front-end, middle-end, and back-end caches are used to pre-cache and allocate global memory, and priority is given to applying for global memory from the near-end side, and cache is managed through FreeList, combining LRU algorithm and intelligent memory access optimization, adjusting the memory access order to reduce cache misses and delays.

Benefits of technology

It improves the efficiency of memory allocation and data access efficiency, improves computing performance, and has significantly improved compared to OpenCL's original global memory management method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029929A_ABST
    Figure CN120029929A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of parallel computing, and discloses an improved OpenCL heterogeneous programming framework memory management method and device.According to the method, a global memory is cached in a front end, a middle table and a rear end in advance, the global memory of the front end is unique to a single SP, and the global memory of the middle table is shared by the SPs in the same working group; the back-end global memory is shared by all the SPs or applied from the DRAM, when the SPs apply for use of the global memory, the SPs apply for use in the front-end global memory unique to the SPs preferentially, when the front-end global memory cannot meet the application condition, the SPs apply for use in the middle global memory, and when the middle global memory cannot meet the application condition, the SPs apply for use in the front-end global memory. And adopting a rear-end global memory or acquiring from the DRM. According to the method, the allocation of the global memory is optimized in a front-end, middle-end and back-end cache mode, the memory allocation efficiency is improved, and the optimal memory access execution sequence is generated through memory access optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of parallel computing technology, and for example, to an improved OpenCL heterogeneous programming framework memory management method and device. Background Art

[0002] OpenCL is a cross-platform parallel programming specification that allows developers to use multiple processors for programming to achieve efficient parallel computing. It greatly expands the application scope of GPUs and makes them no longer limited to the graphics field. However, the existing memory management mechanism is relatively extensive, with low efficiency in memory allocation and memory access, which affects the final performance.

[0003] At present, OpenCL's memory is dynamically divided into global memory, local memory, private memory and registers according to the characteristics of the memory pyramid. Global memory is shared by all SPs of the entire computing card, local memory is shared by SPs in the same workgroup, and private memory and registers are accessed by a single SP. The memory access speed slows down in the order of registers, private memory, local memory, and global memory, and the memory capacity increases in the order of registers, private memory, local memory, and all memory. Since the global memory capacity is large and is shared by all SPs of the entire computing card, it is inevitable that multiple SPs will operate the global memory at the same time, resulting in performance degradation.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] The disclosed embodiments provide an improved OpenCL heterogeneous programming framework memory management method and device to improve data access efficiency and computing performance.

[0007] In some embodiments, the method includes: pre-caching the global memory in the front end, the middle end and the back end, the global memory of the front end is exclusive to a single SP, the global memory of the middle end is shared by SPs in the same workgroup, and the global memory of the back end is shared by all SPs or applied for from DRAM. When the SP applies for the use of global memory, it gives priority to applying in the front end global memory exclusive to the SP. When the front end global memory cannot meet the application conditions, it applies in the middle end global memory. When the middle end global memory cannot meet the application conditions, the back end global memory is used or obtained from the DRM.

[0008] Furthermore, this method is designed so that the middle station provides cache for the front end, and the back end provides cache for the middle station. When the global memory of the front end is insufficient, an application is made from the cache list of the middle station and the front end cache list is supplemented; when the global memory of the middle station is insufficient, an application is made from the cache list of the back end and the middle station cache list is supplemented. When the global memory in the back end cache list is insufficient, an application is made from the DRM; the middle station allows global memory to be exchanged between the two front ends. For example, if SP No. 1 has excess memory and SP No. 2 has insufficient memory, the middle station allows the front end cache portion of SP No. 1 to be adjusted to SP No. 2.

[0009] Furthermore, the entire global memory is managed according to pages. The front-end, middle-end, and back-end all use FreeList to maintain the cache. The basic structure in FreeList is Chunk. Pages of the same size are suspended under each Chunk, and different Chunks maintain pages of different sizes.

[0010] Furthermore, when allocating global memory, first search in the global memory cached on the front end, traverse the Chunks in the FreeList in order, find the first Chunk that meets the conditions, and obtain the global memory that meets the requirements from the Chunk. If the global memory corresponding to the Chunk is larger than the required global memory, the global memory that meets the requirements is provided to the user, and the remaining global memory is placed in the corresponding free Chunk according to size.

[0011] Furthermore, during the operation of the computing card, it is detected whether the size of the applied global memory is consistent with the current actual Chunk combination. If the size of the applied global memory is N times larger than a fixed value of global memory, where N is a positive integer, the global memory smaller than the fixed value will be merged and placed in a Chunk sequence of the corresponding size after merging.

[0012] Furthermore, during the operation of the computing card, the usage frequency of pages under different chunks is detected, and the cache with a usage frequency less than the set value is released based on the LRU algorithm.

[0013] Furthermore, it also includes intelligent memory access optimization, analyzing the execution mode and data dependency of work items, intelligently adjusting the memory access order, and reducing cache misses and memory access delays; specifically, a complete search space is established for all combinations of the current memory access order, and a cost model is established based on the time overhead of memory access, memory usage, and data transmission bandwidth. The cost model is: C total =C time *T+C mem *M+C band *B, where C total is the total cost, Ctime is the cost per unit of memory access time, C mem is the cost per unit of memory usage, C band is the cost per unit bandwidth usage, T is the time overhead of memory access, M is the memory usage, B is the bandwidth occupancy, and the memory access plan generator is used to generate the corresponding memory access plan, and the evaluator is used to score the generated memory access plan, and the cost model is updated according to the scoring result.

[0014] Furthermore, nonlinear calculation is introduced and the cost model is transformed into:

[0015] C total =∫C time (T)*T+C mem *M+∫C band (B) dB, where C band (B) represents the function of bandwidth cost changing with bandwidth usage, C time (T) represents the function of computational cost varying with memory access time.

[0016] Furthermore, different middle platforms are protected by mutex locks.

[0017] In some embodiments, the device includes a processor and a memory storing program instructions, and the processor is configured to execute the above-mentioned improved OpenCL heterogeneous programming framework memory management method when running the program instructions.

[0018] The improved OpenCL heterogeneous programming framework memory management method and device provided by the disclosed embodiment can achieve the following technical effects: the present invention performs dynamic memory partitioning, optimizes the allocation of global memory by adopting the front-end, middle-end, and back-end caching methods, pre-caches the global memory in the front-end, middle-end, and back-end, and when applying for global memory, gives priority to applying from the proximal side, which improves the efficiency of memory allocation compared to the original global memory management method of OpenCL. A corresponding cost model is established for the memory access sequence, and a memory access plan generator is used to generate the corresponding memory access sequence, the results are evaluated by an evaluator, and after memory access optimization, the optimal memory access execution sequence is generated.

[0019] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments, and in which:

[0021] Figure 1It is a schematic diagram of pre-allocating global memory;

[0022] Figure 2 It is a schematic diagram of the front-end, middle-end, and back-end;

[0023] Figure 3 This is a schematic diagram of the FreeList cache;

[0024] Figure 4 It is a global memory allocation flow chart;

[0025] Figure 5 It is a schematic diagram of intelligent memory access optimization. DETAILED DESCRIPTION

[0026] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0027] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0028] Unless otherwise stated, the term "plurality" means two or more.

[0029] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0030] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0031] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0032] Example 1

[0033] This embodiment discloses an improved OpenCL heterogeneous programming framework memory management method. In view of the conflict between multiple SPs applying for global memory, this method pre-allocates a fixed size of global memory for each SP process, such as Figure 1 As shown, the global memory is pre-cached in the front-end, middle-end and back-end. The global memory of the front-end is exclusive to a single SP, the global memory of the middle-end is shared by SPs in the same workgroup, and the global memory of the back-end is shared by all SPs or applied for from DRAM. When the SP applies for the use of global memory, it first applies to the SP's exclusive front-end global memory. When the front-end global memory cannot meet the application conditions, it applies to the middle-end global memory. When the middle-end global memory cannot meet the application conditions, the back-end global memory is used or obtained from the DRM.

[0034] The global memory of the front-end cache is exclusive to a single SP, and there will be no memory allocation conflict. Therefore, allocation and release are very fast. The global memory provided by the middle platform is shared by a group of SPs. Therefore, adding a mutex lock to protect its use will result in a certain performance loss compared to the global memory allocation through the front-end cache. The global memory provided by the back-end is shared by all SPs or requested from DRM, and the memory allocation efficiency is the lowest.

[0035] like Figure 2 As shown, this method is designed so that the middle platform is responsible for providing cache to the front end, and the back end is responsible for providing cache to the middle platform for use. When the global memory of the front end cache is insufficient, apply from the cache list of the middle platform and supplement the front end cache list; when the global memory of the middle platform cache is insufficient, apply from the cache list of the back end and supplement the middle platform cache list; when the global memory in the back end cache list is insufficient, apply from DRM.

[0036] Taking load balancing into consideration, the middle platform allows memory to be exchanged between the two front-end caches. Assuming that SP No. 1 has excess memory and SP No. 2 has insufficient memory, the middle platform allows part of the pre-allocated memory of SP No. 1 to be adjusted to SP No. 2, so as to adapt to changes in load.

[0037] This embodiment manages the entire global memory in pages (hereinafter referred to as Pages), and the default size of Pages is 4KB. Users are supported to set corresponding parameters to meet different needs. The larger the Page, the faster the global memory application speed is, but the more global memory waste is caused. The smaller the Page, the less global memory waste is, but the efficiency of global memory allocation will decrease. The front-end, middle-end, and back-end all use FreeList to maintain the cache, such as Figure 3 As shown, the basic structure in FreeList is Chunk. Pages of the same size are suspended under each Chunk, and different Chunks maintain Pages of different sizes.

[0038] When allocating global memory, first search in the global memory cached by the front end, traverse the Chunks in the FreeList in order, find the first Chunk that meets the conditions, and obtain the global memory that meets the specified size from the Chunk. If the global memory corresponding to the Chunk is larger than the required global memory, the global memory that meets the requirements is provided to the user, and the remaining global memory is placed in the corresponding Chunk according to the size. The specific process of global memory allocation is as follows: Figure 4 As shown. First, determine the size of the memory application. If the memory applied is within the small memory range specified in advance, determine whether the Front Cache in the current SP is free. If it is free, traverse the Front Cache and determine whether a Chunk that meets the requirements is found. If found, determine whether the found Chunk is larger than the memory applied. If so, apply for memory and put the excess memory into the corresponding free Chunk block. Otherwise, apply for memory and end the process. When the memory applied is within the memory range specified in advance or the Front Cache in the current SP is not free or no Chunk that meets the requirements is found by traversing the Front Cache, determine whether there is any remaining in the Mid Cache. If there is any remaining, traverse the Mid Cache and determine whether a Chunk that meets the requirements is found. If a Chunk that meets the requirements is found, apply for memory and update the Front Cache and end the process. When the memory applied is within the large memory range specified in advance or there is no remaining in the Mid Cache or no Chunk that meets the requirements is found by traversing the Mid Cache, determine whether there is any remaining in the Back Cache. If so, traverse the Back Cache. Cache and determine whether a Chunk that meets the requirements is found. If a Chunk that meets the requirements is found, apply for memory and update Mid Cache. If there is no Chunk left in Back Cache or no Chunk that meets the requirements is found after traversing Back Cache, apply for DRM and determine whether there is any DRM left. If so, apply for memory and update Back Cache. If there is no DRM left, the memory application fails and the process ends.

[0039] During the actual operation, check whether the size of the global memory requested is consistent with the current actual Chunk combination. For example, if the program always (N times, N is greater than the set number threshold) requests a global memory larger than 64KB, the global memory smaller than 64KB will be merged and placed in a Chunk sequence of the corresponding size after merging. 64KB is a variable value and can be set to other values ​​according to the actual situation.

[0040] Considering that the corresponding design wastes a certain degree of global memory, the LRU algorithm is used to release caches that are often not used to reduce the number of memory fragments. The specific implementation is to detect the usage frequency of pages under different Chunks during the operation of the computing card, and release caches with a usage frequency less than the set value based on the LRU algorithm.

[0041] The method also includes intelligent memory access optimization, which analyzes the execution mode and data dependency of work items, intelligently adjusts the memory access order, and reduces cache misses and memory access delays. Figure 5 As shown, a complete search space is established for all combinations of the current memory access sequence, and a cost model is established based on the time overhead of memory access, memory usage, and data transmission bandwidth. A memory access plan generator is used to generate a corresponding memory access plan, and an evaluator is used to score the generated memory access plan, and the cost model is updated according to the scoring result. In this embodiment, the memory access plan generator generates the corresponding memory access plan in the prior art, which is mainly implemented by operating the actual physical memory through the input memory access sequence. After updating the cost model, if the score of the current memory access sequence is the current optimal record, the current memory access sequence is selected as the optimal record, otherwise the optimal record is maintained unchanged.

[0042] Let T be the time cost of memory access; M be the memory usage, B be the bandwidth usage; define the following cost factor: C time is the cost per unit of memory access time, C mem is the cost per unit of memory usage, C band is the cost per unit of bandwidth used, and the total cost C total It can be expressed as the weighted sum of all cost factors, namely C total =C time *T+C mem *M+C band *B, this formula is the cost model described above.

[0043] Since the memory access time and the memory bandwidth usage are both nonlinear with the cost, nonlinear calculation is introduced here. The calculation of the total cost model is described as C total =∫C time (T)*T+C mem *M+∫C band (B) dB, where C band (B) represents the function of bandwidth cost changing with bandwidth usage, C time (T) represents the function of computational cost varying with memory access time.

[0044] Example 2

[0045] This embodiment discloses an improved OpenCL heterogeneous programming framework memory management device, including a processor and a memory. Optionally, the device may also include a communication interface and a bus. Among them, the processor, the communication interface, and the memory can communicate with each other through the bus. The communication interface can be used for information transmission. The processor can call the logic instructions in the memory to execute the memory management method of Example 1. The internal management method pre-caches the global memory in the front end, the middle stage, and the back end. The global memory of the front end is unique to a single SP, the global memory of the middle stage is shared by SPs in the same work group, and the global memory of the back end is shared by all SPs or applied from DRAM. When the SP applies to use the global memory, it is preferred to apply in the front end global memory unique to the SP. When the global memory of the front end cannot meet the application conditions, it is applied in the global memory of the middle stage. When the global memory of the middle stage cannot meet the application conditions, the global memory of the back end is used or obtained from the DRM.

[0046] The global memory of the front-end cache is exclusive to a single SP, and there will be no memory allocation conflict. Therefore, allocation and release are very fast. The global memory provided by the middle platform is shared by a group of SPs. Therefore, adding a mutex lock to protect its use will result in a certain performance loss compared to the global memory allocation through the front-end cache. The global memory provided by the back-end is shared by all SPs or requested from DRM, and the memory allocation efficiency is the lowest.

[0047] like Figure 2 As shown, the method is designed that the middle station is responsible for providing cache to the front end, and the back end is responsible for providing cache to the middle station for use. When the global memory of the front end cache is insufficient, apply from the cache list of the middle station and supplement the front end cache list; when the global memory of the middle station cache is insufficient, apply from the cache list of the back end and supplement the middle station cache list; when the global memory in the back end cache list is insufficient, apply from DRM.

[0048] Taking load balancing into consideration, the middle platform allows memory to be exchanged between the two front-end caches. Assuming that SP No. 1 has excess memory and SP No. 2 has insufficient memory, the middle platform allows part of the pre-allocated memory of SP No. 1 to be adjusted to SP No. 2, so as to adapt to changes in load.

[0049] This embodiment manages the entire global memory in pages (hereinafter referred to as Pages), and the default size of Pages is 4KB. Users are supported to set corresponding parameters to meet different needs. The larger the Page, the faster the global memory application speed is, but the more global memory waste is caused. The smaller the Page, the less global memory waste is, but the efficiency of global memory allocation will decrease. The front-end, middle-end, and back-end all use FreeList to maintain the cache, such as Figure 3 As shown, the basic structure in FreeList is Chunk. Pages of the same size are suspended under each Chunk, and different Chunks maintain Pages of different sizes.

[0050] When allocating global memory, first search in the global memory cached by the front end, traverse the Chunks in the FreeList in order, find the first Chunk that meets the conditions, and obtain the global memory that meets the specified size from the Chunk. If the global memory corresponding to the Chunk is larger than the required global memory, the global memory that meets the requirements is provided to the user, and the remaining global memory is placed in the corresponding Chunk according to the size. The specific process of global memory allocation is as follows: Figure 4 shown.

[0051] During the actual operation, check whether the size of the global memory requested is consistent with the current actual Chunk combination. For example, if the program always (N times, N is greater than the set number threshold) requests a global memory larger than 64KB, the global memory smaller than 64KB will be merged and placed in a Chunk sequence of the corresponding size after merging. 64KB is a variable value and can be set to other values ​​according to the actual situation.

[0052] Considering that the corresponding design wastes a certain degree of global memory, the LRU algorithm is used to release caches that are often not used to reduce the number of memory fragments. The specific implementation is to detect the usage frequency of pages under different Chunks during the operation of the computing card, and release caches with a usage frequency less than the set value based on the LRU algorithm.

[0053] The method also includes intelligent memory access optimization, which analyzes the execution mode and data dependency of work items, intelligently adjusts the memory access order, and reduces cache misses and memory access delays. Figure 5As shown, a complete search space is established for all combinations of the current memory access sequence, and a cost model is established based on the time overhead of memory access, memory usage, and data transmission bandwidth. A memory access plan generator is used to generate a corresponding memory access plan, and an evaluator is used to score the generated memory access plan, and the cost model is updated according to the scoring result. In this embodiment, the memory access plan generator generates the corresponding memory access plan, which is a prior art, and is mainly implemented by operating the actual physical memory through the input memory access sequence. After updating the cost model, if the score of the current memory access sequence is the current optimal record, the current memory access sequence is selected as the optimal record, otherwise the optimal record is maintained unchanged.

[0054] Let T be the time cost of memory access; M be the memory usage, B be the bandwidth usage; define the following cost factor: C time is the cost per unit of memory access time, C mem is the cost per unit of memory usage, C band is the cost per unit of bandwidth used, and the total cost C total It can be expressed as the weighted sum of all cost factors, namely C total =C time *T+C mem *M+C band *B, this formula is the cost model described above.

[0055] Since the memory access time and the memory bandwidth usage are both nonlinear with the cost, nonlinear calculation is introduced here. The calculation of the total cost model is described as C total =∫C time (T)*T+C mem *M+∫C band (B) dB, where C band (B) represents the function of bandwidth cost changing with bandwidth usage, C time (T) represents the function of computational cost varying with memory access time.

[0056] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0057] The memory, as a computer-readable storage medium, can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor executes functional applications and data processing by running the program instructions / modules stored in the memory, that is, implementing the memory management method in the above embodiment.

[0058] Example 3

[0059] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned memory management method.

[0060] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0061] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes, or a transient storage medium.

[0062] The above description and accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent possible changes only. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the scope of protection. As used in the description in the text, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0063] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians may clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0064] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

Claims

1. An improved OpenCL heterogeneous programming framework memory management method, characterized by: This method pre-caches the global memory in the front-end, middle-end and back-end. The global memory of the front-end is exclusive to a single SP, the global memory of the middle-end is shared by SPs in the same workgroup, and the global memory of the back-end is shared by all SPs or applied for from DRAM. When the SP applies for the use of global memory, it first applies to the SP's exclusive front-end global memory. When the front-end global memory cannot meet the application conditions, it applies to the middle-end global memory. When the middle-end global memory cannot meet the application conditions, the back-end global memory is used or obtained from the DRM.

2. The improved OpenCL heterogeneous programming framework memory management method according to claim 1, characterized in that: This method is designed in that the middle platform provides cache for the front end, and the back end provides cache for the middle platform. When the global memory of the front end is insufficient, it applies from the cache list of the middle platform and supplements the front end cache list; when the global memory of the middle platform is insufficient, it applies from the cache list of the back end and supplements the middle platform cache list. When the global memory in the back end cache list is insufficient, it applies from the DRM; the middle platform allows global memory to be exchanged between the two front ends. For example, if SP No. 1 has excess memory and SP No. 2 has insufficient memory, the middle platform allows the front end cache part of SP No. 1 to be adjusted to SP No.

2.

3. The improved OpenCL heterogeneous programming framework memory management method according to claim 1, characterized in that: The entire global memory is managed according to pages. The front-end, middle-end, and back-end all use FreeList to maintain the cache. The basic structure in FreeList is Chunk. Pages of the same size are suspended under each Chunk, and different Chunks maintain pages of different sizes.

4. The improved OpenCL heterogeneous programming framework memory management method according to claim 3, characterized in that: When allocating global memory, first search in the global memory cached on the front end, traverse the Chunks in the FreeList in order, find the first Chunk that meets the conditions, and obtain the global memory that meets the requirements from the Chunk. If the global memory corresponding to the Chunk is larger than the required global memory, the global memory that meets the requirements is provided to the user, and the remaining global memory is placed in the corresponding free Chunk according to size.

5. The improved OpenCL heterogeneous programming framework memory management method according to claim 4, characterized in that: During the operation of the computing card, it is checked whether the size of the applied global memory is consistent with the current actual Chunk combination. If the size of the applied global memory is N times larger than a fixed value of global memory, where N is a positive integer, the global memory smaller than the fixed value will be merged and placed in a Chunk sequence of the corresponding size after merging.

6. The improved OpenCL heterogeneous programming framework memory management method according to claim 3, characterized in that: During the operation of the computing card, the usage frequency of pages under different chunks is detected, and the cache with a usage frequency less than the set value is released based on the LRU algorithm.

7. The improved OpenCL heterogeneous programming framework memory management method according to claim 1, characterized in that: It also includes intelligent memory access optimization, analyzing the execution mode and data dependency of work items, intelligently adjusting the memory access order, and reducing cache misses and memory access latency. Specifically, a complete search space is established for all combinations of the current memory access order, and a cost model is established based on the time overhead of memory access, memory usage, and data transmission bandwidth. The cost model is: C total =C time *T+C mem *M+C band *B, where C total is the total cost, C time is the cost per unit of memory access time, C mem is the cost per unit of memory usage, C band is the cost per unit bandwidth usage, T is the time overhead of memory access, M is the memory usage, B is the bandwidth occupancy, and the memory access plan generator is used to generate the corresponding memory access plan, and the evaluator is used to score the generated memory access plan, and the cost model is updated according to the scoring result.

8. The improved OpenCL heterogeneous programming framework memory management method according to claim 7, characterized in that: Introducing nonlinear calculation, the cost model is transformed into: C total =∫C time (T)*T+C mem *M+∫C band (B) dB, where C band (B) represents the function of bandwidth cost changing with bandwidth usage, C time (T) represents the function of computational cost varying with memory access time.

9. The improved OpenCL heterogeneous programming framework memory management method according to claim 1, characterized in that: Different middle platforms are protected by mutex locks.

10. An improved OpenCL heterogeneous programming framework memory management device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the improved OpenCL heterogeneous programming framework memory management method according to any one of claims 1 to 9 when running the program instructions.