High-Bandwidth Memory Expansion Method, Device, Equipment, and Storage Medium

By connecting the storage space of the expansion card device to the high-bandwidth memory, and performing data replication and parallel computing in the kernel function of the graphics processor, the problem of insufficient capacity of high-bandwidth memory is solved, and the expansion of high-bandwidth memory and the improvement of GPU computing efficiency is achieved.

CN119759296BActive Publication Date: 2025-06-20SHENZHEN QUANXING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510262090.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing high-bandwidth memory capacity is insufficient, making it difficult to meet the needs of large-scale data model processing.

Method used

The expansion card device is connected to the graphics processor through the PCIe interface on the graphics processor, so that the storage space of the expansion card device is used as a shared storage space for high-bandwidth memory. In the kernel function of the graphics processor, the shared keyword is used to declare the shared storage space and generate an extended declaration. Then, the memory data of the high-bandwidth memory is copied to the shared storage space of the expansion card device through the threads in the thread block, and thread synchronization is performed on the thread block through the synchronization function. Finally, parallel computing processing is performed on the shared data set on the high-bandwidth memory and the expansion card device to achieve capacity expansion of the high-bandwidth memory.

Benefits of technology

It effectively expands the capacity of high-bandwidth memory, meets the needs of big data and large model processing, improves GPU computing efficiency, and reduces expansion costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759296B_ABST
    Figure CN119759296B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment and storage medium for expanding the capacity of a high-bandwidth memory. The method includes: connecting an expansion card device to a graphics processing unit through a PCIe interface on the graphics processing unit, so that the storage space of the expansion card device serves as a shared storage space for a high-bandwidth memory (HBM). First, in the kernel function of the graphics processing unit, the shared storage space is declared using the shared keyword to generate an extended declaration. Then, the HBM memory data is copied to the shared storage space of the expansion card device through the threads in the thread block, and a synchronization function is used to synchronize the threads of the thread block. On this basis, parallel computing processing is performed, and the calculation results are allocated and copied between the HBM and the shared storage space of the expansion card device to achieve storage expansion. The present invention can effectively expand the capacity of the HBM, meet the processing requirements of big data and large models, improve the GPU computing efficiency, and reduce the expansion cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graphics processors, and particularly to a method, device, equipment, and storage medium for expanding the capacity of a high-bandwidth memory. Background Art

[0002] With the increasingly wide application of various types of big data and large models, data storage and computing performance have become key factors affecting the computing efficiency of graphics processors (GPUs). As a high-speed cache storage medium for core GPU computing, high-bandwidth memory (HBM) greatly affects the amount of data that can be processed during artificial intelligence training and inference. However, the HBM integrated in existing GPUs has limited capacity and usually can only support relatively small data models, unable to meet the computing requirements of current artificial intelligence models. Although expanding the HBM capacity by parallelly connecting multiple GPUs can solve some problems, this method has a high investment cost and poor adaptability to various application scenarios. Summary of the Invention

[0003] The main purpose of the present invention is to solve the technical problem that the existing high-bandwidth memory has insufficient capacity and is difficult to meet the processing requirements of large-scale data models;

[0004] In a first aspect of the present invention, a method for expanding the capacity of a high-bandwidth memory is provided. A high-bandwidth memory is integrated on the graphics processor, and the method for expanding the capacity of the high-bandwidth memory includes:

[0005] Connect an expansion card device to the graphics processor through a PCIe interface on the graphics processor, so as to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory;

[0006] In the kernel function of the graphics processor, perform a declaration process on the shared storage space according to a preset shared keyword to obtain an extended declaration of the expansion card device;

[0007] In the kernel function, according to the extended declaration, copy the memory data of the high-bandwidth memory to the shared storage space of the expansion card device through threads in a preset thread block, and perform thread synchronization on the thread block through a preset synchronization function;

[0008] Perform parallel computing processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain calculation result data, and allocate and copy the calculation result data between the shared storage spaces of the high-bandwidth memory and the expansion card device to achieve the expansion of the high-bandwidth memory.

[0009] Optionally, in the first implementation manner of the first aspect of the present invention, connecting the expansion card device to the graphics processor through the PCIe interface on the graphics processor to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory includes:

[0010] Connect the expansion card device to the graphics processor through the PCIe interface on the graphics processor, and perform initialization processing on the root complex of the graphics processor to obtain a system topology structure connecting the CPU, memory subsystem, and PCIe device;

[0011] According to the system topology structure, perform enumeration processing on the PCIe bus of the PCIe interface to obtain a function list containing different combinations of bus numbers and device numbers;

[0012] Based on the function list, perform a read process on the configuration space of the expansion card device to obtain the identity identification information and hardware characteristics of the expansion card device;

[0013] According to the identity identification information and hardware characteristics, perform configuration processing on the base address register of the expansion card device to obtain the mapping position of the expansion card device in the memory address space of the graphics processor;

[0014] Based on the mapping position, update the memory-mapped I / O system of the graphics processor to obtain an extended memory access mechanism including the storage space of the expansion card device, so as to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory.

[0015] Optionally, in the second implementation manner of the first aspect of the present invention, the declaration process of using the shared storage space according to a preset shared keyword in the kernel function of the graphics processor to obtain the extended declaration of the expansion card device includes:

[0016] Perform preprocessing on the kernel function code of the graphics processor to obtain a code segment of the shared keyword, and insert the identifier of the expansion card device into the code segment to obtain a shared memory declaration statement for extended memory;

[0017] Perform syntax parsing processing on the shared memory declaration statement to obtain a parameter list including the storage space size of the expansion card device;

[0018] According to the parameter list, detect the firmware version of the expansion card device, and based on the firmware version, perform conditional compilation processing on the extended shared memory declaration statement to obtain the extended declaration of the expansion card device adapted to the firmware version.

[0019] Optionally, in the third implementation manner of the first aspect of the present invention, in the kernel function, copying the memory data of the high-bandwidth memory to the shared storage space of the expansion card device through the threads in a preset thread block according to the extended declaration, and performing thread synchronization on the thread block through a preset synchronization function includes:

[0020] In the kernel function, determining the thread block index and thread index in a preset thread block according to the extended declaration;

[0021] According to the thread block index and thread index, determining the global memory access address corresponding to each thread;

[0022] Through the threads in the thread block, copying the memory data in the high-bandwidth memory to the shared storage space of the expansion card device according to the global memory access address, and performing thread synchronization on the thread block through a preset synchronization function.

[0023] Optionally, in the fourth implementation manner of the first aspect of the present invention, performing thread synchronization on the thread block through a preset synchronization function includes:

[0024] Performing a marking process on the execution status of each thread in the thread block through a preset synchronization function to obtain a status identifier for the thread reaching the synchronization point;

[0025] According to the status identifier, performing an update process on the thread counter in the thread block to obtain the number of threads that have reached the synchronization point;

[0026] Performing a judgment process on the number of threads that have reached the synchronization point to obtain a judgment result on whether all threads in the thread block have reached the synchronization point, and performing thread synchronization on the thread block according to the judgment result.

[0027] Optionally, in the fifth implementation manner of the first aspect of the present invention, performing parallel computing processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain calculation result data, and allocating and copying the calculation result data between the shared storage spaces of the high-bandwidth memory and the expansion card device to implement the expansion of the high-bandwidth memory includes:

[0028] Performing a task partitioning process on the shared data set on the high-bandwidth memory and the expansion card device to obtain data subsets suitable for parallel computing and corresponding calculation instruction sets;

[0029] According to the data subsets and calculation instruction sets, performing parallel scheduling processing on the computing units of the graphics processor to obtain parallel computing results executed simultaneously on the high-bandwidth memory and the expansion card device;

[0030] Perform data distribution analysis on the parallel computing results to obtain an optimal allocation scheme for the computing result data in the high-bandwidth memory and the shared storage space of the expansion card device;

[0031] According to the optimal allocation scheme, perform selective copying on the computing result data to realize the expansion of the high-bandwidth memory.

[0032] Optionally, in the sixth implementation manner of the first aspect of the present invention, the task partitioning of the shared data set on the high-bandwidth memory and the expansion card device to obtain a data subset suitable for parallel computing and the corresponding computing instruction set includes:

[0033] Analyze the storage characteristics of the high-bandwidth memory and the expansion card device to obtain a storage performance model including access latency and bandwidth parameters;

[0034] According to the storage performance model, perform data dependence analysis on the shared data set to obtain parallel-executable data blocks and the corresponding computing task graph;

[0035] Based on the data blocks and the computing task graph, perform optimized allocation on the computing tasks to obtain a data subset evenly distributed on the high-bandwidth memory and the expansion card device and the corresponding computing instruction set.

[0036] The second aspect of the present invention provides a high-bandwidth memory expansion device, which is applied to a graphics processor, and the graphics processor is integrated with a high-bandwidth memory. The high-bandwidth memory expansion device includes:

[0037] A connection module for connecting the expansion card device to the graphics processor through the PCIe interface on the graphics processor to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory;

[0038] A declaration module for using a preset shared keyword in the kernel function of the graphics processor to perform declaration processing on the shared storage space to obtain an extended declaration of the expansion card device;

[0039] A copy synchronization module for, in the kernel function, copying the memory data of the high-bandwidth memory to the shared storage space of the expansion card device through the threads in a preset thread block according to the extended declaration, and performing thread synchronization on the thread block through a preset synchronization function;

[0040] A computing module for performing parallel computing on the shared data set on the high-bandwidth memory and the expansion card device to obtain computing result data, and allocating and copying the computing result data between the shared storage spaces of the high-bandwidth memory and the expansion card device to realize the expansion of the high-bandwidth memory.

[0041] In the third aspect of the present invention, a high-bandwidth memory expansion device is provided, including: a memory and at least one processor. Instructions are stored in the memory, and the memory and the at least one processor are interconnected through a line; the at least one processor invokes the instructions in the memory to cause the high-bandwidth memory expansion device to execute the steps of the above-mentioned high-bandwidth memory expansion method.

[0042] In the fourth aspect of the present invention, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the steps of the above-mentioned high-bandwidth memory expansion method.

[0043] For the above-mentioned high-bandwidth memory expansion method, device, equipment and storage medium, the expansion card device is connected to the graphics processor through the PCIe interface on the graphics processor, so that the storage space of the expansion card device serves as the shared storage space of the high-bandwidth memory (HBM). First, in the kernel function of the graphics processor, the shared keyword is used to declare the shared storage space to generate an extended declaration. Then, the HBM memory data is copied to the shared storage space of the expansion card device through the threads in the thread block, and a synchronization function is used to synchronize the threads of the thread block. On this basis, parallel computing processing is performed, and the calculation results are allocated and copied between the shared storage spaces of the HBM and the expansion card device to achieve storage expansion. The present invention can effectively expand the capacity of the HBM, meet the processing requirements of big data and large models, improve the GPU computing efficiency, and reduce the expansion cost.

[0044] Other features and advantages of the present invention will be described in the subsequent description, and part of them will become obvious from the description, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, claims and drawings.

[0045] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. Description of the Drawings

[0046] Figure 1 It is a schematic diagram of the first embodiment of the high-bandwidth memory expansion method in the embodiment of the present invention;

[0047] Figure 2 It is a schematic diagram of an embodiment of the high-bandwidth memory expansion device in the embodiment of the present invention;

[0048] Figure 3 It is a schematic diagram of an embodiment of the high-bandwidth memory expansion equipment in the embodiment of the present invention. Detailed Embodiments

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] As used in the embodiments of the present invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0051] For ease of understanding of this embodiment, a high-bandwidth memory expansion method disclosed in the embodiments of the present invention will be introduced in detail first. As Figure 1 shown, this method is applied to a graphics processing unit, and a high-bandwidth memory is integrated on the graphics processing unit, and includes the following steps:

[0052] 101. Connect an expansion card device to the graphics processing unit through the PCIe interface on the graphics processing unit, so as to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory;

[0053] In an embodiment of the present invention, the step of connecting an expansion card device to the graphics processing unit through the PCIe interface on the graphics processing unit and using the storage space in the expansion card device as the shared storage space of the high-bandwidth memory includes: connecting an expansion card device to the graphics processing unit through the PCIe interface on the graphics processing unit, and initializing the root complex of the graphics processing unit to obtain a system topology structure connecting the CPU, memory subsystem, and PCIe device; according to the system topology structure, performing an enumeration process on the PCIe bus of the PCIe interface to obtain a function list including different combinations of bus numbers and device numbers; based on the function list, performing a read process on the configuration space of the expansion card device to obtain the identity information and hardware characteristics of the expansion card device; according to the identity information and hardware characteristics, performing a configuration process on the base address register of the expansion card device to obtain the mapping position of the expansion card device in the memory address space of the graphics processing unit; based on the mapping position, updating the memory-mapped I / O system of the graphics processing unit to obtain an extended memory access mechanism including the storage space of the expansion card device, so as to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory.

[0054] Connect the expansion card device to the graphics processor through the PCIe interface on the graphics processor, and initialize the root complex of the graphics processor to obtain a system topology connecting the CPU, memory subsystem, and PCIe devices. First, it involves hardware connection, physically inserting the expansion card device into the PCIe slot of the graphics processor. Subsequently, at system startup, the BIOS or UEFI firmware will recognize the newly inserted PCIe device and allocate basic resources to it. When the operating system is loaded, the driver of the graphics processor is activated and starts initializing the root complex. The root complex is a core component in the PCIe architecture, responsible for connecting the CPU and PCIe devices. During the initialization process, the driver will scan all connected PCIe devices, including the newly inserted expansion card device. It will read the configuration space of each device to obtain information such as device type, function, and resource requirements. Based on this information, the driver constructs a detailed system topology that describes the physical and logical connection relationships between the CPU, memory subsystem, and all PCIe devices. This topology is the basis for subsequent resource allocation and device management, ensuring that the system can correctly identify and utilize the expansion card device.

[0055] Specifically, according to the system topology, enumerate the PCIe buses of the PCIe interface to obtain a function list containing different combinations of bus numbers and device numbers. PCIe bus enumeration is a system-level operation, usually performed by the operating system or a dedicated device driver. The enumeration process starts from the root complex and probes down along the PCIe bus structure level by level. For each possible combination of bus number and device number, the system attempts to read the configuration space of the device, especially the vendor ID (VID) and device ID (DID). If the read is successful, it indicates that there is a valid PCIe device at that location. The system will record the bus number, device number, and function number (BDF) of this device, as well as other key information such as device type and interrupt line. During the probing process, the system will also identify PCIe bridges, which connect different PCIe bus segments. By recursively probing these bus segments, the system can discover all connected PCIe devices, including those deeply nested in a complex topology. Finally, the system generates a comprehensive function list that contains all discovered PCIe devices and their detailed information. This list is an important basis for subsequent device configuration and resource allocation.

[0056] Specifically, based on the function list, the configuration space of the expansion card device is read to obtain the identity information and hardware characteristics of the expansion card device. The configuration space is a standardized memory area for PCIe devices, containing 256 bytes (for traditional devices) or 4096 bytes (for extended-capability devices) of information. The reading process starts from the standard configuration header, which contains the basic information of the device. The system first reads the vendor ID and device ID, and these two values uniquely identify the manufacturer and specific model of the device. Then, the system reads the class code to determine the functional type of the device (such as graphics card, network card, etc.). Next, the system reads the status register of the device to understand the current status and supported functions of the device. For the expansion card device, it is particularly important to read its memory requirement information, which is usually stored in the base address register (BAR). The system also reads the interrupt line information, maximum payload size, power management capabilities of the device, etc. If the device supports extended functions, the system further reads the extended capability pointer and traverses all extended capability structures to obtain more advanced feature information. This comprehensive reading process ensures that the system obtains all the key information required to configure and use the expansion card device.

[0057] Specifically, based on the identity information and hardware characteristics, the base address register of the expansion card device is configured to obtain the mapping position of the expansion card device in the memory address space of the graphics processor. This process first involves analyzing the memory requirements of the expansion card device. The system checks each base address register (BAR) of the device to determine the type (memory-mapped or I / O-mapped) and size of the memory space required. For the memory-mapped type, the system also needs to consider factors such as whether prefetch support is required and whether 64-bit addressing is required. Then, the system searches for an idle area in the global memory address space of the graphics processor that meets these requirements. This search process needs to consider alignment requirements and possible address conflicts. After finding a suitable address range, the system writes the start address of this range to the corresponding BAR of the device. If the device has multiple BARs, this process is repeated multiple times. For 64-bit BARs, the system needs to configure two consecutive 32-bit registers simultaneously. After the configuration is completed, the system reads the BAR again to verify whether the configuration is successful. If the verification passes, the system updates the internal data structure to record this newly allocated address mapping. This configuration process essentially establishes the mapping relationship between the physical memory of the expansion card device and the virtual address space accessible by the graphics processor.

[0058] Specifically, based on the mapping location, the memory-mapped I / O system of the graphics processing unit is updated to obtain an extended memory access mechanism that includes the storage space of the expansion card device. This step mainly involves updating the memory management unit (MMU) of the graphics processing unit and related software data structures. First, the system needs to add the memory mapping information of the expansion card device to the page table of the graphics processing unit. This includes creating new page table entries to map the virtual address range to the physical address of the expansion card device. At the same time, the system may need to adjust the configuration of the MMU, such as updating the TLB (Translation Lookaside Buffer), etc. Next, the system updates the memory management-related code in the device driver so that it can recognize and use these newly mapped memory areas. This may involve modifying the memory allocator, updating the DMA controller configuration, etc. The system also needs to establish an interrupt handling mechanism to ensure that the expansion card device can correctly trigger and process interrupts. In addition, the system may need to implement a memory barrier or cache coherence mechanism to ensure data consistency between the graphics processing unit and the expansion card device. Finally, the system may need to update the power management and thermal management policies to accommodate the addition of the expansion card device. Through these updates, the storage space of the expansion card device is successfully integrated into the memory system of the graphics processing unit, forming a seamless extended memory access mechanism.

[0059] 102. In the kernel function of the graphics processing unit, use a preset shared keyword to declare and process the shared storage space to obtain an extended declaration of the expansion card device;

[0060] In one embodiment of the present invention, the using a preset shared keyword to declare and process the shared storage space in the kernel function of the graphics processing unit to obtain the extended declaration of the expansion card device includes: preprocessing the kernel function code of the graphics processing unit to obtain a code segment of the shared keyword, and inserting the identifier of the expansion card device into the code segment to obtain a shared memory declaration statement for the extended memory; performing syntax parsing processing on the shared memory declaration statement to obtain a parameter list including the storage space size of the expansion card device; according to the parameter list, detecting the firmware version of the expansion card device, and based on the firmware version, performing conditional compilation processing on the extended shared memory declaration statement to obtain the extended declaration of the expansion card device adapted to the firmware version.

[0061] Specifically, the kernel function code of the graphics processor is preprocessed to obtain a code segment with shared keywords, and the identifier of the expansion card device is inserted into the code segment to obtain a shared memory declaration statement for extended memory. This first involves lexical and syntactic analysis of the CUDA kernel function code. The preprocessor scans the entire kernel function code and identifies all code segments that use the __shared__ keyword. These code segments typically represent statements for declaring variables in shared memory. Next, the preprocessor inserts a specific identifier of the expansion card device into these identified code segments. This identifier may be a predefined macro or a special comment used to mark that this part of the shared memory declaration is for the expansion card device. The insertion process needs to consider syntactic correctness to ensure that the inserted code still conforms to the CUDA syntax specification. In addition, the preprocessor also needs to handle conditional compilation directives that may exist, such as #ifdef or #if, to ensure that the declaration of the expansion card device can be correctly processed under different compilation conditions. Finally, this step will generate a modified code segment that contains the original shared memory declaration and the newly added expansion card device-related declarations.

[0062] Specifically, syntactic parsing is performed on the shared memory declaration statement to obtain a parameter list containing the storage space size of the expansion card device, which involves deeper syntactic analysis. The parser analyzes the modified shared memory declaration statement in detail and extracts key information. First, the parser needs to identify the variable type, variable name, and array size (if applicable) of the declaration. For the declaration of the expansion card device, the parser needs to pay special attention to the inserted identifier and extract relevant additional information based on this identifier. This may include the model of the expansion card device, the expected storage space size, etc. During the parsing process, CUDA-specific syntactic elements such as template parameters and const qualifiers also need to be processed. The parser constructs an abstract syntax tree (AST) to represent the structure and relationships of these declarations. From the AST, the parser extracts a parameter list that contains detailed information about each shared memory declaration, especially the storage space size of the expansion card device. This parameter list contains not only the statically declared size but also may contain dynamically calculated size information, which is crucial for subsequent memory allocation and management.

[0063] 103. In the kernel function, according to the extended declaration, the memory data of the high-bandwidth memory is copied to the shared storage space of the expansion card device through the threads in the preset thread block, and thread synchronization is performed on the thread block through the preset synchronization function;

[0064] In one embodiment of the present invention, in the kernel function, according to the extended declaration, the memory data of the high bandwidth memory is copied to the shared storage space of the expansion card device through the threads in the preset thread block, and the thread block is thread-synchronized through the preset synchronization function, including: in the kernel function, according to the extended declaration, determining the thread block index and the thread index in the preset thread block; according to the thread block index and the thread index, determining the global memory access address corresponding to each thread; according to the global memory access address, by the threads in the thread block, copying the memory data in the high bandwidth memory to the shared storage space of the expansion card device, and thread-synchronizing the thread block through the preset synchronization function.

[0065] Specifically, in the kernel function, the thread block index and thread index in the preset thread block are determined according to the extension declaration. This process first involves parsing the thread block and thread structure defined in the extension declaration. In the CUDA programming model, each thread has a unique identifier, which consists of a thread block index and a thread index. The thread block index is obtained through the built-in variable blockIdx, which is a three-dimensional vector containing three components x, y, and z, indicating the position of the current thread block in the entire grid. The thread index is obtained through the built-in variable threadIdx, which is also a three-dimensional vector, indicating the position of the thread in the thread block to which it belongs. The kernel function calculates the global unique identifier of each thread through these built-in variables. In addition, the kernel function also needs to obtain the dimensions of the thread block (through the blockDim variable) and the dimensions of the grid (through the gridDim variable), which are used for subsequent address calculations. In actual implementation, the kernel function may store this index information in local variables for subsequent use. The key to this step is to ensure that each thread can correctly identify its own position and prepare for subsequent data copy operations.

[0066] Specifically, based on the thread block index and thread index, determine the global memory access address corresponding to each thread. This step involves mapping the logical index of the thread to the actual memory address. The calculation usually follows a certain pattern, and the most common one is one-dimensional mapping: converting the three-dimensional thread and block indices into a one-dimensional global thread ID. This can be achieved through the formula: globalThreadId = blockIdx.x * blockDim.x + threadIdx.x (for the one-dimensional case). For the multi-dimensional case, the calculation is more complex and needs to consider the y and z dimensions. After obtaining the global thread ID, map it to the actual memory address. This mapping relationship depends on the storage method and access pattern of the data in the global memory. For example, if the data is linearly stored, then the memory address can be simply calculated by: address = baseAddress + globalThreadId * elementSize, where baseAddress is the starting address of the data in the global memory and elementSize is the size of each data element. In some complex cases, more complex addressing algorithms may be required, especially when the data layout involves padding or alignment requirements. The goal of this step is to ensure that each thread can accurately calculate the global memory address it needs to access, preparing for the next data copy step.

[0067] Finally, the threads in the thread block copy the memory data in the high-bandwidth memory to the shared storage space of the expansion card device according to the global memory access address, and perform thread synchronization on the thread block through a preset synchronization function. This process first involves the actual data copy operation. Each thread reads data from the high-bandwidth memory using the previously calculated global memory address. The read operation usually uses global memory load instructions, such as the ld.global instruction used in CUDA. The read data is then written to the shared storage space of the expansion card device. The write operation needs to consider the bank conflict problem of the shared memory, and a reasonable data layout can significantly improve the access efficiency. After the data copy is completed, a preset synchronization function needs to be used to synchronize the thread block. In CUDA, this is usually achieved by calling the __syncthreads() function. This synchronization operation ensures that all threads in the thread block have completed the data copy before proceeding to the next step. Synchronization is necessary because the execution speeds of different threads may vary, and subsequent calculations may depend on all data being correctly loaded into the shared memory. In addition, synchronization also ensures memory consistency and prevents data race problems.

[0068] Further, the thread synchronization of the thread block through a preset synchronization function includes: marking the execution status of each thread in the thread block through a preset synchronization function to obtain a status identifier of the thread reaching the synchronization point; updating the thread counter within the thread block according to the status identifier to obtain the number of threads that have reached the synchronization point; making a judgment process on the number of threads that have reached the synchronization point to obtain a judgment result on whether all threads within the thread block have reached the synchronization point, and performing thread synchronization on the thread block according to the judgment result.

[0069] Specifically, marking the execution status of each thread in the thread block through a preset synchronization function to obtain a status identifier of the thread reaching the synchronization point. In the CUDA programming model, this is usually achieved through the __syncthreads() function. When a thread executes to the __syncthreads() call, it triggers a hardware-level synchronization mechanism. In a specific implementation, each thread sets its own status to "reached the synchronization point". This status identifier is usually stored in a special hardware register shared by the thread block, with one bit corresponding to each thread. When a thread executes to __syncthreads(), it sets the corresponding bit to 1. This operation is atomic to ensure correctness in a multi-threaded environment. At the same time, the hardware records the address of the last instruction executed by each thread, and this address serves as the identifier of the synchronization point. If some threads in the thread block do not execute to __syncthreads() due to conditional branches, then their corresponding status bits will not be set, which will be detected in the subsequent steps.

[0070] Specifically, updating the thread counter within the thread block according to the status identifier to obtain the number of threads that have reached the synchronization point. This step is automatically performed at the hardware level. The thread block scheduler of CUDA maintains a counter for tracking the number of threads that have reached the synchronization point. Whenever the status of a thread is marked as "reached the synchronization point", this counter automatically increments. The increment operation is atomic to ensure accuracy in a highly concurrent situation. The initial value of the counter is 0, and the maximum value is equal to the total number of threads in the thread block. This counting process is efficient because it is implemented at the hardware level without additional overhead at the software level. At the same time, the hardware also records the order in which each thread reaches the synchronization point, which is very important in some scenarios where the execution order needs to be guaranteed. If there are threads in the thread block that do not reach the synchronization point, then the value of the counter will not reach the total number of threads in the thread block, which will be detected in the next step.

[0071] Specifically, the number of threads that have reached the synchronization point is judged and processed to obtain the judgment result on whether all threads within the thread block have reached the synchronization point, and thread synchronization of the thread block is performed according to the judgment result. This step is also automatically executed at the hardware level. The thread block scheduler of CUDA continuously monitors the thread counter that has reached the synchronization point. When the value of the counter is equal to the total number of threads in the thread block, it indicates that all threads have reached the synchronization point. At this time, the hardware will trigger an internal signal indicating that the synchronization is completed. If the counter does not reach the total number of threads within the predetermined time, the hardware will consider that a deadlock or other error has occurred and may trigger an exception. Once it is confirmed that all threads have reached the synchronization point, the hardware will reset the execution states of all threads, allowing them to continue executing subsequent instructions. This process is atomic, ensuring that all threads start executing again at exactly the same moment. At the same time, the hardware also flushes the cache to ensure that the shared memory state seen by all threads is consistent. This hardware-level synchronization mechanism greatly improves efficiency and minimizes the overhead of thread synchronization.

[0072] 104. Perform parallel computing processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain the computed result data, and allocate and copy the computed result data between the shared storage spaces of the high-bandwidth memory and the expansion card device to achieve the expansion of the high-bandwidth memory.

[0073] In an embodiment of the present invention, the performing parallel computing processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain the computed result data, and allocating and copying the computed result data between the shared storage spaces of the high-bandwidth memory and the expansion card device to achieve the expansion of the high-bandwidth memory includes: performing task partitioning processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain data subsets suitable for parallel computing and corresponding computing instruction sets; performing parallel scheduling processing on the computing units of the graphics processor according to the data subsets and computing instruction sets to obtain the parallel computing results executed simultaneously on the high-bandwidth memory and the expansion card device; performing data distribution analysis processing on the parallel computing results to obtain the optimal allocation scheme of the computed result data in the shared storage spaces of the high-bandwidth memory and the expansion card device; and performing selective copying processing on the computed result data according to the optimal allocation scheme to achieve the expansion of the high-bandwidth memory.

[0074] Specifically, the shared data set on the high-bandwidth memory and the expansion card device is subjected to task partitioning processing to obtain data subsets suitable for parallel computing and corresponding computing instruction sets. This process first involves analyzing the structure and size of the shared data set. The system evaluates the distribution characteristics of the data, such as data continuity, access patterns, and dependencies. Based on this information, the system adopts specific task partitioning algorithms, such as block partitioning, loop partitioning, or dynamic partitioning, etc. For the data in the high-bandwidth memory, the system will consider its high-speed access characteristics and tend to allocate computationally intensive tasks to this part of the data. For the data on the expansion card device, the system will consider its lower access speed, which is more suitable for processing I / O-intensive tasks. The partitioning process also needs to consider load balancing to ensure that each computing unit can receive an appropriate amount of work. At the same time, the system will generate corresponding computing instruction sets, which contain the detailed steps on how to process each data subset. The generation of the instruction set needs to consider the location of the data to minimize data movement. Finally, a series of optimized data subsets and their corresponding computing instruction sets are produced.

[0075] Specifically, according to the data subsets and computing instruction sets, the computing units of the graphics processor are subjected to parallel scheduling processing to obtain parallel computing results that are executed simultaneously on the high-bandwidth memory and the expansion card device. This step first involves the work of the task scheduler. The scheduler analyzes each data subset and its corresponding computing instruction set, evaluates its computing complexity and resource requirements. Based on this information, the scheduler formulates an execution plan to decide which tasks are allocated to which computing units. For the data subsets in the high-bandwidth memory, the scheduler will preferentially allocate them to the stream processors on the GPU to utilize their high-speed data access capabilities. For the data subsets on the expansion card device, the scheduler may choose to allocate tasks to dedicated asynchronous computing units to mask the data access latency. The scheduling process also needs to consider the dependencies between tasks to ensure data consistency. During execution, each computing unit will process the tasks allocated to it in parallel, and at the same time, the system will monitor the execution progress and perform dynamic load balancing if necessary. During the computing process, the system will use the CUDA stream and event mechanisms to manage concurrent execution and synchronization. Finally, this step produces parallel computing results that are executed simultaneously on the high-bandwidth memory and the expansion card device, making full use of the computing resources of the entire system.

[0076] Specifically, perform data distribution analysis on the parallel computing results to obtain the optimal allocation plan for the computing result data in the shared storage space of the high-bandwidth memory and the expansion card device. This process first involves collecting and analyzing the characteristics of the computing results. The system will evaluate factors such as the size, access frequency, and data correlation of the result data. For hot data that is frequently accessed, the system tends to allocate it to the high-bandwidth memory to improve the efficiency of subsequent operations. For cold data with a lower access frequency, it may be allocated to the storage space of the expansion card device. The analysis process also considers the data life cycle, and data that needs to be frequently used in the short term is preferentially retained in the high-bandwidth memory. At the same time, the system will evaluate the storage capacity and usage of the current high-bandwidth memory and expansion card device to ensure that there is no overload situation in a certain storage device. In addition, the analysis will also consider the correlation between data and try to allocate related data in the same storage device to reduce cross-device data access. Based on these analysis results, the system will use specific optimization algorithms, such as the greedy algorithm or dynamic programming, to generate an optimal data allocation plan. This plan will specify in detail which device and which location each piece of computing result data should be stored in to achieve the optimal overall performance.

[0077] Specifically, according to the optimal allocation plan, perform selective replication processing on the computing result data to achieve the expansion of the high-bandwidth memory. This step first involves the planning of data movement operations. The system will determine which data needs to be moved from the original location to the new location according to the optimal allocation plan. For data that needs to be moved from the high-bandwidth memory to the expansion card device, the system will use DMA (Direct Memory Access) technology for efficient transmission to minimize CPU intervention. Conversely, for data that needs to be moved from the expansion card device to the high-bandwidth memory, the system will utilize the high-bandwidth characteristics of the PCIe bus to optimize the transmission path. During the data movement process, the system will adopt pipeline technology to perform data transmission and calculation simultaneously to mask the latency of data movement. For data that does not need to be moved, the system will update the memory mapping table to ensure that subsequent accesses can correctly locate the data. In addition, the system will also maintain a global data location table to record the current location of each piece of data for subsequent quick access. Through this selective data replication and redistribution, the system effectively expands the capacity of the high-bandwidth memory, enabling large-scale data sets to be efficiently managed and accessed between the limited high-bandwidth memory and expansion card device, thus achieving the logical expansion of the high-bandwidth memory.

[0078] Further, the task division process for the shared data set on the high-bandwidth memory and the expansion card device to obtain data subsets suitable for parallel computing and corresponding computing instruction sets includes: analyzing the storage characteristics of the high-bandwidth memory and the expansion card device to obtain a storage performance model including access latency and bandwidth parameters; based on the storage performance model, performing data dependence analysis on the shared data set to obtain data blocks that can be executed in parallel and corresponding computing task graphs; based on the data blocks and computing task graphs, performing optimization allocation on the computing tasks to obtain data subsets evenly distributed on the high-bandwidth memory and the expansion card device and corresponding computing instruction sets.

[0079] Specifically, analyze the storage characteristics of the high-bandwidth memory and the expansion card device to obtain a storage performance model including access latency and bandwidth parameters. This process first involves measuring the performance characteristics of the high-bandwidth memory (HBM). The system evaluates the read and write latency, peak bandwidth, and sustained bandwidth of the HBM through a series of microbenchmarks. These tests include sequential read and write, random read and write, and access patterns for different data sizes. At the same time, the system also evaluates the capacity and cache hierarchy of the HBM. For the expansion card device, the system also conducts a series of performance tests, focusing on its data transfer characteristics through the PCIe interface. This includes measuring the effective bandwidth, latency, and transfer efficiency of the PCIe interface at different data sizes. In addition, the system also evaluates the internal storage characteristics of the expansion card device, such as the access speed and capacity of its internal memory. Based on these test results, the system constructs a comprehensive storage performance model that includes not only static latency and bandwidth parameters but also dynamic performance characteristics under different loads. This model will provide a key decision-making basis for subsequent task division and data allocation.

[0080] Specifically, according to the storage performance model, perform data dependence analysis on the shared data set to obtain data blocks that can be executed in parallel and the corresponding computational task graph. This step first involves in-depth analysis of the structure and access pattern of the shared data set. The system constructs a data dependence graph to represent the read and write relationships between data elements. This process uses static code analysis techniques to identify data access patterns and potential parallel opportunities. At the same time, the system considers the latency and bandwidth parameters in the storage performance model to evaluate the performance impact of different data access strategies. Based on this information, the system adopts specific data chunking algorithms, such as block partitioning or loop partitioning, to divide the shared data set into multiple data blocks that can be processed in parallel. The size and boundary of each data block are optimized according to the characteristics of the storage device to maximize data access efficiency. In addition, the system also generates a computational task graph to describe the processing order and dependence relationships between these data blocks. This task graph not only considers data dependence but also computational complexity and storage access patterns to ensure that computational resources can be fully utilized in subsequent parallel execution while minimizing data movement overhead.

[0081] Specifically, based on the data blocks and the computational task graph, perform optimized allocation processing on the computational tasks to obtain data subsets evenly distributed on the high-bandwidth memory and the expansion card device and the corresponding computational instruction sets. This process first involves evaluating the computational requirements and storage characteristics of each data block. The system considers the size, access frequency, computational complexity of the data block, and its dependence relationship with other data blocks. At the same time, the system refers to the previously constructed storage performance model to balance the high-speed access characteristics of the high-bandwidth memory and the large-capacity advantages of the expansion card device. Based on these factors, the system uses heuristic algorithms or dynamic programming methods to optimize task allocation. For data blocks that are computationally intensive and require frequent access, the system tends to allocate them to the high-bandwidth memory; while for I / O-intensive or less frequently accessed data blocks, they may be allocated to the expansion card device. The system also considers load balancing to ensure that both storage devices are fully utilized. At the same time, the system generates the corresponding computational instruction sets, which contain detailed steps on how to process each data subset, including data loading, computational operations, and result storage, etc. The generation of the instruction sets takes into account the location of the data, optimizes the data movement path, and utilizes advanced CUDA features such as shared memory and warp-level operations to improve efficiency.

[0082] In this embodiment, the expansion card device is connected to the graphics processor through the PCIe interface on the graphics processor, so that the storage space of the expansion card device serves as the shared storage space of the high-bandwidth memory (HBM). First, in the kernel function of the graphics processor, the shared storage space is declared using the shared keyword to generate an extended declaration. Then, the HBM memory data is copied to the shared storage space of the expansion card device through the threads in the thread block, and a synchronization function is used to synchronize the threads of the thread block. On this basis, parallel computing processing is performed, and the calculation results are allocated and copied between the HBM and the shared storage space of the expansion card device to achieve storage expansion. The present invention can effectively expand the capacity of the HBM, meet the processing requirements of big data and large models, improve the GPU computing efficiency, and reduce the expansion cost.

[0083] The method for expanding the high-bandwidth memory in the embodiment of the present invention is described above. Next, the high-bandwidth memory expansion device in the embodiment of the present invention will be described. The high-bandwidth memory expansion device is applied to a graphics processor, and the high-bandwidth memory is integrated on the graphics processor. Please refer to Figure 2 , an embodiment of the high-bandwidth memory expansion device in the embodiment of the present invention includes:

[0084] A connection module 201, configured to connect the expansion card device to the graphics processor through the PCIe interface on the graphics processor, so as to use the storage space in the expansion card device as the shared storage space of the high-bandwidth memory;

[0085] A declaration module 202, configured to perform a declaration process on the shared storage space using a preset shared keyword in the kernel function of the graphics processor to obtain an extended declaration of the expansion card device;

[0086] A copy synchronization module 203, configured to copy the memory data of the high-bandwidth memory to the shared storage space of the expansion card device through the threads in a preset thread block according to the extended declaration in the kernel function, and perform thread synchronization on the thread block through a preset synchronization function;

[0087] A calculation module 204, configured to perform parallel computing processing on the shared data sets on the high-bandwidth memory and the expansion card device to obtain calculation result data, and allocate and copy the calculation result data between the high-bandwidth memory and the shared storage space of the expansion card device to achieve expansion of the high-bandwidth memory.

[0088] In an embodiment of the present invention, the high-bandwidth memory expansion device runs the above high-bandwidth memory expansion method. The high-bandwidth memory expansion device connects the expansion card device to the graphics processing unit through the PCIe interface on the graphics processing unit, so that the storage space of the expansion card device serves as the shared storage space of the high-bandwidth memory (HBM). First, in the kernel function of the graphics processing unit, the shared keyword is used to declare the shared storage space to generate an extended declaration. Then, the HBM memory data is copied to the shared storage space of the expansion card device by the threads in the thread block, and the thread synchronization function is used to synchronize the threads in the thread block. On this basis, parallel computing processing is performed, and the calculation results are allocated and copied between the shared storage spaces of the HBM and the expansion card device to achieve storage expansion. The present invention can effectively expand the capacity of the HBM, meet the processing requirements of big data and large models, improve the GPU computing efficiency, and reduce the expansion cost.

[0089] above Figure 2 The medium and high-bandwidth memory expansion device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the high-bandwidth memory expansion device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0090] Figure 3 FIG. 9 is a schematic structural diagram of a high-bandwidth memory expansion device provided by an embodiment of the present invention. The high-bandwidth memory expansion device 300 may vary greatly due to configuration or performance, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the high-bandwidth memory expansion device 300. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the high-bandwidth memory expansion device 300 to implement the steps of the above high-bandwidth memory expansion method.

[0091] The high-bandwidth memory expansion device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand, Figure 3The structure of the high-bandwidth memory expansion device shown does not constitute a limitation on the high-bandwidth memory expansion device provided by the present invention, and may include more or fewer components than shown, or combine certain components, or have a different component arrangement.

[0092] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the high-bandwidth memory expansion method.

[0093] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0094] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0095] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A high bandwidth memory expansion method, characterized in that: Applied to a graphics processor, the graphics processor is integrated with a high-bandwidth memory, and the high-bandwidth memory expansion method includes: The expansion card device is connected to the graphics processor through the PCIe interface on the graphics processor, and the root complex of the graphics processor is initialized to obtain a system topology structure connecting the CPU, the memory subsystem and the PCIe device; according to the system topology structure, the PCIe bus of the PCIe interface is enumerated to obtain a function list containing different bus number and device number combinations; based on the function list, the configuration space of the expansion card device is read to obtain the identity information and hardware characteristics of the expansion card device; according to the identity information and the hardware characteristics, the base address register of the expansion card device is configured to obtain the mapping position of the expansion card device in the memory address space of the graphics processor; based on the mapping position, the memory mapping I / O system of the graphics processor is updated to obtain an extended memory access mechanism containing the storage space of the expansion card device, so that the storage space in the expansion card device is used as the shared storage space of the high bandwidth memory; In the kernel function of the graphics processor, a declaration process is performed on the shared storage space according to a preset shared keyword to obtain an extended declaration of the expansion card device; In the kernel function, a thread block index and a thread index in a preset thread block are determined according to the extended declaration; a global memory access address corresponding to each thread is determined according to the thread block index and the thread index; memory data in a high bandwidth memory is copied to a shared storage space of the expansion card device through threads in the thread block according to the global memory access address, and thread synchronization is performed on the thread block through a preset synchronization function; Parallel computing is performed on the shared data set on the high bandwidth memory and the expansion card device to obtain computing result data, and the computing result data is distributed and copied between the shared storage space of the high bandwidth memory and the expansion card device to achieve capacity expansion of the high bandwidth memory.

2. The high bandwidth memory expansion method according to claim 1, characterized in that: The process of declaring the shared storage space using the preset shared keyword in the kernel function of the graphics processor to obtain the extended declaration of the expansion card device includes: Preprocessing the kernel function code of the graphics processor to obtain a code segment of a shared keyword, and inserting an identifier of the expansion card device into the code segment to obtain a shared memory declaration statement of the extended memory; Performing syntax parsing on the shared memory declaration statement to obtain a parameter list including the storage space size of the expansion card device; According to the parameter list, the firmware version of the expansion card device is detected, and based on the firmware version, the extended shared memory declaration statement is conditionally compiled to obtain an extended declaration of the expansion card device adapted to the firmware version.

3. The high bandwidth memory expansion method according to claim 1, characterized in that: The thread block is synchronized by a preset synchronization function, comprising: The execution status of each thread in the thread block is marked by a preset synchronization function to obtain a status identifier of the thread execution reaching the synchronization point; According to the state identifier, updating the thread counter in the thread block to obtain the number of threads that have reached the synchronization point; The number of threads that have reached the synchronization point is judged to obtain a judgment result of whether all threads in the thread block have reached the synchronization point, and the thread block is synchronized according to the judgment result.

4. The high bandwidth memory expansion method according to claim 1, characterized in that: The step of performing parallel computing on the shared data set on the high bandwidth memory and the expansion card device to obtain computing result data, and distributing and copying the computing result data between the high bandwidth memory and the shared storage space of the expansion card device to achieve the expansion of the high bandwidth memory includes: Performing task division processing on the shared data set on the high bandwidth memory and the expansion card device to obtain a data subset suitable for parallel computing and a corresponding computing instruction set; According to the data subset and the computing instruction set, the computing units of the graphics processor are parallelly scheduled to obtain parallel computing results that are simultaneously executed on the high-bandwidth memory and the expansion card device; Performing data distribution analysis on the parallel computing results to obtain an optimal allocation scheme for the computing result data in the shared storage space of the high-bandwidth memory and the expansion card device; According to the optimal allocation scheme, the calculation result data is selectively copied to achieve the expansion of the high bandwidth memory.

5. The high bandwidth memory expansion method according to claim 4, characterized in that: The task division processing of the shared data set on the high bandwidth memory and the expansion card device to obtain a data subset suitable for parallel computing and a corresponding computing instruction set includes: Analyze and process the storage characteristics of high-bandwidth memory and expansion card devices to obtain a storage performance model that includes access delay and bandwidth parameters; According to the storage performance model, data dependency analysis is performed on the shared data set to obtain data blocks that can be executed in parallel and corresponding computing task graphs; Based on the data blocks and the computing task graph, the computing tasks are optimized and allocated to obtain data subsets and corresponding computing instruction sets that are evenly distributed on the high-bandwidth memory and expansion card device.

6. A high bandwidth memory expansion device, characterized in that: Applied to a graphics processor, the graphics processor is integrated with a high-bandwidth memory, and the high-bandwidth memory expansion device comprises: A connection module is used to connect the expansion card device to the graphics processor through the PCIe interface on the graphics processor, and initialize the root complex of the graphics processor to obtain a system topology structure connecting the CPU, the memory subsystem and the PCIe device; according to the system topology structure, enumerate the PCIe bus of the PCIe interface to obtain a function list containing different bus number and device number combinations; based on the function list, read the configuration space of the expansion card device to obtain the identity information and hardware characteristics of the expansion card device; according to the identity information and hardware characteristics, configure the base address register of the expansion card device to obtain the mapping position of the expansion card device in the memory address space of the graphics processor; based on the mapping position, update the memory mapping I / O system of the graphics processor to obtain an extended memory access mechanism containing the storage space of the expansion card device, so as to use the storage space in the expansion card device as the shared storage space of the high bandwidth memory; A declaration module, used to perform declaration processing on the shared storage space according to a preset shared keyword in a kernel function of the graphics processor to obtain an extended declaration of the expansion card device; A copy synchronization module is used to determine, in the kernel function, a thread block index and a thread index in a preset thread block according to the extended declaration; determine a global memory access address corresponding to each thread according to the thread block index and the thread index; copy memory data in a high bandwidth memory to a shared storage space of the expansion card device through threads in the thread block according to the global memory access address, and perform thread synchronization on the thread block through a preset synchronization function; The computing module is used to perform parallel computing processing on the shared data set on the high-bandwidth memory and the expansion card device to obtain computing result data, and distribute and copy the computing result data between the shared storage space of the high-bandwidth memory and the expansion card device to achieve the expansion of the high-bandwidth memory.

7. A high bandwidth memory expansion device, characterized in that: The high bandwidth memory expansion device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory to enable the high bandwidth memory expansion device to perform the steps of the high bandwidth memory expansion method according to any one of claims 1 to 5.

8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the high bandwidth memory expansion method as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Online memory expansion method and device, equipment and storage medium

    CN117806570A