Data storage location dynamic optimization method, device, equipment and storage medium
By connecting the expansion card device on the GPU, the storage space configuration information acquisition, kernel function code preprocessing, data block division and dynamic migration are solved, and the problem of insufficient optimization of data storage distribution in the existing technology is achieved, and efficient data storage and computing performance improvement is achieved.
Patent Information
- Application Number
- CN202510294712.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing technology lacks an effective mechanism to dynamically optimize the distribution of data between internal storage and external storage of GPU, resulting in inefficient data access when processing large-scale data, inability to make full use of external storage resources, affecting overall computing performance.
Connect the expansion card device through the PCIe interface on the GPU, perform BDF structure scanning and space mapping, and obtain storage space configuration information. Then, the kernel function code of the GPU is preprocessed and the capacity expansion card identifier is inserted, the data transmission parameters are configured, and the CUDA shared storage mechanism is initialized. Next, the storage space of the expansion card device is partitioned by the operation capacity area and the cache area, and the data to be calculated by the GPU is thread synchronization and parallel computing through the SLC data block division method to obtain the operation status information of the storage space. Finally, based on the operating status information, the distribution of the to-be-calculated data between the operating capacity area and the cache area is dynamically migrated to achieve dynamic optimization of the data storage location.
By dynamically optimizing data storage locations, the utilization of storage resources is improved and data access delay is reduced, thereby accelerating the GPU computing process and improving overall system performance.
Smart Images

Figure CN119806434B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graphics processors, and in particular to a method, device, equipment and storage medium for dynamically optimizing a data storage location. Background Art
[0002] With the rapid development of graphics processing unit (GPU) technology, its application in the field of high-performance computing is becoming more and more extensive. At present, GPU usually implements parallel computing through CUDA (Compute Unified Device Architecture) architecture. In the traditional GPU computing model, data storage mainly relies on the shared memory and global memory inside the GPU. In order to improve data processing capabilities and storage capacity, some systems use the method of connecting external storage devices through the PCIe interface to expand the storage space of the GPU.
[0003] However, the existing technology has a major flaw: when the external storage device is connected to the GPU through the PCIe interface, the system lacks an effective mechanism to dynamically optimize the distribution of data between the GPU internal storage and external storage. This leads to low data access efficiency when processing large-scale data, and the external storage resources cannot be fully utilized, thus affecting the overall computing performance. Summary of the invention
[0004] The main purpose of the present invention is to solve the technical problem that the prior art lacks an effective mechanism to dynamically optimize the distribution of data between GPU internal storage and external storage;
[0005] A first aspect of the present invention provides a method for dynamically optimizing a data storage location, the method comprising:
[0006] Connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information;
[0007] Preprocessing the shared keyword code segment of the GPU kernel function code and inserting the expansion card identifier according to the storage space configuration information, and configuring the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism;
[0008] According to the initialization configuration result, the storage space of the expansion card device is partitioned into an operation capacity area and a cache area, and thread synchronization and parallel computing are performed on the to-be-calculated data of the GPU by SLC data block partitioning to obtain operation status information of the storage space;
[0009] According to the operation status information, the data access frequency in the GPU operation process is detected and analyzed, and according to the detection result, the distribution of the data to be operated between the operation capacity area and the cache area is dynamically migrated to achieve dynamic optimization of the data storage location.
[0010] Optionally, in a first implementation of the first aspect of the present invention, the expansion card device is connected to the graphics processor through a PCIe interface on the GPU, and a BDF structure scan and space mapping process is performed on the GPU and the expansion card device to obtain the storage space configuration information, including:
[0011] Connect the expansion card device to the GPU through the PCIe interface on the GPU, and read the VID / DID and scan the BDF structure of the GPU device and the expansion card device to obtain the device identification code;
[0012] Perform an HBM firmware version check on the expansion card device according to the device identification code to obtain firmware information, and perform a memory space size check on the expansion card device according to the firmware information through a BAR register in a configuration space in a PCIe interface to obtain available storage space information;
[0013] The CPU allocates an address range to the expansion card device according to the available storage space information to obtain storage space configuration information required for the GPU operation.
[0014] Optionally, in a second implementation of the first aspect of the present invention, preprocessing the shared keyword code segment of the GPU kernel function code and inserting the expansion card identifier according to the storage space configuration information, and configuring the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism includes:
[0015] Scanning the kernel function code of the GPU according to the storage space configuration information, and identifying and extracting the shared keywords in the kernel function code to obtain the shared keyword code segment;
[0016] Performing syntax parsing and parameter parsing processing on the shared keyword code segment, and inserting the identifier information of the expansion card device into the parsed code segment to obtain a shared memory declaration statement;
[0017] The BME bit and the MSE bit of the command register are configured according to the shared memory declaration statement, and the data transmission channel and priority of the expansion card device are set to obtain the initialization configuration result of the CUDA shared memory mechanism.
[0018] Optionally, in a third implementation of the first aspect of the present invention, the partitioning of the storage space of the expansion card device into a running capacity area and a cache area according to the initialization configuration result, and thread synchronization and parallel computing processing of the to-be-calculated data of the GPU by SLC data block partitioning, to obtain the running status information of the storage space includes:
[0019] According to the initialization configuration result, the storage space of the expansion card device is partitioned into a running capacity area and a cache area to obtain a storage area partition result;
[0020] According to the storage area division result, the GPU data to be calculated is grouped into data blocks, and the data blocks are allocated for storage by SLC data block calculation to obtain a data storage state;
[0021] The SLC data blocks are processed in parallel according to the data storage status, and synchronization status marking processing is performed on the threads in the thread block through a synchronization function to obtain the operation status information of the storage space.
[0022] Optionally, in a fourth implementation of the first aspect of the present invention, performing data block grouping processing on the GPU's data to be calculated according to the storage area division result, and performing storage allocation processing on the data blocks by SLC data block calculation, to obtain the data storage state includes:
[0023] Calculate the available space size of the running capacity area and the cache area according to the storage area division result to obtain the capacity parameters of each storage area;
[0024] According to the capacity parameter, the GPU data to be calculated is divided into SLC data blocks and grouped and numbered to obtain a data block organization scheme;
[0025] According to the data block organization scheme, storage location allocation processing is performed on each SLC data block between the running capacity area and the cache area to obtain the data storage state.
[0026] Optionally, in a fifth implementation of the first aspect of the present invention, the detecting, analyzing and processing the data access frequency in the GPU computing process according to the running status information, and dynamically migrating the distribution of the data to be computed between the running capacity area and the cache area according to the detection result, to achieve dynamic optimization of the data storage location includes:
[0027] Performing access counting processing on the data to be calculated in the GPU calculation according to the operation status information, and performing statistical analysis processing on the access frequency and time interval of the data to obtain data access characteristics;
[0028] According to the data access characteristics, data distribution evaluation processing is performed on the GPU memory space and the storage space of the expansion card device, and the distribution of the operation data between the running capacity area and the cache area is calculated to obtain a data migration plan;
[0029] According to the data migration scheme, the position of the data to be operated is adjusted between the running capacity area and the cache area, and the data access efficiency of the GPU operation process is statistically analyzed to obtain the optimal configuration result of the data storage position.
[0030] Optionally, in a sixth implementation of the first aspect of the present invention, performing data distribution evaluation processing on the video memory space of the GPU and the storage space of the expansion card device according to the data access characteristics, and calculating and processing the distribution of the data to be operated between the running capacity area and the cache area, to obtain the data migration plan includes:
[0031] Monitor and collect the working status of the GPU's video memory space and the storage space of the expansion card device to obtain load status information including space usage and bandwidth occupancy;
[0032] Performing correlation analysis on the load status information and the data access characteristics to obtain a data migration threshold parameter for each storage area;
[0033] The storage location of the data to be operated is evaluated according to the data migration threshold parameter to obtain a location allocation result of the data to be operated;
[0034] The execution order and resource allocation of data migration are planned and processed according to the position allocation result to obtain the data migration plan.
[0035] A second aspect of the present invention provides a data storage location dynamic optimization device, the data storage location dynamic optimization device comprising:
[0036] The device initialization module is used to connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information;
[0037] A storage configuration module is used to pre-process the shared keyword code segment of the GPU kernel function code and insert the expansion card identifier according to the storage space configuration information, and configure the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism;
[0038] A space partitioning module is used to partition the storage space of the expansion card device into an operating capacity area and a cache area according to the initialization configuration result, and perform thread synchronization and parallel computing processing on the to-be-calculated data of the GPU by SLC data block partitioning to obtain operating status information of the storage space;
[0039] The dynamic optimization module is used to detect and analyze the data access frequency during the GPU operation process according to the operation status information, and dynamically migrate the distribution of the data to be operated between the operation capacity area and the cache area according to the detection results to achieve dynamic optimization of the data storage location.
[0040] The third aspect of the present invention provides a data storage location dynamic optimization device, comprising: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via lines; the at least one processor calls the instructions in the memory so that the data storage location dynamic optimization device executes the steps of the above-mentioned data storage location dynamic optimization method.
[0041] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the above-mentioned method for dynamically optimizing data storage locations.
[0042] The above-mentioned data storage location dynamic optimization method, device, equipment and storage medium connect the expansion card device to the GPU through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping to obtain storage space configuration information; preprocess the shared keyword code segment of the GPU kernel function code and insert the expansion card identifier, and configure the data transmission parameters to obtain the initialization configuration of the CUDA shared storage mechanism; partition the expansion card device storage space into the running capacity area and the cache area, and perform thread synchronization and parallel calculation on the GPU to be calculated data through SLC data block division to obtain storage space operation status information; according to the operation status information, the distribution of the to-be-calculated data between the running capacity area and the cache area is dynamically migrated. The present invention improves the storage resource utilization rate and reduces the data access delay by dynamically optimizing the data storage location, thereby accelerating the GPU calculation process and improving the overall system performance.
[0043] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0044] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of a first embodiment of a method for dynamically optimizing data storage locations according to an embodiment of the present invention;
[0046] Figure 2 A schematic diagram of an embodiment of a device for dynamically optimizing data storage locations in an embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of an embodiment of a device for dynamically optimizing data storage locations in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0050] To facilitate understanding of this embodiment, a method for dynamically optimizing data storage locations disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps:
[0051] 101. Connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information;
[0052] In one embodiment of the present invention, the expansion card device is connected to the graphics processor through the PCIe interface on the GPU, and the GPU and the expansion card device are subjected to BDF structure scanning and space mapping processing to obtain storage space configuration information, including: connecting the expansion card device to the GPU through the PCIe interface on the GPU, and performing VID / DID reading and BDF structure scanning on the GPU device and the expansion card device to obtain a device identification code; performing an HBM firmware version check on the expansion card device according to the device identification code to obtain firmware information, and performing a memory space size detection on the expansion card device according to the firmware information through the BAR register of the configuration space in the PCIe interface to obtain available storage space information; and performing an address range allocation on the expansion card device according to the available storage space information by the CPU to obtain the storage space configuration information required for the GPU operation.
[0053] Specifically, first, the expansion card device is connected to the GPU through the PCIe interface on the GPU, and the VID / DID of the GPU device and the expansion card device are read and the BDF structure is scanned to obtain the device identification code. This process begins with the physical insertion of the expansion card device into the PCIe slot of the GPU. When the system starts, the BIOS or UEFI firmware will identify the newly inserted PCIe device and allocate basic resources to it. Subsequently, the GPU driver is activated and begins to initialize the root complex. During the initialization process, the driver scans all connected PCIe devices, including the newly inserted expansion card device. The system reads the VID (vendor ID) and DID (device ID) of the GPU device and the expansion card device through the PCIe bus. This reading process is implemented by accessing the configuration space of the device. Specifically, the Configuration Address Register is set at address 0xCF8 of the I / O space to determine the BDF combination to be accessed, and then the actual reading operation is performed at address 0xCFC. At the same time, the system will also scan the BDF (Bus, Device, Function) structure of these devices, that is, determine the unique identifier of each PCIe device node through the combination of bus, device and function. As the root node of the entire I / O architecture, the root complex dominates the device identification process. It completes the identification and management of PCIe devices under the control of the CPU. In this way, the system can accurately identify each device connected to the PCIe bus and ensure the correct matching of device types through the combination of VID / DID. These device identification information and BDF structure information together constitute the device identification code, which provides the basis for subsequent device configuration and resource allocation.
[0054] Specifically, the expansion card device is then checked for the HBM firmware version according to the device identification code to obtain the firmware information, and the memory space size of the expansion card device is detected according to the firmware information through the BAR register in the configuration space of the PCIe interface to obtain the available storage space information. In this step, the system first checks the version of the HBM (high bandwidth memory) firmware of the expansion card device according to the previously obtained device identification code. This check process is completed by reading specific registers in the configuration space of the expansion card device. The check result contains the specific version number and functional characteristics of the HBM firmware, which is recorded as firmware information and provides an important reference for subsequent memory space configuration. Next, the system operates through the BAR (Base Address Register) register in the PCIe configuration space. This process is achieved by writing all 1s to the BAR register and then reading the return value. Since some bits of the BAR register are read-only, by observing which bits remain unchanged, the size of the storage space required by the device can be known. For example, if the lower 12 bits do not change, it means that 4KB of space is required. This method of determining the size of the storage space through read and write operations is a standard mechanism specified by the PCIe protocol. In addition, the lower bits of the BAR register also contain some attribute information, such as whether the storage space is memory space or I / O space, whether it supports prefetching and other features. The system will decide how to access the device's storage space based on these features. In this way, the system obtains available storage space information including space size and access characteristics, laying the foundation for subsequent memory management and data transmission.
[0055] Specifically, finally, the CPU allocates the address range to the expansion card device according to the available storage space information, and obtains the storage space configuration information required for GPU computing. At this stage, the CPU will allocate a detailed address range to the expansion card device according to the previously obtained available storage space information. This allocation process is achieved by setting the BAR register in the PCIe configuration space. The CPU will allocate a continuous address range in the system address space to the expansion card device according to the storage space required by the expansion card device. This address range will be written into the BAR register, thereby establishing a mapping relationship between the physical address and the device storage space. In the PCIe system, due to the existence of the root complex, when the CPU needs to read the data of the PCIe device, it will first let the root complex read the data from the PCIe device to the CPU memory, and then the CPU will read the data from the memory. Conversely, when the CPU needs to write data to the device, it will also first write the data to the memory, and then write it to the PCIe device through the root complex. This memory-mapped I / O method enables the CPU to access the device's storage space like accessing memory. In this way, the storage space configuration information required for GPU computing is fully established. These configuration information include important parameters such as device identification information, storage capacity information, address mapping relationship and access characteristics.
[0056] 102. Preprocess the shared keyword code segment of the GPU kernel function code and insert the expansion card identifier according to the storage space configuration information, and configure the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism;
[0057] In one embodiment of the present invention, the shared keyword code segment of the GPU kernel code is preprocessed and the expansion card identifier is inserted according to the storage space configuration information, and the data transmission parameters are configured through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism, including: scanning the GPU kernel code according to the storage space configuration information, and identifying and extracting the shared keywords in the kernel code to obtain the shared keyword code segment; parsing the shared keyword code segment and inserting the identifier information of the expansion card device into the parsed code segment to obtain a shared memory declaration statement; configuring the BME bit and the MSE bit of the command register according to the shared memory declaration statement, and setting the data transmission channel and priority of the expansion card device to obtain the initialization configuration result of the CUDA shared storage mechanism.
[0058] Specifically, the kernel function code of the GPU is first scanned and processed according to the storage space configuration information, and the shared keywords in the kernel function code are identified and extracted to obtain the shared keyword code segment. This step uses lexical analysis and syntax analysis technology to comprehensively scan the GPU kernel function code. During the scanning process, the system will identify all code segments using the __shared__ keyword, which usually represent statements declaring variables in shared memory. The identification process includes not only the direct use of the __shared__ keyword, but also the indirect use through macro definitions or templates. The extraction process retains the context information of the shared keyword code segment, including variable type, variable name, array size (if applicable), etc. The purpose of this is to prepare for the subsequent expansion card device adaptation and ensure that the semantics and functions of the original code are not destroyed when the shared memory is expanded. The obtained shared keyword code segment is a collection of all shared memory usages, which provides a basis for subsequent processing.
[0059] Next, the shared keyword code segment is parsed and parameter parsed, and the identifier information of the expansion card device is inserted into the parsed code segment to obtain the shared memory declaration statement. The syntax parsing process uses the abstract syntax tree (AST) technology to convert the shared keyword code segment into a structured syntax representation. This process analyzes the syntax structure of each shared memory declaration in detail, including modifiers, type specifiers, variable names, array dimensions, etc. Parameter parsing focuses on the specific parameters in the shared memory declaration, such as memory size, alignment requirements, etc. The parsing process also processes CUDA-specific syntax elements such as template parameters and const qualifiers that may exist. After the parsing is completed, the system inserts the identifier information of the expansion card device into the parsed code segment. This identifier may be a predefined macro or a special comment to mark that this part of the shared memory declaration is for the expansion card device. The insertion process needs to consider syntax correctness to ensure that the inserted code still conforms to the CUDA syntax specification. The final generated shared memory declaration statement contains not only the original shared memory declaration information, but also the specific properties of the expansion card device, providing necessary information for subsequent memory allocation and management.
[0060] Finally, the BME bit and MSE bit of the command register are configured according to the shared memory declaration statement, and the data transmission channel and priority of the expansion card device are set to obtain the initialization configuration result of the CUDA shared memory mechanism. The process of configuring the command register first involves setting the BME (Bus Master Enable) bit, which allows the expansion card device to initiate DMA transmission as a bus master. At the same time, the MSE (Memory Space Enable) bit is set to allow access to the device's memory space. The setting of these two bits provides the basic conditions for the establishment of the data transmission channel. Then, the system creates independent data transmission channels for the expansion card device based on the information in the shared memory declaration statement. These channels use PCIe differential signal pairs for data transmission, and each channel has its own independent transmission instruction management mechanism. The system also sets the priority of data transmission to ensure that key data can be transmitted in time. This process takes into account the hardware characteristics of the expansion card device and the requirements of the CUDA shared memory mechanism, and optimizes the data transmission efficiency. The final CUDA shared memory mechanism initialization configuration result contains key parameters such as memory mapping information, data transmission channel configuration, and priority setting of the expansion card device, providing an efficient memory access and data transmission mechanism for subsequent GPU computing processes.
[0061] 103. Perform partitioning of the storage space of the expansion card device into an operating capacity area and a cache area according to the initialization configuration result, and perform thread synchronization and parallel computing processing on the GPU's to-be-calculated data by SLC data block partitioning to obtain operating status information of the storage space;
[0062] In one embodiment of the present invention, the storage space of the expansion card device is partitioned into an operating capacity area and a cache area according to the initialization configuration result, and thread synchronization and parallel calculation processing are performed on the data to be calculated of the GPU through the SLC data block partitioning method to obtain the operating status information of the storage space, including: the storage space of the expansion card device is partitioned into an operating capacity area and a cache area according to the initialization configuration result to obtain a storage area partition result; the data to be calculated of the GPU is grouped into data blocks according to the storage area partition result, and storage allocation processing is performed on the data blocks through the SLC data block calculation method to obtain a data storage status; the SLC data blocks are parallelly calculated according to the data storage status, and synchronization status marking processing is performed on the threads in the thread block through a synchronization function to obtain the operating status information of the storage space.
[0063] Specifically, first, the storage space of the expansion card device is partitioned into the running capacity area and the cache area according to the initialization configuration result to obtain the storage area partition result. This process starts with analyzing the initialization configuration result, which contains key information such as the total capacity size, storage characteristics and access mode of the expansion card device. Based on this information, the system uses a specific partitioning algorithm to partition the storage space of the expansion card device. Specifically, the system will divide a part of the total capacity into the running capacity area for storing actual computing data; the remaining part will be divided into the cache area for improving data access speed. The determination of the partition ratio takes into account the computing characteristics and data access mode of the GPU. Usually, the running capacity area will occupy a larger proportion to ensure sufficient space to store large-scale data sets. This partitioning is achieved by setting the corresponding address range register in the root complex. The root complex will write these partitioned address ranges into the PCIe configuration space, thereby completing the space partitioning at the hardware level. The partitioning process also includes the setting of access rights and attributes for each area, such as read and write permissions, cache strategy, etc. These settings are recorded in the storage area division result, which contains detailed information such as the starting address, size, access attributes, etc. of each area, providing a basis for subsequent data allocation and access.
[0064] Next, the data blocks to be calculated by the GPU are grouped according to the storage area division results, and the data blocks are allocated and processed by the SLC data block calculation method to obtain the data storage status. At this stage, the system first determines the specific boundaries and sizes of the running capacity area and the cache area according to the storage area division results. Then, the system begins to analyze and process the data to be calculated by the GPU. This process involves in-depth analysis of the structure, size and access pattern of the data. Based on these analysis results, the system uses the SLC (Single-Level Cell) data block division method to group the data to be calculated. The SLC method is selected because it can provide faster read and write speeds and lower latency, which is suitable for the needs of GPU high-performance computing. The size of the data block is determined by considering the characteristics of the SLC storage unit and the processing power of the GPU. Usually, a size that can balance storage efficiency and access speed is selected. During the grouping process, the system will number and classify the data blocks according to the data access mode and dependency, and organize the data blocks with similar access characteristics or computational dependencies together. At the same time, the system will establish an association mapping table between data blocks to record the dependency and access order between data blocks. This information has important guiding significance for subsequent parallel processing. After completing the data block grouping, the system allocates data storage based on the characteristics of each data block group, combined with the storage characteristics of the operating capacity area and the cache area. This allocation process takes into account the access frequency of data blocks, the correlation between data blocks, and the bandwidth characteristics of the storage area, and reasonably allocates data blocks to different storage areas. Ultimately, the system generates a complete data storage state, including detailed information such as the physical location, logical location, access attributes, and topological relationships between data blocks for each data block.
[0065] Finally, the SLC data blocks are processed in parallel according to the data storage status, and the threads in the thread block are marked with synchronization status through the synchronization function to obtain the operation status information of the storage space. In this stage, the system starts to process the SLC data blocks in parallel based on the data storage status. During the processing, the system finds the corresponding data blocks according to the location information recorded in the data storage status, and then starts the parallel processing unit of the GPU to operate on these data blocks. During the operation, the system tracks the processing progress of each data block and records the intermediate results and data dependencies of the operation. The operation of each data block will go through multiple processing stages, including data loading, calculation execution, result storage, etc. The system will record detailed execution information at each stage, including calculation time, resource occupation, data access mode, etc. When the operation of a group of data blocks is completed, the system will store the operation results to the specified storage location and update the corresponding status mark. These operation results will be organized into a structured data set, which contains key data such as the value, precision information, and timestamp of the calculation results. Next, the system uses the synchronization function to mark the threads in the thread block with synchronization status. This process uses the __syncthreads() function or similar mechanisms in CUDA to ensure that all threads in the thread block have completed the current stage of operation. When performing synchronization operations, the system will record the specific situation of each synchronization point, including the number of threads involved in synchronization, the execution status of each thread, the synchronization waiting time, etc. These synchronization operations not only ensure data consistency, but also provide important performance monitoring data. At the same time, the system will also record the data exchange within the thread block, including information such as the access mode of shared memory and the amount of data moved. All these execution records are integrated to form complete storage space operation status information, which not only reflects the current execution status, but also contains detailed performance indicators and resource utilization data, providing an important basis for subsequent performance optimization and resource scheduling.
[0066] Furthermore, the data to be calculated of the GPU is grouped into data blocks according to the storage area division result, and the data blocks are stored and allocated through SLC data block calculation to obtain the data storage status, including: calculating the available space size of the running capacity area and the cache area according to the storage area division result to obtain the capacity parameters of each storage area; dividing the data to be calculated by the GPU into SLC data blocks and grouping them according to the capacity parameters to obtain a data block organization scheme; and allocating the storage position of each SLC data block between the running capacity area and the cache area according to the data block organization scheme to obtain the data storage status.
[0067] Specifically, first, the available space size of the running capacity area and the cache area is calculated and processed according to the storage area division results to obtain the capacity parameters of each storage area. This process starts with analyzing the storage area division results, which contain detailed partition information of the storage space of the expansion card device. The system first extracts the starting and ending addresses of the running capacity area and the cache area, and then calculates the total capacity of each area. Next, the system considers the reserved space of each area, such as space for metadata storage or system management, and subtracts these reserved spaces from the total capacity to obtain the actual available storage space size. For the running capacity area, the system also considers the data alignment requirements to ensure that the allocated space can meet the alignment requirements of GPU access, which is usually in units of a specific number of bytes (such as 256 bytes). The calculation of the cache area needs to consider the cache line size and cache organization to ensure that the cache can operate efficiently. The system also calculates the bandwidth parameters of each area, which involves the transmission rate of the PCIe interface and the internal memory characteristics of the expansion card device. Finally, the system generates a detailed capacity parameter table, including key parameters such as the available space size, bandwidth, and access latency of the running capacity area and the cache area. These capacity parameters provide important basic information for subsequent data block division and storage allocation, ensuring that data allocation can fully utilize the storage resources of the expansion card device and achieve optimal performance.
[0068] Next, the GPU's data to be calculated is divided into SLC data blocks and grouped according to the capacity parameters to obtain a data block organization scheme. At this stage, the system first analyzes the characteristics of the GPU's data to be calculated, including the total amount of data, data structure, access mode, etc. Then, based on the previously obtained capacity parameters, especially the available space size of the running capacity area and the cache area, the system begins to determine the size of the SLC data block. The selection of the SLC (Single-Level Cell) data block size requires balancing multiple factors: on the one hand, a larger block size can reduce management overhead and improve storage efficiency; on the other hand, a smaller block size is conducive to finer-grained data management and more flexible storage allocation. The system usually chooses a size that can strike a balance between the two, such as 4KB or 8KB. After determining the block size, the system begins to divide the data to be calculated. The division process takes into account the natural boundaries of the data to avoid dividing related data into different blocks. At the same time, the system assigns a unique number to each data block, which is not only used to identify the data block, but also contains the attribute information of the data block, such as data type, priority, etc. During the grouping process, the system analyzes the correlation between data blocks and organizes data blocks with similar access patterns or computational dependencies to form data block groups. This organization helps improve cache hit rates and reduce data movement. In addition, the system also builds a dependency graph between data blocks, which is very important for subsequent parallel computing and data prefetching. Ultimately, the system generates a complete data block organization scheme that includes information such as the size, number, attributes, and organizational relationships of each data block. This organization scheme provides detailed guidance for subsequent storage location allocation, ensuring that the organization of data on the expansion card device can maximize support for efficient GPU computing.
[0069] Finally, according to the data block organization scheme, the storage location of each SLC data block is allocated between the running capacity area and the cache area to obtain the data storage status. This process first analyzes the characteristics of each data block in the data block organization scheme, including the size, access frequency, priority, etc. of the data block. Then, the system starts to allocate storage locations. For the running capacity area, the system will give priority to those data blocks that need long-term storage and are frequently accessed. When allocating, the size of the data block and the available space in the running capacity area will be considered, and the specific storage location will be determined by algorithms such as the best match or first match. At the same time, the system will try to keep the physical proximity of the associated data blocks to reduce the latency of data access. For the cache area, the system will give priority to those data blocks that are frequently accessed but do not need long-term storage. The allocation of the cache area also needs to consider the cache replacement strategy, such as LRU (least recently used) or FIFO (first in first out), to ensure that the cache can operate efficiently. During the allocation process, the system will dynamically track the space usage of each area to ensure that one area is not overused while another area is idle. For some special data blocks, such as data used across multiple computing stages, the system may adopt a dynamic migration strategy to dynamically adjust between the running capacity area and the cache area. After the allocation is completed, the system will generate a detailed data storage status table to record the specific storage location, access rights, cache strategy and other information of each SLC data block. This data storage status not only contains static storage allocation information, but also dynamic data access control information, providing complete guidance for the GPU to efficiently access and manage this data during the computing process.
[0070] 104. According to the operation status information, the data access frequency in the GPU operation process is detected and analyzed, and according to the detection result, the distribution of the operation data between the operation capacity area and the cache area is dynamically migrated to achieve dynamic optimization of the data storage location.
[0071] In one embodiment of the present invention, the data access frequency in the GPU operation process is detected and analyzed based on the operation status information, and the distribution of the data to be operated between the operation capacity area and the cache area is dynamically migrated based on the detection result to achieve dynamic optimization of the data storage location, including: performing access counting processing on the data to be operated in the GPU operation according to the operation status information, and performing statistical analysis processing on the access frequency and time interval of the data to obtain data access characteristics; performing data distribution evaluation processing on the GPU's video memory space and the storage space of the expansion card device according to the data access characteristics, and calculating the distribution of the data to be operated between the operation capacity area and the cache area to obtain a data migration plan; performing position adjustment processing on the data to be operated between the operation capacity area and the cache area according to the data migration plan, and performing statistical analysis processing on the data access efficiency of the GPU operation process to obtain the optimized configuration result of the data storage location.
[0072] Specifically, first, the access count processing is performed on the data to be calculated in the GPU operation according to the operation status information, and the access frequency and time interval of the data are statistically analyzed to obtain the data access characteristics. This process starts with analyzing the operation status information, which contains detailed access records of each data block during the GPU operation process. The system first sets an access counter for each data block, and the corresponding counter increases whenever the data block is accessed. This counting process is completed by the hardware counter implemented in the GPU memory controller to ensure high efficiency and not affect the normal operation of the GPU. At the same time, the system also records the timestamp of each access for subsequent time interval analysis. After collecting enough access data, the system starts statistical analysis. First, the access frequency of each data block is calculated, that is, the number of accesses per unit time. Then the time distribution of the access is analyzed, including the average access interval, the periodicity of the access, etc. The system also identifies access patterns, such as sequential access, random access, or specific access patterns. In addition, the system analyzes the access correlation between data blocks and identifies groups of data blocks that are often accessed together. These analysis results are integrated into a complete data access feature description, including the access frequency, access mode, time distribution characteristics of each data block, and the access correlation with other data blocks. These characteristics provide key decision-making basis for subsequent data distribution optimization, enabling the system to adjust the storage location of data according to actual access conditions.
[0073] Next, the data distribution of the GPU memory space and the expansion card device storage space is evaluated based on the data access characteristics, and the distribution of the data to be operated between the running capacity area and the cache area is calculated to obtain the data migration plan. At this stage, the system first comprehensively evaluates the current usage of the GPU memory and expansion card device storage space. This includes analyzing the space occupancy, data block distribution, and access bandwidth utilization of each storage area. The physical characteristics of different storage areas, such as access latency, bandwidth limitations, and concurrency capabilities, are considered during the evaluation. Then, the system combines this information with the previously obtained data access characteristics and begins to formulate a data migration plan. For data blocks with high access frequency, the system tends to migrate them to storage areas with faster access speeds, such as the GPU memory or the cache area of the expansion card device. For data blocks with low access frequency, they may be migrated to the running capacity area of the expansion card device. When formulating the migration plan, the system also needs to consider the access correlation between data blocks, and try to put data blocks that are frequently accessed together in the same storage area to reduce data movement overhead. At the same time, the system will evaluate the cost of each data migration, including the migration time and the impact on overall performance, to ensure that the performance benefits of the migration are greater than the migration costs. In addition, the system will predict future access patterns and reserve appropriate storage space for data that may be frequently accessed in the future. Finally, the system generates a detailed data migration plan that includes information such as the target storage location of each data block, migration priority, and expected performance improvement. This plan not only guides subsequent data migration operations, but also provides the system with a policy framework for dynamically optimizing data storage locations.
[0074] Finally, the data to be processed is adjusted between the running capacity area and the cache area according to the data migration plan, and the data access efficiency of the GPU computing process is statistically analyzed to obtain the optimized configuration result of the data storage location. In this stage, the system first starts to perform data migration operations according to the migration plan. The migration process is completed through the DMA (direct memory access) controller to minimize the impact on GPU computing. The system schedules the migration task according to the migration priority to ensure that the most important data blocks can be migrated first. During the migration process, the system monitors the execution of the migration operation in real time, including the migration speed, resource usage, and the impact on normal computing. If it is found that some migration operations may cause performance degradation, the system will dynamically adjust the migration plan. After the migration is completed, the system will update the data location mapping table to ensure that the GPU can correctly access the new storage location. Next, the system starts to perform statistical analysis on the optimized data access efficiency. This includes measuring multiple performance indicators such as data access latency, bandwidth utilization, and cache hit rate. The system will compare these indicators with the data before optimization to evaluate the effect of the optimization. At the same time, the system will also analyze whether the optimized data distribution meets expectations and whether there are new performance bottlenecks. Based on these analysis results, the system will generate a detailed optimization configuration result report, including the final storage location of each data block, access performance improvement, overall system performance improvement, etc. This report not only reflects the effect of this optimization, but also provides important reference information for the next round of optimization. Through this continuous monitoring, analysis and optimization cycle, the system can continuously adjust the data storage location to adapt to the computing needs of different stages, thereby achieving dynamic optimization of the data storage location and maximizing the computing performance of the GPU.
[0075] Furthermore, the data distribution evaluation process is performed on the video memory space of the GPU and the storage space of the expansion card device according to the data access characteristics, and the distribution of the data to be calculated between the running capacity area and the cache area is calculated to obtain a data migration plan, including: monitoring and collecting the working status of the video memory space of the GPU and the storage space of the expansion card device to obtain load status information including space utilization rate and bandwidth occupancy rate; performing correlation analysis on the load status information and the data access characteristics to obtain data migration threshold parameters of each storage area; evaluating the storage location of the data to be calculated according to the data migration threshold parameters to obtain a location allocation result of the data to be calculated; planning the execution order and resource allocation of data migration according to the location allocation result to obtain the data migration plan.
[0076] Specifically, first, the working status of the GPU's video memory space and the storage space of the expansion card device is monitored, collected and processed to obtain load status information including space usage and bandwidth occupancy. This process involves real-time monitoring of the GPU and expansion card devices. The system obtains the video memory usage through the API interface provided by the GPU driver, including the used video memory size, free video memory size, and the degree of video memory fragmentation. At the same time, the system also monitors the GPU's computing unit utilization, which reflects the current GPU computing load. For the expansion card device, the system reads the device's status register through the PCIe interface to obtain the usage of the storage space. This includes information such as the space occupancy of the running capacity area and the cache area, the frequency of read and write operations, etc. The monitoring of bandwidth occupancy involves the statistics of the data transmission volume of the PCIe bus. The system sets a counter in the PCIe controller to record the data transmission volume per unit time, and compares it with the theoretical bandwidth of the PCIe interface to obtain the bandwidth occupancy. These monitoring data are collected at a certain time interval (usually milliseconds) to capture the dynamic changes in the system load. The collected raw data undergoes preliminary processing, such as removing outliers and smoothing, to ultimately form a comprehensive load status information including space usage and bandwidth occupancy. This load status information provides an important real-time basis for subsequent data migration decisions, ensuring that the system can make optimal data distribution adjustments based on current hardware resource usage.
[0077] Next, the load status information and data access characteristics are analyzed for correlation to obtain the data migration threshold parameters for each storage area. At this stage, the system first analyzes the load status information for correlation with the previously obtained data access characteristics. This analysis process uses a variety of statistical and machine learning techniques. The system calculates the correlation coefficients between various load status indicators (such as space utilization, bandwidth occupancy) and data access characteristics (such as access frequency, access mode) to identify which factors have the greatest impact on data access efficiency. At the same time, the system also establishes a multivariate regression model between these factors to predict data access performance under different load conditions. Based on these analysis results, the system begins to determine the threshold parameters for data migration for each storage area (GPU memory, operating capacity area of expansion card devices, and cache area). These threshold parameters include the upper limit of space utilization, the upper limit of bandwidth occupancy, and the lower limit of data access frequency. For example, when the utilization rate of GPU memory exceeds a certain threshold (such as 90%), the system will consider migrating some infrequently used data to the expansion card device. Similarly, when the access frequency of certain data on the expansion card device exceeds the set threshold, the system will consider migrating it to the GPU memory. These thresholds are not fixed, but are dynamically adjusted based on the current load status and historical data. The system uses machine learning algorithms, such as reinforcement learning, to continuously optimize these thresholds to adapt to different computing tasks and data access patterns. Ultimately, the system obtains a set of dynamic data migration threshold parameters that can guide the system to make optimal data migration decisions under different load conditions.
[0078] Then, the storage location of the data to be operated is evaluated according to the data migration threshold parameters to obtain the location allocation result of the data to be operated. In this step, the system will conduct a detailed evaluation of each data block to be operated. The evaluation process first considers the current storage location and access characteristics of the data block and compares them with the previously determined migration threshold parameters. For example, if a data block is currently stored in the operating capacity area of the expansion card device, but its access frequency exceeds the set threshold, the system will mark it as a potential migration object. At the same time, the system will also consider factors such as the size of the data block, its association with other data blocks, and the current storage space usage. For each potential migration object, the system will calculate the expected performance improvement after migration. This calculation is based on the previously established performance prediction model and takes into account factors such as reduced access latency and improved bandwidth utilization after migration. The system will also evaluate the cost of the migration operation itself, including migration time, impact on current operations, etc. By comparing the performance improvement brought by migration and the migration cost, the system can decide whether it is worth migrating. In addition, the system will also consider the global optimization effect to avoid frequent small-scale migrations that cause system instability. Based on these comprehensive evaluations, the system assigns an optimal storage location to each data block. This location may remain unchanged or be migrated to another storage area. Finally, the system generates a detailed location allocation result, including the target storage location of each data block, expected performance improvement, migration priority, etc. This location allocation result provides clear guidance for subsequent data migration operations.
[0079] Finally, the execution order and resource allocation of data migration are planned and processed according to the location allocation results to obtain a data migration plan. In this stage, the system first prioritizes the data blocks to be migrated. The ranking takes into account multiple factors, including the expected performance improvement, the importance of the data blocks, and the current system load status. For data blocks with large performance improvement and high importance, the system will give higher migration priority. Next, the system will formulate a detailed migration schedule based on the current resource status, especially the PCIe bandwidth and the read and write speed of the storage device. This schedule will spread the migration operations into different time periods to avoid instantaneous high loads affecting system performance. The system will also consider data dependencies to ensure that the associated data blocks can be migrated in the correct order. In terms of resource allocation, the system will allocate appropriate DMA channels and buffers for each migration operation. For large data blocks, the system may adopt a block transfer strategy to divide them into multiple small blocks and migrate them step by step to reduce the impact on normal operations. At the same time, the system will also formulate a rollback plan to deal with errors or unexpected situations that may occur during the migration process. Finally, the system generates a complete data migration plan, which includes a detailed list of migration operations, the execution time of each operation, the required resources, the expected completion time, and other information. This migration plan not only guides the actual data migration process, but also provides the system with a dynamically adjustable execution framework, enabling the system to flexibly adjust the migration strategy according to real-time conditions, thereby optimizing the data storage location.
[0080] In this embodiment, the expansion card device is connected to the GPU through the PCIe interface on the GPU, and the BDF structure scanning and space mapping are performed to obtain storage space configuration information; the shared keyword code segment of the GPU kernel function code is preprocessed and the expansion card identifier is inserted, and the data transmission parameters are configured to obtain the initialization configuration of the CUDA shared storage mechanism; the expansion card device storage space is partitioned into the running capacity area and the cache area, and the GPU to-be-calculated data is thread-synchronized and parallelized through the SLC data block partition to obtain the storage space running status information; according to the running status information, the distribution of the to-be-calculated data between the running capacity area and the cache area is dynamically migrated. The present invention improves the storage resource utilization rate and reduces the data access delay by dynamically optimizing the data storage location, thereby accelerating the GPU computing process and improving the overall system performance.
[0081] The above describes the method for dynamically optimizing the data storage location in the embodiment of the present invention. The following describes the device for dynamically optimizing the data storage location in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a device for dynamically optimizing data storage locations includes:
[0082] The device initialization module 201 is used to connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information;
[0083] The storage configuration module 202 is used to pre-process the shared keyword code segment of the GPU kernel function code and insert the expansion card identifier according to the storage space configuration information, and configure the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism;
[0084] The space partitioning module 203 is used to partition the storage space of the expansion card device into an operating capacity area and a cache area according to the initialization configuration result, and perform thread synchronization and parallel computing processing on the to-be-calculated data of the GPU by SLC data block partitioning to obtain the operating status information of the storage space;
[0085] The dynamic optimization module 204 is used to detect and analyze the data access frequency during the GPU operation process according to the operation status information, and dynamically migrate the distribution of the data to be operated between the operation capacity area and the cache area according to the detection results to achieve dynamic optimization of the data storage location.
[0086] In an embodiment of the present invention, the data storage location dynamic optimization device runs the above-mentioned data storage location dynamic optimization method, the data storage location dynamic optimization device connects the expansion card device to the GPU through the PCIe interface on the GPU, and performs BDF structure scanning and space mapping to obtain storage space configuration information; preprocesses the shared keyword code segment of the GPU kernel function code and inserts the expansion card identifier, and configures the data transmission parameters to obtain the initialization configuration of the CUDA shared storage mechanism; partitions the expansion card device storage space into the running capacity area and the cache area, and performs thread synchronization and parallel calculation on the GPU to-be-calculated data through SLC data block division to obtain storage space running status information; according to the running status information, the distribution of the to-be-calculated data between the running capacity area and the cache area is dynamically migrated. The present invention improves the storage resource utilization rate and reduces the data access delay by dynamically optimizing the data storage location, thereby accelerating the GPU computing process and improving the overall system performance.
[0087] above Figure 2 The apparatus for dynamically optimizing the data storage location in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The apparatus for dynamically optimizing the data storage location in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0088] Figure 3It is a structural schematic diagram of a data storage location dynamic optimization device provided by an embodiment of the present invention. The data storage location dynamic optimization device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the data storage location dynamic optimization device 300. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the data storage location dynamic optimization device 300 to implement the steps of the above-mentioned data storage location dynamic optimization method.
[0089] The data storage location dynamic optimization device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that Figure 3 The structure of the data storage location dynamic optimization device shown does not constitute a limitation on the data storage location dynamic optimization device provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0090] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the method for dynamically optimizing the data storage location.
[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0093] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for dynamically optimizing data storage location, characterized in that: The method for dynamically optimizing data storage location comprises: Connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information; Preprocessing the shared keyword code segment of the GPU kernel function code and inserting the expansion card identifier according to the storage space configuration information, and configuring the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism; According to the initialization configuration result, the storage space of the expansion card device is partitioned into an operation capacity area and a cache area, and thread synchronization and parallel computing are performed on the to-be-calculated data of the GPU by SLC data block partitioning to obtain operation status information of the storage space; According to the operation status information, the data access frequency in the GPU operation process is detected and analyzed, and according to the detection result, the distribution of the data to be operated between the operation capacity area and the cache area is dynamically migrated to achieve dynamic optimization of the data storage location.
2. The method for dynamically optimizing data storage locations according to claim 1, characterized in that: The expansion card device is connected to the graphics processor through the PCIe interface on the GPU, and the GPU and the expansion card device are subjected to BDF structure scanning and space mapping processing to obtain storage space configuration information, including: Connect the expansion card device to the GPU through the PCIe interface on the GPU, read the VID / DID and scan the BDF structure of the GPU device and the expansion card device to obtain the device identification code; Perform an HBM firmware version check on the expansion card device according to the device identification code to obtain firmware information, and perform a memory space size check on the expansion card device according to the firmware information through a BAR register in a configuration space in a PCIe interface to obtain available storage space information; The CPU allocates an address range to the expansion card device according to the available storage space information to obtain storage space configuration information required for the GPU operation.
3. The method for dynamically optimizing data storage location according to claim 1, characterized in that: The preprocessing of the shared keyword code segment of the GPU kernel function code and the insertion of the expansion card identifier according to the storage space configuration information, and the configuration of the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism include: Scanning the kernel function code of the GPU according to the storage space configuration information, and identifying and extracting the shared keywords in the kernel function code to obtain the shared keyword code segment; Performing syntax parsing and parameter parsing processing on the shared keyword code segment, and inserting the identifier information of the expansion card device into the parsed code segment to obtain a shared memory declaration statement; The BME bit and the MSE bit of the command register are configured according to the shared memory declaration statement, and the data transmission channel and priority of the expansion card device are set to obtain the initialization configuration result of the CUDA shared memory mechanism.
4. The method for dynamically optimizing data storage location according to claim 1, characterized in that: The step of partitioning the storage space of the expansion card device into a running capacity area and a cache area according to the initialization configuration result, and performing thread synchronization and parallel computing processing on the to-be-calculated data of the GPU by SLC data block partitioning to obtain the running status information of the storage space includes: According to the initialization configuration result, the storage space of the expansion card device is partitioned into a running capacity area and a cache area to obtain a storage area partition result; According to the storage area division result, the GPU data to be calculated is grouped into data blocks, and the data blocks are allocated for storage by SLC data block calculation to obtain a data storage state; The SLC data blocks are processed in parallel according to the data storage status, and synchronization status marking processing is performed on the threads in the thread block through a synchronization function to obtain the operation status information of the storage space.
5. The method for dynamically optimizing data storage location according to claim 4, characterized in that: The step of performing data block grouping processing on the GPU data to be calculated according to the storage area division result, and performing storage allocation processing on the data blocks by SLC data block calculation to obtain the data storage state includes: Calculate the available space size of the running capacity area and the cache area according to the storage area division result to obtain the capacity parameters of each storage area; According to the capacity parameter, the GPU data to be calculated is divided into SLC data blocks and grouped and numbered to obtain a data block organization scheme; According to the data block organization scheme, storage location allocation processing is performed on each SLC data block between the running capacity area and the cache area to obtain the data storage state.
6. The method for dynamically optimizing data storage location according to claim 1, characterized in that: The detecting and analyzing the data access frequency in the GPU computing process according to the running status information, and dynamically migrating the distribution of the data to be computed between the running capacity area and the cache area according to the detection result to achieve dynamic optimization of the data storage location includes: Performing access counting processing on the data to be calculated in the GPU calculation according to the operation status information, and performing statistical analysis processing on the access frequency and time interval of the data to obtain data access characteristics; According to the data access characteristics, data distribution evaluation processing is performed on the GPU memory space and the storage space of the expansion card device, and the distribution of the operation data between the running capacity area and the cache area is calculated to obtain a data migration plan; According to the data migration scheme, the position of the data to be operated is adjusted between the running capacity area and the cache area, and the data access efficiency of the GPU operation process is statistically analyzed to obtain the optimal configuration result of the data storage position.
7. The method for dynamically optimizing data storage location according to claim 6, characterized in that: The data distribution evaluation process is performed on the GPU memory space and the storage space of the expansion card device according to the data access characteristics, and the distribution of the operation data between the running capacity area and the cache area is calculated and processed to obtain the data migration plan, which includes: Monitor and collect the working status of the GPU's video memory space and the storage space of the expansion card device to obtain load status information including space usage and bandwidth occupancy; Performing correlation analysis on the load status information and the data access characteristics to obtain a data migration threshold parameter for each storage area; The storage location of the data to be operated is evaluated according to the data migration threshold parameter to obtain a location allocation result of the data to be operated; The execution order and resource allocation of data migration are planned and processed according to the position allocation result to obtain the data migration plan.
8. A data storage location dynamic optimization device, characterized in that: The data storage location dynamic optimization device comprises: The device initialization module is used to connect the expansion card device to the graphics processor through the PCIe interface on the GPU, and perform BDF structure scanning and space mapping processing on the GPU and the expansion card device to obtain storage space configuration information; A storage configuration module is used to pre-process the shared keyword code segment of the GPU kernel function code and insert the expansion card identifier according to the storage space configuration information, and configure the data transmission parameters through the command register to obtain the initialization configuration result of the CUDA shared storage mechanism; A space partitioning module is used to partition the storage space of the expansion card device into an operating capacity area and a cache area according to the initialization configuration result, and perform thread synchronization and parallel computing processing on the to-be-calculated data of the GPU by SLC data block partitioning to obtain operating status information of the storage space; The dynamic optimization module is used to detect and analyze the data access frequency during the GPU operation process according to the operation status information, and dynamically migrate the distribution of the data to be operated between the operation capacity area and the cache area according to the detection result to achieve dynamic optimization of the data storage location.
9. A data storage location dynamic optimization device, characterized in that: The data storage location dynamic optimization device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the data storage location dynamic optimization device to perform the steps of the data storage location dynamic optimization method as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the steps of the method for dynamically optimizing data storage locations as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
GPU resource elastic scheduling method based on heterogeneous application platform
CN112698947A
Streaming data heterogeneous computing memory optimization method based on dynamic telescopic memory pool
CN114048025A