Method and system for automatically configuring CUDA environment and testing GPU pressure

By automatically configuring the CUDA environment and GPU stress testing method, the problem of CUDA environment configuration and GPU performance testing in the prior art relying on manual operation, and automated installation and all-round testing are realized, and efficiency and accuracy are improved.

CN119922183AActive Publication Date: 2025-05-02POWERLEADER COMPUTER SYST CO LTD

Patent Information

Application Number
CN202510416068.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-02
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the prior art, the configuration and GPU performance testing of the CUDA environment rely on manual operation, the process is cumbersome and error-prone, and the existing testing tools cannot conduct comprehensive GPU stress testing.

Method used

Provides automatic configuration of CUDA environment and GPU stress testing methods, and builds a mapping relationship table of device model and architecture, CUDA version, and cuDNN version, enter account information, automatically download and install CUDA and cuDNN, configure environment variables, and finally conduct comprehensive GPU testing.

Benefits of technology

It realizes automatic download and installation of CUDA and cuDNN, and automatically configures environment variables, reduces the need for manual operations, improves the efficiency and accuracy of the installation process, and judges whether the hardware is damaged through comprehensive testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922183A_ABST
    Figure CN119922183A_ABST
Patent Text Reader

Abstract

The invention relates to a CUDA environment automatic configuration and GPU pressure test method and system. The method comprises the following steps: constructing a mapping relation table of the model and architecture of equipment, cuda version and cudnn version; account information is input; according to the account number and the password of the configured operating system, the Linux server is accessed through SSH protocol connection, and then the Shell script is transmitted through the SFTP protocol; acquiring the model of the equipment, returning the model of the equipment, querying a predefined mapping relation table according to a query requirement for acquiring the model of the equipment so as to determine versions of a framework, cuda and cudnn corresponding to the model of the equipment, and automatically executing downloading, installation and verification so as to acquire an automatically configured environment variable; memory is distributed, a matrix is initialized, a matrix operation algorithm is started, and the computing power and the error correction capability of the GPU are tested in an omnibearing mode. According to the method, automatic cuda downloading and cuda environment configuration are realized, and the GPU is subjected to omnibearing testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an automatic configuration CUDA environment and GPU stress testing method and system, belonging to the communication field. Background Art

[0002] In today's high-performance computing and artificial intelligence fields, the configuration of the CUDA environment and GPU performance testing are two crucial tasks. However, these two tasks currently mostly rely on manual operations, which have many problems and shortcomings. Among them, for the configuration of the CUDA environment, users are usually required to manually download the corresponding version of the CUDA toolkit from the official website according to their own needs, and then configure it according to the complex installation steps. This process is not only cumbersome, but also prone to configuration failure due to improper operation. In terms of GPU performance testing, although there are a variety of testing tools on the market, most tools can only perform single-dimensional testing and cannot perform comprehensive stress testing on the GPU. To perform a comprehensive test on it, it is often necessary to manually combine multiple testing tools and manually set the test parameters, which not only increases the complexity of the operation, but also makes it difficult to accurately determine whether the hardware has potential damage. Summary of the invention

[0003] The present invention provides a method and system for automatically configuring a CUDA environment and a GPU stress test, aiming to solve at least one of the technical problems existing in the prior art.

[0004] The technical solution of the present invention relates to a method for automatically configuring a CUDA environment and a GPU stress test. The method according to the present invention comprises the following steps: A method for automatically configuring a CUDA environment and a GPU stress test, the method comprising the following steps: S100, build the mapping table between device model and architecture, cuda version, and cudnn version; S200, enter the account information; according to the configured operating system account and password, connect to the Linux server through the SSH protocol, and then transfer the Shell script through the SFTP protocol; S300, obtaining the model of the device and returning it, querying the predefined mapping relationship table according to the query need of obtaining the device model to determine the architecture, cuda and cudnn versions corresponding to the device model, and automatically downloading, installing and verifying to obtain automatically configured environment variables; S400, allocates memory and initializes the matrix, starts the matrix operation algorithm, and comprehensively tests the GPU's computing power and error correction capabilities.

[0005] Further, the step S200 includes: S210, obtaining input information of the front-end web server; wherein the input information includes an IP address, an account number and a password; S220, passing the input information to the backend server, and starting the SSH service through the Python server thread; S230. Use the configured operating system account and password to log in through SSH, connect to the Linux system through the SSH protocol, and transfer the automatic installation script through the SFTP protocol.

[0006] Further, the step S300 includes: S310, SSH executes Shell, which connects to a remote server via SSH and executes a Shell script; S320, in the Shell script, call exec_command to execute command nvidia-smi to output model information; S330, querying a predefined mapping relationship table according to the extracted model to obtain CUDA and cuDNN version information corresponding to different models; S340, automatically downloading and installing CUDA according to the queried CUDA version, and automatically downloading and installing cuDNN according to the queried cuDNN version; S350, after installation, automatically configures environment variables to ensure that CUDA and cuDNN can be correctly identified and used; S350, verify whether the CUDA environment is configured correctly.

[0007] Further, the step S340 includes: The supported CUDA version range is found from the predefined mapping table, and the minimum and maximum CUDA version numbers of the corresponding entries in the mapping table are read by setting the internal field separator IFS to a space, and the maximum version number is output to determine the CUDA version range suitable for the computing architecture.

[0008] Further, in step S340, Assign the incoming cuDNN version number and CUDA version number to the local variables cudnn_version and cuda_version respectively, output the cuDNN version information being downloaded, build the cuDNN download link, store it in the variable cudnn_url, and then download the link to the specified file; Use the curl command and the cookies file to log in and download. After logging in to the NVIDIA developer website through the curl command and saving the cookies information to the cookie_file file, use the saved cookies information to download the cuDNN library from the built link to the specified cudnn_path path; In the install_cudnn function, assign the passed cuDNN version number and CUDA version number to the local variables cudnn_version and cuda_version respectively, build the decompression path of cuDNN, use the tar command to decompress the downloaded cuDNN compressed package to the specified directory, and complete the installation of the cuDNN library.

[0009] Further, the step S400 includes: S410, allocating memory on the host and initializing matrices A and B; S420, randomly inject some errors into matrices A and B; S430, allocating memory for matrices A and B and result matrix C on the GPU; S440, copying the matrix data on the host from the host memory to the device memory; S450, starting a first kernel function of a matrix operation to multiply matrices A and B, and storing the result in matrix C; S460, after the matrix operation is completed, using the second kernel function to randomly inject some errors into the result matrix C; S460, check whether there is any error in CUDA call; S470. Release memory resources on the device and the host.

[0010] Furthermore, in step S420, the randomly injected error is implemented by an injectError function, which adds random values ​​to certain elements in the matrix with a certain probability.

[0011] Further, in step S400, Calculate the starting row row and column col of the current thread block in the global matrix to locate the matrix element access; The row is calculated as follows: multiply the thread block's y-dimension index blockIdx.y by the thread block's y-dimension size blockDim.y, and then add the thread's y-dimension index threadIdx.y; where col is calculated by multiplying the thread block's x-dimension index blockIdx.x by the thread block's x-dimension size blockDim.x, and then adding the thread's x-dimension index threadIdx.x.

[0012] The technical solution of the present invention also relates to a computer-readable storage medium on which program instructions are stored, and the above-mentioned method is implemented when the program instructions are executed by a processor.

[0013] The technical solution of the present invention also relates to an automatic configuration CUDA environment and GPU stress testing system, wherein the system comprises a computer device, and the computer device comprises the above-mentioned computer-readable storage medium.

[0014] The beneficial effects of the present invention are as follows: The method and system for automatically configuring the CUDA environment and GPU stress testing of the present invention can realize automatic downloading of cuda and configuration of the cuda environment, and perform a comprehensive test on the GPU. The present invention realizes the automatic downloading and installation of CUDA and cuDNN, and automatically configures the environment variables, ensuring that the environment is immediately available after installation, without the need for users to manually edit the configuration files, greatly reducing the need for manual operations, and improving the efficiency and accuracy of the installation process. Combined with the GPU stress testing function, it can perform a comprehensive test on the GPU, which introduces artificial damage at the software level, tests the robustness and error handling capabilities of the system, and is conducive to accurately judging whether the hardware is damaged. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flow chart of the method according to the present invention.

[0016] Figure 2 It is a schematic diagram of a mapping relationship table between device model and architecture, cuda version, and cudnn version according to the method of the present invention.

[0017] Figure 3 4 is a flowchart of account information input according to the method of the present invention.

[0018] Figure 4 It is a Shell execution flow chart according to the method of the present invention.

[0019] Figure 5 It is a flow chart of the matrix operation algorithm according to the method of the present invention. DETAILED DESCRIPTION

[0020] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effect of the present invention.

[0021] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to another feature, or it may be indirectly fixed or connected to another feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. The terms used in this specification are intended only to describe specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0022] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.

[0023] Reference Figures 1 to 5 In some embodiments, the method for automatically configuring the CUDA environment and GPU stress testing according to the present invention includes at least the following steps: S100, build the mapping table between device model and architecture, cuda version, and cudnn version; S200, enter the account information; according to the configured operating system account and password, connect to the Linux server through the SSH protocol, and then transfer the Shell script through the SFTP protocol; S300, obtaining the model of the device and returning it, querying the predefined mapping relationship table according to the query need of obtaining the device model to determine the architecture, cuda and cudnn versions corresponding to the device model, and automatically downloading, installing and verifying to obtain automatically configured environment variables; S400, allocates memory and initializes the matrix, starts the matrix operation algorithm, and comprehensively tests the GPU's computing power and error correction capabilities.

[0024] In some embodiments, see Figure 2 The present invention first constructs a mapping relationship table between device model and architecture, cuda version, and cudnn version. The above mapping relationship table will be updated with the iteration of GPU.

[0025] In some embodiments, in the account information input process of the present invention, according to the configuration of the operating system account and password, the Linux server is connected through the SSH protocol, and then the Shell script is transmitted through the SFTP protocol. Figure 3 , obtain the input information of the front-end web server, including the IP address, account and password, pass the input information to the back-end server, start the SSH service through the Python server thread, use the configured operating system account and password to log in through SSH, connect to the Linux system through the SSH protocol, and transfer the automatic installation script through the SFTP protocol.

[0026] Specifically, first import the paramiko module. Then create an SSH client instance named ssh to create an SSH connection. Next, set the host key of the SSH client and use paramiko.AutoAddPolicy() to automatically add unknown host keys to the local known host list. Next, connect to the remote server through the connect function, whose parameters include: the remote server's IP address remote_ip, port number remote_port, username remote_user and password remote_password.

[0027] Then, define a string variable named env_content to store the environment variable configuration content. Next, open (or create) a file named env.sh and write the content in the env_content variable to the file in write mode ("w"). Next, create an SFTP client instance through ssh.open_sftp() and assign it to the variable sftp. Next, upload the local env.sh file to the directory of the remote server. Finally, delete the local env.sh file.

[0028] In some embodiments, during the execution of Shell in the present invention, the gpu model is determined by the shell script and returned to the control end. The control end queries the architecture and the corresponding cuda and cudnn versions according to the model query needs, and automatically performs download, installation and nvcc verification. After completion, a cuda environment is configured. See Figure 4 First, SSH executes Shell, which connects to the remote server through SSH and executes a Shell script. Then in the Shell script, call exec_command to execute the command nvidia-smi to output model information. Then, according to the extracted model, query the predefined mapping relationship table (see Figure 2), get the CUDA and cuDNN version information corresponding to different models, automatically download and install CUDA according to the queried CUDA version, and automatically download and install cuDNN according to the queried cuDNN version. After the installation is complete, automatically configure the environment variables to ensure that CUDA and cuDNN can be correctly identified and used. Then use the nvcc (NVIDIA CUDA Compiler) command to verify whether the CUDA environment is configured correctly. If the nvcc command can be executed successfully, it means that the CUDA environment has been correctly configured.

[0029] Specifically, by executing the command nvidia-smi --query-gpu=name --format=csv,noheader,nounits, the model information of the NVIDIA graphics card in the current system is obtained and stored in the variable gpu_name. Using the queried graphics card model, the predefined mapping relationship table (see Figure 2 ) to retrieve the corresponding computing architecture code and output it. Then, according to the retrieved computing architecture code, select the predefined mapping relationship table (see Figure 2 ) to find the supported CUDA version range, and by setting the internal field separator IFS to a space, read the minimum and maximum CUDA version numbers of the corresponding entry in the mapping table, and output the maximum version number to determine the CUDA version range suitable for the computing architecture. Then, according to the determined CUDA version number, select the CUDA version from the predefined mapping relationship table (see Figure 2 ) to retrieve the compatible cuDNN version number and output it so that the correct version of the cuDNN library can be downloaded and installed.

[0030] Next, through the defined install_cuda function, the incoming CUDA version number is assigned to the local variable cuda_version, the CUDA version information being installed is output, the download link of the CUDA installer is constructed, and it is stored in the variable installer_url, and the link is downloaded to the temporary path file. Then, the wget command is used to download the CUDA installation package from the constructed link to the specified installer_path path, and after setting the file to executable permissions, the sudo command is used to install the CUDA toolkit in silent mode. Then, in the download_cudnn function, the incoming cuDNN version number and CUDA version number are assigned to the local variables cudnn_version and cuda_version respectively, the cuDNN version information being downloaded is output, the download link of cuDNN is constructed, and it is stored in the variable cudnn_url, and the link is downloaded to the specified file. Then, use the curl command and the cookies file to log in and download. After logging in to the NVIDIA developer website through the curl command and saving the cookies information to the cookie_file file, use the saved cookies information to download the cuDNN library from the constructed link to the specified cudnn_path path. Then, in the install_cudnn function, assign the incoming cuDNN version number and CUDA version number to the local variables cudnn_version and cuda_version respectively, then build the cuDNN decompression path, use the tar command to decompress the downloaded cuDNN compressed package to the specified directory, and complete the installation of the cuDNN library.

[0031] Next, define a function named configure_environment in the provided script to configure the CUDA environment variables. It assigns the incoming CUDA version number to the local variable cuda_version, outputs the path of the environment variable being configured, and uses the echo command to add CUDA-related environment variables to the user's .bashrc file, where the environment variables include PATH and LD_LIBRARY_PATH. Use the source ~ / .bashrc command to reload the .bashrc file so that the environment variables take effect immediately. Then, the main function of the script calls the get_gpu_info function to obtain the GPU model, stores it in the variable gpu_name, and outputs the detected GPU model information. Then, call the get_gpu_architecture function to obtain the GPU computing architecture and store it in the variable architecture. If the architecture is not detected, output the unknown GPU model and exit the program. Then, call the get_cuda_version_for_architecture function to get the recommended CUDA version according to the architecture, store it in the variable recommended_cuda_version, and output the recommended CUDA version. Also, call the get_cudnn_version_for_cuda function to get the recommended cuDNN version according to the recommended CUDA version, store it in the variable recommended_cudnn_version, and output the recommended cuDNN version. Then, call the install_cuda function to install the recommended CUDA version, call the download_cudnn function to download the recommended cuDNN version, and call the install_cudnn function to install the version. Then, call the configure_environment function to configure the environment variables, and use the nvcc --version and nvidia-smi commands to verify the installation of CUDA and cuDNN, and output the information that the installation and configuration are completed.

[0032] It should be noted that the present invention can automatically identify the GPU model and architecture through the install_cuda and install_cudnn functions, download and install compatible CUDA and cuDNN versions, achieve full automation, no user intervention, and significantly reduce the complexity and error rate of operation. In addition, the configure_environment function automatically configures the environment variables, ensuring that the installed environment is immediately available without the need for the user to manually edit the configuration file.

[0033] In some embodiments, the present invention executes a custom matrix operation algorithm in multiple steps, including memory allocation, matrix initialization, GPU memory allocation, data copying to memory, thread block and grid setting, matrix multiplication startup, error checking and resource cleaning, so as to comprehensively test the GPU's computing power and error correction capabilities. If the hardware is damaged, this program will report an error. Figure 4 , which includes the steps: S410, initialize the matrix. The program first allocates memory on the host and initializes two matrices A and B, where the size of each matrix is ​​N x N, where N is the defined matrix size. Further, in this example, the size of the matrix is ​​2048.

[0034] S420, inject errors. In order to simulate hardware failure or memory corruption, the program randomly injects some errors into the initialized matrices A and B. Among them, these errors are implemented by the injectError function, which adds random values ​​to certain elements in the matrix with a certain probability. The above probability is defined by INJECTION_RATE, which is 10% in this example.

[0035] S430, allocate device memory. The program allocates memory for matrices A, B and result matrix C on the GPU.

[0036] S440, data transfer. Copy the matrix data on the host from the host memory to the device memory.

[0037] S450, execute the matrix multiplication kernel. The program starts an optimized matrix multiplication kernel function matMulOptimized, which uses shared memory to improve calculation efficiency. Further, the kernel function multiplies matrices A and B and stores the result in matrix C.

[0038] S460, simulate memory corruption. To further simulate memory corruption, the program uses another kernel simulateMemoryCorruption to randomly inject some errors into the result matrix C after the matrix multiplication is completed.

[0039] S470, Error check. The program uses cudaGetLastError to check whether an error occurs in the CUDA call.

[0040] S480, clean up resources. The program releases memory resources on the device and host.

[0041] Specifically, a CUDA kernel function named matMulOptimized is first defined to optimize matrix multiplication operations. The function defines a shared memory array sharedMem, which is used to store the sub-block data of matrices A and B during the execution of the kernel function. Then, three floating-point pointers sharedA, sharedB and C are defined to point to the corresponding positions of the sub-blocks of matrices A and B and the result matrix C in the shared memory, respectively. Then, the starting row row and column col of the current thread block in the global matrix are calculated to prepare for the subsequent matrix element access positioning, where row is calculated by multiplying the thread block's y-dimension index blockIdx.y by the thread block's y-dimension size blockDim.y, and then adding the thread's y-dimension index threadIdx.y, and col is calculated by multiplying the thread block's x-dimension index blockIdx.x by the thread block's x-dimension size blockDim.x, and then adding the thread's x-dimension index threadIdx.x.

[0042] It should be noted that the present invention can introduce artificial corruption at the software level by simulating the memory corruption function simulateMemoryCorruption to test the robustness and error handling capability of the system.

[0043] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or implemented by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level process or object-oriented programming language to communicate with a computer system. However, if necessary, the program can be implemented in an assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed ASIC for this purpose.

[0044] In addition, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.

[0045] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, an RSM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0046] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including specific visual depictions of physical and tangible objects produced on the display.

[0047] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same means, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical scheme and / or implementation method may have various modifications and changes.

Claims

1. A method for automatically configuring a CUDA environment and a GPU stress test, characterized in that: The method comprises the following steps: S100, build the mapping table between device model and architecture, cuda version, and cudnn version; S200, enter the account information; according to the configured operating system account and password, connect to the Linux server through the SSH protocol, and then transfer the Shell script through the SFTP protocol; S300, obtaining the model of the device and returning it, querying the predefined mapping relationship table according to the query need of obtaining the device model to determine the architecture, cuda and cudnn versions corresponding to the device model, and automatically downloading, installing and verifying to obtain automatically configured environment variables; S400, allocates memory and initializes the matrix, starts the matrix operation algorithm, and comprehensively tests the GPU's computing power and error correction capabilities.

2. The method according to claim 1, characterized in that The step S200 includes: S210, obtaining input information of the front-end web server; wherein the input information includes an IP address, an account number and a password; S220, passing the input information to the backend server, and starting the SSH service through the Python server thread; S230. Use the configured operating system account and password to log in through SSH, connect to the Linux system through the SSH protocol, and transfer the automatic installation script through the SFTP protocol.

3. The method according to claim 1, characterized in that The step S300 includes: S310, SSH executes Shell, which connects to a remote server via SSH and executes a Shell script; S320, in the Shell script, call exec_command to execute command nvidia-smi to output model information; S330, querying a predefined mapping relationship table according to the extracted model to obtain CUDA and cuDNN version information corresponding to different models; S340, automatically downloading and installing CUDA according to the queried CUDA version, and automatically downloading and installing cuDNN according to the queried cuDNN version; S350, after installation, automatically configures environment variables to ensure that CUDA and cuDNN can be correctly identified and used; S350, verify whether the CUDA environment is configured correctly.

4. The method according to claim 3, characterized in that The step S340 includes: Find the supported CUDA version range from the predefined mapping table, set the internal field separator IFS to a space, read the minimum and maximum CUDA version numbers of the corresponding entry in the mapping table, and output the maximum version number to determine the CUDA version range suitable for the computing architecture.

5. The method according to claim 4, characterized in that In the step S340, Assign the incoming cuDNN version number and CUDA version number to the local variables cudnn_version and cuda_version respectively, output the cuDNN version information being downloaded, build the cuDNN download link, store it in the variable cudnn_url, and then download the link to the specified file; Use the curl command and the cookies file to log in and download. After logging in to the NVIDIA developer website through the curl command and saving the cookies information to the cookie_file file, use the saved cookies information to download the cuDNN library from the built link to the specified cudnn_path path; In the install_cudnn function, assign the passed cuDNN version number and CUDA version number to the local variables cudnn_version and cuda_version respectively, build the decompression path of cuDNN, use the tar command to decompress the downloaded cuDNN compressed package to the specified directory, and complete the installation of the cuDNN library.

6. The method according to claim 1, characterized in that The step S400 includes: S410, allocating memory on the host and initializing matrices A and B; S420, randomly inject some errors into matrices A and B; S430, allocating memory for matrices A and B and result matrix C on the GPU; S440, copying the matrix data on the host from the host memory to the device memory; S450, starting a first kernel function of a matrix operation to multiply matrices A and B, and storing the result in matrix C; S460, after the matrix operation is completed, using the second kernel function to randomly inject some errors into the result matrix C; S460, check whether there is any error in CUDA call; S470. Release memory resources on the device and the host.

7. The method according to claim 5, characterized in that In step S420, the randomly injected error is implemented by the injectError function, which adds random values ​​to certain elements in the matrix with a certain probability.

8. The method according to claim 6, characterized in that In the step S400, Calculate the starting row row and column col of the current thread block in the global matrix to locate the matrix element access; The row is calculated as follows: multiply the thread block's y-dimension index blockIdx.y by the thread block's y-dimension size blockDim.y, and then add the thread's y-dimension index threadIdx.y; where col is calculated by multiplying the thread block's x-dimension index blockIdx.x by the thread block's x-dimension size blockDim.x, and then adding the thread's x-dimension index threadIdx.x.

9. A computer-readable storage medium, characterized in that: Program instructions are stored thereon, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. An automated configuration CUDA environment and GPU stress testing system, characterized in that: include: A computer device comprising a computer readable storage medium according to claim 9.

Citation Information

Patent Citations

  • Linux-based man-machine interaction NVIDIA GPU (Graphics Processing Unit) automatic testing method

    CN104268046A

  • A method and a system for testing the floating-point operation performance of a GPU

    CN108958999A

  • Deployment method and device of deep learning platform cluster, medium and electronic equipment

    CN111459506A

  • Automatic test method for GPU systematization

    CN116089181A

  • Pod creation method and device, computer equipment and storage medium

    CN119002937A

Cited By

  • GPU performance test method and storage medium

    CN120179483A