Method and device for testing GPU (Graphics Processing Unit) graphics card and related product

By testing GPU graphics cards under power and temperature fluctuations, the problem of insignificant test results in the existing technology is solved, the test effect and accuracy are improved, and the background engineering technology for model training is optimized.

CN120371615APending Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410109533.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the test effect of GPU graphics cards is not obvious, and it is impossible to effectively simulate power and temperature fluctuations in actual operation, resulting in poor test results.

Method used

By increasing and adjusting the test fixed power and temperature, simulating the power and temperature fluctuation environment in actual operation, testing the GPU graphics card, obtaining test abnormal points, and generating a result table that fails the power and temperature test.

Benefits of technology

It improves the testing effect and accuracy of GPU graphics cards, can comprehensively test in power and temperature fluctuations, discover potential test abnormalities, optimize the background engineering technology for model training, and improve model training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371615A_ABST
    Figure CN120371615A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU graphics card testing method and device and a related product. The method comprises the following steps: testing a to-be-tested GPU display card under a test increased power to obtain a test abnormal point of the to-be-tested GPU display card under the test increased power, and testing the to-be-tested GPU display card under a test increased temperature to obtain a test abnormal point of the to-be-tested GPU display card under the test increased temperature; and according to the test abnormal points of the to-be-tested GPU display card under the test increased power and the test abnormal points of the to-be-tested GPU display card under the test increased temperature, obtaining a test result table that the to-be-tested GPU display card does not pass the power test and a test result table that the to-be-tested GPU display card does not pass the temperature test. Therefore, according to the method and the device, the test of the to-be-tested GPU display card is realized by increasing the test power and the test temperature, and the test effect of the GPU display card is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of GPU testing, and particularly to a testing method, device and related products for a GPU graphics card. Background Art

[0002] GPU graphics cards are mainly used to support the development of services such as model training. Before carrying out the model training service, it is necessary to test the GPU graphics card to ensure its availability during the model training process. In the related art, a simulation test program related to the model training service is mainly constructed to test whether the GPU graphics card can run normally through this simulation test program. However, the simulation test program mimics the environmental stable state during the model training process. Therefore, the load and temperature of the GPU graphics card are in a stable state during the test by this simulation test program, resulting in an unclear test effect of the GPU graphics card. Therefore, how to improve the test effect of the GPU graphics card has become an urgent technical problem in the current field. Summary of the Invention

[0003] Embodiments of this application provide a testing method, device and related products for a GPU graphics card, aiming to improve the test effect of the GPU graphics card.

[0004] The first aspect of this application provides a testing method for a GPU graphics card, including:

[0005] Obtain a GPU graphics card to be tested, a test fixed power and a test fixed temperature;

[0006] Increase and adjust the test fixed power and the test fixed temperature respectively to obtain a test increased power and a test increased temperature;

[0007] Test the GPU graphics card to be tested at the test increased power to obtain a test abnormal point of the GPU graphics card to be tested at the test increased power, and test the GPU graphics card to be tested at the test increased temperature to obtain a test abnormal point of the GPU graphics card to be tested at the test increased temperature;

[0008] According to the test abnormal point of the GPU graphics card to be tested at the test increased power and the test abnormal point of the GPU graphics card to be tested at the test increased temperature, obtain a test result table of the GPU graphics card to be tested that fails the power test and a test result table of the GPU graphics card to be tested that fails the temperature test.

[0009] The second aspect of this application provides a testing device for a GPU graphics card, including:

[0010] A GPU graphics card acquisition unit, configured to obtain a GPU graphics card to be tested, a test fixed power and a test fixed temperature;

[0011] A power temperature adjustment unit for respectively increasing and adjusting the test fixed power and the test fixed temperature to obtain a test increased power and a test increased temperature;

[0012] A GPU graphics card test unit for testing the to-be-tested GPU graphics card under the test increased power to obtain a test abnormal point of the to-be-tested GPU graphics card under the test increased power, and testing the to-be-tested GPU graphics card under the test increased temperature to obtain a test abnormal point of the to-be-tested GPU graphics card under the test increased temperature;

[0013] A test result table obtaining unit for obtaining a test result table of the to-be-tested GPU graphics card failing the power test and a test result table of the to-be-tested GPU graphics card failing the temperature test according to the test abnormal point of the to-be-tested GPU graphics card under the test increased power and the test abnormal point of the to-be-tested GPU graphics card under the test increased temperature.

[0014] The third aspect of the present application provides a computer device, which includes a processor and a memory:

[0015] The memory is used for storing a computer program and transmitting the computer program to the processor;

[0016] The processor is used for executing the steps of the GPU graphics card test method provided in the first aspect according to the instructions in the computer program.

[0017] The fourth aspect of the present application provides a computer-readable storage medium, which is used for storing a computer program, and when the computer program is executed by a computer device, the steps of the GPU graphics card test method provided in the first aspect are implemented.

[0018] The fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a GPU graphics card test device, the steps of the GPU graphics card test method provided in the first aspect are implemented.

[0019] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0020] In the technical solution of this application, first, the GPU graphics card to be tested, the test fixed power, and the test fixed temperature are obtained. After that, the test fixed power and the test fixed temperature are respectively increased and adjusted to obtain the test increased power and the test increased temperature. Then, the GPU graphics card to be tested is tested under the test increased power to obtain the test abnormal points of the GPU graphics card to be tested under the test increased power, and the GPU graphics card to be tested is tested under the test increased temperature to obtain the test abnormal points of the GPU graphics card to be tested under the test increased temperature. Finally, according to the test abnormal points of the GPU graphics card to be tested under the test increased power and the test abnormal points of the GPU graphics card to be tested under the test increased temperature, the test result table of the GPU graphics card to be tested that fails the power test and the test result table of the GPU graphics card to be tested that fails the temperature test are obtained.

[0021] It can be seen that in this application, by increasing the test power and the test temperature, the power fluctuation environment and the temperature fluctuation environment that the GPU graphics card to be tested may encounter in the actual operating environment are simulated, so that the GPU graphics card to be tested can be tested under the power fluctuation environment and the temperature fluctuation environment. In this way, the GPU graphics card to be tested can be tested more comprehensively, thereby improving the test effect of the GPU graphics card. Description of the Drawings

[0022] Figure 1 It is a scene architecture diagram of a test method for a GPU graphics card provided by an embodiment of this application;

[0023] Figure 2 It is an application scenario diagram of a GPU graphics card in a test method for a GPU graphics card provided by an embodiment of this application;

[0024] Figure 3 It is a flowchart of a test method for a GPU graphics card provided by an embodiment of this application;

[0025] Figure 4 It is a test flowchart of the GPU graphics card to be tested in a test method for a GPU graphics card provided by an embodiment of this application;

[0026] Figure 5 It is a flowchart of testing the GPU graphics card to be tested in a test method for a GPU graphics card provided by an embodiment of this application;

[0027] Figure 6 It is a flowchart of testing the GPU graphics card to be tested in another test method for a GPU graphics card provided by an embodiment of this application;

[0028] Figure 7 It is a test flowchart of the GPU graphics card to be tested in another test method for a GPU graphics card provided by an embodiment of this application;

[0029] Figure 8Schematic diagram of a test result table in a GPU graphics card test method provided by an embodiment of the present application;

[0030] Figure 9 Full flow chart for testing a GPU graphics card in a GPU graphics card test method provided by an embodiment of the present application;

[0031] Figure 10 Schematic structural diagram of a GPU graphics card test device provided by an embodiment of the present application;

[0032] Figure 11 Schematic structural diagram of a server in an embodiment of the present application;

[0033] Figure 12 Schematic structural diagram of a terminal device in an embodiment of the present application. Detailed implementation manners

[0034] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0035] First, several noun terms that may be involved in the following embodiments of the present application are explained.

[0036] GPU graphics card: Refers to an independent slot-type hardware installed in a computer, which is specifically used to process graphics and computing tasks.

[0037] The GPU graphics card is mainly used to support the development of services such as model training. Before carrying out the model training service, it is necessary to test the GPU graphics card to ensure its availability during the model training process. It can be understood that the GPU graphics card can be used as the underlying hardware data to provide computing services for services such as model training. That is, a normally operating GPU graphics card must be used during the model training process. Therefore, the testing of the GPU graphics card has become a crucial preparatory work before model training.

[0038] In the related art, a simulation test program related to the model training service is mainly constructed to test whether the GPU graphics card can operate normally through this simulation test program. However, the simulation test program mimics the environmental stable state during the model training process. It can be understood that the simulation test program only mimics the stable environment of model training, that is, the GPU graphics card is tested in an environment where both the test load and the test temperature are relatively stable.

[0039] Although the GPU graphics card is tested in this way, when the GPU graphics card is actually running, it may encounter an environment where the running load and running temperature are in a relatively fluctuating environment. At this time, the GPU graphics card obtained through the above test process cannot provide certain support for model training. Therefore, testing the GPU graphics card through the test method in the relevant technology may lead to poor test results of the GPU graphics card. Therefore, how to improve the test effect of the GPU graphics card has become a technical problem that needs to be solved urgently in the current field.

[0040] In view of the above problems, a test method, device and related products for a GPU graphics card are provided in the present application, the purpose of which is to improve the test effect of the GPU graphics card. In the technical solution provided in the present application, firstly, the test fixed power and the test fixed temperature are increased and adjusted respectively to obtain the test increased power and the test increased temperature, after which the GPU graphics card to be tested is tested under the test increased power to obtain the test abnormal point of the GPU graphics card to be tested under the test increased power, and the GPU graphics card to be tested is tested under the test increased temperature to obtain the test abnormal point of the GPU graphics card to be tested under the test increased temperature, and finally, according to the test abnormal point of the GPU graphics card to be tested under the test increased power and the test abnormal point of the GPU graphics card to be tested under the test increased temperature, a test result table of the GPU graphics card to be tested failing the power test and a test result table of the GPU graphics card to be tested failing the temperature test are obtained.

[0041] It can be seen that in the present application, by increasing and adjusting the test fixed power and the test fixed temperature, the test increased power and the test increased temperature obtained after the adjustment can be used to simulate the power fluctuation environment and the temperature fluctuation environment that the GPU graphics card to be tested may encounter in the actual operating environment, so that the GPU graphics card to be tested can be tested in the power fluctuation environment and the temperature fluctuation environment. In this way, the GPU graphics card to be tested can be tested more comprehensively, thereby improving the test effect of the GPU graphics card. In the present application, the test abnormal points that may exist in the power fluctuation environment and the temperature fluctuation environment of the GPU graphics card to be tested can also be output, further improving the test accuracy of the GPU graphics card.

[0042] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. In the embodiments of this application, artificial intelligence technology can use machines to adjust the test fixed power and test fixed temperature, and can test the GPU graphics card to be tested under the adjusted power and temperature to obtain the test result table of the GPU graphics card to be tested, achieving better testing of the GPU graphics card to be tested.

[0043] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0044] The GPU graphics card testing method provided in this application mainly involves machine learning. Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0045] Furthermore, the GPU graphics card testing method provided in this application also involves cloud computing. Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources on demand and in an easily scalable manner through a network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services on demand and in an easily scalable manner through a network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as Grid Computing, Distributed Computing, Parallel Computing, Utility Computing, Network Storage Technologies, Virtualization, and Load Balance. In the embodiments of this application, cloud computing technology is mainly used to test the GPU graphics card to be tested at the adjusted power and temperature to obtain the test result table of the GPU graphics card to be tested, thus realizing the testing of the GPU graphics card to be tested in an environment with power and temperature fluctuations.

[0046] The execution subject of the GPU graphics card testing method provided in the embodiments of this application can be a terminal device. For example, obtain the GPU graphics card to be tested, the test fixed power, and the test fixed temperature on the terminal device. As an example, the terminal device can specifically include, but is not limited to, mobile phones, desktop computers, tablet computers, laptop computers, handheld computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. The execution subject of the GPU graphics card testing method provided in the embodiments of this application can also be a server, that is, the GPU graphics card to be tested, the test fixed power, and the test fixed temperature can be obtained on the server. In addition, the GPU graphics card testing method provided in the embodiments of this application can also be executed jointly by the terminal device and the server. Among them, the terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. Therefore, the implementation subject of executing the technical solution of this application in the embodiments of this application is not limited.

[0047] Figure 1 Exemplarily shows a scenario architecture diagram of a GPU graphics card testing method. The figure includes a server and various forms of terminal devices. Figure 1The server shown can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Additionally, the server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0048] See Figure 2 , which is an application scenario diagram of a GPU graphics card in a GPU graphics card testing method provided by an embodiment of the present application. In Figure 2 , it is shown that the GPU graphics card to be tested can be tested through the testing solution provided by the present application, and the GPU graphics card to be tested that passes the test can be used as the underlying data support for model training. It should be noted that the GPU graphics card that passes the test through the technical solution of the present application can be used to support model training, and this model training includes, but is not limited to, the training of voice AI models, visual CV models, game AI models, and medical AI models. It can be seen that the present application optimizes the background engineering technology of model training through the testing method of GPU graphics cards and improves the model training efficiency.

[0049] See Figure 3 , which is a flowchart of a GPU graphics card testing method provided by an embodiment of the present application. As Figure 3 shown in the GPU graphics card testing method, the following steps are included:

[0050] S301: Obtain the GPU graphics card to be tested, the test fixed power, and the test fixed temperature.

[0051] In this step, the GPU graphics card to be tested can be understood as the GPU graphics card that needs to be used to support model training. That is, it is necessary to test the GPU graphics card in advance to improve the usability of the GPU graphics card during model training. It should be noted that the number of GPU graphics cards to be tested in the present application mainly depends on the number of GPU graphics cards required during model training, and the number of GPU graphics cards to be tested is not limited herein.

[0052] Furthermore, before obtaining the GPU graphics card to be tested, the test fixed power, and the test fixed temperature, the present application can also obtain the total test data set, where the total test data set includes the total data set that is input into the GPU graphics card to be tested and runs in the GPU graphics card to be tested. It can be understood that the GPU graphics card is a hardware device, and in actual application, it is necessary to input the model training data into the GPU graphics card to run the model training data in the GPU graphics card to support the construction of the model.

[0053] In an implementable embodiment, before obtaining the GPU graphics card to be tested, the test fixed power, and the test fixed temperature, the present application may also obtain a historical test table, where the historical test table is constructed based on the test power and the test temperature corresponding to the historical test abnormal points during the historical test process. It can be understood that the historical test table is mainly composed of multiple historical test environments where abnormal points may be tested. For example, the GPU graphics card can be tested in the A historical test environment to obtain the test results obtained in the A historical test environment (such as the GPU graphics card fails the test in the A historical test environment, and the test abnormal point is test abnormal point A). The A historical test environment may include the A test power and the A test temperature, that is, the test results for the GPU graphics card can be obtained at the A test power and the A test temperature. Further, the historical test abnormal point can be understood as the abnormal point where the GPU graphics card fails the test in the A historical test environment.

[0054] It should be noted that the historical test table includes the test fixed power and the test fixed temperature, that is, a historical test environment can be determined through the test fixed power and the test fixed temperature. The historical test table may include the test fixed power and the test fixed temperature corresponding to different parameters. For example, the test fixed power corresponding to the A1 power parameter and the test fixed temperature corresponding to the A1 temperature parameter can form a historical test environment A, and the test fixed power corresponding to the B1 power parameter and the test fixed temperature corresponding to the B1 temperature parameter can form a historical test environment B.

[0055] It should also be noted that the total test data set in the present application can be separately stored in a data unit for quickly extracting data when the GPU graphics card needs to be tested. The total test data set can be classified according to the test requirements. For example, when testing the A GPU graphics card, the A total test data set is required to implement the test, and when testing the B GPU graphics card, the B total test data set is required to implement the test. The total test data set can also be stored in the historical test table, that is, each historical test environment in the historical test table may include one or more total test data sets. It can be understood that different test environments may require different test data to run on the GPU graphics card to be tested. The main body for storing the total test data set is not limited here and can also be applied according to actual needs.

[0056] As shown in Table 1, Table 1 is the historical test table of a test method for a GPU graphics card provided in an embodiment of the present application. Table 1 shows the parameter information included in the historical test table when the total test data set is stored in the historical test table. The historical test table is mainly composed of the historical test environment and the parameter information under the historical test environment. For example, the historical test table includes the historical test environment A, and mainly includes the A total test data set, the A test fixed power, and the A test fixed temperature under the historical test environment A.

[0057] It should be noted that the parameter information in the historical test environment of the historical test table in this application may also include the image name (i.e., the computing program and the running library files), data throughput (i.e., the amount of computing data throughput per second), execution time (i.e., the time range of computing operation), frequency (i.e., the frequency interval of fluctuations), and number of times (i.e., the count of the number of fluctuations), etc. There is no specific limitation here, and the parameter information in the historical test table can also be increased or deleted according to needs during the actual test process.

[0058] Table 1

[0059]

[0060] In some examples, before executing step S302 in this application, the total set of test data can also be input into the GPU graphics card to be tested at a fixed test power and a fixed test temperature, so as to run the total set of test data in the GPU graphics card to be tested, and realize the test of the GPU graphics card to be tested at a fixed test power and a fixed test temperature, so as to obtain an initial test result. It can be seen that since the historical test table proposed in this application includes the fixed test power and fixed test temperature that may detect abnormal points determined in advance through the historical test process, testing the GPU graphics card to be tested in advance according to the fixed test power and fixed test temperature can improve the test efficiency for the GPU graphics card.

[0061] Furthermore, if the initial test result is that the GPU graphics card to be tested fails the test at a fixed test power and a fixed test temperature, record the test abnormal point; if the initial test result is that the GPU graphics card to be tested passes the test at a fixed test power and a fixed test temperature, step S302 can be further executed to test the GPU graphics card to be tested in an environment with fluctuating power and temperature, so as to improve the test accuracy for the GPU graphics card.

[0062] S302: Increase and adjust the fixed test power and the fixed test temperature respectively to obtain an increased test power and an increased test temperature.

[0063] In this step, the fixed test power includes the test power generated by inputting the total set of test data into the GPU graphics card to be tested in fixed test batches. It can be understood that the fixed test batches are composed of multiple single batches, and the data input volume of each single batch is the same, and the fixed test power can be determined according to the amount of data calculated each time input into the GPU graphics card to be tested in the fixed test batches. Among them, if the data input volume is small, the test power is also small. Correspondingly, if the data input volume is large, the test power is also large. For example: The fixed test batches can be set to five times, and the total set of test data needs to be input into the GPU graphics card to be tested within five times.

[0064] Since it has been determined before this step that the GPU graphics card to be tested passes the test at the test fixed power and test fixed temperature, the test fixed power can be increased and adjusted next to test the GPU graphics card to be tested in an environment with a larger power. Specifically, the adjustment of the test fixed power can be achieved by adjusting the test fixed batch, that is, reducing the test fixed batch to obtain a test reduced batch, so as to increase the data input volume of each single batch, and then achieve the increase adjustment of the power.

[0065] It should be noted that the number of times of the test reduced batch is less than the number of times of the test fixed batch, and the data volume input to the GPU graphics card to be tested each time in the test reduced batch is greater than the data volume input to the GPU graphics card to be tested each time in the test fixed batch. Therefore, the test increased power can be determined according to the data calculation volume input to the GPU graphics card to be tested each time in the test reduced batch. As an example, if the total set of data to be tested includes 50 pieces of data and the test fixed batch is five times, then the data volume input to the GPU graphics card to be tested each time in the five times is 10 pieces of data. Since only 10 pieces of data are output each time, the input calculation volume is not large, and the corresponding test fixed power is not large either; at this time, by reducing the test fixed batch to two times, then the data volume input to the GPU graphics card to be tested each time in the two times is 25 pieces of data. Since 25 pieces of data are output each time, the input calculation volume is relatively increased, and correspondingly, the obtained test increased power is also increased compared with the test fixed power.

[0066] In this step, what also needs to be introduced is the test fixed temperature. The test fixed temperature includes the temperature generated by encrypting and inputting a partial data subset in the test data total set with a test fixed data volume into the GPU graphics card to be tested. Among them, the partial data subset is randomly determined, which can be the subset at the front of the queue in the test data total set or the subset at the back of the queue in the test data total set, and no specific limitation is made here.

[0067] It can be understood that the test fixed temperature is determined according to the fixed data volume input to the GPU graphics card to be tested. Since the encryption calculation needs to use the CPU and memory together, correspondingly, if the data volume is small, the temperature generated by inputting the data corresponding to this data volume into the GPU graphics card to be tested is also small; if the data volume is large, the temperature generated by inputting the data corresponding to this data volume into the GPU graphics card to be tested is also large. It should be noted that in this application, in addition to encrypting and inputting a partial data subset into the GPU graphics card to be tested to determine the temperature, the temperature can also be determined by performing encryption calculation on the partial data subset input to the GPU graphics card to be tested.

[0068] Further, in this application, the test fixed data volume can be adjusted to increase the data volume to obtain a test increased data volume. Since the data volume that needs to be encrypted and input is large, the corresponding test fixed temperature will also increase to a test increased temperature, where the test increased data volume is greater than the test fixed data volume. As an example, if the total set of data to be tested is 40 data entries and the test fixed data volume is 15 data entries, then the test fixed temperature is determined based on the encrypted calculation input for 15 data entries; at this time, the test fixed data volume (15 data entries) can be increased to 30 data entries (i.e., obtain the test increased data volume) to determine the test increased temperature based on the test increased data volume. In this way, by adjusting the data volume that needs to be encrypted and input for calculation, the increase adjustment from the test fixed temperature to the test increased temperature is achieved.

[0069] S303: Test the to-be-tested GPU graphics card at the test increased power to obtain the test anomaly points of the to-be-tested GPU graphics card at the test increased power, and test the to-be-tested GPU graphics card at the test increased temperature to obtain the test anomaly points of the to-be-tested GPU graphics card at the test increased temperature.

[0070] As Figure 4 shown, Figure 4 is the test flow chart of the to-be-tested GPU graphics card in a method for testing a GPU graphics card provided by an embodiment of this application. In Figure 4 (a), the power test anomaly points determined after adjusting the test fixed power and testing the to-be-tested GPU graphics card at the adjusted power are shown; in Figure 4 (b), the temperature test anomaly points determined after adjusting the test fixed temperature and testing the to-be-tested GPU graphics card at the adjusted temperature are shown. It should be noted that in this application, the test sequence of the power environment and the temperature environment is not limited. For example, the to-be-tested GPU graphics card can be tested in the temperature environment first, and then in the power environment.

[0071] In the embodiments of this application, for the implementation methods of testing the GPU graphics card at the test increased power and testing the GPU graphics card at the test increased temperature mentioned in the above step S303, the following will be introduced separately. It should be noted that the implementation methods given in the following introduction are only for exemplary illustration and do not represent all the implementation methods of the embodiments of this application.

[0072] The first implementation method mentioned in step S303 is: the implementation method of testing the to-be-tested GPU graphics card at the test increased power. As Figure 5 shown, Figure 5Flowchart for testing a GPU graphics card to be tested in a testing method provided by an embodiment of the present application Figure 5 The method includes the following steps:

[0073] S501: Divide the total set of test data into multiple data subsets according to the test reduction batches, and input the multiple data subsets into the GPU graphics card to be tested to drive the GPU graphics card to be tested to run the multiple data subsets, and obtain power test results.

[0074] In this step, one test reduction batch corresponds to the division of one data subset. It can be understood that the number of test reduction batches is the same as the number of data subset divisions, that is, the input of one data subset needs to be completed within a single batch of the test reduction batch. Since the smaller the test batch, the larger the amount of data input into the GPU graphics card to be tested, and at this time, the power generated by the multiple data subsets input into the GPU graphics card to be tested also becomes larger. After that, by driving the GPU graphics card to be tested to run these multiple data subsets, power test results for the GPU graphics card to be tested in a power fluctuation environment are obtained, realizing the excavation of large power fluctuations and improving the efficiency of subsequent process model training.

[0075] It should be noted that after this step, the power test results can also be processed to determine whether the GPU graphics card to be tested is operating abnormally. If the power test results show that the GPU graphics card to be tested is operating abnormally, then step S402 is executed; if the power test results show that the GPU graphics card to be tested is not operating abnormally, then the test reduction batches can be adjusted to continue simulating the power fluctuation environment to achieve the testing of the GPU graphics card hardware in different frequency fluctuation environments.

[0076] Specifically, in the embodiment of the present application, if the power test results show that the GPU graphics card to be tested is operating normally, then the test reduction batches can be adjusted to increase the batch size to obtain test increase batches, where the number of test increase batches is less than the number of test fixed batches and greater than the number of test reduction batches.

[0077] It should be noted that the test reduction batch in this application can be the minimum batch acceptable to the system. That is, when this application makes a reduction adjustment to the test fixed batch, the test fixed power can be adjusted to the maximum power acceptable to the system (test increased power). If the GPU card to be tested operates normally under the test increased power, the power can be adjusted downward on the basis of the test increased power. Among them, the fluctuation range of the adjusted power can span power adjustments (such as from power 5 to power 3), which is not specifically limited here and can be determined according to the actual situation. In this way, an environment with power fluctuations is simulated, so that the test of the GPU card to be tested can be realized in any power fluctuation environment.

[0078] After that, this application can re-divide the total test data set into multiple data target subsets according to the test increase batch, and input the multiple data target subsets into the GPU card to be tested. Then, the GPU card to be tested is driven to run the multiple data target subsets to obtain the power test target result. It should be noted that one test increase batch corresponds to the division of one data target subset, that is, the number of test increase batches is the same as the number of divisions of the data target subsets, and the input of one data target subset needs to be completed within a single batch of the test increase batch.

[0079] As an example, it is assumed that the total test data set includes 50 pieces of data, and the test reduction batch is two times. The amount of data input into the GPU card to be tested each time in the two times is 25 pieces of data; at this time, the test reduction batch can be increased to four times, then the amount of data input into the GPU card to be tested each time in the four times is 12.5 pieces of data, and the power is reduced on the basis of the test increased power to test the GPU card to be tested under the test reduced power.

[0080] Furthermore, this application can process the power test target result to determine whether the GPU card to be tested operates abnormally under the test increased power. If the power test target result is that the GPU card to be tested is abnormal, the power test target result can be analyzed to obtain the test abnormal point of the GPU card to be tested under the test reduced power; if the power test target result is that the GPU card to be tested is normal, the test increase batch can be adjusted until the GPU card to be tested operates normally at each power or the GPU card to be tested operates abnormally at a certain power.

[0081] In another implementable embodiment, when making a decreasing adjustment to a fixed test batch, the test fixed power can be gradually adjusted to the maximum power acceptable to the system. For example, from power 1 to power 2, and then from power 2 to power 3, until the maximum power (power 5) is reached or the test result of the GPU graphics card to be tested is abnormal at a certain power, thereby simulating each power fluctuation environment and solving the singularity of only testing the GPU graphics card in one environment in the related art.

[0082] S502: If the power test result shows that the GPU graphics card to be tested is operating abnormally, analyze the power test result to obtain the test abnormal point of the GPU graphics card to be tested at the increasing test power.

[0083] Among them, the manifestation of the test abnormal point is usually that problems occur with the GPU graphics card, such as GPU Lost errors, GPU ECC errors, or xid errors in the corresponding nvidia ecosystem stack. It should be noted that the test abnormal points may be different in any environment with power fluctuations. Therefore, through the technical solution of the present application, the environment in which the GPU graphics card may have abnormal operation can be found, as well as the test abnormal points that may occur in the environment with abnormal operation.

[0084] The second implementation method mentioned in step S303 is: the implementation method of testing the GPU graphics card to be tested at an increasing test temperature. As Figure 6 shown, Figure 6 is the flowchart for testing the GPU graphics card to be tested in another GPU graphics card testing method provided by the embodiment of the present application, Figure 6 including the following steps:

[0085] S601: Process the total test data set according to the increasing test data volume to obtain the target test data subset corresponding to the increasing test data volume in the total test data set.

[0086] In this step, the total test data set can be extracted according to the increasing test data volume to obtain the target test data subset corresponding to the increasing test data volume in the total test data set. For example: assuming that the total test data set includes 50 pieces of data and the increasing test data volume is 20 pieces of data, at this time, 20 pieces of data can be randomly obtained from the total test data set as the target test data subset corresponding to the increasing test data volume.

[0087] S602: Encrypt and input the target test data subset into the GPU graphics card to be tested to drive the GPU graphics card to be tested to run the target test data subset and obtain the temperature test result.

[0088] Further, the present application can perform encryption calculation on the obtained target test data subset to obtain the encrypted target test data subset, and input the encrypted target test data subset into the GPU card to be tested, so as to finally obtain the temperature test result by driving the GPU card to be tested to run the target test data subset. It can be understood that since the encryption calculation needs to run in coordination with the CPU and memory, the larger the data volume of the target test data subset, the faster the running speed of the CPU and memory, and it is possible to run the target test data subset in a high-temperature environment to obtain the temperature test result.

[0089] It should be noted that the present application can also process the temperature test result to determine whether the GPU card to be tested is operating abnormally. If the temperature test result shows that the GPU card to be tested is operating abnormally, then step S603 is executed; if the temperature test result shows that the GPU card to be tested is not operating abnormally, then the encryption calculation for the target test data subset can be stopped to simulate the test of the GPU card hardware in an environment where the temperature gradually decreases.

[0090] It can be understood that the increased data volume in the test of the present application can be the maximum encrypted data volume acceptable to the system, that is, when the present application increases and adjusts the fixed test data volume, the fixed test temperature can be adjusted to the maximum temperature acceptable to the system (increased test temperature). If the GPU card to be tested operates normally at the increased test temperature, then the encryption calculation can be gradually stopped on the basis of the increased test temperature, so as to ensure that data with different data volumes can be encrypted and calculated in each temperature fluctuation environment to obtain the test of the GPU card to be tested in different temperature environments.

[0091] S603: If the temperature test result shows that the GPU card to be tested is abnormal, analyze the temperature test result to obtain the test abnormal point of the GPU card to be tested at the increased test temperature.

[0092] In the present application, the test abnormal points existing in the temperature environment usually also manifest as problems with the GPU card, such as GPULost errors, GPU ECC errors, or xid errors in the corresponding nvidia ecosystem stack. It can be seen that in the present application, the specific temperature environment where the test abnormal points exist can be determined, further improving the test accuracy of the GPU card.

[0093] It should be noted that the terminal device can implement the above two implementation methods separately. Next, it is introduced that the present application can also test the GPU card to be tested simultaneously in the power environment and the temperature environment. As Figure 7 shown, Figure 7 is the test flow chart of the GPU card to be tested in another test method of the GPU card provided by the embodiment of the present application. InFigure 7 It shows that after power adjustment of the test fixed power, the test increased power is obtained, and after temperature adjustment of the test fixed temperature, the test increased temperature is obtained.

[0094] After that, the present application can bridge the test increased temperature and the test increased power, or bridge the test increased power and the test increased temperature, so as to simultaneously test the GPU graphics card to be tested under the test increased power and under the test increased temperature, and obtain the overlapping test anomaly points in the superimposed test environment of power and temperature. In this way, the present application can explore more test environments where the GPU graphics card to be tested may have test anomaly points to avoid test omission.

[0095] S304: Obtain the test result table of the GPU graphics card to be tested that fails the power test and the test result table of the GPU graphics card to be tested that fails the temperature test according to the test anomaly points of the GPU graphics card to be tested under the test increased power and the test anomaly points of the GPU graphics card to be tested under the test increased temperature.

[0096] As shown in Table 2, Table 2 is the test result table of a GPU graphics card testing method provided in an embodiment of the present application. The test anomaly points obtained by the GPU graphics card to be tested in different test power environments are shown in Table 2. For example: W1 is the test fixed power, and the GPU graphics card to be tested runs without anomaly under this test fixed power. W2 is the test increased power, and the GPU graphics card to be tested runs anomalously under this test increased power, and the test anomaly point is GPU ECC error.

[0097] Table 2

[0098] Test Power W1 W2 Test Abnormal Point -- ECC

[0099] As shown in Table 3, Table 3 is the test result table of another GPU graphics card testing method provided in an embodiment of the present application. The test anomaly points obtained by the GPU graphics card to be tested in different test temperature environments are shown in Table 3. For example: M1 is the test fixed temperature, and the GPU graphics card to be tested runs without anomaly under this test fixed temperature. M2 is the test increased temperature, and the GPU graphics card to be tested runs anomalously under this test increased temperature, and the test anomaly point is GPU Lost error.

[0100] Table 3

[0101] Test Temperature M1 M2 Test Abnormal Point -- Lost

[0102] Further, as shown in Table 4, Table 4 is a test result table of another GPU graphics card test method provided in the embodiments of the present application. The overlapping test result table obtained by the GPU graphics card to be tested in the overlapping test temperature environment of temperature and power is shown in Table 4, and the overlapping test abnormal points are reflected in the overlapping test result table. The overlapping test abnormal points obtained in this overlapping test temperature environment can be GPU ECC errors and GPU Lost errors, or xid errors in the nvidia ecosystem stack.

[0103] Table 4

[0104] Test Power W1 W2 Test Temperature M1 M2 Test Abnormal Point -- ECC, Lost / Xid Error

[0105] Further, after the present application tests the GPU graphics card to be tested under increasing power to obtain the test abnormal points of the GPU graphics card to be tested under increasing power, and tests the GPU graphics card to be tested under increasing temperature to obtain the test abnormal points of the GPU graphics card to be tested under increasing temperature, the historical test table can also be updated according to the test abnormal points under increasing power and the test abnormal points under increasing temperature.

[0106] It can be understood that the test abnormal points under increasing power and the possible abnormal operation of the GPU graphics card to be tested under increasing temperature are not recorded in the historical test table. At this time, the power environment and the temperature environment, as well as the possible test abnormal points corresponding to each environment, can be recorded in the historical test table, so that when it is necessary to test a new GPU graphics card to be tested in the subsequent process, the environmental parameters can be directly called, reducing the test time.

[0107] In an implementable embodiment, since the simulation test program in the related art needs to test whether the GPU graphics card can run normally at the day level, that is, it takes several days to complete the test of a GPU graphics card using the simulation test program, and the time consumption is too long.

[0108] Therefore, in the embodiments of the present application, it is also proposed that multiple data subsets can be input into the GPU graphics card to be tested within a preset minute interval, multiple data target subsets can be input into the GPU graphics card to be tested within a preset minute interval, and the target test data subset can be encrypted and input into the GPU graphics card to be tested within a preset minute interval. The preset minute interval can be understood as completing the input of data within three minutes. The preset minute interval is not specifically limited here and can also be set according to the actual input duration required by the data. It should be noted that the preset minute intervals corresponding to each data input process in the present application can be the same or different.

[0109] As an example, multiple data subsets can be input into the GPU graphics card to be tested within two minutes to obtain power test results. Multiple data target subsets can also be input into the GPU graphics card to be tested within three minutes to obtain power test target results. Additionally, the target test data subsets can be encrypted and input into the GPU graphics card to be tested within four minutes to obtain temperature test results. In this way, the present application can achieve data input and complete testing at the minute level, improving the testing efficiency. Moreover, continuous adjustment of power or temperature can be achieved in a short time, realizing high-frequency fluctuation testing.

[0110] Further, as Figure 8 shown, Figure 8 is a schematic diagram of a test result table in a test method for a GPU graphics card provided by an embodiment of the present application. In Figure 8 it shows the test result table determined according to the test time granularity at different test powers and different test temperatures. The test time granularity can be understood as the time for inputting data and testing the GPU graphics card to be tested within a preset minute interval. For example: in the test power W1 environment or test temperature M1 environment corresponding to time T1, there are no test abnormal points; in the test power W2 environment or test temperature M2 environment corresponding to time T2, there are test abnormal points such as GPU ECC errors or GPU Lost errors. It should be noted that the test time granularity in the present application can be configured by itself. For example, continuous simulation of power and temperature fluctuations can be carried out within a two-hour time granularity to achieve hardware fault detection.

[0111] Next, in combination with Figure 9 it is further described the test process for test abnormal points under the adjustment of the test fixed power and test fixed temperature. As Figure 9 shown, Figure 9 is a full flow chart for testing a GPU graphics card in a test method for a GPU graphics card provided by an embodiment of the present application. In Figure 9 the test reduced batch can be obtained by reducing the adjustment of the test fixed batch, and the total set of test data is input into the GPU graphics card to be tested according to the test reduced batch within a preset minute interval, and the time point and power value of this input process are recorded. Since the batches of data input are different, the power values of the GPU graphics card to be tested are also different. At this time, it can be determined whether the GPU graphics card to be tested is operating abnormally at this power value. If the GPU graphics card to be tested is operating abnormally, the test abnormal points existing in the GPU graphics card in this power environment are recorded, and the historical test table is updated.

[0112] And if the GPU card to be tested does not run abnormally, the test batch can be increased and adjusted, obtaining an increased test batch. Then, within a preset minute interval, the total test data set is input into the GPU card to be tested according to the increased test batch, and the time point and power value of this input process are recorded. At this power value, it is determined whether the GPU card to be tested runs abnormally. If the GPU card to be tested still runs normally, this increase process is looped until the GPU card to be tested runs normally in each power environment.

[0113] Furthermore, the fixed test data volume can be increased and adjusted (that is, the fixed test data volume is adjusted to the maximum encrypted data volume that the system can accept), obtaining an increased test data volume. Then, within a preset minute interval, the data subset corresponding to the increased test data volume in the total test data set is encrypted and calculated and input into the GPU card to be tested, and the time point and temperature value of this input process are recorded. Since the data volume of data encryption and input is at the maximum data volume, the temperature value of the GPU card to be tested will also be at the maximum temperature. At this time, it can be determined whether the GPU card to be tested runs abnormally at this temperature value. If the GPU card to be tested runs abnormally, the test abnormal points existing in the GPU card to be tested in this temperature environment are recorded, and the historical test table is updated.

[0114] And if the GPU card to be tested does not run abnormally, the encryption calculation of the data set corresponding to the increased test data volume in the total test data set can be gradually stopped, that is, the increased test data volume is decreased and adjusted, so that the temperature value gradually decreases. Then, the GPU card to be tested is tested in the gradually decreasing temperature environment, and the time point and temperature value are recorded. If the GPU card to be tested runs normally in each temperature environment, the test is stopped and the test result is output. In some examples, the present application can record a graphics card repair table for the GPU card required for model training, that is, what is recorded in this graphics card repair table are all the GPU cards with abnormalities in the GPU cards required for this model training. As shown in Table 5, Table 5 is a graphics card repair table for a GPU card test method provided in an embodiment of the present application. In Table 5, the device IP corresponding to the GPU card with abnormalities in the GPU cards required for model training is shown, and the abnormal points existing in each device IP are recorded, such as IP1 has abnormal point 2 and IP2 has abnormal point 1, so as to repair the abnormal GPU cards subsequently.

[0115] Table 5

[0116] Device IP Abnormal Point 1 (e.g., ECC Error) Abnormal Point 2 (e.g., a certain xid error) ...... IP1 -- Yes ...... IP2 Yes -- ...... IP(N) ...... ...... ......

[0117] Further, after determining that the GPU graphics card required for model training has no abnormal operation in the test power environment or the test temperature environment, the present application can regularize the acceptance result table of the GPU graphics card to be tested in terms of hardware dimension, software dimension, network dimension, and power dimension according to the test results. As shown in Table 6, Table 6 is the acceptance result table of a GPU graphics card test method provided in an embodiment of the present application. In Table 6, the device IPs of the GPU graphics cards required for model training that have no abnormal operation are shown, and it is recorded that each GPU graphics card with no abnormal operation has no problems in terms of hardware dimension, software dimension, network dimension, and power dimension, so as to facilitate the viewing and acceptance of subsequent model training personnel.

[0118] Table 6

[0119] Device IP Hardware Software Network Power IP1 Yes Yes Yes Yes IP2 Yes Yes Yes Yes IP(N) Yes Yes Yes Yes

[0120] In summary, in the present application, the GPU graphics card to be tested can be tested in a power fluctuation environment and a temperature fluctuation environment, so that the GPU graphics card to be tested can be tested more comprehensively, thereby improving the test effect of the GPU graphics card. In addition, in the present application, the possible test abnormal points of the GPU graphics card to be tested in the power fluctuation environment and the temperature fluctuation environment can also be output, further improving the test accuracy of the GPU graphics card. In addition, compared with the traditional test time-consuming at the day level, the technical solution of the present application greatly reduces the test time-consuming, achieving the purpose of reducing costs and increasing efficiency.

[0121] Based on the GPU graphics card test method provided in the foregoing embodiments, the present application also correspondingly provides a GPU graphics card test device. The GPU graphics card test device provided in the embodiments of the present application will be specifically introduced below.

[0122] See Figure 10 , which is a schematic structural diagram of a GPU graphics card test device provided in an embodiment of the present application. As Figure 10 shown, the GPU graphics card test device specifically includes:

[0123] A GPU graphics card acquisition unit 1001, configured to acquire a GPU graphics card to be tested, a test fixed power, and a test fixed temperature;

[0124] A power and temperature adjustment unit 1002, configured to respectively increase and adjust the test fixed power and the test fixed temperature to obtain a test increased power and a test increased temperature;

[0125] The GPU graphics card test unit 1003 is used to test the GPU graphics card to be tested under the increased test power to obtain the test abnormal points of the GPU graphics card to be tested under the increased test power, and to test the GPU graphics card to be tested under the increased test temperature to obtain the test abnormal points of the GPU graphics card to be tested under the increased test temperature;

[0126] The test result table obtaining unit 1004 is used to obtain the test result table of the GPU graphics card to be tested that fails the power test and the test result table of the GPU graphics card to be tested that fails the temperature test according to the test abnormal points of the GPU graphics card to be tested under the increased test power and the test abnormal points of the GPU graphics card to be tested under the increased test temperature.

[0127] Optionally, the device further includes:

[0128] The total data set obtaining unit is used to obtain the total test data set, where the total test data set includes the total data set input to the GPU graphics card to be tested and running in the GPU graphics card to be tested;

[0129] The power and temperature adjustment unit 1002 includes:

[0130] The reduced batch obtaining unit is used to perform batch reduction adjustment on the fixed test batch to obtain a reduced test batch, where the number of times of the reduced test batch is less than the number of times of the fixed test batch, and the amount of data input to the GPU graphics card to be tested each time in the reduced test batch is greater than the amount of data input to the GPU graphics card to be tested each time in the fixed test batch;

[0131] The increased power determination unit is used to determine the increased test power according to the reduced test batch.

[0132] Optionally, the GPU graphics card test unit 1003 includes:

[0133] The power test result obtaining unit is used to divide the total test data set into multiple data subsets according to the reduced test batch, and input the multiple data subsets into the GPU graphics card to be tested to drive the GPU graphics card to run the multiple data subsets to obtain the power test result, where one reduced test batch corresponds to the division of one data subset;

[0134] The power test result analysis unit is used to analyze the power test result if the power test result is that the GPU graphics card to be tested runs abnormally, and obtain the test abnormal points of the GPU graphics card to be tested under the increased test power.

[0135] Optionally, the device further includes:

[0136] An increased batch obtaining unit, configured to, if the power test result indicates that the GPU to be tested is not operating abnormally, perform batch increase adjustment on the reduced test batch to obtain an increased test batch, where the number of times of the increased test batch is less than the number of times of the fixed test batch and greater than the number of times of the reduced test batch;

[0137] A test target result obtaining unit, configured to divide the total test data set into multiple data target subsets according to the increased test batch, and input the multiple data target subsets into the GPU to be tested to drive the GPU to be tested to run the multiple data target subsets, so as to obtain a power test target result, where one increased test batch corresponds to the division of one data target subset;

[0138] A test target result parsing unit, configured to, if the power test target result indicates that the GPU to be tested is abnormal, parse the power test target result to obtain the test abnormal point of the GPU to be tested at the reduced test power.

[0139] Optionally, the power and temperature adjustment unit 1002 includes:

[0140] An increased data volume obtaining unit, configured to perform data volume increase adjustment on the fixed test data volume to obtain an increased test data volume, where the increased test data volume is greater than the fixed test data volume;

[0141] An increased temperature determining unit, configured to determine an increased test temperature according to the increased test data volume.

[0142] Optionally, the GPU test unit 1003 is specifically configured to:

[0143] A test data subset obtaining unit, configured to process the total test data set according to the increased test data volume to obtain a target test data subset corresponding to the increased test data volume in the total test data set;

[0144] A temperature test result obtaining unit, configured to encrypt and input the target test data subset into the GPU to be tested to drive the GPU to be tested to run the target test data subset, so as to obtain a temperature test result;

[0145] A temperature test result parsing unit, configured to, if the temperature test result indicates that the GPU to be tested is abnormal, parse the temperature test result to obtain the test abnormal point of the GPU to be tested at the increased test temperature.

[0146] Optionally, the power test result obtaining unit is specifically configured to: input the multiple data subsets into the GPU to be tested within a preset minute interval;

[0147] The test target result obtaining unit is specifically configured to: input the multiple data target subsets into the GPU to be tested within the preset minute interval;

[0148] The temperature test result obtaining unit is specifically configured to: encrypt and input the target test data subset into the GPU to be tested within the preset minute interval.

[0149] Optionally, the apparatus further includes:

[0150] The overlapping test anomaly point obtaining unit is configured to simultaneously test the GPU to be tested under the increased test power and under the increased test temperature, and obtain the overlapping test anomaly points of the GPU to be tested in the overlapping test environment of the increased test power and the increased test temperature;

[0151] The overlapping test result table obtaining unit is configured to obtain the overlapping test result table of the GPU to be tested in the overlapping test environment according to the overlapping test anomaly points.

[0152] Optionally, the apparatus further includes:

[0153] The historical test table obtaining unit is configured to obtain a historical test table, which is constructed by the test power and test temperature corresponding to the historical test anomaly points in the historical test process, and the historical test table includes a test fixed power and a test fixed temperature;

[0154] The apparatus further includes:

[0155] The historical test table updating unit is configured to update the historical test table according to the test anomaly points under the increased test power and the test anomaly points under the increased test temperature.

[0156] An embodiment of the present application provides a computer device, and the computer device may be a server. Figure 11FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 922 (for example, one or more processors) and a memory 932, and one or more storage media 930 (for example, one or more mass storage devices) for storing application programs 942 or data 944. Among them, the memory 932 and the storage media 930 may be transient storage or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 922 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 900.

[0157] The server 900 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0158] Among them, the CPU 922 is used to execute the following steps:

[0159] Obtain the GPU graphics card to be tested, the test fixed power, and the test fixed temperature;

[0160] Increase and adjust the test fixed power and the test fixed temperature respectively to obtain a test increased power and a test increased temperature;

[0161] Test the GPU graphics card to be tested at the test increased power to obtain the test abnormal points of the GPU graphics card to be tested at the test increased power, and test the GPU graphics card to be tested at the test increased temperature to obtain the test abnormal points of the GPU graphics card to be tested at the test increased temperature;

[0162] According to the test abnormal points of the GPU graphics card to be tested at the test increased power and the test abnormal points of the GPU graphics card to be tested at the test increased temperature, obtain the test result table of the GPU graphics card to be tested that fails the power test and the test result table of the GPU graphics card to be tested that fails the temperature test.

[0163] An embodiment of the present application further provides another computer device, and this computer device may be a terminal device. As Figure 12As shown, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. Taking the terminal device as a mobile phone as an example:

[0164] Figure 12 What is shown is a block diagram of a part of the structure of the mobile phone provided by the embodiments of the present application. Refer to Figure 12 , the mobile phone includes: a radio frequency (full English name: Radio Frequency, English abbreviation: RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (full English name: wirelessfidelity, English abbreviation: WiFi) module 1070, a processor 1080, and a power supply 1090 and other components. Those skilled in the art can understand that Figure 12 the structure of the mobile phone shown in

[0165] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 12 The following will specifically introduce each component of the mobile phone:

[0166] The RF circuit 1010 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1080 for processing; in addition, the designed uplink data is sent to the base station. Usually, the RF circuit 1010 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (full English name: LowNoise Amplifier, English abbreviation: LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (full English name: Global System of Mobile communication, English abbreviation: GSM), General Packet Radio Service (full English name: General Packet Radio Service, GPRS), CodeDivision Multiple Access (full English name: CodeDivision Multiple Access, English abbreviation: CDMA), Wideband CodeDivision Multiple Access (full English name: Wideband CodeDivision Multiple Access, English abbreviation: WCDMA), Long TermEvolution (full English name: Long TermEvolution, English abbreviation: LTE), email, Short Messaging Service (full English name: Short Messaging Service, SMS), etc.

[0167] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0168] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1031), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute the commands sent by the processor 1080. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1031. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0169] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of, for example, a liquid crystal display (LCD for short), an organic light-emitting diode (OLED for short), etc. Further, a touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 12 the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and output functions of the mobile phone, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0170] The mobile phone may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0171] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and then converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent to, for example, another mobile phone through the RF circuit 1010, or the audio data is output to the memory 1020 for further processing.

[0172] WiFi belongs to short - range wireless transmission technology. Through the WiFi module 1070, a mobile phone can help users send and receive emails, browse the web, and access streaming media, etc. It provides users with wireless broadband Internet access. Although Figure 12 the WiFi module 1070 is shown, it can be understood that it is not an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.

[0173] The processor 1080 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1020, and by calling data stored in the memory 1020, it performs various functions of the mobile phone and processes data, thereby collecting overall data and information of the mobile phone. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1080 either.

[0174] The mobile phone also includes a power source 1090 (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the processor 1080 through a power management system, thereby realizing functions such as management of charging, discharging, and power consumption management through the power management system.

[0175] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0176] In the embodiment of this application, the processor 1080 included in the mobile phone further has the following functions:

[0177] Obtain the GPU graphics card to be tested, the test fixed power, and the test fixed temperature;

[0178] Increase and adjust the test fixed power and the test fixed temperature respectively to obtain the test increased power and the test increased temperature;

[0179] Test the GPU graphics card to be tested under the test increased power to obtain the test abnormal points of the GPU graphics card to be tested under the test increased power, and test the GPU graphics card to be tested under the test increased temperature to obtain the test abnormal points of the GPU graphics card to be tested under the test increased temperature;

[0180] Based on the test anomaly points of the GPU graphics card to be tested under the increasing power in the test and the test anomaly points of the GPU graphics card to be tested under the increasing temperature in the test, a test result table for the GPU graphics card to be tested failing the power test and a test result table for the GPU graphics card to be tested failing the temperature test are obtained.

[0181] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, it causes the computer device to execute any one of the implementation manners of the test method for a GPU graphics card described in the foregoing various embodiments.

[0182] An embodiment of the present application also provides a computer program product including a computer program. When it runs on a computer device, it causes the computer device to execute any one of the implementation manners of the test method for a GPU graphics card described in the foregoing various embodiments.

[0183] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0184] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the system is only a logical function division. In actual implementation, there may be other division methods. For example, multiple systems can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0185] The system described as a separated component may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0186] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0187] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store computer programs, such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs.

[0188] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of this module or unit.

[0189] The above embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A test method for a GPU graphics card, characterized in that, Including: Obtain the GPU graphics card to be tested, the test fixed power, and the test fixed temperature; Perform an increasing adjustment on the test fixed power and the test fixed temperature respectively to obtain the test increasing power and the test increasing temperature; Test the GPU graphics card to be tested under the test increasing power to obtain the test abnormal points of the GPU graphics card to be tested under the test increasing power, and test the GPU graphics card to be tested under the test increasing temperature to obtain the test abnormal points of the GPU graphics card to be tested under the test increasing temperature; According to the test abnormal points of the GPU graphics card to be tested under the test increasing power and the test abnormal points of the GPU graphics card to be tested under the test increasing temperature, obtain the test result table of the GPU graphics card to be tested that fails the power test and the test result table of the GPU graphics card to be tested that fails the temperature test.

2. The method according to claim 1, wherein Before obtaining the GPU graphics card to be tested, the test fixed power, and the test fixed temperature, it further includes: Obtain the total test data set, where the total test data set includes the total data set input into the GPU graphics card to be tested and running in the GPU graphics card to be tested; The test fixed power includes the test power generated by inputting the total test data set into the GPU graphics card to be tested in a test fixed batch; performing an increasing adjustment on the test fixed power to obtain the test increasing power includes: Perform a batch decreasing adjustment on the test fixed batch to obtain a test decreasing batch, where the number of times of the test decreasing batch is less than the number of times of the test fixed batch, and the amount of data input into the GPU graphics card to be tested each time in the test decreasing batch is greater than the amount of data input into the GPU graphics card to be tested each time in the test fixed batch; Determine the test increasing power according to the test decreasing batch.

3. The method according to claim 2, wherein The testing the GPU graphics card to be tested under the test increasing power to obtain the test abnormal points of the GPU graphics card to be tested under the test increasing power includes: Divide the total test data set into multiple data subsets according to the test decreasing batch, and input the multiple data subsets into the GPU graphics card to be tested to drive the GPU graphics card to run the multiple data subsets to obtain the power test result, where one test decreasing batch corresponds to the division of one data subset; If the power test result is that the GPU graphics card to be tested runs abnormally, analyze the power test result to obtain the test abnormal points of the GPU graphics card to be tested under the test increasing power.

4. The method according to claim 3, wherein After driving the GPU graphics card to be tested to run the multiple data subsets to obtain the power test result, it further includes: If the power test result is that the GPU graphics card to be tested does not run abnormally, perform a batch increasing adjustment on the test decreasing batch to obtain a test increasing batch, where the number of times of the test increasing batch is less than the number of times of the test fixed batch, and the number of times of the test increasing batch is greater than the number of times of the test decreasing batch; Divide the total set of test data into multiple data target subsets according to the increasing test batches, and input the multiple data target subsets into the GPU card to be tested, so as to drive the GPU card to be tested to run the multiple data target subsets and obtain the power test target result, where one increasing test batch corresponds to the division of one data target subset; If the power test target result indicates that the GPU card to be tested is abnormal, analyze the power test target result to obtain the test abnormal points of the GPU card to be tested at the reduced test power.

5. The method according to claim 2, wherein The fixed test temperature includes the temperature generated by encrypting and inputting a partial data subset in the total set of test data into the GPU card to be tested with a fixed test data volume; Increasingly adjust the fixed test temperature to obtain an increased test temperature, including: Increasingly adjust the fixed test data volume to obtain an increased test data volume, where the increased test data volume is greater than the fixed test data volume; Determine the increased test temperature according to the increased test data volume.

6. The method according to claim 5, wherein Testing the GPU card to be tested at the increased test temperature to obtain the test abnormal points of the GPU card to be tested at the increased test temperature, including: Process the total set of test data according to the increased test data volume to obtain the target test data subset corresponding to the increased test data volume in the total set of test data; Encrypt and input the target test data subset into the GPU card to be tested to drive the GPU card to be tested to run the target test data subset and obtain the temperature test result; If the temperature test result indicates that the GPU card to be tested is abnormal, analyze the temperature test result to obtain the test abnormal points of the GPU card to be tested at the increased test temperature.

7. The method according to any one of claims 3-6, characterized in that, Inputting the multiple data subsets into the GPU card to be tested includes: inputting the multiple data subsets into the GPU card to be tested within a preset minute interval; Inputting the multiple data target subsets into the GPU card to be tested includes: inputting the multiple data target subsets into the GPU card to be tested within the preset minute interval; Encrypting and inputting the target test data subset into the GPU card to be tested includes: encrypting and inputting the target test data subset into the GPU card to be tested within the preset minute interval.

8. The method according to claim 1, characterized in that After increasingly adjusting the fixed test power and the fixed test temperature respectively to obtain the increased test power and the increased test temperature, it further includes: Testing the GPU card to be tested simultaneously at the increased test power and at the increased test temperature to obtain the overlapping test abnormal points of the GPU card to be tested in the overlapping test environment of the increased test power and the increased test temperature; Obtain the overlapping test result table of the GPU card to be tested in the overlapping test environment according to the overlapping test abnormal points.

9. The method according to claim 1, wherein Before obtaining the GPU graphics card to be tested, the test fixed power, and the test fixed temperature, it further includes: Obtaining a historical test table, which is constructed by the test power and test temperature corresponding to the historical test abnormal points in the historical test process, and the historical test table includes the test fixed power and the test fixed temperature; After testing the GPU graphics card to be tested at the test increased power to obtain the test abnormal points of the GPU graphics card to be tested at the test increased power, and testing the GPU graphics card to be tested at the test increased temperature to obtain the test abnormal points of the GPU graphics card to be tested at the test increased temperature, it further includes: Updating the historical test table according to the test abnormal points at the test increased power and the test abnormal points at the test increased temperature.

10. A test device for a GPU graphics card, characterized in that, It includes: A GPU graphics card acquisition unit, configured to acquire a GPU graphics card to be tested, a test fixed power, and a test fixed temperature; A power and temperature adjustment unit, configured to respectively perform an increase adjustment on the test fixed power and the test fixed temperature to obtain a test increased power and a test increased temperature; A GPU graphics card test unit, configured to test the GPU graphics card to be tested at the test increased power to obtain the test abnormal points of the GPU graphics card to be tested at the test increased power, and test the GPU graphics card to be tested at the test increased temperature to obtain the test abnormal points of the GPU graphics card to be tested at the test increased temperature; A test result table acquisition unit, configured to obtain a test result table in which the GPU graphics card to be tested fails the power test and a test result table in which the GPU graphics card to be tested fails the temperature test according to the test abnormal points of the GPU graphics card to be tested at the test increased power and the test abnormal points of the GPU graphics card to be tested at the test increased temperature.

11. A computer device, characterized in that, The device includes a processor and a memory: The memory is configured to store a computer program and transmit the computer program to the processor; The processor is configured to execute the steps of the test method of the GPU graphics card according to any one of claims 1 to 9 based on the instructions in the computer program.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store a computer program, and when the computer program is executed by a computer device, it implements the steps of the test method of the GPU graphics card according to any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a computer device, it implements the steps of the test method of the GPU graphics card according to any one of claims 1 to 9.