Task transfer method and device based on double-GPU architecture and medium

By monitoring and predicting the running status of the GPU in real time, building task transfer coefficients and strategies, and optimizing task allocation using genetic algorithms, solving the problem of insufficient flexibility and dynamics in resource allocation and task scheduling of the dual GPU architecture, achieving more efficient and stable task execution.

CN120011069APending Publication Date: 2025-05-16SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102668.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional dual GPU architecture lacks flexibility and dynamicity in computing resource allocation and task scheduling, resulting in the inability to maximize the overall performance of the system, which may cause idle or overload resources, affecting the efficiency and quality of task execution.

Method used

By monitoring the running status of each GPU in real time, building transmission task transfer coefficients, determining reference task transfer strategies, using linear regression models for load prediction, and optimizing task allocation through genetic algorithms to adjust task transfer strategies to achieve more effective task allocation and load balancing.

Benefits of technology

It improves the overall performance and stability of the system, realizes more efficient and intelligent task scheduling, avoids idle or overload of resources, and improves the efficiency and quality of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011069A_ABST
    Figure CN120011069A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a task transfer method and device based on a double-GPU architecture and a medium, belongs to the technical field of data communication, and solves the problem that the task execution efficiency and quality are affected due to the fact that a traditional double-GPU architecture lacks flexibility and dynamics in computing resource allocation and task scheduling. Comprising the steps of obtaining current operation data and historical operation data of a first GPU and a second GPU based on a preset interval duration; constructing a transmission task transfer coefficient according to the historical operation data of the first GPU and the second GPU, and determining a reference task transfer strategy based on the transmission task transfer coefficient and the current operation data of the first GPU and the second GPU; load prediction is carried out through a preset linear regression model, the current operation data and the historical operation data; and based on the load prediction result, optimizing task allocation through a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data communication technology, and in particular to a task transfer method, device and medium based on a dual GPU architecture. Background Art

[0002] In today's era of rapid development of information technology, big data, cloud computing, metaverse and higher-level artificial intelligence applications are booming at an unprecedented speed, profoundly changing our lifestyle, work mode and even the operating logic of the entire society. The rise of these cutting-edge technologies has not only greatly enriched the dimensions and complexity of data processing, but also put forward unprecedented high requirements on the real-time processing capabilities and data throughput of computing platforms. As the core infrastructure supporting these high-tech applications, computing platforms must be able to process massive amounts of data efficiently and accurately to ensure instant feedback of information and accurate decision-making.

[0003] Traditionally, in order to cope with the growing computing needs, dual GPU (graphics processing unit) architecture has been widely used in the field of high-performance computing. This architecture has achieved a significant improvement in computing power to a certain extent by integrating two independent GPU units, providing more powerful computing support for complex data analysis and model training tasks. However, with the continuous advancement of technology and the increasing diversification of application scenarios, especially in the face of extremely complex or large-scale parallel processing tasks, the collaboration between GPUs is often limited by fixed hardware connections and established software protocols, resulting in a lack of sufficient flexibility and dynamism in computing resource allocation and task scheduling. This not only limits the maximization of the overall performance of the system, but may also cause imbalances such as idle or overloaded resources, thereby affecting the efficiency and quality of task execution. Summary of the invention

[0004] The embodiments of the present application provide a task transfer method, device and medium based on a dual-GPU architecture, which are used to solve the following technical problems: the traditional dual-GPU architecture lacks flexibility and dynamics in computing resource allocation and task scheduling, which not only limits the maximization of the overall performance of the system, but may also cause imbalances such as resource idleness or overload, thereby affecting the efficiency and quality of task execution.

[0005] The present application embodiment adopts the following technical solutions:

[0006] The embodiment of the present application provides a task transfer method based on a dual GPU architecture. The method includes: based on a preset interval duration, obtaining the current operation data of the first GPU and the current operation data of the second GPU, and obtaining the historical operation data of the GPU and the historical operation data of the second GPU; constructing a transmission task transfer coefficient according to the historical operation data of the first GPU and the historical operation data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, the current operation data of the first GPU and the current operation data of the second GPU; performing load prediction through a preset linear regression model, current operation data and historical operation data; optimizing task allocation through a genetic algorithm based on the load prediction result, so as to adjust the reference task transfer strategy and obtain an adjusted task transfer strategy.

[0007] The embodiment of the present application monitors the operating status of each GPU in real time and feeds back data in a timely manner, so that the system can make reasonable task allocation decisions based on the current load situation. The linear regression model is used to predict future task requirements, so that the system can predict the load changes of tasks in advance, so that it is more efficient and intelligent in task scheduling. By adjusting and improving the reference task transfer strategy, the adjusted strategy is more in line with the current load status and performance requirements, and can more effectively balance the task allocation between the two GPUs, thereby improving the overall performance and stability of the system.

[0008] In one implementation of the present application, a reference task transfer strategy is determined based on a transmission task transfer coefficient, current operating data of the first GPU, and current operating data of the second GPU, specifically including: determining a current load corresponding to the first GPU by the ratio of current power consumption in the current operating data of the first GPU to the maximum power consumption corresponding to the first GPU; determining a current load corresponding to the second GPU by the ratio of current power consumption in the current operating data of the second GPU to the maximum power consumption corresponding to the second GPU; obtaining a load balancing target based on the current load corresponding to the first GPU and the current load corresponding to the second GPU; determining a reference task transfer strategy based on the current load corresponding to the first GPU, the current load corresponding to the second GPU, and the load balancing target.

[0009] In one implementation of the present application, a reference task transfer strategy is determined based on a current load corresponding to the first GPU, a current load corresponding to the second GPU, and a load balancing target, specifically including: comparing the current load corresponding to the first GPU and the current load corresponding to the second GPU with the load balancing target respectively; based on the comparison result, determining a reference load corresponding to the GPU that is greater than the load balancing target; based on the difference between the reference load and the balancing target, and the transmission task transfer coefficient, determining a reference task transfer strategy.

[0010] In one implementation of the present application, a transmission task transfer coefficient is constructed based on the historical operation data of the first GPU and the historical operation data of the second GPU, specifically including: based on a reference load, determining reference historical operation data from the historical operation data of the first GPU and the historical operation data of the second GPU; determining multiple historical transfer coefficients corresponding to the reference load in the reference historical operation data; based on the reference load and the average of the current operation data of the first GPU and the current data of the second GPU, screening multiple historical transfer coefficients, and sorting the screened historical transfer coefficients based on the difference between the average and the average; based on the sorting order, weighting the multiple historical transfer coefficients, and obtaining the transmission task transfer coefficient through the result after weighting processing.

[0011] In one implementation of the present application, based on the difference between the reference load and the balancing target, and the transmission task transfer coefficient, a reference task transfer strategy is determined, specifically including:

[0012] Function-based:

[0013] T transfer = k·(L u -L avg );

[0014] Determine the load to be transferred corresponding to the reference load, so as to determine the reference task transfer strategy based on the load to be transferred; wherein, T transfer is the load to be transferred; k is the transmission task transfer coefficient; L u is the reference load; L avg For the goal of balance.

[0015] In one implementation of the present application, load prediction is performed by presetting a linear regression model, current operation data, and historical operation data, specifically including: obtaining historical load data of a GPU corresponding to a reference load; based on the historical load data and the linear regression model:

[0016] L pred =β 0 +β 1 ·T+β 2 H;

[0017] Determine the predicted load of the GPU corresponding to the reference load; where L pred is the predicted load; T is the time variable; H is the historical load data; β 0 is the first parameter; β 1 is the second parameter; β 2 Is the third parameter.

[0018] In one implementation of the present application, based on the load prediction result, the task allocation is optimized by a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy, which specifically includes: optimizing the task allocation by a genetic algorithm and setting the fitness function to:

[0019]

[0020] By minimizing the fitness function, the reference task transfer strategy is adjusted to obtain the adjusted task transfer strategy.

[0021] In one implementation of the present application, current operating data of a first GPU and current operating data of a second GPU are acquired based on a preset interval duration, specifically including: comparing the acquired current operating data of the first GPU and current operating data of the second GPU with a preset frequency table; wherein the preset frequency table includes a plurality of operating data ranges, and also includes data acquisition frequencies corresponding to the plurality of operating data ranges; determining corresponding data acquisition frequencies based on comparison results of the current operating data of the first GPU and the current operating data of the second GPU with the plurality of operating data ranges respectively; determining a new preset interval duration based on the data acquisition frequency, so as to acquire the current operating data of the first GPU and the current operating data of the second GPU based on the new preset interval duration.

[0022] An embodiment of the present application provides a task transfer method and device based on a dual-GPU architecture, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: based on a preset interval duration, obtain current operating data of a first GPU and current operating data of a second GPU, and obtain historical operating data of a GPU and historical operating data of a second GPU; construct a transmission task transfer coefficient according to the historical operating data of the first GPU and the historical operating data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, the current operating data of the first GPU and the current operating data of the second GPU; perform load prediction by means of a preset linear regression model, current operating data and historical operating data; based on the load prediction result, optimize task allocation by means of a genetic algorithm, so as to adjust the reference task transfer strategy and obtain an adjusted task transfer strategy.

[0023] A non-volatile computer storage medium provided in an embodiment of the present application stores computer executable instructions, wherein the computer executable instructions are configured to: based on a preset interval duration, obtain current operating data of a first GPU and current operating data of a second GPU, and obtain historical operating data of a GPU and historical operating data of a second GPU; construct a transmission task transfer coefficient according to the historical operating data of the first GPU and the historical operating data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, the current operating data of the first GPU and the current operating data of the second GPU; perform load prediction by using a preset linear regression model, current operating data and historical operating data; and optimize task allocation by using a genetic algorithm based on the load prediction result, so as to adjust the reference task transfer strategy and obtain an adjusted task transfer strategy.

[0024] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: The embodiments of the present application monitor the operating status of each GPU in real time and feedback data in a timely manner, so that the system can make reasonable task allocation decisions based on the current load conditions. The linear regression model is used to predict future task requirements, so that the system can predict the load changes of tasks in advance, so that it is more efficient and intelligent when scheduling tasks. By adjusting and improving the reference task transfer strategy, the adjusted strategy is more in line with the current load conditions and performance requirements, and can more effectively balance the task allocation between the two GPUs, thereby improving the overall performance and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:

[0026] Figure 1 A flowchart of a task transfer method based on a dual GPU architecture provided in an embodiment of the present application;

[0027] Figure 2 A high-performance computing system with a novel dual-GPU architecture provided in an embodiment of the present application;

[0028] Figure 3 A schematic diagram of the structure of a task transfer device based on a dual GPU architecture provided in an embodiment of the present application.

[0029] Reference numerals:

[0030] 200: Task transfer device based on dual GPU architecture, 201: Processor, 202: Memory. DETAILED DESCRIPTION

[0031] The embodiments of the present application provide a task transfer method, device and medium based on a dual GPU architecture.

[0032] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0033] The technical solution proposed in the embodiment of the present invention is described in detail below with reference to the accompanying drawings.

[0034] Figure 1 A flowchart of a task transfer method based on a dual GPU architecture is provided in an embodiment of the present application, such as Figure 1 As shown, the task transfer method based on dual GPU architecture includes the following steps:

[0035] S101 . Based on a preset interval duration, current operation data of a first GPU and current operation data of a second GPU are acquired, and historical operation data of the first GPU and historical operation data of the second GPU are acquired.

[0036] In one embodiment of the present application, the acquired current operating data of the first GPU and the current operating data of the second GPU are compared with a preset frequency table; wherein the preset frequency table includes a plurality of operating data ranges, and also includes data acquisition frequencies corresponding to the plurality of operating data ranges. Based on the comparison results of the current operating data of the first GPU and the current operating data of the second GPU with the plurality of operating data ranges, the corresponding data acquisition frequency is determined. Based on the data acquisition frequency, a new preset interval duration is determined, so as to acquire the current operating data of the first GPU and the current operating data of the second GPU based on the new preset interval duration.

[0037] Specifically, the preset frequency table in the embodiment of the present application is a table containing multiple operating data ranges and their corresponding data acquisition frequencies, each operating data range represents a specific GPU operating state or performance range. The corresponding data acquisition frequency indicates how often the GPU operating data should be acquired within the operating data range.

[0038] Furthermore, when the current operating data of the first GPU and the second GPU are obtained, these data are compared with the operating data range in the preset frequency table, and the operating data range of each GPU is determined according to the comparison result. According to the operating data range of each GPU, the corresponding data acquisition frequency is searched in the preset frequency table, and the interval duration for subsequent adjustment of data acquisition is determined by the data acquisition frequency. Based on the determined data acquisition frequency, a new preset interval duration is calculated, and the current operating data of the first GPU and the second GPU are acquired according to the new preset interval duration at the new frequency to ensure that when the GPU performance changes, its operating data can be captured more accurately, so as to make more reasonable decisions.

[0039] And, historical operation data of the first GPU and historical operation data of the second GPU are obtained.

[0040] Among them, the operating data in the embodiment of the present application at least includes information such as GPU utilization, temperature, power consumption, etc.

[0041] S102: construct a transmission task transfer coefficient according to historical operation data of the first GPU and historical operation data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, current operation data of the first GPU and current operation data of the second GPU.

[0042] In one embodiment of the present application, the current load corresponding to the first GPU is determined by the ratio between the current power consumption in the current operation data of the first GPU and the maximum power consumption corresponding to the first GPU. The current load corresponding to the second GPU is determined by the ratio between the current power consumption in the current operation data of the second GPU and the maximum power consumption corresponding to the second GPU. A load balancing target is obtained based on the current load corresponding to the first GPU and the current load corresponding to the second GPU. A reference task transfer strategy is determined based on the current load corresponding to the first GPU, the current load corresponding to the second GPU, and the load balancing target.

[0043] Specifically, the system can collect the GPU's operating status every 100 milliseconds, including utilization, temperature, power consumption, etc. The current load is calculated using the following formula:

[0044]

[0045] Among them, P i is the current power consumption of the GPU, C i is the maximum power consumption of the GPU, i is the serial number of the GPU, and in the embodiment of the present application, it can be the first GPU or the second GPU.

[0046] Furthermore, the target formula for load balancing is set as:

[0047]

[0048] Where L 1 Indicates the first GPU; L 2 Indicates the current load of the second GPU.

[0049] In one embodiment of the present application, the current load corresponding to the first GPU and the current load corresponding to the second GPU are compared with the load balancing target respectively. Based on the comparison result, a reference load corresponding to the GPU greater than the load balancing target is determined. Based on the difference between the reference load and the balancing target and the transmission task transfer coefficient, a reference task transfer strategy is determined.

[0050] Specifically, the current load value of the first GPU is obtained, and this value is compared with the preset load balancing target. If the current load of the first GPU is greater than the load balancing target, the first GPU may be overloaded. Similarly, the current load value of the second GPU is obtained, and this value is compared with the load balancing target. According to the comparison result, it is determined whether the second GPU is overloaded. In the comparison result, the GPU whose current load is greater than the load balancing target is found. This GPU is overloaded and needs to transfer some tasks to reduce the load. For these overloaded GPUs, their current load values ​​are the reference loads.

[0051] Further, for each overloaded GPU, the difference between its reference load and the load balancing target is calculated, and the difference represents the amount of tasks that need to be transferred or the degree of load reduction. The transmission task transfer coefficient in the embodiment of the present application is a parameter for adjusting the amount of task transfer. Based on the calculated difference and the transmission task transfer coefficient, the number or proportion of tasks that need to be transferred is determined, and appropriate tasks are selected and transferred to GPUs whose current load is lower than the load balancing target.

[0052] For example, when L 1 >L avg When the system transfers part of the tasks to the second GPU, and vice versa.

[0053] In one embodiment of the present application, based on the reference load, reference historical operation data is determined from the historical operation data of the first GPU and the historical operation data of the second GPU. In the reference historical operation data, multiple historical transfer coefficients corresponding to the reference load are determined. Based on the reference load and the mean of the current operation data of the first GPU and the current data of the second GPU, multiple historical transfer coefficients are screened, and the screened historical transfer coefficients are sorted based on the difference between the mean and the reference load. Based on the sorting order, multiple historical transfer coefficients are weighted, and the transfer task transfer coefficient is obtained through the result of weight processing.

[0054] Specifically, in the reference historical operation data, for each historical load point close to the reference load, the transmission task transfer coefficient used at that time is found. These coefficients constitute a set of historical transfer coefficients. The mean of the current operation data of the first GPU and the second GPU is calculated. This mean reflects the current overall load state of the system. The means corresponding to these historical transfer coefficients are compared with the current data mean, and the historical transfer coefficient close to the current mean is determined for further analysis.

[0055] Furthermore, for each historical transfer coefficient selected, the difference between the coefficient and the current data mean is calculated. According to these differences, the historical transfer coefficients are sorted, and the sorting method can be that the coefficient with the smallest difference (i.e., the coefficient closest to the current mean) is placed in front. Based on the sorting order, multiple historical transfer coefficients are weighted, and the coefficients with higher rankings are assigned higher weights. The final transmission task transfer coefficient is calculated by weighted average.

[0056] In one embodiment of the present application, based on the function:

[0057] T transfer = k·(L u -L avg );

[0058] The load to be transferred corresponding to the reference load is determined, so as to determine the reference task transfer strategy based on the load to be transferred. transfer is the load to be transferred; k is the transmission task transfer coefficient; L u is the reference load; L avg For the goal of balance.

[0059] S103, performing load forecasting by using a preset linear regression model, current operation data, and historical operation data.

[0060] In one embodiment of the present application, historical load data of a GPU corresponding to a reference load is obtained, and based on the historical load data and a linear regression model:

[0061] L pred =β 0 +β 1 T+β 2 H;

[0062] Determine the predicted GPU load corresponding to the reference load. pred is the predicted load; T is the time variable; H is the historical load data; β 0 is the first parameter; β 1 is the second parameter; β 2 Is the third parameter.

[0063] The embodiment of the present application combines machine learning technology and uses a linear regression model to predict future task requirements, so that the system can predict task load changes in advance, thereby being more efficient and intelligent in task scheduling.

[0064] S104. Based on the load prediction result, optimize task allocation by using a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy.

[0065] In one embodiment of the present application, the task scheduling is adjusted according to the prediction results to ensure the load balance of the GPU. The task allocation is optimized by genetic algorithm, and the fitness function is set as:

[0066]

[0067] By minimizing the fitness function, the reference task transfer strategy is adjusted to obtain the adjusted task transfer strategy.

[0068] The embodiment of the present application introduces a genetic algorithm as an optimization tool for task scheduling, which can find the optimal solution among multiple task scheduling schemes, minimize load imbalance, and thus improve overall performance.

[0069] The algorithm in the embodiment of the present application is implemented by the computing unit in the FPGA, and the task allocation strategy is updated regularly. The embodiment of the present application integrates an FPGA module on the system motherboard, which is specifically used to implement dynamic load balancing and intelligent scheduling algorithms. The parallel processing capability of the FPGA enables the algorithm to run in real time, providing faster response time and higher processing efficiency. The circuit design of the FPGA module is implemented in the Verilog language, which can be customized according to specific needs, providing convenience for future algorithm upgrades and function expansions.

[0070] Figure 2 A high-performance computing system with a new dual-GPU architecture is provided in the embodiments of the present application, such as Figure 2 As shown, the high-performance computing system with a new dual-GPU architecture includes GPUs (a first GPU and a second GPU), a high-speed bus interconnection matrix, a dynamic load balancing algorithm module, an intelligent scheduling module, a low-speed bus interconnection matrix, and a power management module.

[0071] 1. Hardware architecture

[0072] (1) GPU configuration: The system includes two high-performance GPUs (e.g., NVIDIA RTX 3090), which are connected via a high-speed data transmission interface that supports the PCIe 5.0 standard and has a theoretical bandwidth of 32GT / s.

[0073] (2) FPGA module: An FPGA module is integrated on the system motherboard to implement the dynamic load balancing algorithm and intelligent scheduling system. It is designed in Verilog language and contains multiple parallel processing units to improve data processing capabilities.

[0074] (3) Power management: Design an intelligent power management module to dynamically adjust power consumption and ensure a balance between high performance and energy efficiency.

[0075] The high-performance computing system with a novel dual-GPU architecture in the embodiment of the present application also uses a high-bandwidth data transmission channel:

[0076] Multi-channel parallel transmission technology: The embodiment of the present application adopts a high-speed data transmission interface with 8 data channels, which significantly improves the bandwidth of data transmission to 50GB / s through parallel transmission technology. This design reduces the delay of data transmission, improves the efficiency of data exchange between GPUs, and ensures that computing tasks can be carried out quickly and efficiently.

[0077] Customized data transmission protocol: A new data transmission protocol is introduced, which supports confirmation response mechanism to ensure the reliability and integrity of data packets during transmission, and effectively cope with the data transmission challenges under high load conditions.

[0078] High-bandwidth data transmission channels:

[0079] Using multi-channel parallel transmission technology, a high-speed data transmission interface with 8 data channels is designed. The bandwidth of each channel is 6.25GB / s, and the total bandwidth can reach 50GB / s.

[0080] The data transmission circuit uses high-frequency differential signal transmission technology to reduce signal interference and delay. The data input port is connected to multiple data channels through a multiplexer (MUX), and the data channels are connected to the target GPU through a demultiplexer (DEMUX). A FIFO (first-in-first-out) queue is introduced to buffer data traffic and ensure the order and integrity of data.

[0081] Figure 3 A schematic diagram of a task transfer device based on a dual GPU architecture provided in an embodiment of the present application. Figure 3As shown, a task transfer device 200 based on a dual-GPU architecture includes: at least one processor 201; and a memory 202 in communication with the at least one processor 201; wherein the memory 202 stores instructions that can be executed by the at least one processor 201, and the instructions are executed by the at least one processor 201 so that the at least one processor 201 can: based on a preset interval duration, obtain the current operating data of the first GPU and the current operating data of the second GPU, and obtain the historical operating data of the GPU and the historical operating data of the second GPU; construct a transmission task transfer coefficient according to the historical operating data of the first GPU and the historical operating data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, the current operating data of the first GPU and the current operating data of the second GPU; perform load prediction through a preset linear regression model, the current operating data and the historical operating data; based on the load prediction result, optimize task allocation through a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy.

[0082] A non-volatile computer storage medium provided in an embodiment of the present application stores computer executable instructions, wherein the computer executable instructions are configured to: based on a preset interval duration, obtain current operating data of a first GPU and current operating data of a second GPU, and obtain historical operating data of a GPU and historical operating data of a second GPU; construct a transmission task transfer coefficient according to the historical operating data of the first GPU and the historical operating data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, the current operating data of the first GPU and the current operating data of the second GPU; perform load prediction by using a preset linear regression model, current operating data and historical operating data; and optimize task allocation by using a genetic algorithm based on the load prediction result, so as to adjust the reference task transfer strategy and obtain an adjusted task transfer strategy.

[0083] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0084] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the embodiments of the present application may have various changes and variations. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A task transfer method based on dual GPU architecture, characterized in that: The method comprises: Based on a preset interval duration, current operation data of the first GPU and current operation data of the second GPU are acquired, and historical operation data of the first GPU and historical operation data of the second GPU are acquired; constructing a transmission task transfer coefficient according to historical operation data of the first GPU and historical operation data of the second GPU, so as to determine a reference task transfer strategy based on the transmission task transfer coefficient, current operation data of the first GPU and current operation data of the second GPU; Load forecasting is performed through preset linear regression models, current operating data, and historical operating data; Based on the load prediction result, the task allocation is optimized by a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy.

2. The task transfer method based on dual GPU architecture according to claim 1, characterized in that: The step of determining a reference task transfer strategy based on the transmission task transfer coefficient, the current operation data of the first GPU, and the current operation data of the second GPU specifically includes: determining a current load corresponding to the first GPU by using a ratio between current power consumption in current operation data of the first GPU and a maximum power consumption corresponding to the first GPU; determining a current load corresponding to the second GPU by using a ratio between current power consumption in the current operation data of the second GPU and a maximum power consumption corresponding to the second GPU; Obtaining a load balancing target based on a current load corresponding to the first GPU and a current load corresponding to the second GPU; The reference task transfer strategy is determined based on a current load corresponding to the first GPU, a current load corresponding to the second GPU, and the load balancing target.

3. The task transfer method based on dual GPU architecture according to claim 2, characterized in that: The determining the reference task transfer strategy based on the current load corresponding to the first GPU, the current load corresponding to the second GPU, and the load balancing target specifically includes: Comparing the current load corresponding to the first GPU and the current load corresponding to the second GPU with the load balancing target respectively; Based on the comparison result, determining a reference load corresponding to a GPU that is greater than the load balancing target; The reference task transfer strategy is determined based on the difference between the reference load and the balancing target, and the transmission task transfer coefficient.

4. The task transfer method based on dual GPU architecture according to claim 3, characterized in that: The constructing a transmission task transfer coefficient according to the historical operation data of the first GPU and the historical operation data of the second GPU specifically includes: Based on the reference load, determining reference historical operation data from the historical operation data of the first GPU and the historical operation data of the second GPU; In the reference historical operating data, a plurality of historical transfer coefficients corresponding to the reference load are determined; Filtering the plurality of historical transfer coefficients based on the reference load and an average of current operating data of the first GPU and current data of the second GPU, and sorting the filtered historical transfer coefficients based on differences between the historical transfer coefficients and the average; Based on the sorting order, weights are assigned to the plurality of historical transfer coefficients, and the transmission task transfer coefficient is obtained through the result after weight processing.

5. The task transfer method based on dual GPU architecture according to claim 3, characterized in that: The determining of the reference task transfer strategy based on the difference between the reference load and the balancing target and the transmission task transfer coefficient specifically includes: Function-based: T transfer =k·(L u -L avg ); Determine the load to be transferred corresponding to the reference load, so as to determine the reference task transfer strategy based on the load to be transferred; Among them, T transfer is the load to be transferred; k is the transfer coefficient of the transmission task; L u is the reference load; L avg is the equilibrium target.

6. The task transfer method based on dual GPU architecture according to claim 5, characterized in that: The load forecasting is performed by using a preset linear regression model, current operation data and historical operation data, specifically including: Obtain historical load data of the GPU corresponding to the reference load; Based on the historical load data and the linear regression model: L pred =β0+β1·T+β2·H; Determining a predicted load of the GPU corresponding to the reference load; Among them, L pred is the predicted load; T is the time variable; H is the historical load data; β0 is the first parameter; β1 is the second parameter; β2 is the third parameter.

7. The task transfer method based on dual GPU architecture according to claim 6, characterized in that: The step of optimizing task allocation based on the load prediction result by using a genetic algorithm to adjust the reference task transfer strategy to obtain an adjusted task transfer strategy specifically includes: The task allocation is optimized by genetic algorithm, and the fitness function is set as: The reference task transfer strategy is adjusted by minimizing the fitness function to obtain the adjusted task transfer strategy.

8. The task transfer method based on dual GPU architecture according to claim 1, characterized in that: The acquiring of the current operation data of the first GPU and the current operation data of the second GPU based on the preset interval duration specifically includes: Comparing the acquired current operation data of the first GPU and the acquired current operation data of the second GPU with a preset frequency table; wherein the preset frequency table includes a plurality of operation data ranges and a plurality of data acquisition frequencies corresponding to the operation data ranges; Determine a corresponding data acquisition frequency based on comparison results of the current operation data of the first GPU and the current operation data of the second GPU with the plurality of operation data ranges; Based on the data acquisition frequency, a new preset interval duration is determined, so as to acquire the current operation data of the first GPU and the current operation data of the second GPU based on the new preset interval duration.

9. A task transfer method and device based on dual GPU architecture, characterized in that: The device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.