GPU sharing method, system and medium for integrated neural network model

CN122433820BActive Publication Date: 2026-09-18NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610886342.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-18
Estimated Expiration
2046-06-18

AI Technical Summary

Technical Problem

如何发挥集成神经网络模型的性能优势并实现高效的GPU共享存在下述阻碍:首先,多个相关的子模型的并行执行可能会相互影响,因此需要重新考虑整体模型执行效率的评价指标和优化方向

Benefits of technology

[0015] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: (1) Constructing an efficient execution mechanism that supports GPU sharing and solving the performance bottleneck of parallel execution: The present invention designs and implements an integrated neural network online execution framework that supports multi-process parallel deployment and GPU sharing on the PyTorch platform. By dividing the integrated model into sub-models, allocating devices and managing resource ratios, the execution engine coordinates the GPU resource ratio of each sub-process (through MPS) and accurately records the execution time. In order to solve the latency and memory waste caused by inter-process data interaction, the present invention introduces a shared memory and event controller mechanism, which effectively reduces communication costs and achieves stable process control, thereby significantly improving the overall execution efficiency. (2) Proposing a two-layer optimization strategy for GPU resources to improve search efficiency and deployment quality: The present invention proposes a single GPU resource partitioning optimization strategy and a multi-GPU device allocation algorithm based on the optimization goal of "shortest completion time". For single-GPU optimization, the resource allocation starting point is predicted by combining the computational amount and intensity of the sub-model, and the iterative search is accelerated by a dynamic step size adjustment strategy. In multi-GPU allocation, reasonable device allocation candidate schemes are generated based on the static features of the sub-model, and high-efficiency sub-models are scheduled first to avoid the computational burden of exhaustive schemes, thereby achieving a comprehensive improvement in integrated deployment efficiency. (3) Significant advantages of this invention in execution efficiency and search performance: In the scenario of integrated neural network deployment, experimental results show that the method of this invention is significantly better than the existing execution scheme, which not only accelerates the model execution process, but also improves the search convergence efficiency of GPU resources. Compared with the traditional static or average partitioning strategy, the scheme of this invention has achieved excellent performance in key indicators such as task completion time, GPU utilization and system response latency, and can realize efficient resource scheduling and deployment in various real scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433820B_ABST
    Figure CN122433820B_ABST
Patent Text Reader

Abstract

The application discloses a GPU sharing method and system for an integrated neural network model and a medium. The method comprises the following steps: performing GPU partitioning on sub-models in the integrated neural network model, determining an initial GPU resource allocation ratio of each sub-model according to the calculation intensity and calculation load of the sub-models to obtain an initial GPU allocation scheme; performing online deployment on the current GPU allocation scheme to obtain the execution time of each sub-model and the total execution time of the integrated neural network model under the current GPU allocation scheme; adjusting the GPU resource allocation ratio of each sub-model according to the total execution time to obtain a new GPU allocation scheme; and outputting the finally obtained new GPU allocation scheme if a preset search termination condition is reached. The application aims to realize GPU shared resource allocation optimization of the integrated neural network model, improve the execution efficiency of the integrated neural network model, and balance the parallel execution time of the sub-models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of GPU sharing technology for neural network models, and specifically to a GPU sharing method, system, and medium for integrated neural network models. Background Technology

[0002] Ensemble neural network models improve prediction accuracy by integrating multiple neural network models, and their excellent performance has led to their widespread application in various artificial intelligence competitions and large-scale commercial deployments. Meanwhile, the growing demand for interactive services based on ensemble neural network models has led researchers to focus more on service quality, especially how to fully utilize the performance of computing devices. For online interactive services, the low GPU utilization caused by mini-batch input data hinders the full performance of computing devices. To address this obstacle, GPU sharing technology has been proposed, allowing multiple tasks to be executed in parallel on a single device, effectively solving the problem of low GPU utilization caused by mini-batch input data. Therefore, it is widely used in various online model inference frameworks to improve device utilization and overall task execution efficiency.

[0003] However, existing inference execution frameworks based on GPU sharing technology are mostly applied to multi-tenant models, i.e., parallel inference execution of multiple unrelated neural network models; they rarely focus on ensemble neural network models, i.e., parallel inference execution of multiple related sub-models. Existing GPU sharing technologies also fall short when applied to ensemble neural network models. The following obstacles hinder the realization of the performance advantages of ensemble neural network models and the achievement of efficient GPU sharing: First, the parallel execution of multiple related sub-models may interfere with each other, thus requiring a reconsideration of the evaluation metrics and optimization directions for overall model execution efficiency. Second, existing resource search frameworks based on GPU sharing can only perform simple searches without considering model parameters and characteristics, leading to insufficient efficiency. Finally, due to issues such as process blocking and data interaction, the parallel execution of sub-models during the execution of ensemble models suffers from process blocking and data interaction, making efficiency metrics ambiguous. The different completion times of each sub-model also disrupt the optimization direction of computational resource allocation, making current GPU sharing-based execution frameworks unsuitable for ensemble neural network models. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a GPU sharing method, system and medium for integrated neural network models, which addresses the above-mentioned problems in the prior art. This invention aims to optimize the allocation of GPU shared resources for integrated neural network models, improve the execution efficiency of integrated neural network models and improve the balance of parallel execution time of sub-models.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A GPU sharing method for integrated neural network models includes the following steps: S101 partitions the sub-models in the integrated neural network model into GPU partitions, allocating each sub-model to the GPU; S102, for each sub-model allocated to a GPU, determine the initial GPU resource allocation ratio of each sub-model based on the computational intensity and computational load of the sub-model, thereby obtaining the initial GPU allocation scheme; S103, deploys online for the current GPU allocation scheme, and obtains the execution time of each sub-model and the total execution time of the integrated neural network model when the current GPU allocation scheme is executed on the execution engine; S104, adjust the GPU resource allocation ratio of each sub-model according to the execution time of each sub-model and the total execution time of the integrated neural network model, so as to obtain a new GPU allocation scheme; S105, determine whether the preset search termination condition has been met. If the preset search termination condition has been met, output the new GPU allocation scheme obtained at the end; otherwise, jump to step S103 to continue the iteration.

[0006] Optionally, step S101 includes: (1) generating a set of candidate GPU partitioning schemes for partitioning the sub-models in the integrated neural network model onto GPUs, wherein each candidate scheme in the set of candidate GPU partitioning schemes is used to partition the sub-models in the integrated neural network model onto GPUs; (2) obtaining the efficiency evaluation value of each GPU in each candidate GPU partitioning scheme through simulation modeling, and calculating the variance of the efficiency evaluation values ​​of all GPUs in each candidate GPU partitioning scheme; (3) selecting the candidate GPU partitioning scheme with the smallest variance as the final GPU partitioning scheme.

[0007] Optionally, in step S102, the computational intensity of the sub-model is the ratio of the computational amount of the sub-model to the memory access overhead, wherein the computational amount of the sub-model is the sum of all floating-point calculation operations during the execution of the sub-model, and the memory access overhead of the sub-model is the sum of all memory moves during the execution of the sub-model, including input data moves, weight data moves, and intermediate data moves.

[0008] Optionally, step S102, determining the initial GPU resource allocation ratio for each sub-model based on its computational intensity and computational load, includes: normalizing the computational intensity and computational load of the sub-model to obtain normalized computational intensity and normalized computational load, and then calculating the efficiency evaluation value of each sub-model according to the following formula: ; in, Let be the efficiency evaluation value of the i-th sub-model. The normalized strength is calculated for the i-th sub-model. The normalized computational load for the i-th sub-model is... and For the calibration parameters of the i-th sub-model, determine the initial GPU resource allocation ratio for each sub-model based on the efficiency evaluation values ​​of each sub-model: ; in, ~ These represent the initial GPU resource allocation ratios for the 1st to nth sub-models, respectively. ~ These are the efficiency evaluation values ​​for the 1st to nth sub-models, respectively.

[0009] Optionally, step S104, adjusting the GPU resource allocation ratio of each sub-model based on the execution time of each sub-model and the total execution time of the integrated neural network model, includes: ① Sub-model grouping: Calculate the average execution time of each sub-model based on its execution time. Using the average execution time combined with positive and negative thresholds as the average execution interval, group sub-models whose execution time is greater than the upper limit of the average execution interval into the increased GPU resource group, sub-models whose execution time is within the average execution interval into the maintained GPU resource group, and sub-models whose execution time is less than the lower limit of the average execution interval into the decreased GPU resource group; ② Grouping control: Increase the GPU resource allocation ratio of sub-models in the increased GPU resource group; decrease the GPU resource allocation ratio of sub-models in the decreased GPU resource group; and keep the GPU resource allocation ratio unchanged for sub-models in the maintained GPU resource group.

[0010] Optionally, increasing the GPU resource allocation ratio means increasing the GPU resource allocation ratio by a preset ratio or multiplying it by a coefficient greater than 1; decreasing the GPU resource allocation ratio means increasing or decreasing the GPU resource allocation ratio by a preset ratio or multiplying it by a coefficient less than 1.

[0011] Optionally, the preset search termination condition in step S105 is that the number of iterations is equal to a preset threshold or the maximum execution time difference between sub-models is lower than the preset maximum execution time difference threshold Tmax, where the maximum execution time difference between sub-models refers to the maximum value of the time difference between the execution times of each sub-model.

[0012] The present invention also provides a GPU sharing system for integrated neural network models, including interconnected microprocessors and memory, wherein the microprocessors are programmed or configured to execute the GPU sharing method for integrated neural network models.

[0013] The present invention also provides a computer-readable storage medium storing a computer program or instructions that are programmed or configured to execute the GPU-sharing method for integrated neural network models via a processor.

[0014] The present invention also provides a computer program product, including a computer program or instructions that are programmed or configured to execute the GPU sharing method for integrated neural network models via a processor.

[0015] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: (1) Constructing an efficient execution mechanism that supports GPU sharing and solving the performance bottleneck of parallel execution: The present invention designs and implements an integrated neural network online execution framework that supports multi-process parallel deployment and GPU sharing on the PyTorch platform. By dividing the integrated model into sub-models, allocating devices and managing resource ratios, the execution engine coordinates the GPU resource ratio of each sub-process (through MPS) and accurately records the execution time. In order to solve the latency and memory waste caused by inter-process data interaction, the present invention introduces a shared memory and event controller mechanism, which effectively reduces communication costs and achieves stable process control, thereby significantly improving the overall execution efficiency. (2) Proposing a two-layer optimization strategy for GPU resources to improve search efficiency and deployment quality: The present invention proposes a single GPU resource partitioning optimization strategy and a multi-GPU device allocation algorithm based on the optimization goal of "shortest completion time". For single-GPU optimization, the resource allocation starting point is predicted by combining the computational amount and intensity of the sub-model, and the iterative search is accelerated by a dynamic step size adjustment strategy. In multi-GPU allocation, reasonable device allocation candidate schemes are generated based on the static features of the sub-model, and high-efficiency sub-models are scheduled first to avoid the computational burden of exhaustive schemes, thereby achieving a comprehensive improvement in integrated deployment efficiency. (3) Significant advantages of this invention in execution efficiency and search performance: In the scenario of integrated neural network deployment, experimental results show that the method of this invention is significantly better than the existing execution scheme, which not only accelerates the model execution process, but also improves the search convergence efficiency of GPU resources. Compared with the traditional static or average partitioning strategy, the scheme of this invention has achieved excellent performance in key indicators such as task completion time, GPU utilization and system response latency, and can realize efficient resource scheduling and deployment in various real scenarios. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention.

[0017] Figure 2 This is a schematic diagram of the overall framework structure of the method in an embodiment of the present invention.

[0018] Figure 3This is a schematic diagram illustrating the principle of multi-GPU computing resource allocation in an embodiment of the present invention.

[0019] Figure 4 This is a schematic diagram illustrating the principle of single GPU computing resource allocation in an embodiment of the present invention.

[0020] Figure 5 This is a schematic diagram illustrating the working principle of the "execution engine" in an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] like Figure 1 As shown, the GPU sharing method for integrated neural network models in this embodiment includes the following steps: S101 partitions the sub-models in the integrated neural network model into GPU partitions, allocating each sub-model to the GPU; S102, for each sub-model allocated to a GPU, determine the initial GPU resource allocation ratio of each sub-model based on the computational intensity and computational load of the sub-model, thereby obtaining the initial GPU allocation scheme; S103, deploys online for the current GPU allocation scheme, and obtains the execution time of each sub-model and the total execution time of the integrated neural network model when the current GPU allocation scheme is executed on the execution engine; S104, adjust the GPU resource allocation ratio of each sub-model according to the execution time of each sub-model and the total execution time of the integrated neural network model, so as to obtain a new GPU allocation scheme; S105, determine whether the preset search termination condition has been met. If the preset search termination condition has been met, output the new GPU allocation scheme obtained at the end; otherwise, jump to step S103 to continue the iteration.

[0023] like Figure 2As shown, the input to the method in this embodiment includes an integrated neural network model and a set of GPUs in the computing device. The overall framework of the method in this embodiment includes two parts: a "GPU optimization framework" and an "execution engine". The input integrated neural network model supports mainstream model representation methods such as PyTorch, ONNX, and TVM. The "GPU optimization framework" is used to execute steps S101 to S105, and the "execution engine" is used to execute the current GPU allocation scheme. "Sub-model parameter extraction" is to obtain information about the sub-model, including computational cost and memory access cost, as well as the ratio of computational cost to memory access cost. "Multi-GPU computing resource allocation" is the GPU partitioning in step S101, which divides each sub-model onto GPUs. "Single GPU computing resource allocation" is the optimization step in steps S103 to S104. When the "GPU optimization framework" and "execution engine" are working, they first optimize the allocation of multiple GPU devices based on the input integrated neural network model and hardware execution device to obtain a multi-GPU allocation scheme, that is, a scheme for deploying sub-models on each GPU. Then, the framework performs single GPU resource allocation optimization, allocating computing resources to the sub-models on each GPU. Finally, the overall computing resource scheme needs to be deployed on the execution engine for online inference services. Furthermore, the search framework iteration process in the optimization of individual GPU resource allocation requires the online execution engine to provide the execution time for each sub-model and the overall model. Addressing the resource waste and execution imbalance issues in shared execution of ensemble neural network models on shared GPU devices, this embodiment, unlike traditional multi-tenant GPU sharing frameworks, fully considers the data interaction and execution dependencies between sub-models in the ensemble model and designs a two-level resource partitioning mechanism suitable for both single-GPU and multi-GPU environments. This embodiment's method models sub-model parameters (such as computational load and memory access) to guide device and resource allocation, thereby achieving load balancing and minimizing idle time, improving overall execution efficiency. This embodiment's method includes GPU partitioning and dynamically adjusting the GPU resource allocation ratio of each sub-model based on the execution time of each sub-model and the total execution time of the ensemble neural network model under the current GPU allocation scheme. Through optimization in two dimensions—multi-GPU device partitioning and single-GPU resource allocation—it can optimize the shared GPU resource allocation of the ensemble neural network model, improve the execution efficiency of the ensemble neural network model, and improve the balance of parallel execution time of sub-models.

[0024] In this embodiment, step S101 includes: (1) generating a set of candidate GPU partitioning schemes for partitioning the sub-models in the integrated neural network model onto GPUs, wherein each candidate scheme in the set of candidate GPU partitioning schemes is used to partition the sub-models in the integrated neural network model onto GPUs; (2) obtaining the efficiency evaluation value of each GPU in each candidate GPU partitioning scheme through simulation modeling, and calculating the variance of the efficiency evaluation values ​​of all GPUs in each candidate GPU partitioning scheme; (3) selecting the candidate GPU partitioning scheme with the smallest variance as the final GPU partitioning scheme.

[0025] In step S102 of this embodiment, the computational intensity of the sub-model obtained through parameter modeling is the ratio of the computational cost to the memory access overhead of the sub-model. The computational cost of the sub-model is the sum of all floating-point calculations during its execution, and the memory access overhead is the sum of all memory moves during its execution, including input data moves, weight data moves, and intermediate data moves. In step S102 of this embodiment, determining the initial GPU resource allocation ratio for each sub-model based on its computational intensity and computational load includes: normalizing the computational intensity and computational load of the sub-model to obtain normalized computational intensity and normalized computational load, and then calculating the efficiency evaluation value of each sub-model according to the following formula: ; in, Let be the efficiency evaluation value of the i-th sub-model. The normalized strength is calculated for the i-th sub-model. The normalized computational load for the i-th sub-model is... and For the calibration parameters of the i-th sub-model, determine the initial GPU resource allocation ratio for each sub-model based on the efficiency evaluation values ​​of each sub-model: ; in, ~ These represent the initial GPU resource allocation ratios for the 1st to nth sub-models, respectively. ~ These are the efficiency evaluation values ​​for the 1st to nth sub-models, respectively.

[0026] In this embodiment, step S104, adjusting the GPU resource allocation ratio of each sub-model based on the execution time of each sub-model and the total execution time of the integrated neural network model, includes: ① Sub-model grouping: Calculating the average execution time of each sub-model based on its execution time, and using a positive and negative threshold as the average execution interval, sub-models with execution times greater than the upper limit of the average execution interval are grouped into the "increased GPU resource" group, sub-models with execution times within the average execution interval are grouped into the "maintained GPU resource" group, and sub-models with execution times less than the lower limit of the average execution interval are grouped into the "decreased GPU resource" group; ② Grouping control: For sub-models in the "increased GPU resource" group, increasing their GPU resource allocation ratio; for sub-models in the "decreased GPU resource" group, decreasing their GPU resource allocation ratio; and for sub-models in the "maintained GPU resource" group, keeping their GPU resource allocation ratio unchanged. In this embodiment, increasing the GPU resource allocation ratio means increasing it by a preset percentage or multiplying it by a coefficient greater than 1; decreasing the GPU resource allocation ratio means increasing or decreasing it by a preset percentage or multiplying it by a coefficient less than 1.

[0027] In step S105 of this embodiment, the preset search termination condition is that the number of iterations is equal to a preset threshold or the maximum execution time difference between sub-models is lower than the preset maximum execution time difference threshold Tmax. The maximum execution time difference between sub-models refers to the maximum value of the time difference between the execution times of each sub-model. Figure 3This is a schematic diagram of the principle of multi-GPU computing resource allocation in this embodiment. When allocating multi-GPU computing resources, the process includes determining the initial GPU resource allocation ratio for each sub-model allocated to each GPU based on the computational intensity and computational load of the sub-model. This includes parameter modeling of the sub-model parameters to obtain the computational intensity of the sub-model. Then, using a computing resource search framework, a preset resource search optimization strategy is adopted to perform resource search iteration to obtain the initial GPU allocation scheme. Using a computing resource search framework, by setting the optimization search direction, search parameters, search starting point and search dynamic step size, the required resource search optimization strategy can be adopted to perform resource search iteration to obtain the initial GPU allocation scheme. (1) Search direction: Under GPU sharing technology, the key to optimizing the execution efficiency of the integrated neural network model is to make the completion time of each sub-model executed in parallel as close as possible, thereby shortening the overall execution time. To this end, this invention proposes an iterative fine-tuning GPU resource allocation strategy, which gradually approaches the optimal allocation state by dynamically adjusting the computational resource ratio of each sub-model on the GPU. (2) Search parameters: In each iteration, the resource allocation scheme is updated based on the actual execution time feedback of the previous round. The strategy sets the maximum number of iterations N and the maximum execution time difference threshold Tmax allowed by the sub-model before the search begins, as the search termination condition to balance the search effect and resource consumption. The specific search process includes initializing the GPU partition ratio and then performing up to N iterations. In each iteration, the sub-model is deployed and executed according to the current GPU resource allocation scheme, its execution time is recorded, and the maximum execution time difference between the sub-models is calculated. When the difference is lower than the threshold Tmax, the current allocation is considered to be close to the optimal and the search ends early; otherwise, the partition ratio is adjusted according to the actual execution of the sub-model and the iteration continues. (3) Prediction of the search starting point: The length of the search path, that is, the distance from the search starting point to the ending point, directly affects the efficiency of resource allocation optimization. The selection of the starting point is crucial to the search quality: an excellent starting point can not only significantly reduce the number of iterations, but sometimes even obtain a good allocation effect without searching. This invention models the static parameters of each sub-network to predict the optimal GPU allocation ratio, and uses this as the starting point of the search, thereby improving the overall model execution efficiency and shortening the total execution time (JCT). The key to achieving this goal is to identify and model the core factors that affect the execution efficiency of the sub-model. This embodiment constructs a prediction model from two dimensions: computational cost and computational intensity. Computational cost measures the total number of operation instructions required to execute a task, mainly corresponding to multiply-accumulate operations (FLOPs) in neural networks. It is positively correlated with execution time; that is, the greater the computational cost, the longer the execution time. Computational intensity reflects the efficiency of the task's utilization of computing power and is defined as the amount of computation completed per unit of memory access. The computational cost and memory access overhead (MAC) of each sub-model are obtained using tools (such as torch summary).By combining these two indicators, a performance prediction model based on the characteristics of sub-model parameters can be established, thereby obtaining a reasonable initial GPU resource allocation ratio, reducing the search path length, and thus accelerating the search process and improving resource optimization efficiency. (4) Dynamic search step size: In GPU resource optimization search, the setting of the search step size has an important impact on the efficiency and accuracy of the search process. Too large a step size will lead to insufficient accuracy, while too small a step size will increase unnecessary iteration overhead. Therefore, this invention proposes to adopt a dynamic step size adjustment mechanism to achieve fast convergence and maintain sufficient accuracy. In each iteration, the search framework adaptively adjusts the step size of the next round according to the execution time difference between the current sub-models and the GPU allocation status, so that the completion time of each sub-model gradually approaches, thereby achieving the goal of minimizing the overall execution time. In order to further improve search efficiency, this embodiment also introduces a sub-model grouping mechanism: the sub-models are divided into three categories according to the execution time: increasing resources, decreasing resources, or keeping resources unchanged, thereby guiding the new round of GPU resource allocation strategy. In addition, in order to reduce search complexity, isomorphic sub-models always maintain the same GPU allocation ratio throughout the search process.

[0028] like Figure 4 As shown, the optimization of multiple GPU device allocation begins with candidate scheme generation. Considering that device allocation is an NP-hard problem, directly enumerating all possible allocation candidate schemes is not feasible when the number of sub-models and devices is large. Therefore, this embodiment adopts a scenario-based processing strategy: when the number of sub-models and devices is small, all possible device allocation schemes are traversed; while when the scale is large, a candidate scheme set is constructed based on the evaluation efficiency of the sub-models. For sub-models with similar evaluation efficiencies, the allocation scheme with load balancing on each device is given priority; while when the efficiency difference is large, the sub-model with high evaluation efficiency is given priority to be independently allocated to different devices to improve overall execution performance. Then, this embodiment designs an efficient device allocation evaluation method by modeling the sub-model parameters (such as computational load, memory access, etc.), aiming to minimize idle time during execution and thus improve overall system efficiency. Finally, using the obtained allocation candidate schemes and the evaluation method based on sub-model parameter modeling, the optimal GPU allocation scheme is found, i.e., the corresponding sub-model deployed on each GPU device. Resource search iteration is the specific resource search iteration process, which requires activating the execution engine and using the execution results for search iteration. During each search iteration, the search framework provides the execution engine with a GPU resource allocation scheme and a set of sub-models. After execution, the execution engine provides the search framework with the overall execution time and the execution time of each sub-model.

[0029] After obtaining the GPU allocation scheme, the current GPU allocation scheme can be deployed online through step S103. The current GPU allocation scheme is executed through the "execution engine" to obtain the execution time of each sub-model and the total execution time of the integrated neural network model when the current GPU allocation scheme is executed on the execution engine. Figure 5 This is a schematic diagram illustrating the working principle of the "execution engine" in this embodiment of the invention. The "execution engine" includes a process set consisting of multiple sub-processes. Each sub-process simulates a GPU. An event controller controls the process set to execute the current GPU allocation scheme, reads user requests from shared memory, outputs the results to shared memory, and outputs a response to achieve the purpose of simulated execution. Each sub-model is assigned to an independent sub-process and bound to a specific GPU in a multi-GPU environment through a device allocation algorithm. The GPU utilization ratio of each sub-process is dynamically set by the MPS (Multi-Process Service) mechanism. The execution engine is responsible for scheduling the entire inference process, while recording the earliest completion time and actual execution time of each sub-model. In each iteration, it adjusts resource allocation based on the search results, driving the overall system towards optimization towards the shortest task completion time (JCT). To address the problem of frequent data interaction between parallel sub-models in the ensemble model, this invention introduces a shared memory and event controller mechanism, effectively reducing the data migration overhead between processes. Input request data and output result data are stored in dedicated shared memory areas; the event controller uniformly manages the sub-process state, only unblocking the process and entering the next round of inference after all sub-processes have completed the previous round of computation and the shared data has been updated. Each subprocess immediately blocks after completing data processing until the next round of input arrives. Finally, the execution engine integrates the results and generates a response after all sub-models have finished outputting. Compared to traditional GPU inference frameworks, the architecture in this embodiment significantly optimizes inter-process collaboration and GPU partition search efficiency while improving execution efficiency, making it more suitable for deploying latency-sensitive integrated neural network online services.

[0030] In summary, this embodiment of the GPU sharing method for integrated neural network models comprises two parts: resource allocation for a single GPU and device allocation across multiple GPUs. The specific steps are as follows: Device allocation across multiple GPUs: When the number of deployed devices is greater than one, the sub-models of the integrated model are first allocated to different GPU devices using a device allocation algorithm. By modeling the parameters of the sub-models, device allocation can be guided to achieve the optimization goal of minimizing idle time. This process ensures load balancing for each GPU and minimizes idle time, thereby avoiding excessive resource concentration or waste. Resource allocation for a single GPU: Then, corresponding computing resources are allocated for each GPU. Computational resource allocation is based on GPU utilization (GPU utilization %) to ensure that the resources of each sub-model on the GPU are used reasonably. In the optimization of single GPU resource allocation, the optimization direction is: the goal is to make the execution times of parallel sub-models as close as possible. Fine-tuning of computing resources can gradually approach the optimal optimization goal in each iteration. By dynamically adjusting the search step size, convergence can be accelerated, ensuring that the execution times of all sub-models gradually converge, thereby maximizing overall efficiency. (2) Optimization starting point: By considering the parameter characteristics of each sub-network model (such as floating-point computation and storage access), a prediction function is constructed to predict the GPU allocation ratio of the optimization search starting point. With a suitable search starting point and dynamically fine-tuned search step size, the search process can be effectively accelerated to achieve higher efficiency. (3) Search step size: In each iteration, the search step size is dynamically adjusted according to the execution time difference of the current sub-model and the GPU allocation. In this way, convergence can be accelerated and GPU allocation can be optimized, ultimately maximizing the overall execution efficiency. In multi-GPU device allocation optimization, the goal is to optimize device allocation by modeling the parameters of the sub-model to achieve the goal of minimum idle time. By adjusting the GPU allocation ratio, the load of each GPU is more balanced, thereby improving the execution efficiency of the ensemble model. By modeling based on the sub-model parameters, this invention finds a suitable sub-model device allocation scheme for the ensemble model. At the same time, combined with the optimization partitioning algorithm of a single GPU, an effective resource allocation scheme is provided for the deployment of the ensemble neural network model. Through this resource allocation scheme, users can obtain better online interactive services and achieve low-latency, high-precision ensemble model application and optimization. By using the methods described above, integrated neural network models can efficiently allocate computing resources and improve overall execution efficiency, thereby providing a better interactive experience for online services.

[0031] Those skilled in the art will understand that the technical solutions provided by this invention can take the form of methods, systems, or computer program products. For example, this invention can provide a GPU sharing system for integrated neural network models, including interconnected microprocessors and memory, wherein the microprocessors are programmed or configured to execute the GPU sharing method for integrated neural network models. This invention can provide a computer-readable storage medium storing a computer program or instructions programmed or configured to execute the GPU sharing method for integrated neural network models via a processor. This invention can provide a computer program product including computer programs or instructions programmed or configured to execute the GPU sharing method for integrated neural network models via a processor. Furthermore, this invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this invention can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of a flowchart and / or block diagram, and combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0032] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A GPU sharing method for integrated neural network models, characterized in that, Includes the following steps: S101 partitions the sub-models in the integrated neural network model into GPU partitions, allocating each sub-model to the GPU; S102, for each sub-model allocated to a GPU, determine the initial GPU resource allocation ratio of each sub-model based on the computational intensity and computational load of the sub-model, thereby obtaining the initial GPU allocation scheme; the computational intensity of the sub-model is the ratio of the computational amount of the sub-model to the memory access overhead, wherein the computational amount of the sub-model is the sum of all floating-point calculation operations during the execution of the sub-model. S103, deploys online for the current GPU allocation scheme, and obtains the execution time of each sub-model and the total execution time of the integrated neural network model when the current GPU allocation scheme is executed on the execution engine; S104, adjust the GPU resource allocation ratio of each sub-model according to the execution time of each sub-model and the total execution time of the integrated neural network model, so as to obtain a new GPU allocation scheme; S105, determine whether the preset search termination condition has been met. If the preset search termination condition has been met, output the final new GPU allocation scheme. Otherwise, proceed to step S103 and continue iterating; Step S102, determining the initial GPU resource allocation ratio for each sub-model based on its computational intensity and load, includes: normalizing the computational intensity and load of the sub-model to obtain normalized computational intensity and load, and then calculating the efficiency evaluation value of each sub-model according to the following formula: ; in, For the first Efficiency evaluation value of each sub-model For the first Normalization computational intensity of each sub-model For the first Normalized computational load of each sub-model and For the first The calibration parameters for each sub-model are used to determine the initial GPU resource allocation ratio for each sub-model based on its efficiency evaluation value. ; in, ~ These represent the initial GPU resource allocation ratios for the 1st to nth sub-models, respectively. ~ These are the efficiency evaluation values ​​for the 1st to nth sub-models, respectively.

2. The GPU sharing method for integrated neural network models according to claim 1, characterized in that, Step S101 includes: (1) generating a set of candidate GPU partitioning schemes for partitioning the sub-models in the integrated neural network model onto GPUs, wherein each candidate scheme in the set of candidate GPU partitioning schemes is used to partition the sub-models in the integrated neural network model onto GPUs; (2) obtaining the efficiency evaluation value of each GPU in each candidate GPU partitioning scheme through simulation modeling, and calculating the variance of the efficiency evaluation values ​​of all GPUs in each candidate GPU partitioning scheme; (3) selecting the candidate GPU partitioning scheme with the smallest variance as the final GPU partitioning scheme.

3. The GPU sharing method for integrated neural network models according to claim 1, characterized in that, In step S102, the memory access overhead of the sub-model is the sum of all memory moves during the execution of the sub-model, including input data moves, weight data moves, and intermediate data moves.

4. The GPU sharing method for integrated neural network models according to claim 1, characterized in that, Step S104 involves adjusting the GPU resource allocation ratio of each sub-model based on its execution time and the total execution time of the integrated neural network model. This includes: ① Sub-model grouping: Calculating the average execution time of each sub-model based on its execution time, and using positive and negative thresholds as the average execution interval. Sub-models with execution times greater than the upper limit of the average execution interval are grouped into the "increased GPU resource" group, those with execution times within the average execution interval are grouped into the "maintained GPU resource" group, and those with execution times less than the lower limit of the average execution interval are grouped into the "decreased GPU resource" group; ② Grouping adjustment: For sub-models in the "increased GPU resource" group, increasing their GPU resource allocation ratio; for sub-models in the "decreased GPU resource" group, decreasing their GPU resource allocation ratio; and for sub-models in the "maintained GPU resource" group, keeping their GPU resource allocation ratio unchanged.

5. The GPU sharing method for integrated neural network models according to claim 4, characterized in that, Increasing the GPU resource allocation ratio means increasing the GPU resource allocation ratio by a preset ratio or multiplying it by a coefficient greater than 1; decreasing the GPU resource allocation ratio means increasing or decreasing the GPU resource allocation ratio by a preset ratio or multiplying it by a coefficient less than 1.

6. The GPU sharing method for integrated neural network models according to claim 1, characterized in that, The preset search termination condition in step S105 is that the number of iterations is equal to a preset threshold or the maximum execution time difference between sub-models is lower than the preset maximum execution time difference threshold Tmax. The maximum execution time difference between sub-models refers to the maximum value of the time difference between the execution times of each sub-model.

7. A GPU-shared system for integrated neural network models, comprising interconnected microprocessors and memory, characterized in that, The microprocessor is programmed or configured to execute the GPU sharing method for integrated neural network models as described in any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the GPU-sharing method for integrated neural network models as described in any one of claims 1 to 6.

9. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the GPU-sharing method for integrated neural network models as described in any one of claims 1 to 6.