Model resource management method and computing device
By monitoring model performance indicators and dynamically adjusting the memory resources and the number of model instances, the problem of how to reasonably deploy multiple neural network models under limited hardware resources is solved, and the efficient operation of the system and user experience improvement is achieved.
Patent Information
- Application Number
- CN202510262561.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-18
AI Technical Summary
Under limited hardware resources, how to reasonably deploy and manage multiple neural network models to ensure efficient operation of the system and improve user experience.
By monitoring the performance indicators of the model, dynamically adjust the video memory resources, including increasing or decreasing the number of video memory resources and model instances, load balancing is achieved and the efficient operation of the model is ensured.
It improves the operating efficiency and performance of the model, improves the user experience, and achieves the accurate matching of user needs and model capabilities.
Smart Images

Figure CN120339034A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model management, and particularly to a method for managing model resources and a computing device. Background Art
[0002] In the current field of artificial intelligence, neural network models (hereinafter referred to as models) have made remarkable progress, especially in the fields of natural language processing, image recognition and generation, demonstrating powerful capabilities and extensive application potential. In the process of model training and inference, the graphics card plays a crucial role, providing indispensable computing power. Video memory is the dedicated memory of the graphics card, used to store data that the graphics card needs to access frequently during the calculation process. The capacity and bandwidth of the video memory directly affect the performance of the graphics card, thereby affecting the response speed and processing ability of the model.
[0003] With the increasingly wide application of models, a single model may not be able to meet the needs of some complex scenarios, and multiple models need to be deployed simultaneously to achieve more comprehensive applications. For example, in software development tasks, multiple models may need to be deployed: the code generation model can automatically generate complete code based on the developer's comments; the natural language processing model converts the requirements described in natural language into corresponding code implementations; the test generation model automatically generates unit test code; the code interpretation model provides natural language explanations for code snippets; and the question-and-answer model answers questions related to the code. These models work together to greatly assist developers in coding efficiently and smoothly.
[0004] Since each model needs to use a large amount of video memory, how to reasonably deploy and manage these models under limited hardware resources to ensure the efficient operation of the system and a good user experience has become an urgent problem to be solved. Summary of the Invention
[0005] The embodiments of this application provide a method for managing model resources and a computing device, which can manage video memory resources according to the running requirements of the model, improve the resource utilization efficiency of the system and enhance the user experience.
[0006] In a first aspect, the embodiments of this application provide a method for managing model resources, which is applied to a server. The server deploys multiple models for processing different tasks. The method includes: in response to receiving a user request, determining a target model from the multiple models according to the type of task requested by the user request to process the user request; collecting performance metric statistical data of the target model for processing multiple user requests; and adjusting the video memory resources for the target model according to the performance metric statistical data.
[0007] Thus, the video memory resources can be managed according to the running requirements of the model, improving the resource utilization efficiency of the system and enhancing the user experience.
[0008] In a possible implementation, performance metric statistics of a target model for processing multiple user requests are obtained, including: taking a preset time as a period, and within each period, counting the number of user requests processed by the target model; and obtaining the average response time within each period based on the number and the response time corresponding to each user request within each period, as the performance metric statistics.
[0009] Thus, the required performance metric statistics can be obtained to identify the performance bottleneck of the model during the process of processing user requests.
[0010] In a possible implementation, video memory resources are adjusted for the target model according to the performance metric statistics, including: if the average response time is greater than a first time, increasing the video memory resources for the target model; and if the average response time is less than a second time, reducing the video memory resources for the target model.
[0011] In a possible implementation, video memory resources are adjusted for the target model according to the performance metric statistics, including: in response to increasing the video memory resources of a preset size for the target model at the end of the current period, calculating the response time reduction rate based on the preset time, the average response time of the current period, and the average response time of the next period; if the reduction rate is less than a threshold, and the average response time of the next period is greater than a third time, increasing the number of instances for deploying the target model.
[0012] In a possible implementation, video memory resources are adjusted for the target model according to the performance metric statistics, including: if the average response time of a continuous preset number of periods is less than a fourth time, determining the number of instances for currently deploying the target model; if the current number of deployed instances is greater than 1, reducing the number of instances for deploying the target model; and if the current number of deployed instances is 1, and the video memory resources currently allocated to the target model are greater than the preset size, reducing the video memory resources of the preset size for the target model.
[0013] Thus, adjusting the video memory resources for the target model specifically according to the performance metric statistics can improve the running efficiency and performance of the model.
[0014] In a possible implementation, determining the target model from multiple models to process user requests, including: if the number of instances for currently deploying the target model is greater than 1, evenly distributing multiple user requests to multiple instances of the target model for processing.
[0015] Thus, it is possible to avoid overloading of some instances while idling of other instances, ensuring the overall inference performance and stability of the system.
[0016] In a possible implementation, the preset size is determined according to the number of parameters of the target model, and the number of parameters is determined according to the task objective of processing user requests.
[0017] Thus, when deploying multiple models, appropriate model parameter quantities can be selected according to the specific task objectives of different types of models for processing user requests, so as to achieve the best balance between model inference performance and response time.
[0018] In a possible implementation manner, determining a target model from multiple models to process a user request includes: using the model specified by the user request as the target model to process the user request; or determining, from multiple models, a model that matches the usage scenario of the user request as the target model to process the user request.
[0019] Thus, precise matching between user requirements and model capabilities can be achieved, enhancing the user experience.
[0020] In a possible implementation manner, the types of tasks include one or more of code generation, code completion, code annotation, code modification, or code interpretation.
[0021] Thus, multiple models can be deployed simultaneously to achieve a more comprehensive application.
[0022] In a second aspect, an embodiment of the present application provides a computing device, including:
[0023] Multiple memories for storing programs;
[0024] Multiple processors for executing the programs stored in the memories. When the programs stored in the memories are executed, the processors are used to execute the methods described in the first aspect or any possible implementation manner of the first aspect.
[0025] In a third aspect, an embodiment of the present application provides a computer storage medium. Instructions are stored in the computer storage medium. When the instructions run on a computer, the computer is caused to execute the methods described in the first aspect or any possible implementation manner of the first aspect.
[0026] In a fourth aspect, an embodiment of the present application provides a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute the methods described in the first aspect or any possible implementation manner of the first aspect.
[0027] It can be understood that the beneficial effects of the above second aspect to the fourth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is an architecture diagram of a model resource management platform provided by an embodiment of the present application;
[0030] Figure 2 It is a flowchart of a model resource management method provided by an embodiment of the present application;
[0031] Figure 3a It is a schematic diagram of a video memory resource adjustment strategy provided by an embodiment of the present application;
[0032] Figure 3b It is a schematic diagram of a video memory resource adjustment strategy provided by an embodiment of the present application;
[0033] Figure 4 It is a schematic diagram of a model resource management process provided by an embodiment of the present application;
[0034] Figure 5 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0035] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will describe the technical solutions in the embodiments of the present application with reference to the drawings.
[0036] In the description of the embodiments of the present application, any embodiment or design solution described as "exemplary", "for example", or "for instance" should not be understood as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.
[0037] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "A plurality" can be one or more, where a plurality is two or more.
[0038] Model deployment refers to running a trained model on dedicated computing resources so that it can operate efficiently and reliably in an independent operating environment and provide inference services for business applications. According to the way of inference use, it can be divided into CPU deployment and GPU deployment. GPU deployment includes selecting a GPU with sufficient video memory and computing power according to the scale and computing requirements of the model, and reasonably configuring the video memory resources to ensure the efficient operation of the model.
[0039] In the embodiments of the present application, multiple models for processing different tasks are deployed on computing devices such as servers and terminals. The target model is determined from the multiple models according to the type of task requested by the user request. The user request is processed using the target model, which can achieve an accurate match between the user's needs and the model's capabilities and improve the user experience. Further, performance metric statistical data for the target model processing multiple user requests is obtained, and the video memory resources are adjusted for the target model according to the performance metric statistical data, which can improve the operating efficiency and performance of the model.
[0040] Exemplarily, Figure 1 FIG. shows an architecture diagram of a model resource management platform provided by an embodiment of the present application.
[0041] As Figure 1 shown, the model resource management platform 100 is an application software or tool for managing model resources, and it can run on a server for deploying models. The model resource management platform 100 includes a deployment module 110, a load distribution module 120, a monitoring module 130, a statistics module 140, and a resource adjustment module 150. Each module has a connection and call relationship as Figure 1 shown.
[0042] Specifically, the deployment module 110 is used for model deployment. It determines multiple task types to be processed by the models according to the application domain requirements. Taking the software development application domain as an example, it can be determined that the task types include: code generation, code completion, code annotation, code modification, and code interpretation. Correspondingly, it determines the types of multiple models according to the task types, including: code generation model, natural language processing model, test generation model, code interpretation model, and question answering model. Each type of model has its specific task type. For example, the code interpretation model is used to process code interpretation tasks.
[0043] The deployment module 110 is also responsible for loading multiple models into the server and performing necessary initialization operations for each model, such as allocating initial computing resources, setting model parameters, etc. And according to the running requirements of the models, configure the corresponding software environment and hardware environment, including installing a deep learning framework, setting the version of the dependency library, setting the GPU and video memory resources, etc., to make the multiple models in a runnable state.
[0044] It can be understood that for each type of model, at least one instance is initially deployed. A model instance refers to a specific and independently running copy of the model, which contains all the parameters, structures, and configuration information of the model and can perform inference or training tasks in a specific hardware and software environment.
[0045] The load distribution module 120 receives multiple user requests. For each of these user requests, the load distribution module 120 reasonably distributes the user request to the target model for processing according to the type of task requested by the user request. The target model is a certain type of model among multiple models that best matches the task requested by the user request. This distribution strategy can ensure that each task is supported by the most optimized model, thereby improving the overall inference performance and task completion quality of the system.
[0046] In addition, when deploying multiple instances of a certain type of model on the server, the load distribution module 120 also adjusts the load of each instance through a load balancing algorithm to avoid overloading some instances while other instances are idle, ensuring the overall inference performance and stability of the system.
[0047] The monitoring module 130 monitors the running performance of the model in real time, including key indicators such as response time, processing speed, accuracy, etc., and collects relevant running data on the model processing multiple user requests on the server, such as request volume, response time, resource occupancy, etc.
[0048] The statistics module 140 statistically analyzes the performance metric data collected by the monitoring module 130 to obtain the performance metric statistical data of the model processing multiple user requests, and based on this, identifies the performance bottlenecks of the model in the process of processing user requests.
[0049] The resource adjustment module 150 formulates a resource allocation strategy for adjusting the resources used by the model according to the analysis results of the statistics module 140. For example, if the model response time is long, increase the allocated video memory resources to improve the inference speed. Conversely, reduce the allocated video memory resources to optimize the system resource utilization rate. The resource adjustment module 150 also sends the resource allocation strategy to the deployment module 110 to adjust the model deployment on the server.
[0050] Thus, these modules work together to jointly build an efficient and stable model resource management platform 100, which can achieve efficient model deployment, load balancing, performance monitoring, data analysis, and resource optimization, ensuring that the model achieves the optimal inference efficiency, thereby improving the user experience.
[0051] Based on the above content, a model resource management method proposed in this application will be introduced in detail.
[0052] Exemplarily, Figure 2The figure shows a flowchart of a model resource management method provided by an embodiment of the present application. This method is applied to computing devices such as servers, terminals, etc. Taking a server as an example, the server is deployed with multiple models for processing different tasks.
[0053] It can be understood that in order to process user requests, the server needs to pre-deploy multiple models. Taking the software development application field as an example, the server can deploy different types of code models to meet the needs of users for different types of tasks. For each type of model, at least one instance is initially deployed.
[0054] At the initial deployment, the server allocates respective preset-size video memory resources for different types of models. The video memory resources mainly refer to the video memory capacity, and the video memory capacity determines the data scale that the model can store. The preset size can be the minimum video memory resource amount required for the model to run normally (i.e., the minimum video memory), which ensures the basic functions and performance of the model. At the initial deployment, the video memory allocation cannot be lower than this minimum video memory, otherwise the model may not be able to load or run. The minimum video memory values of different types of models are different. The minimum video memory can be determined according to the number of parameters of the model. The larger the number of parameters, the larger the minimum video memory. The number of parameters can be determined according to the task objectives of processing user requests.
[0055] For example, the code completion function requires the model to be able to respond to the user's input in real time and quickly provide accurate code suggestions without affecting the user's programming process. Therefore, the number of parameters of the model should not be too large, and it is suitable to deploy lightweight code completion models. Such as small models based on Transformer or models based on RNN.
[0056] Functions such as intelligent question answering require the model to be able to accurately understand natural language, which usually requires the model to have strong language understanding capabilities in order to accurately extract the programming intent from the user's natural language description and generate corresponding code. Such models often require a large number of parameters to support complex language understanding tasks. Such as GPT series, BERT, etc.
[0057] Unit test generation requires the model to be able to generate corresponding test cases for the code, which usually requires fine-tuning training of the model for unit testing to improve the accuracy and effectiveness of generating test cases. Models such as sequence-to-sequence based models or generative adversarial network based models can be selected.
[0058] It can be understood that the number of model parameters is negatively correlated with the response time. This is because more parameters mean more complex calculations and higher memory requirements, which will increase the calculation time and memory access latency required for each inference. That is, the larger the number of model parameters, the longer the response time (usually, the shorter the response time is better); while the number of model parameters is positively correlated with the inference performance. This is because more parameters allow the model to learn more complex features and patterns, thereby improving the data fitting ability and generalization ability. That is, the larger the number of model parameters, the stronger the inference performance (usually, the stronger the inference performance is better).
[0059] Therefore, when deploying multiple models, the appropriate number of model parameters can be selected according to the specific task objectives of different types of models to process user requests, so as to achieve the best balance between model inference performance and response time.
[0060] After determining the number of model parameters, estimate the minimum video memory of the model. In practical applications, a common rule of thumb is that the video memory requirement is about 2 to 3 times, or even higher, of the number of model parameters, depending on factors such as model architecture, batch size, and computational complexity. Taking 2 times as an example, assume that the number of parameters of a certain type of model is 1 billion (10^9). Each parameter is usually represented by a floating-point number. Assuming 32-bit floating-point numbers (i.e., 4 bytes) are used, the video memory capacity required for parameter storage is: 1 billion parameters × 4 bytes = 4 billion bytes = 4GB. Then the minimum allocated video memory is 8GB.
[0061] Exemplarily, the server uses the deployment module 110 to deploy multiple models.
[0062] After deploying multiple models on the server, implement Figure 2 The model resource management method shown mainly includes the following execution steps:
[0063] Step 8201, in response to receiving a user request, determine a target model from multiple models to process the user request according to the type of task requested by the user request.
[0064] In one embodiment, after deploying multiple models on the server, the load distribution module 120 receives multiple user requests. A user request is an instruction or information sent by a user to the server, hoping to obtain a certain intelligent service or solution through the model. A user request contains the problems that the user needs to solve, the tasks to be completed, or the information to be received, and is the starting point of the interaction between the user and the model. The types and forms of user requests include voice, text, images, data, or interactive operations, etc. For each user request among multiple user requests, the target model is a certain type of model among multiple models that best matches the task requested by the user request.
[0065] Exemplarily, the load distribution module 120 also determines a target model from multiple models to process the user request according to the type of task requested to be executed by the user request.
[0066] In one implementation, the load distribution module 120 uses the model specified by the user request as the target model to process the user request. Different types of models can be predefined in the deployment module 110, and each type of model corresponds to processing a specific type of user task or request. When a user request is received, the user request specifies the type of model to be accessed, and the server directly assigns the user request to the specified type of model for processing according to the model type label requested to be accessed.
[0067] In another implementation, the load distribution module 120 is used to determine a model that matches the usage scenario of the user request from multiple models as the target model to process the user request. For example, if the semantics of the user request is analyzed as a code interpretation task, it will be assigned to a code interpretation model dedicated to code interpretation; if the semantics of the user request is analyzed as a code generation task described in natural language, it will be assigned to a natural language processing model that generates code based on natural language descriptions, etc.
[0068] In addition, the user request can also be assigned to the model with the fastest current response time according to the priority of the user request to process the user request as soon as possible.
[0069] Thus, determining the target model that is most suitable for processing the user request from multiple models can achieve an accurate match between user needs and model capabilities.
[0070] Step S202, obtain the performance metric statistical data of the target model for processing multiple user requests.
[0071] In one embodiment, the monitoring module 130 uses a performance monitoring tool (such as Prometheus, Grafana, etc.) to monitor the running performance of the model in real time, including key metrics such as response time, processing speed, accuracy, etc., and collect relevant running data of the model for processing multiple user requests, such as request volume, response time, resource occupancy, etc. It can be understood that the time in the embodiments of the present application can be a duration or a time period, etc. For example, the response time is the response duration, and the first time, the second time, the third time, the fourth time... are respectively the first duration, the second duration, the third duration, the fourth duration...
[0072] Optionally, collect the response time of the target model in processing user requests. The response time is the time required for the target model to receive a user request and return a result. It is closely related to the video memory resources allocated to the model. For the model, the video memory resources need to be large enough to accommodate the model's parameters, intermediate calculation results, input data, etc. If the video memory resources are insufficient, the model may not be able to be fully loaded and run, resulting in a too long response time or even unable to complete the task.
[0073] In addition, it is also possible to collect performance metric data such as the processing speed and accuracy of the target model in processing multiple user requests.
[0074] Furthermore, the statistics module 140 performs statistics and analysis on the performance metric data collected by the monitoring module 130 to obtain performance metric statistics data. Based on this, identify the performance bottlenecks in the process of the target model processing user requirements, such as too long response time.
[0075] The resource adjustment module 150 also formulates a resource allocation strategy for adjusting the resources used by the target model according to the performance metric statistics data, such as a video memory resource allocation strategy.
[0076] Thus, the model resource management platform 100 collects and analyzes the operation data of the target model in processing multiple user requests, obtains the performance metric statistics data of the model in processing multiple user requests, and formulates a resource allocation strategy according to the performance metric statistics data, so that the server adjusts the model deployment according to these strategies, such as adjusting the video memory resources allocated to the model, to optimize the model operation efficiency.
[0077] Step S203, adjust the video memory resources for the target model according to the performance metric statistics data.
[0078] Exemplarily, the server may receive multiple user requests in a short period of time, and the target model needs to centrally process these user requests. For example, the R & D personnel of a team may need to generate similar code structures or API interface codes for multiple projects. The code generation model needs to process multiple generation requests sent by developers at the same time, automatically generate complete code according to the developers' comments, and quickly output the required code. For multiple user requests, the response time is an important indicator to measure the model processing speed.
[0079] Optionally, if the number of instances of the currently deployed target model is greater than 1, the load distribution module 120 is used to evenly distribute the received multiple user requests to multiple instances of the target model for processing.
[0080] Thus, it is possible to avoid overloading a single instance and ensure that each instance can operate efficiently within its processing capacity.
[0081] In one embodiment, taking a preset time T (such as one day or several hours) as a period, the number of user requests processed by the target model is counted within each period. According to the number and the response time corresponding to each user request within each period, the average response time m within each period is obtained as the performance metric statistical data. Then, according to the average response time m, a video memory resource adjustment strategy is determined for the target model. These operations can be executed by the resource adjustment module 150.
[0082] Exemplarily, Figure 3a FIG. shows a schematic diagram of a video memory resource adjustment strategy provided by an embodiment of the present application.
[0083] As Figure 3a shown, first, execute step 1: Obtain the average response time of the current period, denoted as m t . Execute step 2: Determine whether m t >t1. If m t >t1, execute step 3: Increase the video memory resources for the target model. Then, by executing step 5: Let t = t + 1, and continue to execute step 1.
[0084] And, if m t ≤t1, execute step 4: Determine whether m t <t2. If m t <t2, execute step 6: Reduce the video memory resources for the target model. Then, continue to execute step 1 by executing step 5. If m t ≥t2, directly continue to execute step 1 by executing step 5.
[0085] Among them, the amount of resources increased or decreased in step 3 or step 6 is a preset size. The first time t1 and the second time t2 are usually determined by empirical values. For example, t1 = 1000ms is set based on the response time expectations of most users for real-time interaction applications. Research shows that when the response time exceeds 1000ms, users will start to feel waiting and the user experience will be affected. Therefore, setting t1 to 1000ms means that when the average response time exceeds this value, it is necessary to increase the video memory resources for the target model to improve its processing speed and response ability and ensure the smoothness of the user experience.
[0086] For another example, considering scenarios such as high-frequency trading and real-time recommendation, a small difference in response time may lead to completely different results. Setting t2 to 200ms means that when the average response time is lower than this value, the video memory resources of the target model can be appropriately reduced to improve the utilization efficiency of resources while ensuring that the response speed of the model can still meet the business requirements.
[0087] Optionally, considering the effect of reducing the response time brought by adding a model instance, set the maximum value of increasing or decreasing the preset size of the video memory resource to the minimum value G of the video memory allocation in the above initial video memory allocation. This is because increasing the video memory resource can alleviate the video memory bottleneck, improve data transmission and computing efficiency, and thus reduce the response time. However, this effect will gradually weaken as the video memory resource increases, and there is a boundary effect phenomenon.
[0088] Specifically, considering being limited by the video memory bandwidth or the processing power of the graphics card, sometimes even if the video memory resource (mainly referring to the video memory capacity) of the target model is increased, the response time requirement cannot be met. For example, even if the video memory capacity is sufficient, if the video memory bandwidth is limited, the data transmission speed will become a bottleneck. When the model performs large-scale calculations, it needs to frequently read and write data in the video memory. Insufficient video memory bandwidth will cause the data transmission speed to not keep up with the calculation speed, thus increasing the response time. At this time, simply increasing the video memory resource unilaterally cannot effectively reduce the response time, and this problem can be solved by increasing the number of instances of deploying this model.
[0089] It can be understood that in this solution, increasing the model instance not only involves increasing the video memory, but also includes other software or hardware resources required for deploying the model instance. Therefore, increasing the model instance is actually a comprehensive expansion of the video memory and other necessary software and hardware resources. For the sake of clear distinction, this solution separately describes the operation of increasing the model instance and separately increasing the video memory (i.e., increasing the video memory) for the existing model instance. Similarly, there are similar differences between reducing the model instance and reducing the video memory, which will not be elaborated here.
[0090] Exemplarily, Figure 3b FIG. shows a schematic diagram of a video memory resource adjustment strategy provided by an embodiment of the present application.
[0091] As Figure 3b shown, considering the above boundary effect phenomenon, a more detailed video memory resource adjustment strategy is provided.
[0092] First, similar to Figure 3a , perform steps 1 to 4.
[0093] Secondly, under the branch where m t > t1, after performing step 3, calculate the response time reduction rate according to the preset time, the average response time of the current cycle, and the average response time of the next cycle. If the reduction rate is less than the threshold K, and m (t+1) of the next cycle is greater than the third time t3, it indicates that the boundary effect occurs, and simply increasing the video memory resource of the target model can no longer meet the response time requirement. At this time, the number of instances of deploying the target model should be increased. The threshold K can be set to 0.1 according to experience values, and it can be reflected in the image as the slope of the average response time curve relative to time within the preset time T. Based on the above settings, when m(t+1) > At time t3, it is judged whether If so, it indicates that boundary effect occurs; otherwise, it does not occur. The third time t3 can also be set to 1000 ms considering the impact of user experience.
[0094] Specifically, in response to adding video memory resources of a preset size to the target model at the end of the current cycle, step 7 is executed: Obtain the average response time m of the next cycle (t+1) . Step 8 is executed: Judge whether m (t+1) > t3. If m (t+1) ≤t3, it indicates that the ideal response effect can be achieved by increasing the video memory. Then, step 1 is continued by executing step 5. If m (t+1) > t3, it indicates that the ideal response effect cannot be achieved by increasing the video memory. Then step 9 is executed: Judge whether boundary effect occurs. If boundary effect occurs, step 10 is executed: Increase the number of instances of the deployed target model; otherwise, step 11 is executed: Add video memory resources of the preset size again. After step 10 or 11, step 1 is continued by executing step 5. It can be understood that if there are multiple model instances in model deployment, video memory resources of the preset size are added to the most recently deployed model instance in step 3 and step 11.
[0095] Further, under the branch where m t <t1, if the average response time of a continuous preset number (such as N) of cycles is less than the fourth time t4, judge the number of instances of the currently deployed target model. The counting function of the preset number can be implemented by setting a counter with an initial value of 0. If the current number of deployed instances is greater than 1, reduce the number of instances of the deployed target model; and if the current number of deployed instances is 1 and the video memory resources currently allocated to the target model are greater than the above preset size, reduce the video memory resources of the target model. The fourth time t4 can also be set to 200 ms considering ensuring the model response speed while still meeting the business requirements.
[0096] For example, when the number of user users decreases and the received user requests decrease, and the response time m < t4 in multiple consecutive unit cycles, at this time, it is necessary to reduce the resources deployed for the model. First, try to reduce the number of model deployment instances until only one model deployment instance remains. If it is judged that the response time m when there is a unique deployed model instance is still less than t4, then reduce the deployment resources of the unique deployed model instance until the model occupies the minimum video memory G.
[0097] Specifically, if the execution result of step 4 is m t ≥t2, step 12 is executed: Clear the counter. Then, step 1 is continued by executing step 5: Let t = t + 1. If the execution result of step 4 is m t<t2, perform step 13: increment the current counter by 1. Perform step 14: determine whether the calculator ≥ N - 1. If the calculator < N - 1, then, by performing step 5: set t = t + 1, and continue to perform step 1. If the calculator ≥ N - 1, it means that the average response time for N consecutive cycles is less than the fourth time t4. Perform step 15: determine whether the number of instances of the currently deployed target model > 1. If so, perform step 16: reduce the number of instances of the deployed target model. If not, perform step 17: determine whether the video memory resources currently allocated to the target model > the preset size. If the video memory resources > the preset size, perform step 18: reduce the video memory resources for the target model. If the video memory resources ≤ the preset size, or after steps 16 and 18 are executed, perform step 12: clear the counter, and then, by performing step 5: set t = t + 1, and continue to perform step 1.
[0098] Thus, by adjusting the video memory resources of the target model, the response time of the target model can be optimized, and the user experience can be improved.
[0099] In summary, in the embodiment of the present application, multiple models for processing different tasks are deployed on the server, and the target model most suitable for processing the user request is determined from the multiple models according to the type of task requested by the user request, so as to achieve an accurate match between the user requirements and the model capabilities. And obtaining the performance index statistical data of the target model processing multiple user requests, and adjusting the video memory resources for the target model in a targeted manner to optimize its resource configuration, improve the operation efficiency and performance of the model, and improve the user experience.
[0100] Exemplarily, Figure 4 shows a schematic diagram of a model resource management process provided by an embodiment of the present application. As Figure 4 shown, this flowchart shows how the server manages and allocates resources to optimize the model performance when the user uses multiple code functions. This process mainly includes the following steps:
[0101] Step S401, receive multiple user requests for using multiple code functions.
[0102] Exemplarily, the user uses different code functions, which may include code completion, intelligent question answering, unit test generation, etc.
[0103] Step S402, statistically analyze the performance index data.
[0104] Exemplarily, the server uses the model resource management platform 100 to collect the performance index data of the model processing multiple user requests, and performs statistical analysis to obtain the performance index statistical data, so as to understand the usage frequency and resource requirements of different types of models.
[0105] Step S403, make a judgment based on the results of statistical analysis.
[0106] Exemplarily, the server makes a judgment based on the performance metric statistical data of statistical analysis to decide whether to adjust the resource allocation of the model.
[0107] Specifically, process steps S401 to S403 according to the scheme described in steps S201 to S202.
[0108] Step S404, increase the video memory resources for the model.
[0109] Exemplarily, if the judgment result shows that the current video memory resources of the target model are insufficient to meet the requirements of the user request, the server will increase resources for the model. This may involve allocating more video memory resources for the model to improve its processing ability.
[0110] Step S405, increase the model instances.
[0111] Exemplarily, if it is still judged that the requirements of the user request cannot be met after increasing the video memory resources, or in order to improve the processing ability of the target model, the server may increase the number of model instances. This means deploying more model copies to process user requests.
[0112] Specifically, according to Figure 3a or the video memory resource adjustment strategy shown in 3b, execute steps S404 and S405.
[0113] Step S406, achieve load balancing.
[0114] Exemplarily, after increasing the video memory resources or model instances, the server will use a load balancer to distribute multiple user requests to multiple instances of the target model. This can ensure that each instance will not be overloaded, while improving the overall response speed and system stability.
[0115] After that, re-enter step S401, which means that the server will continuously monitor the user usage situation and continuously adjust the resource allocation as needed. This is a dynamic adjustment process to ensure that the system can always efficiently process user requests.
[0116] Thus, system performance can be improved and user experience can be enhanced through dynamic resource management and load balancing.
[0117] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to actual situations, can be partially executed, or can be fully executed, which is not limited herein. Additionally, all or part of any feature in the above embodiments can be freely combined in any way without contradiction. The combined technical solution is also within the scope of the present application.
[0118] Exemplarily, an embodiment of the present application further provides a computing device 1000. As Figure 5 shown, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000.
[0119] The bus 1002 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 a single line is used in the figure, but it does not mean that there is only one bus or one type of bus. The bus 1004 can include a path for transmitting information between various components of the computing device 1000 (for example, the memory 1006, the processor 1004, and the communication interface 1008).
[0120] The processor 1004 can include any one or more of a central processing unit, a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a baseboard management controller, and other processors.
[0121] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0122] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 or a cluster of multiple computing devices 1000 and other devices or a communication network.
[0123] The computing device 1000 includes internal network devices, or the computing device 1000 is externally connected to multiple network devices. The internal network devices communicate with the processor 1004, the memory 1006, and the communication interface 1008 through the bus 1002, and the external network devices communicate with the computing device 1000 through interfaces such as Ethernet, Fibre Channel, and InfiniBand.
[0124] The memory 1006 stores executable program code / instructions, and the processor 1004 executes the executable program code / instructions to implement Figure 2 the process shown in [the figure], thereby implementing all or part of the steps of the method in the above embodiments. In other words, the memory 1006 stores a program / instructions for executing all or part of the steps of the method in the above embodiments.
[0125] An embodiment of the present application provides a computing device, including: a memory and a processor; the memory and the processor are coupled; the memory is used for storing a program; the processor is used for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used for executing the method in the above embodiments.
[0126] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0127] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product, and when the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0128] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0129] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0130] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application.
Claims
1. A model resource management method, characterized in that Applied to a server, where multiple models for processing different tasks are deployed on the server, the method includes: In response to receiving a user request, determine a target model from multiple models to process the user request according to the type of task requested by the user request; Obtain performance metric statistical data of the target model for processing multiple user requests; Adjust the video memory resources for the target model according to the performance metric statistical data.
2. The method according to claim 1, wherein The obtaining of the performance metric statistical data of the target model for processing multiple user requests includes: Taking a preset time as a cycle, and counting the number of user requests processed by the target model within each cycle; According to the number and the response time corresponding to each user request within each cycle, obtain the average response time within each cycle as the performance metric statistical data.
3. The method according to claim 2, characterized in that The adjusting of the video memory resources for the target model according to the performance metric statistical data includes: If the average response time is greater than a first time, increase the video memory resources for the target model; and If the average response time is less than a second time, reduce the video memory resources for the target model.
4. The method according to claim 2, wherein The adjusting of the video memory resources for the target model according to the performance metric statistical data includes: In response to increasing the video memory resources of a preset size for the target model at the end of the current cycle, calculate the response time reduction rate according to the preset time, the average response time of the current cycle, and the average response time of the next cycle; If the reduction rate is less than a threshold and the average response time of the next cycle is greater than a third time, increase the number of instances of the target model deployed.
5. The method according to claim 4, wherein The adjusting of the video memory resources for the target model according to the performance metric statistical data includes: If the average response times of a continuous preset number of cycles are all less than a fourth time, determine the current number of instances of the target model deployed; If the current number of deployed instances is greater than 1, reduce the number of instances of the target model deployed; and If the current number of deployed instances is 1 and the video memory resources currently allocated to the target model are greater than the preset size, reduce the preset size of the video memory resources for the target model.
6. The method according to claim 4, characterized in that The determining of a target model from multiple models to process the user request includes: If the current number of instances of the target model deployed is greater than 1, evenly distribute the multiple user requests to multiple instances of the target model for processing.
7. The method according to claim 4, characterized in that The preset size is determined according to the number of parameters of the target model, and the number of parameters is determined according to the task objective of processing the user request.
8. The method according to claim 1, characterized in that The determining of a target model from multiple models to process the user request includes: Designate the model specified by the user request as the target model to process the user request; or Determine a model that matches the usage scenario of the user request from multiple models as the target model to process the user request.
9. The method according to claim 1, characterized in that The type of the task includes one or more of code generation, code completion, code annotation, code modification, or code interpretation.
10. A computing device, characterized in that, Includes: Multiple memories for storing programs; A plurality of processors for executing the program stored in the memory, and when the program stored in the memory is executed, the processors are used to execute the method according to any one of claims 1-9.