Model scheduling method and device based on multi-level cache, equipment and medium
Through the multi-level cache architecture and model popularity scheduling method, the problem of limited video memory resources during model deployment is solved, the model loading efficiency and resource utilization are improved, and intelligent management is realized.
Patent Information
- Application Number
- CN202510482562.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology has high model loading delay and resource waste caused by limited video memory resources during the model deployment process. Especially when large-scale model deployment, video memory resources are quickly full, and process memory management lacks resource sharing mechanism, resulting in cumbersome loading of models and low memory utilization.
The multi-level cache architecture is adopted, including the video memory cache layer, the process memory cache layer, the shared memory cache layer and the persistent cache layer. The model popularity is updated through model access requests, the model is reasonably scheduled to the appropriate cache layer for inference operations, and the model popularity is used for model scheduling and cache management.
This improves model loading efficiency, reduces model loading delay, optimizes resource utilization, and realizes efficient management and intelligent scheduling of model resources.
Smart Images

Figure CN120353553A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a model scheduling method, device, equipment and medium based on multi-level cache. Background Art
[0002] In today's era of booming artificial intelligence, AI models are increasingly being used in various fields. However, in the process of model deployment, there are many severe technical challenges.
[0003] From the perspective of video memory cache, frameworks such as TensorRT (i.e. Turing Tensor R-Engine, a high-performance deep learning inference optimizer and runtime acceleration library) use a cache mechanism. When faced with large-scale model deployment, limited video memory resources will soon be occupied. When a new model is loaded, the original model often needs to be replaced, and subsequent inference requests for these models have to be reloaded from the disk, resulting in a significant increase in loading delays. In terms of process memory management, the process-level model cache of frameworks such as Python memory pool and Torch (a deep learning framework) uses a simple memory key-value pair structure. When a model inference request arrives, if there is no target model in the process memory, it must first be loaded from the disk to the memory and then sent to the video memory, which is a cumbersome process. In addition, there is a lack of resource sharing mechanism between different processes, and the same model is frequently loaded repeatedly across processes, resulting in resource waste and extremely low memory utilization. However, when it comes to model deployment, shared memory technology fails to meet the characteristics and requirements of AI models, and it is difficult to effectively solve key problems such as tight video memory, high loading latency, and memory waste in large-scale model deployment.
[0004] Therefore, how to avoid inefficiencies in model loading and resource allocation due to limited video memory resources during model deployment is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a model scheduling method, device, equipment and medium based on multi-level cache, which can avoid the inefficiency of model loading and resource allocation caused by limited video memory resources during model deployment. The specific scheme is as follows:
[0006] In a first aspect, the present application provides a model scheduling method based on multi-level cache, comprising:
[0007] During the model deployment process, the current model heat of the corresponding model to be inferred in the preset cache architecture is updated based on the obtained model access request; wherein the cache layers of the preset cache architecture from top to bottom are respectively the video memory cache layer, the process memory cache layer, the shared memory cache layer and the persistent cache layer; the model heat is used to indicate the frequency of the model being accessed within a preset time period;
[0008] Determine the target cache layer of the to-be-inferred model in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger a preset inference operation on the to-be-inferred model;
[0009] When the target cache layer is not the video memory cache layer, schedule the to-be-inferred model to the video memory cache layer based on the current model heat in other cache layers above the target cache layer, so as to trigger the preset inference operation on the to-be-inferred model in the video memory cache layer.
[0010] Optionally, updating the current model heat of the corresponding to-be-inferred model in the preset cache architecture based on the obtained model access request includes:
[0011] Obtain the model access request sent by the client, and determine the to-be-inferred model in the preset cache architecture based on the model access request;
[0012] Determine the time difference based on the current time and the model access time corresponding to the model access request, and use the time difference and the heat decay period to determine the heat decay exponent;
[0013] Update the current model heat of the to-be-inferred model by using a preset heat decay factor, the heat decay exponent, the size of the to-be-inferred model, and the target model access times;
[0014] Wherein, the target model access times are the number of times the to-be-inferred model is accessed within the preset time period, and the target model access times do not include the model access times corresponding to the current model access request.
[0015] Optionally, when the target cache layer is not the video memory cache layer, scheduling the to-be-inferred model to the video memory cache layer based on the current model heat in other cache layers above the target cache layer includes:
[0016] When the target cache layer is the process memory cache layer, determine the current model heat in the video memory cache layer, and arrange the current model heat in the video memory cache layer in ascending order of model heat to obtain a first heat arrangement result;
[0017] Determine the to-be-evicted model in the video memory cache layer based on the first heat arrangement result, and perform a first preset model eviction operation on the to-be-evicted model in the video memory cache layer to obtain the video memory cache layer after the model is evicted;
[0018] Schedule the to-be-inferred model to the video memory cache layer after the model is evicted.
[0019] Optionally, when the target cache layer is not the video memory cache layer, scheduling the model to be inferred to the video memory cache layer based on the current model popularity in other cache layers above the target cache layer includes:
[0020] When the target cache layer where the target model data corresponding to the model to be inferred is located is the shared memory cache layer, constructing the model to be inferred based on the target model data through a file lock mechanism, a memoryview function, and a preset deserialization operation;
[0021] Determining the current model popularity in the process memory cache layer, comparing the current model popularity of the model to be inferred with the current model popularity in the process memory cache layer, and obtaining a first comparison result;
[0022] If the first comparison result indicates that there is a model in the process memory cache layer whose current model popularity is not greater than that of the model to be inferred, arranging the models in the process memory cache layer whose current model popularity is not greater than that of the model to be inferred in ascending order of model popularity, and obtaining a second popularity arrangement result;
[0023] Determining the model to be evicted in the process memory cache layer based on the second popularity arrangement result, and performing a second preset model eviction operation on the model to be evicted in the process memory cache layer to obtain the process memory cache layer after evicting the model;
[0024] Scheduling the model to be inferred to the process memory cache layer after evicting the model, and jumping to the step of determining the current model popularity in the video memory cache layer until the model to be inferred is scheduled to the video memory cache layer after evicting the model.
[0025] Optionally, the constructing the model to be inferred based on the target model data through a file lock mechanism, a memoryview function, and a preset deserialization operation includes:
[0026] Obtaining the target storage information of the target model data in the shared memory cache layer based on the file lock mechanism;
[0027] Locating the position of the target model data using the target storage information and the memoryview function, and obtaining the target memory block where the target model data is located;
[0028] Reading the target model data from the target memory block based on a preset byte stream reading method, and converting the target model data into the model to be inferred using the preset deserialization operation.
[0029] Optionally, when the target cache layer is not the video memory cache layer, scheduling the model to be inferred to the video memory cache layer based on the current model popularity in other cache layers above the target cache layer includes:
[0030] When the model to be inferred is not in the video memory cache layer, the process memory cache layer, and the shared memory cache layer, under a preset deep learning framework, create a target model class, and construct the model to be inferred based on the target model class and the model weights in the persistent cache layer;
[0031] Compare the current model popularity of the model to be inferred with the current model popularity corresponding to each model data in the shared memory cache layer to obtain a second comparison result;
[0032] If the second comparison result indicates that there is model data in the shared memory cache layer whose current model popularity is not greater than that of the model to be inferred, then arrange the current model popularity corresponding to each model data in the shared memory cache layer in ascending order of model popularity to obtain a third popularity arrangement result;
[0033] Determine the model data to be evicted in the shared memory cache layer based on the third popularity arrangement result, and perform a preset model data eviction operation on the model data to be evicted to obtain the shared memory cache layer after the model data is evicted;
[0034] Schedule the model to be inferred to the shared memory cache layer after the model data is evicted, and jump to the step of determining the current model popularity in the process memory cache layer until the model to be inferred is scheduled to the video memory cache layer after the model is evicted.
[0035] In a second aspect, the present application provides a model scheduling device based on a multi-level cache, including:
[0036] A popularity update module, configured to update the current model popularity of a model to be inferred corresponding in a preset cache architecture based on an obtained model access request during the model deployment process; wherein, the preset cache architecture from top to bottom cache layers are a video memory cache layer, a process memory cache layer, a shared memory cache layer, and a persistent cache layer; the model popularity is used to represent the frequency of a model being accessed within a preset time period;
[0037] A first model inference module, configured to determine the target cache layer of the model to be inferred in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger a preset inference operation on the model to be inferred;
[0038] A second model inference module, configured to, when the target cache layer is not the video memory cache layer, schedule the model to be inferred to the video memory cache layer based on the current model popularity in other cache layers above the target cache layer, so as to trigger the preset inference operation on the model to be inferred in the video memory cache layer.
[0039] In a third aspect, the present application provides an electronic device, including:
[0040] A memory, configured to store a computer program;
[0041] A processor, configured to execute the computer program to implement the foregoing model scheduling method based on multi-level caches.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the foregoing model scheduling method based on multi-level caches is implemented.
[0043] In this application, during the process of model deployment, the current model popularity of the to-be-inferred model corresponding in the preset cache architecture is updated based on the obtained model access request; wherein, the cache layers of the preset cache architecture from top to bottom are the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer; the model popularity is used to represent the frequency of the model being accessed within a preset time period; determine the target cache layer of the to-be-inferred model in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger the preset inference operation for the to-be-inferred model; when the target cache layer is not the video memory cache layer, schedule the to-be-inferred model to the video memory cache layer based on the current model popularities in other cache layers above the target cache layer, so as to trigger the preset inference operation for the to-be-inferred model in the video memory cache layer. As can be seen from the above, during model deployment in this application, according to the obtained model access request, the current model popularity of the to-be-inferred model corresponding in the preset cache architecture is updated. Here, the preset cache architecture, its cache layers are sequentially the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer from top to bottom. And the model popularity refers to the frequency of the model being accessed within a preset time period. Next, determine the target cache layer of the to-be-inferred model in the preset cache architecture. If the target cache layer is the video memory cache layer, trigger the preset inference operation for the to-be-inferred model. If the target cache layer is not the video memory cache layer, then according to the current model popularities of each model in other cache layers above the target cache layer, schedule the to-be-inferred model to the video memory cache layer, and then trigger the preset inference operation for the to-be-inferred model in the video memory cache layer. In this way, this application can avoid inefficient problems in aspects such as model loading and resource allocation caused by limited video memory resources during the model deployment process, thereby achieving efficient utilization and intelligent management of model resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0045] Figure 1 It is a flowchart of a model scheduling method based on multi-level caching disclosed in this application;
[0046] Figure 2 It is a schematic diagram of a preset cache architecture in a model scheduling based on multi-level caching disclosed in this application;
[0047] Figure 3Flowchart of a specific model scheduling method based on multi - level caches disclosed in this application;
[0048] Figure 4 Schematic structural diagram of a model scheduling device based on multi - level caches disclosed in this application;
[0049] Figure 5 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] Currently, during the model deployment process, many severe technical challenges are faced. From the perspective of video memory caching, frameworks such as TensorRT adopt a caching mechanism. When facing large - scale model deployment, limited video memory resources will soon be occupied. When a new model is loaded, the original model often needs to be replaced. Subsequently, for inference requests for these models, they need to be re - loaded from the disk again, resulting in a significant increase in loading latency. In terms of process memory management, the process of the Python memory pool and the process - level model caching of frameworks such as Torch is cumbersome. Moreover, there is a lack of resource sharing mechanism between different processes, and the situation where the same model is repeatedly loaded across processes frequently occurs, causing resource waste and extremely low memory utilization. And the shared memory technology fails to fit the characteristics and requirements of AI models when facing model deployment, and it is difficult to effectively solve key problems such as video memory tension, high loading latency, and memory waste in large - scale model deployment. Therefore, this application provides a model scheduling method, device, device, and medium based on multi - level caches, which can avoid inefficient problems in aspects such as model loading and resource allocation caused by limited video memory resources during the model deployment process.
[0052] See Figure 1 As shown, the embodiments of the present invention disclose a model scheduling method based on multi - level caches, including:
[0053] Step S11, during the model deployment process, update the current model popularity of the to - be - inferred model corresponding in the preset cache architecture based on the obtained model access request; wherein, the preset cache architecture from top to bottom cache layers are the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer; the model popularity is used to represent the frequency of the model being accessed within a preset time period.
[0054] In this embodiment, first, during the model deployment process, a model access request sent by the user end is obtained. The user end includes but is not limited to an artificial intelligence server cluster, an intelligent terminal device, etc. The user end will send a model access request containing a model reasoning task according to its own business needs. After receiving the model access request, the model to be reasoned is determined in the preset cache architecture based on the relevant information in the model access request, such as the model identifier, task type, etc.
[0055] It should be noted that see Figure 2 As shown in the figure, in the preset cache architecture, the video memory cache layer is at the top layer, where the model objects can be directly used for reasoning. Its capacity depends on the graphics card specifications. The common range for reasoning ranges from 12-40GB (i.e., Gigabyte). The higher the graphics card specifications, the larger the corresponding video memory capacity, and the higher the graphics card cost. For example, in the model reasoning scenario of a large game server, in order to ensure that models such as real-time rendering and character behavior prediction in the game can be quickly reasoned, it is often necessary to configure a graphics card with a high video memory capacity to directly reason about the relevant models contained in the video memory cache layer. The process memory cache layer is below the video memory cache layer. The model objects in this layer need to be switched to the video memory before reasoning the model. Its capacity can be determined by the host memory configuration, such as 50-100GB. The shared memory cache layer is below the process memory cache layer. The model objects in this layer need to be deserialized from the shared memory to the video memory before reasoning the model. The capacity can also be determined by the host memory configuration, such as 200-300GB. The lowest layer of the persistent cache layer (i.e. Figure 2 The persistent model storage layer in the GPU stores model weights. Before reasoning about the model, you need to use the weight information to build a model object and load the model into the graphics memory cache layer.
[0056] After determining the model to be inferred, the time difference is determined based on the current time and the model access time corresponding to the model access request. The current time refers to the time point when the system receives the model access request, and the model access time is the time point when the model to be inferred was last accessed. The time difference is obtained by obtaining these two time points and calculating the difference between them. Then, the heat decay index is determined using the time difference and the heat decay period. The heat decay period is a time interval pre-set according to the actual request scenario, for example, it can be set to half an hour, one hour, etc., and can also be automatically adjusted according to indicators such as cache hit rate during system operation. Divide the time difference by the heat decay period, and the result is the heat decay index.
[0057] Then, the current model popularity of the model to be inferred is updated using a preset popularity decay factor, a popularity decay exponent, the size of the model to be inferred, and the number of accesses to the target model. The preset popularity decay factor is a value greater than zero and less than one, such as 0.8 or 0.95, etc., and can also be set according to the actual request scenario or automatically adjusted according to the cache hit rate. The number of accesses to the target model is the number of accesses to the model to be inferred within a preset time period, and this number of accesses to the target model does not include the number of model accesses corresponding to the current model access request. The specific formula for updating the model popularity is: Model Popularity = Number of Accesses to the Target Model * Preset Popularity Decay Factor ^ (Popularity Decay Exponent) / Size of the Model to be Inferred. Through this formula, the current model popularity of the model to be inferred can be updated based on various factors such as the model's access history, time decay factor, and the size of the model itself, providing a certain basis for subsequent model scheduling and cache management.
[0058] Step S12: Determine the target cache layer of the model to be inferred in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger a preset inference operation on the model to be inferred.
[0059] In this embodiment, after completing the update of the current model popularity of the model to be inferred in the preset cache architecture, it is next necessary to determine the target cache layer of the model to be inferred in the preset cache architecture to ensure that the model can be reasonably scheduled and used.
[0060] When it is determined that the target cache layer is the video memory cache layer, a preset inference operation on the model to be inferred will be immediately triggered. Before performing the inference operation, it is necessary to ensure that the model to be inferred has been correctly loaded into the video memory and that the relevant operating environment and parameters have been configured. The preset inference operation will be executed according to specific business requirements and the characteristics of the model.
[0061] Taking the field of intelligent security as an example, when a model for analyzing the behavior of people in real-time monitoring images is deployed, during the process of determining the target cache layer, if the target cache layer of the model is the video memory cache layer, the inference operation will be immediately triggered. When the target cache layer is the video memory cache layer, a preset inference operation on the model to be inferred will be immediately triggered. In the inference scenario of the real-time captured images by the security monitoring camera, before performing the inference operation, it is necessary to ensure that the model for analyzing the behavior of people has been correctly loaded into the video memory and that the relevant operating environment and parameters for identifying action types and analyzing abnormal behaviors have been configured. The preset inference operation will be executed according to specific business requirements and the characteristics of the model. For example, in the video stream inference task in the field of security, the model needs to identify abnormal behaviors such as running and fighting based on features such as the human body posture and movement trajectory in the image.
[0062] In addition, during the inference process, the usage of the video memory and the performance metrics of the inference can be monitored in real time to ensure that the inference operation can be carried out efficiently and stably. If abnormal situations such as insufficient video memory occur during the inference process, corresponding measures can be taken, such as pausing the inference, releasing some video memory, etc., to ensure the normal operation of the entire inference process. In this way, the advantages of the video memory cache layer can be fully utilized to improve the efficiency of model inference and provide users with faster and more accurate services.
[0063] Step S13: When the target cache layer is not the video memory cache layer, schedule the model to be inferred to the video memory cache layer based on the current model popularity in other cache layers above the target cache layer, so as to trigger the preset inference operation on the model to be inferred in the video memory cache layer.
[0064] In this embodiment, it should be noted first that when the target cache layer of the model to be inferred is not the video memory cache layer, it needs to be scheduled to the video memory cache layer so that the preset inference operation can be efficiently executed. That is to say, in the preset cache architecture, in addition to the video memory cache layer that can directly execute model inference, for the remaining three cache layers, the model to be inferred needs to be scheduled to the video memory cache layer according to the situation of each cache layer.
[0065] In a specific implementation manner, when the target cache layer is the process memory cache layer, determine the current model popularity in the video memory cache layer, and arrange the current model popularity in the video memory cache layer in ascending order of model popularity to obtain the first popularity arrangement result. And it can be understood that the purpose of this step is to find the models with lower popularity in the video memory cache layer to make room for the model to be inferred.
[0066] Furthermore, determine the model to be evicted in the video memory cache layer based on the first popularity arrangement result, and perform the first preset model eviction operation on the model to be evicted in the video memory cache layer. For example, it can be removed from the video memory cache layer and stored in other cache layers or persistent storage as needed to obtain the video memory cache layer after evicting the model. Finally, schedule the model to be inferred to the video memory cache layer after evicting the model, so as to trigger the preset inference operation on the model to be inferred in the video memory cache layer.
[0067] In another specific implementation, when the target cache layer where the target model data corresponding to the model to be inferred is located is the shared memory cache layer, the model to be inferred is constructed based on the target model data through a file lock mechanism, the memoryview function (i.e., a function in Python that accesses and operates on the memory buffer of an object in a zero-copy manner), and a preset deserialization operation. Specifically, since there may be concurrent read and write situations in the shared memory cache layer, in order to avoid shared memory pollution, the target storage information of the target model data in the shared memory cache layer is obtained based on the file lock mechanism.
[0068] Next, the memoryview function is used with the target storage information to locate the position of the target model data, and the target memory block where the target model data is located is obtained. The memoryview function provides a zero-copy way to access memory, without first reading the byte data of the specified memory block in the shared memory into the process memory, thus improving the efficiency of data access. Then, the target model data is read from the target memory block based on a preset byte stream reading method, and the target model data is converted into the model to be inferred using the preset deserialization operation.
[0069] Furthermore, the current model heat of each model in the process memory cache layer is determined, the current model heat of the model to be inferred is compared with the current model heat of each model in the process memory cache layer, and a first comparison result is obtained. If the first comparison result indicates that there is a model in the process memory cache layer whose current model heat is not greater than the current model heat of the model to be inferred, then the models in the process memory cache layer whose current model heat is not greater than the current model heat of the model to be inferred are arranged in ascending order of model heat, and a second heat arrangement result is obtained.
[0070] Finally, the model to be evicted in the process memory cache layer is determined based on the second heat arrangement result, and a second preset model eviction operation is performed on the model to be evicted in the process memory cache layer to obtain the process memory cache layer after the model is evicted. The model to be inferred is scheduled to the process memory cache layer after the model is evicted, and the process jumps to the step of determining the current model heat of each model in the video memory cache layer until the model to be inferred is scheduled to the video memory cache layer after the model is evicted.
[0071] In a third specific implementation, when the model to be inferred is not in the video memory cache layer, the process memory cache layer, and the shared memory cache layer, a target model class is created under a preset deep learning framework, and the model to be inferred is constructed based on the target model class and the model weights in the persistent cache layer.
[0072] Further, compare the current model popularity of the to-be-inferred model with the current model popularities corresponding to each model data in the shared memory cache layer to obtain a second comparison result. If the second comparison result indicates that there is model data in the shared memory cache layer whose current model popularity is not greater than that of the to-be-inferred model, then arrange the current model popularities corresponding to each model data in the shared memory cache layer in ascending order of model popularity to obtain a third popularity arrangement result.
[0073] Next, determine the model data to be evicted in the shared memory cache layer based on the third popularity arrangement result, and perform a preset model data eviction operation on the model data to be evicted to obtain the shared memory cache layer after evicting the model data. Then, schedule the to-be-inferred model to the shared memory cache layer after evicting the model data, and jump to the step of determining the current model popularities in the process memory cache layer until the to-be-inferred model is scheduled to the video memory cache layer after evicting the model.
[0074] As can be seen from the above, during model deployment, this application updates the current model popularity of the to-be-inferred model corresponding in a preset cache architecture according to the obtained model access request. Here, for the preset cache architecture, its cache layers are, in order from top to bottom, the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer. And the model popularity refers to the frequency of the model being accessed within a preset time period. Next, determine the target cache layer of the to-be-inferred model in the preset cache architecture. If the target cache layer is the video memory cache layer, trigger the preset inference operation on the to-be-inferred model. If the target cache layer is not the video memory cache layer, then schedule the to-be-inferred model to the video memory cache layer according to the current model popularities of each model in other cache layers above the target cache layer, and then trigger the preset inference operation on the to-be-inferred model in the video memory cache layer. In this way, this application can avoid inefficient problems such as model loading and resource allocation caused by limited video memory resources during model deployment, thereby achieving efficient utilization and intelligent management of model resources.
[0075] Next, in combination with Figure 3 the schematic diagram shown, and in combination with a specific application scenario, the technical solution of the embodiment of this application will be specifically described.
[0076] In a specific implementation, in the product recommendation system of an e-commerce platform, a model (i.e., the model to be inferred) is required to predict the products that users may be interested in. These models will analyze and recommend based on data such as users' browsing history and purchase records. Specifically, first, based on the access situation of the model within a preset time period, combined with parameters such as the heat decay factor and decay interval, the model heat of the model is updated. For example, during the shopping peak period, the frequency of users accessing product recommendations increases significantly, and the number of accesses to the models related to popular product recommendations increases, so the heat of its model will rise; while during the off-peak period, the heat of some models will gradually decrease over time according to the decay factor.
[0077] After updating the model heat, first determine whether the model is in the video memory. If the model is in the video memory cache layer, since the model in the video memory cache layer can be directly used for inference, the inference operation will be directly triggered. That is to say, if the product recommendation model for a certain popular category is in the video memory cache layer, since the model in the video memory can be directly used for inference, the system will immediately trigger the inference operation to quickly recommend relevant products to users, greatly shortening the user waiting time.
[0078] If the model is not in the video memory cache layer, further determine whether it is in the process memory cache layer. Suppose the recommendation model for a certain niche product is in the process memory. Since the model in this layer needs to be switched to the video memory before inference, the system will calculate the heat of each model in the video memory and sort them from low to high. If the video memory space is insufficient, the models with the lowest heat will be evicted in turn, such as the product recommendation models for some outdated promotional activities, to free up space for the niche product recommendation model, and then schedule it to the video memory for inference.
[0079] If the model is not in the process memory cache layer, continue to determine whether it is in the shared memory cache layer. When the recommendation model for a newly listed product is in the shared memory, the system will first ensure concurrent access security through the file lock mechanism, and use the memoryview function and preset deserialization operations to construct the model based on the target model data. If the heat of the new product recommendation model is higher than that of some models in the process memory, the models with low heat in the process memory will be sorted from low to high and evicted to free up space for the new product recommendation model, schedule it to the process memory, and then try to schedule it to the video memory.
[0080] If the model is not in the shared memory cache layer, it is necessary to load the model weights from the disk (i.e., the persistent cache layer) and construct the model. For example, the newly developed product recommendation algorithm model is initially stored on the disk. After loading, compare its heat with the models in the shared memory. If the heat is high, evict the models with low heat in the shared memory, schedule the new model to the shared memory, and then gradually schedule it to the video memory through the process memory.
[0081] In another specific embodiment, in an intelligent security monitoring system, a model is used to analyze the monitoring video, identify abnormal behaviors, people, or objects. Similarly, the model popularity is updated according to the access situation of the model within a preset time period. For example, during the day when there is a large flow of people, the access frequency of the face recognition model is high and its popularity increases; while at midnight, the perimeter intrusion detection model may be more frequently used and its popularity increases.
[0082] After updating the popularity, if the face recognition model is in the video memory, the system will directly trigger the inference operation to quickly identify the identity of the people in the monitoring video.
[0083] If the abnormal behavior detection model for a specific scenario is in the process memory, when the video memory space is insufficient according to the popularity of the models in the video memory, the model with low popularity will be evicted and the abnormal behavior detection model will be scheduled to the video memory.
[0084] When a new object recognition model is in the shared memory, the model is first constructed through the corresponding mechanism. If its popularity is higher than that of some models in the process memory, the models with low popularity in the process memory will be evicted, and the new model will be scheduled to the process memory and then to the video memory.
[0085] If the model is not in the shared memory, the model weights are loaded from the disk. For example, for a newly deployed intelligent analysis algorithm model, after loading, according to the comparison result of the popularity with the models in the shared memory, it is determined whether to evict the models with low popularity in the shared memory, and the new model is gradually scheduled to the video memory.
[0086] Through such a progressive scheduling process based on model popularity and cache hierarchy, this solution can effectively improve the model inference efficiency, reduce the model loading latency, and optimize the overall performance of the system.
[0087] Correspondingly, as shown in Figure 4 This application embodiment provides a model scheduling device based on a multi-level cache, including:
[0088] A popularity update module 11, configured to update the current model popularity of the to-be-inferred model corresponding in a preset cache architecture based on the obtained model access request during the model deployment process; wherein, the preset cache architecture from top to bottom cache layers are the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer; the model popularity is used to represent the frequency of the model being accessed within a preset time period;
[0089] A first model inference module 12, configured to determine the target cache layer of the to-be-inferred model in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger a preset inference operation on the to-be-inferred model;
[0090] The second model inference module 13 is configured to, when the target cache layer is not the video memory cache layer, schedule the model to be inferred to the video memory cache layer based on the current model heat values in other cache layers above the target cache layer, so as to trigger the preset inference operation on the model to be inferred in the video memory cache layer.
[0091] As can be seen from the above, during model deployment, according to the obtained model access request, the present application updates the current model heat value of the model to be inferred corresponding in the preset cache architecture. Here, for the preset cache architecture, its cache layers are, in order from top to bottom, the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer. And the model heat value refers to the frequency of the model being accessed within a preset time period. Next, the target cache layer of the model to be inferred in the preset cache architecture is determined. If the target cache layer is the video memory cache layer, the preset inference operation on the model to be inferred is triggered. If the target cache layer is not the video memory cache layer, then according to the current model heat values of each model in other cache layers above the target cache layer, the model to be inferred is scheduled to the video memory cache layer, and then the preset inference operation on the model to be inferred is triggered in the video memory cache layer. In this way, the present application can avoid inefficient problems in aspects such as model loading and resource allocation caused by limited video memory resources during the model deployment process, thereby achieving efficient utilization and intelligent management of model resources.
[0092] In some specific embodiments, the heat value update module 11 specifically includes:
[0093] A model determination unit, configured to obtain a model access request sent by a user side, and determine the model to be inferred in the preset cache architecture based on the model access request;
[0094] An exponent determination unit, configured to determine a time difference based on the current time and the model access time corresponding to the model access request, and determine a heat value decay exponent by using the time difference and the heat value decay period;
[0095] A heat value update unit, configured to update the current model heat value of the model to be inferred by using a preset heat value decay factor, the heat value decay exponent, the size of the model to be inferred, and the target model access times;
[0096] Wherein, the target model access times are the number of times the model to be inferred is accessed within the preset time period, and the target model access times do not include the model access times corresponding to the current model access request.
[0097] In some specific embodiments, the second model inference module 13 specifically includes:
[0098] A first heat ranking unit, configured to determine the current model heat in the video memory cache layer when the target cache layer is the process memory cache layer, and rank the current model heat in the video memory cache layer in ascending order of model heat to obtain a first heat ranking result;
[0099] A first model eviction unit, configured to determine the models to be evicted in the video memory cache layer based on the first heat ranking result, and perform a first preset model eviction operation on the models to be evicted in the video memory cache layer to obtain the video memory cache layer after evicting the models;
[0100] A first model scheduling unit, configured to schedule the model to be inferred to the video memory cache layer after evicting the models.
[0101] In some specific embodiments, the second model inference module 13 specifically includes:
[0102] A first model construction unit, configured to construct the model to be inferred based on the target model data through a file lock mechanism, a memoryview function, and a preset deserialization operation when the target cache layer where the target model data corresponding to the model to be inferred is located is the shared memory cache layer;
[0103] A first heat comparison unit, configured to determine the current model heat in the process memory cache layer, compare the current model heat of the model to be inferred with the current model heat in the process memory cache layer, and obtain a first comparison result;
[0104] A second heat ranking unit, configured to rank the models in the process memory cache layer whose current model heat is not greater than the current model heat of the model to be inferred in ascending order of model heat if the first comparison result indicates that there are models in the process memory cache layer whose current model heat is not greater than the current model heat of the model to be inferred, and obtain a second heat ranking result;
[0105] A second model eviction unit, configured to determine the models to be evicted in the process memory cache layer based on the second heat ranking result, and perform a second preset model eviction operation on the models to be evicted in the process memory cache layer to obtain the process memory cache layer after evicting the models;
[0106] A second model scheduling unit, configured to schedule the model to be inferred to the process memory cache layer after evicting the models, and jump to the step of determining the current model heat in the video memory cache layer until the model to be inferred is scheduled to the video memory cache layer after evicting the models.
[0107] In some specific embodiments, the model construction unit specifically includes:
[0108] An information acquisition subunit, configured to obtain target storage information of the target model data in the shared memory cache layer based on a file lock mechanism;
[0109] A memory block acquisition subunit, configured to locate the position of the target model data by using the target storage information and the memoryview function, and obtain a target memory block where the target model data is located;
[0110] A data conversion subunit, configured to read the target model data from the target memory block based on a preset byte stream reading method, and convert the target model data into the model to be inferred by using the preset deserialization operation.
[0111] In some specific embodiments, the second model inference module 13 specifically includes:
[0112] A second model construction unit, configured to create a target model class under a preset deep learning framework and construct the model to be inferred based on the target model class and model weights in the persistent cache layer when the model to be inferred is not in the video memory cache layer, the process memory cache layer, and the shared memory cache layer;
[0113] A second heat comparison unit, configured to compare the current model heat of the model to be inferred with the current model heats corresponding to each model data in the shared memory cache layer to obtain a second comparison result;
[0114] A third heat sorting unit, configured to sort the current model heats corresponding to each model data in the shared memory cache layer in ascending order of model heat if the second comparison result indicates that there is model data in the shared memory cache layer whose current model heat is not greater than the current model heat of the model to be inferred, to obtain a third heat sorting result;
[0115] A model data eviction unit, configured to determine model data to be evicted in the shared memory cache layer based on the third heat sorting result, and perform a preset model data eviction operation on the model data to be evicted to obtain the shared memory cache layer after the model data is evicted;
[0116] A third model scheduling unit, configured to schedule the model to be inferred to the shared memory cache layer after the model data is evicted, and jump to the step of determining the current model heats in the process memory cache layer until the model to be inferred is scheduled to the video memory cache layer after the model is evicted.
[0117] Further, the embodiments of the present application also disclose an electronic device. Figure 5 FIG. 20 is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the model scheduling method based on multi-level caches disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0118] In this embodiment, the power supply 23 is used to provide working voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data loading channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0119] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be transient storage or permanent storage.
[0120] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the model scheduling method based on multi-level caches executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.
[0121] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the model scheduling method based on multi-level caches disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0122] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0123] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0124] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0125] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0126] The above has introduced the technical solutions provided by this application in detail. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A model scheduling method based on multi-level caching, characterized in that, Including: During the process of model deployment, update the current model popularity of the to-be-inferred model corresponding in the preset cache architecture based on the obtained model access request; wherein, the preset cache architecture from top to bottom has cache layers as the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer; the model popularity is used to represent the frequency of the model being accessed within a preset time period; Determine the target cache layer of the to-be-inferred model in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger the preset inference operation for the to-be-inferred model; When the target cache layer is not the video memory cache layer, schedule the to-be-inferred model to the video memory cache layer based on the current model popularities in other cache layers above the target cache layer, so as to trigger the preset inference operation for the to-be-inferred model in the video memory cache layer.
2. The model scheduling method based on multi-level caches according to claim 1, wherein The updating the current model popularity of the to-be-inferred model corresponding in the preset cache architecture based on the obtained model access request includes: Obtain the model access request sent by the client, and determine the to-be-inferred model in the preset cache architecture based on the model access request; Determine the time difference based on the current time and the model access time corresponding to the model access request, and use the time difference and the heat decay period to determine the heat decay exponent; Update the current model popularity of the to-be-inferred model by using a preset heat decay factor, the heat decay exponent, the size of the to-be-inferred model, and the target model access times; Wherein, the target model access times are the number of times the to-be-inferred model is accessed within the preset time period, and the target model access times do not include the model access times corresponding to the current model access request.
3. The model scheduling method based on multi-level caches according to claim 1, wherein When the target cache layer is not the video memory cache layer, the scheduling the to-be-inferred model to the video memory cache layer based on the current model popularities in other cache layers above the target cache layer includes: When the target cache layer is the process memory cache layer, determine the current model popularities in the video memory cache layer, and arrange the current model popularities in the video memory cache layer in ascending order of model popularity to obtain the first popularity arrangement result; Determine the to-be-evicted model in the video memory cache layer based on the first popularity arrangement result, and perform a first preset model eviction operation on the to-be-evicted model in the video memory cache layer to obtain the video memory cache layer after the model is evicted; Schedule the to-be-inferred model to the video memory cache layer after the model is evicted.
4. The model scheduling method based on multi-level cache according to claim 3, characterized in that, When the target cache layer is not the video memory cache layer, the scheduling the to-be-inferred model to the video memory cache layer based on the current model popularities in other cache layers above the target cache layer includes: When the target cache layer where the target model data corresponding to the to-be-inferred model is located is the shared memory cache layer, construct the to-be-inferred model based on the target model data through a file lock mechanism, a memoryview function, and a preset deserialization operation; Determine the current model popularity in the process memory cache layer, compare the current model popularity of the model to be inferred with the current model popularity in the process memory cache layer, and obtain a first comparison result; If the first comparison result indicates that there is a model in the process memory cache layer whose current model popularity is not greater than that of the model to be inferred, then arrange the models in the process memory cache layer whose current model popularity is not greater than that of the model to be inferred in ascending order of model popularity, and obtain a second popularity arrangement result; Based on the second popularity arrangement result, determine the model to be evicted in the process memory cache layer, and perform a second preset model eviction operation on the model to be evicted in the process memory cache layer to obtain the process memory cache layer after the model is evicted; Schedule the model to be inferred to the process memory cache layer after the model is evicted, and jump to the step of determining the current model popularity in the video memory cache layer until the model to be inferred is scheduled to the video memory cache layer after the model is evicted.
5. The model scheduling method based on multi-level caching according to claim 4, wherein, The constructing of the model to be inferred based on the target model data through the file lock mechanism, the memoryview function, and the preset deserialization operation includes: Obtain the target storage information of the target model data in the shared memory cache layer based on the file lock mechanism; Use the target storage information and the memoryview function to locate the position of the target model data, and obtain the target memory block where the target model data is located; Read the target model data from the target memory block based on the preset byte stream reading method, and use the preset deserialization operation to convert the target model data into the model to be inferred.
6. The model scheduling method based on multi-level caches according to claim 4, wherein, When the target cache layer is not the video memory cache layer, scheduling the model to be inferred to the video memory cache layer based on the current model popularity in other cache layers above the target cache layer includes: When the model to be inferred is not in the video memory cache layer, the process memory cache layer, and the shared memory cache layer, then in a preset deep learning framework, create a target model class, and construct the model to be inferred based on the target model class and the model weights in the persistent cache layer; Compare the current model popularity of the model to be inferred with the current model popularity corresponding to each model data in the shared memory cache layer to obtain a second comparison result; If the second comparison result indicates that there is model data in the shared memory cache layer whose current model popularity is not greater than that of the model to be inferred, then arrange the current model popularity corresponding to each model data in the shared memory cache layer in ascending order of model popularity to obtain a third popularity arrangement result; Based on the third popularity arrangement result, determine the model data to be evicted in the shared memory cache layer, and perform a preset model data eviction operation on the model data to be evicted to obtain the shared memory cache layer after the model data is evicted; Schedule the to-be-inferred model to the shared memory cache layer after evicting the model data, and jump to the step of determining the current model heat in the process memory cache layer until the to-be-inferred model is scheduled to the video memory cache layer after evicting the model.
7. A model scheduling device based on a multi-level cache, characterized in that, It includes: A heat update module, which is used to update the current model heat of the to-be-inferred model corresponding to the preset cache architecture based on the obtained model access request during the model deployment process; wherein, the preset cache architecture from top to bottom cache layers are the video memory cache layer, the process memory cache layer, the shared memory cache layer, and the persistent cache layer; the model heat is used to represent the frequency of the model being accessed within a preset time period. A first model inference module, which is used to determine the target cache layer of the to-be-inferred model in the preset cache architecture, and when the target cache layer is the video memory cache layer, trigger the preset inference operation on the to-be-inferred model. A second model inference module, which is used to, when the target cache layer is not the video memory cache layer, schedule the to-be-inferred model to the video memory cache layer based on the current model heats in other cache layers above the target cache layer, so as to trigger the preset inference operation on the to-be-inferred model in the video memory cache layer.
8. The model scheduling device based on multi-level caching according to claim 7, wherein The heat update module includes: A model determination unit, which is used to obtain the model access request sent by the client and determine the to-be-inferred model in the preset cache architecture based on the model access request. An exponent determination unit, which is used to determine the time difference based on the current time and the model access time corresponding to the model access request, and use the time difference and the heat decay period to determine the heat decay exponent. A heat update unit, which is used to update the current model heat of the to-be-inferred model by using the preset heat decay factor, the heat decay exponent, the size of the to-be-inferred model, and the target model access times. Wherein, the target model access times are the number of times the to-be-inferred model is accessed within the preset time period, and the target model access times do not include the model access times corresponding to the current model access request.
9. An electronic device, characterized in that, It includes: A memory, which is used to store computer programs. A processor, which is used to execute the computer program to implement the multi-level cache-based model scheduling method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, It is used to store computer programs; wherein, when the computer program is executed by the processor, it implements the multi-level cache-based model scheduling method according to any one of claims 1 to 6.
Citation Information
Cited By
Storage hierarchical acceleration system for large-model high-concurrency reasoning
CN120848818A
Storage hierarchical acceleration system for large model high concurrency inference
CN120848818B
Large model reasoning method and system based on multi-level cache mechanism, electronic equipment and storage medium
CN120851217A
Edge cloud collaborative distribution method and device for model resources, electronic equipment and medium
CN121531032A