Model scheduling method, storage cluster, device and medium
By storing model image files in layers and dynamically scheduling them, the problems of image file transmission delay and storage space occupation in cloud computing environments are solved, and efficient model deployment and iterative response are achieved.
Patent Information
- Application Number
- CN202511079545.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-12
AI Technical Summary
In a cloud computing environment, traditional model deployment involves large image files, which result in long transmission delays, large storage space usage, low deployment efficiency, and inability to meet rapid iteration requirements.
The model image file is layered into the base layer and the model layer. The base layer is stored in the cloud platform cache space, and the model layer is stored in the network shared space. The model layer image file is pre-loaded according to the historical deployment frequency and popularity frequency, and scheduled through dynamic scheduling strategy and incremental update technology.
It reduces repeated transmission and storage usage, improves deployment efficiency, optimizes network bandwidth utilization, and ensures rapid response and accuracy to model iteration requirements.
Smart Images

Figure CN120639788A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a model scheduling method, storage cluster, device and medium. Background Art
[0002] In a cloud computing environment, the traditional deployment process involves downloading the entire model image file for each deployment task and then copying it to the bare metal server's logical disk. The entire model image file is large, and the download process from the quick-view user interface component (Glance module) to the cloud platform is lengthy, resulting in significant transmission delays. Storing the entire model image file in the cloud platform's corresponding temporary storage area consumes significant storage space. Furthermore, the rapid iteration of models on bare metal servers during deployment makes it difficult to meet deployment requirements, reducing overall deployment efficiency.
[0003] Therefore, how to speed up transmission delay and improve model deployment efficiency while saving storage space is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0004] The purpose of the present invention is to provide a model scheduling method, storage cluster, device and medium to solve the problem that the entire model image file deployment occupies a large storage space during the model deployment process, the deployment requirements are difficult to meet, and the deployment efficiency is reduced.
[0005] To solve the above technical problems, the present invention provides a model scheduling method, comprising:
[0006] The model image file is pre-layered to obtain a base layer image file and a model layer image file; wherein the base layer image file is stored in the corresponding cache space of the cloud platform; the model layer image file is stored in the network shared space;
[0007] In response to the current scheduling request, when the current scheduling request is the first scheduling, the corresponding base layer image file and model layer image file are scheduled from the cache space and the network shared space respectively to be mounted to the target server;
[0008] When the current scheduling request is not the first scheduling, the incremental model layer image file is determined based on the model layer image file stored in the target server and the model layer image file in the network shared space; and the incremental model layer image file is mounted to the target server to complete the scheduling process.
[0009] On the one hand, the pre-loading process of the model layer image file stored in the network shared space includes:
[0010] Obtain the historical deployment frequencies of multiple models in the model scheduling set and the popularity frequencies in different time periods;
[0011] Processing the multiple models based on the historical deployment frequency and / or the heat frequency to obtain a first target model;
[0012] The model layer image file of the first target model is pre-loaded into a network shared space.
[0013] On the other hand, the plurality of models are processed based on the historical deployment frequency and the popularity frequency to obtain a first target model, including:
[0014] Selecting a second target model corresponding to a preset heat frequency from the heat frequencies of multiple models corresponding to different time periods;
[0015] Selecting a third target model corresponding to a frequency exceeding a preset deployment frequency from historical deployment frequencies corresponding to the multiple models;
[0016] Determine a fourth target model of the intersection according to the second target model and the third target model;
[0017] Assigning corresponding weight parameters to the historical deployment frequencies and popularity frequencies of the plurality of fourth target models to determine prediction scores of the plurality of fourth target models;
[0018] A target model having a higher prediction score than the target is selected from the prediction scores of the plurality of fourth target models as the first target model.
[0019] On the other hand, the scheduling process of scheduling the model layer image file to the target server includes:
[0020] Determine the corresponding mounting bandwidth resources based on the functional requirements of multiple target servers;
[0021] Determine the actual bandwidth corresponding to the scheduling process according to the bandwidth resources required by the target model image files corresponding to the multiple target servers and the mounting bandwidth resources;
[0022] Determine the scheduling strategy based on the actual bandwidth and total bandwidth resources of the target model image files corresponding to multiple target servers;
[0023] The target model image files of the multiple target servers are scheduled according to the scheduling strategy.
[0024] On the other hand, the scheduling process of scheduling the model layer image file to the target server includes:
[0025] Obtain the scheduled transmission rates of target model image files corresponding to multiple target servers;
[0026] Adjust the corresponding scheduling transmission rate according to the urgency of the target model image file to determine the scheduling strategy;
[0027] The target model image files of the multiple target servers are scheduled and processed through the flow control mechanism and the scheduling strategy.
[0028] On the other hand, when a scheduling error occurs during the scheduling of the model layer image file, the method further includes:
[0029] Obtaining a scheduling protocol for the model layer image file;
[0030] Determine target storage protocol;
[0031] Pre-configuring the target storage protocol to the target server, so as to install and configure the service of the target storage protocol on the target server;
[0032] Pre-configure the target storage protocol to the client corresponding to the network shared space to facilitate the implementation of shared services;
[0033] The scheduling protocol is switched to the target storage protocol, and scheduling processing is performed on the model layer image file.
[0034] On the other hand, after the incremental model layer image file is mounted to the target server, the method further includes:
[0035] Determine the target model corresponding to the model layer image file stored in the target server;
[0036] Freezing the parameters of the target model;
[0037] Constructing an adapter of the target model and corresponding adapter weights;
[0038] Loading the incremental model layer image file according to the adapter weight to determine an adapter weight file;
[0039] Injecting the adapter weight file into the target model to add a target layer corresponding to the target model;
[0040] The added target layer and the target model are used as the fifth target model;
[0041] Get the test dataset;
[0042] Inputting the test data set into the fifth target model and the target model to determine the output results corresponding to the fifth target model and the target model respectively;
[0043] If the difference between the output results corresponding to the fifth target model and the target model is within a first preset range, it is determined that the fifth target model meets expectations;
[0044] Obtaining the accuracy rates corresponding to the fifth target model and the target model respectively;
[0045] Determining similarity and recall based on overlap between the output results corresponding to the fifth target model and the target model;
[0046] Determine a first scoring result and a second scoring result according to the accuracy, the similarity, the recall rate, and their corresponding weight values;
[0047] If the difference between the first scoring result and the second scoring result is within a second preset range, determining that the test result of the fifth target model is qualified;
[0048] The fifth target model is used as a new target model.
[0049] To solve the above technical problems, the present invention further provides a storage cluster, comprising a master node and a slave node;
[0050] The master node is connected to multiple slave nodes;
[0051] The master node is used to execute the steps of the model scheduling method to schedule the model image file to the slave node.
[0052] To solve the above technical problems, the present invention further provides a model scheduling device, comprising:
[0053] memory for storing computer programs;
[0054] A processor is used to implement the steps of the model scheduling method when executing the computer program.
[0055] In order to solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the model scheduling method as described above are implemented.
[0056] The beneficial effects of the present invention are as follows: first, the model image file is pre-layered to obtain the base layer image file and the model layer image file, breaking the original scheduling of the entire image file and splitting the entire image file to ensure the independence and reusability between the layers. The base layer image file is stored in the corresponding cache space of the cloud platform, which reduces repeated transmission in subsequent deployments and saves the problem of cache space occupation. The model layer image file is stored in the network shared space to achieve high-performance storage cluster storage, ensuring high availability and fast access. There is no need to re-download the process each time deployment, avoiding the long time corresponding to the download process, realizing the hybrid storage of the image file cache space and the network shared space, and also reducing redundant transmission and storage overhead. Secondly, in response to the current scheduling request, when the current scheduling request is the first scheduling, the base layer image file and the model layer image file are mounted from different storage areas respectively. In non-first scheduling, the incremental model layer image file is determined, and there is no need to re-mount the entire image file each time during the model iteration process. After the scheduling of the base layer image file is completed for the first time, the base layer image file does not need to be mounted for non-first scheduling, thereby realizing the reuse of the base layer image file. For model iteration, only the incremental model layer image file is mounted, which reduces transmission overhead while meeting the iteration requirements of each model, reducing network bandwidth consumption and improving deployment efficiency.
[0057] Multiple models are pre-processed based on historical deployment frequency and / or popularity frequency to obtain the first target model for pre-loading. This pre-loading process also reduces deployment latency during mounting. This screening process, based on a combination of these two factors (historical deployment frequency and popularity frequency), improves the accuracy of pre-loading target models while fully utilizing shared cache space, saving space resources and reducing deployment latency during mounting. Model-level image file scheduling on multiple target servers uses a scheduling strategy based on the image file's bandwidth resources, the target server's mounting bandwidth resources, and the total bandwidth resources. This improves network bandwidth utilization while ensuring load balancing. Scheduling, based on flow control mechanisms and scheduled transmission rates, comprehensively considers factors such as thread priority, resource constraints, load balancing, and flow control to improve system performance and resource utilization. Switching to the target storage protocol enables multi-protocol optimization while providing the optimal mounting method to ensure smooth image file scheduling. The incremental file is merged with the original image file to generate a new model file for subsequent deployment and use. In order to ensure the accuracy of the subsequent use of the target model, the new target model is verified and compared with the original target model to ensure that the new target model is qualified and improve the accuracy of model use.
[0058] In addition, the present invention also provides a storage cluster, a device, and a medium, which have the same beneficial effects as the above-mentioned model scheduling method. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0060] Figure 1 A flowchart of a model scheduling method provided by an embodiment of the present invention;
[0061] Figure 2 A schematic diagram of the structure of another storage cluster provided by an embodiment of the present invention;
[0062] Figure 3 A flowchart of another model scheduling method provided by an embodiment of the present invention;
[0063] Figure 4 A structural diagram of a model scheduling device provided by an embodiment of the present invention;
[0064] Figure 5 A structural diagram of a model scheduling device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0066] The core of the present invention is to provide a model scheduling method, storage cluster, device and medium to solve the problem that the entire model image file deployment occupies a large storage space during the model deployment process, the deployment requirements are difficult to meet, and the deployment efficiency is reduced.
[0067] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0068] In cloud computing environments, the OpenStack project (OpenStack Ironic), designed for managing and automating the deployment of bare metal servers, is widely used for bare metal server operating system deployment. Its traditional deployment process relies on downloading images from the Glance module, which manages and stores virtual machine images, to the Ironic module, which manages and automates the deployment of bare metal servers. The images are then copied to the bare metal logical disk by executing various operations (Ironic Python Agent, IPA) on the bare metal servers. However, when deploying artificial intelligence (AI) models, model images are typically large (tens of GB or even terabytes), requiring a complete re-creation of the image. Each AI model version iteration requires re-creating the entire image containing the new model. Even minor adjustments or updates to model parameters require repackaging the entire image. This requires full download and transmission. Each new version deployment requires transferring the entire image file from Glance to Ironic and then to the bare metal server. For example, if a model is updated from v1.0 to v1.1, even if only a 1GB parameter change is made, the entire 50GB image must be transferred. Wasted storage space: Glance must store the complete image for each version; Ironic nodes must download and store the entire new version image. Existing identical infrastructure components cannot be reused. For example: AI model v1.0 image: 50GB (40GB for the base environment + 10GB for the model files); AI model v1.1 image: 50GB (40GB for the base environment + 10GB for the model files). The model files may have only changed by 500MB, but the entire 50GB must still be transferred. The base environment remains identical, yet existing infrastructure components cannot be reused.
[0069] Traditional methods face the following challenges:
[0070] Transmission delay: The process of downloading images from Glance to Ironic takes a long time, especially when the cluster size is large, and the network bandwidth becomes a bottleneck.
[0071] Storage redundancy: Ironic nodes need to temporarily store complete images, resulting in a waste of storage resources.
[0072] Low deployment efficiency: The layer-by-layer transmission of large files cannot meet the requirements of rapid iteration of AI models.
[0073] Resource contention: When multiple nodes are deployed concurrently, there is severe contention for storage and network resources, affecting overall performance.
[0074] Insufficient dynamic scalability: Unable to flexibly cope with large-scale clusters and multi-task scheduling requirements.
[0075] Lack of optimization mechanism: Failure to perform targeted optimization based on the characteristics of AI model deployment (such as frequent iterations).
[0076] The model scheduling method provided by the present invention can solve the above technical problems.
[0077] Figure 1 A flow chart of a model scheduling method provided by an embodiment of the present invention is as follows: Figure 1 As shown, the method includes:
[0078] S11: pre-layering the model image file to obtain a base layer image file and a model layer image file;
[0079] Among them, the base layer image file is stored in the corresponding cache space of the cloud platform; the model layer image file is stored in the network shared space;
[0080] S12: respond to the current scheduling request; determine whether the current scheduling request is the first scheduling, if so, proceed to step S13, if not, proceed to step S14;
[0081] S13: When the current scheduling request is the first scheduling, the corresponding base layer image file and model layer image file are scheduled from the cache space and the network shared space respectively to be mounted to the target server;
[0082] S14: When the current scheduling request is not the first scheduling, determine the incremental model layer image file based on the model layer image file stored in the target server and the model layer image file in the network shared space; and mount the incremental model layer image file to the target server to complete the scheduling process.
[0083] Specifically, conventional solutions use an open source application container engine (Docker) to layer, and build layer by layer according to the file text (Dockerfile) instructions (specify the base image (FROM), execute commands (RUN), etc.). The layering is fine-grained, and there may be more than a dozen layers. The main purpose is to reuse the intermediate layer and reduce storage space. The layering of the present invention is based on functionality, mainly divided into two layers according to business logic and change frequency, the base layer and the model layer. The base layer provides a relatively stable operating environment, such as (operating system + dependent library + runtime environment), such as operating system, general library, runtime environment and other basic components. These contents are usually relatively stable and reused in different tasks. The model layer is AI model files and configurations that are easy to change, such as specific AI model files, training data, configuration files, etc. These contents may change due to task requirements. Here, the base layer image files are stored in the Ironic node's local cache (the corresponding cache space on the cloud platform) to reduce repeated transmissions in subsequent deployments. As a transit and control node for deployment, the Ironic node caches the base layer images locally to avoid repeated downloads; coordinates the Network File System Mount (NFS) and image merging processes; and controls the deployment of bare metal servers through IPA. Model layer image files are stored in a network share, such as a high-performance NFS storage cluster, to ensure high availability and fast access. The Network File System is a file storage (NFS system) with the following features: Distributed storage: The model layer is distributed across multiple NFS nodes; High availability: Supports automatic switching in the event of node failure; Load balancing: Multiple Ironic nodes can concurrently access different NFS nodes; Dynamic mounting: Mounts specific model layers on demand without downloading the entire file.
[0084] The traditional method is that the Ironic node needs to download the entire image locally. This embodiment uses NFS mounting to achieve "on-demand access" rather than "complete download", reducing the storage pressure of the Ironic node; supporting multiple nodes to share the same model layer; and facilitating centralized management and updates of the model layer.
[0085] Regarding the layering process in this embodiment, the image may be split and compressed according to image layering technology (such as Docker or the Open Container Initiative (OCI) standard) to ensure the independence and reusability of each layer.
[0086] The base layer image files are stored in the cache space. In order to reduce the subsequent response to the current scheduling request, they are pre-stored. The model layer image files are stored in the network shared space. Here, the model layers with higher frequency of use can be predicted first to implement the preloading process and pre-load them into the network shared space to reduce the mounting delay during deployment. Regarding the preloading process, the cloud platform receives the deployment request from the bare metal server and parses the request content, such as the required AI model version, hardware configuration, etc. A query request is sent to the prediction and optimization engine. The prediction engine uses machine learning algorithms (such as time series prediction or cluster analysis) based on historical deployment data and current task requirements to predict the frequently used model layers. Determine the preloaded content: Based on the prediction results, the frequently used model layers are pre-loaded into the edge cache of the Ironic node to reduce the mounting delay during deployment.
[0087] Dynamically mount the image files corresponding to the base and model layers. Load the base layer directly from the Ironic node's local cache to avoid duplicate transfers. Model layer mounting: Dynamically mount the required model layer from the NFS storage cluster based on the bare metal server deployment requirements. Use the NFS client protocol to mount the model layer to a specified directory, ensuring that the bare metal server can access the model files.
[0088] Step S12 responds to the current scheduling request. When the current scheduling request is the first scheduling, the corresponding image files of the base layer and the model layer need to be scheduled from the local cache space and the network shared space. What is considered here is that in the first scheduling, which means that the target server has just been mounted in the storage cluster, all the image files corresponding to the target server need to be mounted. Conventionally, they are first downloaded from the node to the local. In this embodiment, the base layer image files stored locally and the model layer image files in the network shared space corresponding to NFS are directly called, saving the download process.
[0089] When scheduling is not the first time, the target server is considered to have already added the base layer image file during the first scheduling process. At this time, because the file version corresponding to the model upgrade needs to be iterated, the model layer image file is changed. Therefore, the model layer image file can be scheduled. However, during the model upgrade process, it may be that only a few requirements or a small number of image files need to be modified for the iteration. Therefore, before scheduling, the model layer image file already stored on the target server is compared with the model layer image file in the network shared space to determine the incremental model layer image file, which further reduces the number of scheduled files and bandwidth resources, thereby improving scheduling efficiency.
[0090] For version iterations of existing images, only transfer the changed parts (such as updated model parameters or newly added configuration files). Use incremental update technologies (such as file synchronization tools (rsync) or block-level differential replication) to merge the changed parts with the existing image to reduce data transfer. Image merging: Merge the base layer, model layer, and incremental updates into a complete logical disk image. Ensure that the merged image is compatible with the bare metal server's hardware configuration.
[0091] During the scheduling process, the bandwidth allocation strategy for the image file will be dynamically adjusted according to the actual situation. When multiple target servers need to schedule the model layer image file, the conventional NFS storage has only one node, which leads to congestion during the scheduling process. This embodiment expands its NFS storage nodes to ensure resource fairness and efficiency when multiple nodes are deployed concurrently. Cache management: Dynamically clean and update the preloaded content in the edge cache to ensure efficient use of cache space. Use the Least Recently Used (LRU) or Most Frequently Used (LFU) algorithm to manage cache content. Load balancing: Deploy a load balancer in the NFS storage cluster to dynamically allocate storage requests and avoid performance bottlenecks of a single node.
[0092] The beneficial effects of the embodiments of the present invention are as follows: first, the model image file is pre-layered to obtain a base layer image file and a model layer image file, breaking the original scheduling of the entire image file and splitting the entire image file to ensure the independence and reusability between the layers. The base layer image file is stored in the corresponding cache space of the cloud platform, which reduces repeated transmission in subsequent deployments and saves cache space occupancy. The model layer image file is stored in a network shared space to achieve high-performance storage cluster storage, ensuring high availability and fast access. There is no need to re-download the process for each deployment, which avoids the long download process. The hybrid storage of the image file cache space and network shared space is achieved, which also reduces redundant transmission and storage overhead. Secondly, in response to the current scheduling request, if the current scheduling request is the first scheduling, the base layer image file and the model layer image file are mounted from different storage areas respectively. In non-first scheduling, the incremental model layer image file is determined, eliminating the need to re-mount the entire image file each time during the model iteration process. After the first scheduling of the base layer image file is completed, the base layer image file does not need to be mounted for non-first scheduling, thus achieving the reuse of the base layer image file. For model iteration, only the incremental model layer image file is mounted, which reduces transmission overhead while meeting the iteration requirements of each model, reducing network bandwidth consumption and improving deployment efficiency.
[0093] In some embodiments, the pre-loading process of storing the model layer image file in a network shared space includes:
[0094] Obtain the historical deployment frequencies of multiple models in the model scheduling set and the popularity frequencies in different time periods;
[0095] Processing the multiple models based on historical deployment frequencies and / or popularity frequencies to obtain a first target model;
[0096] The model layer image file of the first target model is pre-loaded into the network shared space.
[0097] Specifically, to improve scheduling efficiency, we preload the image files corresponding to commonly used models into the network shared space, as multiple target models may require scheduling. This reduces mounting delays during deployment. The definition of commonly used models can be based on factors such as historical deployment frequency, popularity, or a combination of multiple factors.
[0098] Obtain the historical deployment frequencies and the popularity frequencies in different time periods corresponding to the multiple models in the model scheduling set, and screen the multiple models based on the and / or relationship of the two factors to determine the first target model to be pre-loaded into the network shared space.
[0099] The historical deployment frequency is a statistical calculation of the frequency of each model's appearance in historical deployments, which is determined by the number of historical appearances and the total number of deployments. The popularity frequency in different time periods is to analyze the popular models in different time periods of the day (for example, there are more Computer Vision (CV) models during working hours and more Natural Language Processing (NLP) models at night), and give a popularity score adjusted for time factors. The day is divided into hours (0-23h), and for each model, its frequency of appearance within each hour is counted. For the time of the current request (assuming the request timestamp is t), the model frequency of the hour t is taken as the time score: time_score = (number of appearances in the hour at time t) / (total number of deployments in the hour). If the total number of deployments in the hour is 0, the global frequency is used.
[0100] If there is a single factor, a sorting method can be used to select the first N corresponding models as the first target model for pre-loading.
[0101] This embodiment provides a method for pre-processing multiple models based on historical deployment frequency and / or heat frequency to obtain a pre-loaded first target model, and the pre-loading process also reduces the mounting delay during deployment.
[0102] In some embodiments, processing multiple models based on historical deployment frequencies and popularity frequencies to obtain a first target model includes:
[0103] Selecting a second target model corresponding to a preset heat frequency from the heat frequencies of multiple models corresponding to different time periods;
[0104] Selecting a third target model corresponding to a frequency exceeding a preset deployment frequency from historical deployment frequencies corresponding to the multiple models;
[0105] Determine a fourth target model of the intersection according to the second target model and the third target model;
[0106] Assigning corresponding weight parameters to the historical deployment frequencies and popularity frequencies of the plurality of fourth target models to determine prediction scores of the plurality of fourth target models;
[0107] A target model having a higher prediction score than the target is selected from the prediction scores of the plurality of fourth target models as the first target model.
[0108] Specifically, based on the above two factors, the heat frequencies corresponding to different time periods in multiple models are first screened out to obtain the second target model corresponding to the preset heat frequency. Since there are different time periods in the screening here, the time period can also be directly ignored, and all its heat frequencies are screened. In addition, different time periods can also be taken into account, and the preset heat frequencies corresponding to different time periods can be distinguished to obtain the corresponding second target model through screening. The third target model that exceeds the preset deployment frequency is screened out from multiple models corresponding to the historical deployment frequency.
[0109] Since there may be target models with more or less intersections between the second target model and the third target model, the second target model and the third target model can all be used as the first target model, or they can be used as the fourth target model based on the above-mentioned intersection model. In the multiple fourth target models, corresponding weight parameters are assigned to the two factors mentioned above, and the weighted sum is performed to determine the prediction scores of the multiple fourth target models. The multiple prediction models are sorted from high to low, and the target models that exceed the target prediction score are screened out. Alternatively, they are loaded in sequence until the cache space of the shared space is full.
[0110] Weighted fusion: The above two scores are weighted averaged to obtain the final prediction score;
[0111] total_score=w1×freq_score+w2×time_score;
[0112] The weights w1 and w2 can be set based on experience (e.g., w1 = 0.5, w2 = 0.3) or adjusted based on historical performance. Order by total_score from high to low, and then load them sequentially until the cache is full.
[0113] This embodiment provides a screening process based on a comprehensive consideration of two factors (historical deployment frequency and popularity frequency), which improves the screening accuracy of pre-loaded target models while fully utilizing the cache space of the shared space, saving space resources, and reducing mounting delays during deployment.
[0114] In some embodiments, the process of dispatching the model layer image file to the target server includes:
[0115] Determine the corresponding mounting bandwidth resources based on the functional requirements of multiple target servers;
[0116] Determine the actual bandwidth corresponding to the scheduling process based on the bandwidth resources required by the target model image files corresponding to the multiple target servers and the mounting bandwidth resources;
[0117] Determine the scheduling strategy based on the actual bandwidth and total bandwidth resources of the target model image files corresponding to multiple target servers;
[0118] The target model image files of multiple target servers are scheduled and processed according to the scheduling strategy.
[0119] Specifically, the functional requirements of multiple target servers, such as network interfaces and network links, determine the mounted bandwidth resources. Network interface bandwidth here ensures that the server's network interface (such as an Ethernet card) supports the calculated bandwidth requirement. For example, if the calculated bandwidth requirement is 102.6 MB / s, a network interface that supports 1 Gbps (approximately 125 MB / s) can be selected. Network link bandwidth ensures that the network link from the server to the client (such as a switch or router) also supports the calculated bandwidth requirement.
[0120] The bandwidth resources required for target model image files corresponding to multiple target servers include the size, number, download requirements, network latency, and redundancy of the model image files. The download requirements correspond to the average download frequency of each image file and the maximum number of concurrent downloaders that may download the image file during peak hours. The download bandwidth for a single image file = image file size / download time. The total bandwidth requirement for concurrent downloads = single image file download bandwidth × number of concurrent downloads. Network latency may affect download speed. If network latency is high, additional bandwidth may be required to compensate for the impact of latency. To ensure system stability and reliability, it is generally necessary to add a certain amount of redundancy to the calculated bandwidth requirements. For example, a 20% increase in bandwidth can be used for redundancy. The bandwidth requirement (including redundancy) = total bandwidth requirement for concurrent downloads × (1 + redundancy ratio). The bandwidth resources here refer to the bandwidth corresponding to the bandwidth requirement.
[0121] The bandwidth resources required by the target model image file and the mounted bandwidth resources are added together to obtain the actual bandwidth corresponding to the scheduling process.
[0122] The scheduling strategy is determined based on actual bandwidth and total bandwidth resources. The total bandwidth resources here refer to the total bandwidth resources under the threads corresponding to multiple target servers. The scheduling strategy is based on the actual bandwidth of the threads corresponding to each of the target servers and the target model image files. The scheduling strategy can prioritize target model image files with high network bandwidth requirements or prioritize target model image files with the highest degree of urgency to achieve balanced scheduling.
[0123] The model layer image file scheduling process provided in this embodiment under multiple target servers can determine the scheduling strategy based on the bandwidth resources of the image file itself and the mounting bandwidth resources of the target server and the total bandwidth resources for scheduling processing, thereby improving network bandwidth utilization while ensuring load balancing.
[0124] In some embodiments, determining a scheduling strategy based on actual bandwidth and total bandwidth resources of multiple target servers corresponding to target model image files includes:
[0125] Filter the target model image file whose actual bandwidth exceeds the preset bandwidth as the first priority processing file;
[0126] Determining a second priority processing file for parallel scheduling according to total bandwidth resources among the plurality of first priority processing files;
[0127] After the second priority processing files are scheduled, the remaining first priority processing files except the second priority processing files are processed to determine the scheduling strategy.
[0128] Specifically, the files that exceed the preset bandwidth are screened out from the actual bandwidth of multiple target model image files as the first priority processing files. Here, the files that occupy the larger bandwidth are selected as priority processing files. The second priority processing files that can be currently scheduled in parallel are determined by comparing the total bandwidth resources with the bandwidth resources of multiple first priority processing files. It should be noted that the second priority processing files are further selected based on the first priority processing files. Parallel scheduling takes into account the ability to process in parallel within the total bandwidth resources, improves scheduling efficiency, and makes full use of bandwidth. After all the second priority processing files are scheduled, the remaining first priority processing files are processed to form a scheduling strategy. If there are still bandwidth resources after the processing is completed, the remaining processing files except the first priority processing files will be scheduled for processing.
[0129] This embodiment provides a specific process for determining a scheduling strategy based on the actual bandwidth and total bandwidth resources of target model image files corresponding to multiple target servers. Among multiple first-priority processing files, a second-priority processing file is determined for parallel scheduling based on the total bandwidth resources, thereby avoiding performance bottlenecks caused by insufficient bandwidth. By rationally allocating total bandwidth resources, each file can be processed in the shortest possible time. Prioritizing files with higher actual bandwidth can significantly reduce the processing time of these files. This not only improves the processing efficiency of individual files, but also reduces the waiting time for other files, thereby improving the performance of the entire system.
[0130] In other embodiments, the process of dispatching the model layer image file to the target server includes:
[0131] Obtain the scheduled transmission rates of target model image files corresponding to multiple target servers;
[0132] Adjust the corresponding scheduling transmission rate according to the urgency of the target model image file to determine the scheduling strategy;
[0133] The target model image files of multiple target servers are scheduled and processed through flow control mechanism and scheduling strategy.
[0134] Specifically, the scheduling process can also adjust the scheduled transmission rate of the image file based on the urgency of the target model image file through the scheduled transmission rate to achieve reordering.
[0135] Flow control mechanisms, such as token bucket and leaky bucket algorithms, limit the transmission rate of each thread to ensure that the system does not overload due to the high load of a single thread. On this basis, the target model image file is scheduled through a scheduling strategy, ensuring that the scheduled transmission rate of each model image file is strictly scheduled according to the flow control mechanism.
[0136] The scheduling process based on the flow control mechanism and the scheduled transmission rate provided in this embodiment comprehensively considers factors such as thread priority, resource limitations, load balancing and flow control, thereby improving system performance and resource utilization.
[0137] In some embodiments, when a scheduling error occurs during the scheduling of the model layer image file, the method further includes:
[0138] Obtain the scheduling protocol for the model layer image file;
[0139] Determine target storage protocol;
[0140] Pre-configure the target storage protocol to the target server, so as to install and configure the service of the target storage protocol on the target server;
[0141] Pre-configure the target storage protocol to the client corresponding to the network shared space to facilitate the implementation of shared services;
[0142] Switch the scheduling protocol to the target storage protocol and schedule the model layer image file.
[0143] Specifically, a scheduling error occurs when bandwidth resources are limited and the scheduling strategy fails to allocate resources properly, causing some threads or tasks to fail to run properly. Based on this, it is necessary to obtain the scheduling protocol used by the model layer image file and determine the target storage protocol.
[0144] Switching to a storage protocol is specifically for performance improvement and to meet specific storage requirements. Based on the selected storage protocol, configure the storage server (target server). For NFS, install and configure the target storage protocol service on the target server. Configure the appropriate client software on the client to access the storage server. For NFS, mount the NFS share on the client.
[0145] Switch the scheduling protocol to the target storage protocol to continue scheduling processing.
[0146] This embodiment provides switching to the target storage protocol to achieve multi-protocol optimization while providing an optimal mounting method to ensure smooth completion of image file scheduling processing.
[0147] In some embodiments, after mounting the incremental model layer image file to the target server, the method further includes:
[0148] Determine the target model corresponding to the model layer image file stored in the target server;
[0149] Freeze the parameters of the target model;
[0150] Build the adapter of the target model and the corresponding adapter weights;
[0151] Load the incremental model layer image file according to the adapter weight to determine the adapter weight file;
[0152] Inject the adapter weight file into the target model to add the target layer corresponding to the target model;
[0153] The added target layer and target model are used as the fifth target model;
[0154] Get the test dataset;
[0155] Inputting the test data set into the fifth target model and the target model to determine the output results corresponding to the fifth target model and the target model respectively;
[0156] If the difference between the output results corresponding to the fifth target model and the target model is within the first preset range, it is determined that the fifth target model meets expectations;
[0157] Obtaining the accuracy rates corresponding to the fifth target model and the target model;
[0158] Determining similarity and recall based on overlap between the output results corresponding to the fifth target model and the target model;
[0159] Determine the first scoring result and the second scoring result according to the accuracy, similarity and recall rate and their corresponding weight values;
[0160] If the difference between the first scoring result and the second scoring result is within a second preset range, then determining that the test result of the fifth target model is qualified;
[0161] The fifth target model is used as a new target model.
[0162] Specifically, in pure incremental fine-tuning, the parameters of the original base model are frozen, and a series of additional, much smaller "adapter" parameter layers are trained. These adapter layers learn how to adjust the output of the original model to adapt to the new task. The incremental model layer image file is loaded and processed according to the adapter weights to determine the adapter weight file. The adapter weight file is injected into the target model to add the target layer of the target model. The adapter weight is assumed to be a dictionary containing the weights of the adapter layer. The added target layer and target model are used as the fifth target model.
[0163] On this basis, it is necessary to verify whether the fifth target model is correct, which can be verified through performance evaluation. The test data set needs to be input into the fifth target model and the target model to output the corresponding output results. If the difference between the output results of the two target models is within the first preset range, it means that the fifth target model currently meets expectations. It is necessary to further evaluate the model performance. By assigning corresponding weight values to the accuracy, similarity, and recall rate, a weighted sum is performed to determine the first and second scoring results corresponding to each target model, and determine whether the difference between the two scoring results is within the second preset range. If so, it means that the model performance test is qualified. The fifth target model is used as the new target model.
[0164] This embodiment provides for merging the incremental file with the original image file to generate a new model file for subsequent deployment and use. In order to ensure the accuracy of the target model used in subsequent use, the new target model is verified and compared with the original target model to ensure that the new target model is qualified and improve the accuracy of model use.
[0165] Furthermore, the present invention also provides a storage cluster, comprising a master node and a slave node;
[0166] The master node is connected to multiple slave nodes;
[0167] The master node is used to execute the steps of the above-mentioned model scheduling method to schedule the model image file to the slave node.
[0168] For an introduction to a storage cluster provided by the present invention, please refer to the above method embodiment, which will not be described in detail herein. The present invention has the same beneficial effects as the above model scheduling method.
[0169] Figure 2 A schematic diagram of another storage cluster structure provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown in the figure, it includes a storage cluster, a deployment controller, a bare metal node cluster, an image layering management module, and a prediction and optimization engine. The image layering management module is responsible for the layered splitting and storage of AI model images. The NFS storage cluster stores model layer images and provides high-availability access. The Ironic deployment controller extends Ironic functionality to support dynamic mounting and preload scheduling. The prediction and optimization engine uses AI algorithms to predict preload content and optimize bandwidth allocation. The bare metal agent (IPA) supports incremental updates and multi-protocol mounting.
[0170] Image layered storage and dynamic mounting:
[0171] Divide the AI model image into the base layer (such as the operating system and common dependencies) and the model layer (specific AI model files).
[0172] The base layer is pre-stored in the local cache of the Ironic node, and the model layer is stored in NFS and dynamically mounted on demand to reduce redundant transmission.
[0173] Preload prediction mechanism:
[0174] Based on historical deployment data and AI task requirements, machine learning algorithms are used to predict commonly used model layers, and frequently used model layers are preloaded into the Ironic node edge cache in advance.
[0175] Adaptive bandwidth allocation:
[0176] During the NFS mount process, bandwidth allocation is dynamically adjusted based on network load and bare metal performance to avoid resource contention during concurrent deployment on multiple nodes. For example, a concurrent deployment scenario involves: Node A mounting a 50GB image recognition model, Node B mounting a 30GB NLP model, and Node C mounting a 40GB recommendation model. The total network requirement is 120GB of concurrent transfers.
[0177] The problem with the traditional method is that all nodes pull data from NFS at the maximum speed at the same time, which exhausts the network bandwidth, causing all transmissions to slow down and some nodes to timeout and fail.
[0178] The initial total bandwidth is 1000 MB / s, with 400 MB / s allocated to Node A (high performance), 350 MB / s to Node B (medium performance), and 250 MB / s to Node C (low performance). Dynamic adjustment occurs as follows: at time T1, after Node B completes deployment, 350 MB / s is released. At T2, this bandwidth is redistributed to Node A (600 MB / s) and Node C (400 MB / s). At T3, a hardware bottleneck is detected at Node C, and its allocation is reduced to avoid waste.
[0179] However, traditional approaches, such as even allocation or first-come, first-served allocation, can easily lead to resource waste. This embodiment uses an adaptive approach, intelligently allocating resources based on actual needs and capabilities to maximize overall deployment efficiency. For AI model version iterations, incremental image updates are supported, transmitting only the changed parts and merging them with the existing image, reducing data transmission volume.
[0180] Figure 3 A flowchart of another model scheduling method provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown in Figure 1, Step 1: Image Preparation, includes image layering and storage allocation. Image layering involves dividing the AI model image into two layers based on functionality and dependencies: the base layer and the model layer. The base layer includes foundational components such as the operating system, common libraries, and the runtime environment. These components are generally relatively stable and reused across different tasks. The model layer includes specific AI model files, training data, configuration files, and other components, which may vary depending on task requirements. Image layering technologies (such as Docker or the OCI standard) are used to split and compress the image to ensure independence and reusability between layers.
[0181] Storage allocation is to pre-store the base layer in the local cache of the Ironic node to reduce repeated transmission in subsequent deployments. The model layer is uploaded to a high-performance NFS storage cluster to ensure high availability and fast access.
[0182] Step 2: Deployment initialization: First, the deployment request is received, then the prediction is preloaded, and the preloaded content is locked. Receiving the deployment request involves the Ironic deployment controller receiving the deployment request from the bare metal server and parsing the request content (such as the required AI model version and hardware configuration). Preloading the prediction involves the Ironic controller sending a query request to the prediction and optimization engine. The prediction engine uses machine learning algorithms (such as time series prediction or cluster analysis) based on historical deployment data and current task requirements to predict frequently used model layers. Determining the preloaded content involves preloading frequently used model layers into the Ironic node's edge cache based on the prediction results to reduce mounting latency during deployment.
[0183] Step 3: Dynamic mounting, including base layer reuse, model layer mount, and multi-protocol optimization. Base layer reuse: Load the base layer directly from the local cache of the Ironic node to avoid repeated transmission. Model layer mount: Dynamically mount the required model layer from the NFS storage cluster based on the deployment requirements of the bare metal server. Use the NFS client protocol to mount the model layer to the specified directory to ensure that the bare metal server can access the model file. Multi-protocol optimization: If the NFS performance is insufficient or the network load is high, the system dynamically switches to storage protocols such as Internet Small Computer Systems Interface (iSCSI) or Ceph to select the optimal mounting method.
[0184] Step 4: Data Transfer and Merge. This is the image merging process after incremental updates are supported. For version iterations of existing images, only the changed portions (such as model parameter updates or newly added configuration files) are transferred. Incremental update technologies (such as rsync or block-level differential replication) are used to merge the changed portions with the existing image, reducing data transfer. Image Merge: The base layer, model layer, and incremental updates are merged into a complete logical disk image. Ensure that the merged image is compatible with the bare metal server's hardware configuration.
[0185] Step 5: Resource optimization: First, perform bandwidth allocation and cache management to achieve load balancing. During the NFS mount process, monitor the network load and bare metal server performance in real time. Dynamically adjust the bandwidth allocation strategy based on the monitoring data to ensure resource fairness and efficiency when deploying multiple nodes concurrently. Cache management: Dynamically clean and update preloaded content in the edge cache to ensure efficient use of cache space. Use the LRU (least recently used) or LFU (most frequently used) algorithm to manage cache content. Load balancing: Deploy a load balancer in the NFS storage cluster to dynamically allocate storage requests and avoid performance bottlenecks on a single node.
[0186] Step 6: Deployment is complete. During this process, verification and startup are required to implement performance monitoring. Verify the integrity and availability of the logical disk image on the bare metal server. Start the AI model service to ensure system operation. Performance Monitoring: After deployment is complete, continuously monitor the AI model's operational performance and resource usage. Dynamically adjust the deployment strategy based on monitoring data to optimize the execution efficiency of subsequent tasks.
[0187] Step 7: Exception handling, mainly network failures, storage node failures, and cache failures. If a network interruption occurs during the NFS mount process, the system automatically switches to an alternative storage protocol (such as iSCSI) and remounts. Storage node failure: If a node in the NFS storage cluster fails, the system automatically redirects the mount request to another available node to ensure high availability. Cache failure: If the preloaded content in the edge cache fails or is cleared, the system reloads the required content from the NFS storage cluster to ensure the continuity of the deployment task.
[0188] Step 8: Logging and Auditing: During the deployment process, record detailed information about key operations and events (such as image layering, preload prediction, dynamic mounting, and incremental updates). Auditing and Analysis: Analyze log data to optimize prediction algorithms and resource allocation strategies, improving the efficiency and reliability of subsequent deployments.
[0189] The above describes in detail various embodiments corresponding to the model scheduling method. On this basis, the present invention also discloses a model scheduling device corresponding to the above method. Figure 4 This is a structural diagram of a model scheduling device provided by an embodiment of the present invention. Figure 4 As shown, the model scheduling device includes:
[0190] The layered processing module 11 is used to pre-layer the model image file to obtain a base layer image file and a model layer image file; wherein the base layer image file is stored in the corresponding cache space of the cloud platform; the model layer image file is stored in the network shared space;
[0191] The first scheduling module 12 is used to respond to the current scheduling request. When the current scheduling request is the first scheduling, the corresponding base layer image file and model layer image file are scheduled from the cache space and the network shared space respectively to be mounted to the target server;
[0192] The second scheduling module 13 is used to determine the incremental model layer image file based on the model layer image file stored in the target server and the model layer image file in the network shared space when the current scheduling request is not the first scheduling; and mount the incremental model layer image file to the target server to complete the scheduling process.
[0193] Since the embodiments of the device part correspond to the above embodiments, the embodiments of the device part please refer to the description of the embodiments of the method part, and will not be repeated here.
[0194] For an introduction to a model scheduling device provided by the present invention, please refer to the above method embodiment, and the present invention will not be repeated here. It has the same beneficial effects as the above model scheduling method.
[0195] Figure 5 A structural diagram of a model scheduling device provided by an embodiment of the present invention, such as Figure 5 As shown, the device includes:
[0196] Memory 21, for storing computer programs;
[0197] The processor 22 is configured to implement the steps of the model scheduling method when executing a computer program.
[0198] The model scheduling device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0199] The processor 22 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 22 may be implemented in at least one of the following hardware forms: a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array. The processor 22 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 22 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, the processor 22 may also include an AI processor for handling computational operations related to machine learning.
[0200] The memory 21 may include one or more computer-readable storage media, which may be non-transitory. The memory 21 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 21 is at least used to store the following computer program 211, wherein, after the computer program is loaded and executed by the processor 22, it can implement the relevant steps of the model scheduling method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include but is not limited to data involved in the model scheduling method, etc.
[0201] In some embodiments, the model scheduling device may further include a display screen 23 , an input / output interface 24 , a communication interface 25 , a power supply 26 , and a communication bus 27 .
[0202] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation to the model scheduling device, and may include more or fewer components than shown in the figure.
[0203] The processor 22 implements the model scheduling method provided by any of the above embodiments by calling instructions stored in the memory 21 .
[0204] For an introduction to a model scheduling device provided by the present invention, please refer to the above method embodiment, and the present invention will not be repeated here. It has the same beneficial effects as the above model scheduling method.
[0205] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 22, the steps of the above-mentioned model scheduling method are implemented.
[0206] It is understood that if the methods in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0207] For an introduction to a computer-readable storage medium provided by the present invention, please refer to the above method embodiment, which will not be described in detail herein. It has the same beneficial effects as the above model scheduling method.
[0208] The above is a detailed introduction to a model scheduling method, storage cluster, device and medium provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.
[0209] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A model scheduling method, characterized in that: include: The model image file is pre-layered to obtain a base layer image file and a model layer image file; wherein the base layer image file is stored in the corresponding cache space of the cloud platform; the model layer image file is stored in the network shared space; In response to the current scheduling request, when the current scheduling request is the first scheduling, the corresponding base layer image file and model layer image file are scheduled from the cache space and the network shared space respectively to be mounted to the target server; When the current scheduling request is not the first scheduling, the incremental model layer image file is determined based on the model layer image file stored in the target server and the model layer image file in the network shared space; and the incremental model layer image file is mounted to the target server to complete the scheduling process.
2. The model scheduling method according to claim 1, characterized in that: The pre-loading process of the model layer image file stored in the network shared space includes: Obtain the historical deployment frequencies of multiple models in the model scheduling set and the popularity frequencies in different time periods; Processing the multiple models based on the historical deployment frequency and / or the heat frequency to obtain a first target model; The model layer image file of the first target model is pre-loaded into a network shared space.
3. The model scheduling method according to claim 2, characterized in that: Processing the multiple models based on the historical deployment frequency and the popularity frequency to obtain a first target model includes: Selecting a second target model corresponding to a preset heat frequency from the heat frequencies of multiple models corresponding to different time periods; Selecting a third target model corresponding to a frequency exceeding a preset deployment frequency from historical deployment frequencies corresponding to the multiple models; Determine a fourth target model of the intersection according to the second target model and the third target model; Assigning corresponding weight parameters to the historical deployment frequencies and popularity frequencies of the plurality of fourth target models to determine prediction scores of the plurality of fourth target models; A target model having a higher prediction score than the target is selected from the prediction scores of the plurality of fourth target models as the first target model.
4. The model scheduling method according to claim 1, characterized in that: The scheduling process of scheduling the model layer image file to the target server includes: Determine the corresponding mounting bandwidth resources based on the functional requirements of multiple target servers; Determine the actual bandwidth corresponding to the scheduling process according to the bandwidth resources required by the target model image files corresponding to the multiple target servers and the mounting bandwidth resources; Determine the scheduling strategy based on the actual bandwidth and total bandwidth resources of the target model image files corresponding to multiple target servers; The target model image files of the multiple target servers are scheduled according to the scheduling strategy.
5. The model scheduling method according to claim 1, characterized in that: The scheduling process of scheduling the model layer image file to the target server includes: Obtain the scheduled transmission rates of target model image files corresponding to multiple target servers; Adjust the corresponding scheduling transmission rate according to the urgency of the target model image file to determine the scheduling strategy; The target model image files of the multiple target servers are scheduled and processed through the flow control mechanism and the scheduling strategy.
6. The model scheduling method according to claim 4 or 5, characterized in that: When a scheduling error occurs during the scheduling of the model layer image file, the method further includes: Obtaining a scheduling protocol for the model layer image file; Determine target storage protocol; Pre-configuring the target storage protocol to the target server, so as to install and configure the service of the target storage protocol on the target server; Pre-configure the target storage protocol to the client corresponding to the network shared space to facilitate the implementation of shared services; The scheduling protocol is switched to the target storage protocol, and scheduling processing is performed on the model layer image file.
7. The model scheduling method according to claim 1, characterized in that: After mounting the incremental model layer image file to the target server, the method further includes: Determine the target model corresponding to the model layer image file stored in the target server; Freezing the parameters of the target model; Constructing an adapter of the target model and corresponding adapter weights; Loading the incremental model layer image file according to the adapter weight to determine an adapter weight file; Injecting the adapter weight file into the target model to add a target layer corresponding to the target model; The added target layer and the target model are used as the fifth target model; Get the test dataset; Inputting the test data set into the fifth target model and the target model to determine the output results corresponding to the fifth target model and the target model respectively; If the difference between the output results corresponding to the fifth target model and the target model is within a first preset range, it is determined that the fifth target model meets expectations; Obtaining the accuracy rates corresponding to the fifth target model and the target model respectively; Determining similarity and recall based on overlap between the output results corresponding to the fifth target model and the target model; Determine a first scoring result and a second scoring result according to the accuracy, the similarity, the recall rate, and their corresponding weight values; If the difference between the first scoring result and the second scoring result is within a second preset range, determining that the test result of the fifth target model is qualified; The fifth target model is used as a new target model.
8. A storage cluster, characterized in that: Including master node and slave node; The master node is connected to multiple slave nodes; The master node is used to execute the steps of the model scheduling method described in any one of claims 1 to 7 above to schedule the model image file to the slave node.
9. A model scheduling device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the model scheduling method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the model scheduling method according to any one of claims 1 to 7.
Citation Information
Cited By
Mirror image management method suitable for multi-cluster idle computing power scheduling
CN122111566A