Model deployment method, system, electronic device, and computer-readable storage medium
By distributing model files to storage devices in different regions and mounting them to a server cluster in the same region, the problem of low efficiency in deploying large models across regions is solved, achieving high-speed interconnection and efficient loading.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-02
- Publication Date
- 2026-03-27
AI Technical Summary
Large models have large model files, resulting in long loading times when starting the service and low efficiency in cross-regional deployment.
The model files are distributed to multiple storage devices in different geographical locations, and these devices are mounted to a server cluster in the same region to achieve high-speed interconnection between the storage devices and the server cluster, eliminating cross-regional file transfer.
It shortens the loading time of model files, improves the processing efficiency of models in different regions, and enhances deployment efficiency.
Smart Images

Figure CN117008836B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large model technology and model deployment, in particular to a model deployment method and system, an electronic device, and a computer readable storage medium. BACKGROUND
[0002] The current large model stores a large number of parameters and computation graph structures, resulting in a very large model file of the model. Loading the model when starting a service takes a long time, and is limited by factors such as region, network, and hardware. In a cross-region file transmission scenario, the loading time is further prolonged, thereby resulting in low deployment efficiency of the model.
[0003] Currently, no effective solution has been proposed for the above problems. SUMMARY
[0004] Embodiments of the present application provide a model deployment method and system, an electronic device, and a computer readable storage medium, to at least solve the technical problem of low efficiency of model cross-region deployment in related technologies.
[0005] According to an aspect of an embodiment of the present application, a model deployment method is provided, including: obtaining a model file of a to-be-deployed model; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; and mounting the storage devices to a target server cluster in a plurality of server clusters that is in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0006] According to another aspect of an embodiment of the present application, a model deployment method is also provided, including: in response to receiving a model file of a to-be-deployed model, storing the model file to a central repository; in response to receiving a model distribution request, distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location, and different server clusters are also deployed in different regions; and in response to receiving a server deployment request, mounting the storage devices to a target server cluster in a plurality of server clusters that is in the same region as the storage devices based on the server deployment request, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0007] According to another aspect of the embodiments of the present application, a model deployment method is also provided, including: obtaining a model file of a to-be-deployed model by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter is the model file; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; when deploying an inference service of the to-be-deployed model, mounting the storage devices to a target server cluster in a plurality of server clusters that are deployed in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters to obtain a deployment result of the to-be-deployed model; and outputting the deployment result by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter is the deployment result.
[0008] According to another aspect of the embodiments of the present application, a model deployment system is also provided, including: a plurality of server clusters, different server clusters being deployed in different regions in terms of geographical location; a plurality of storage devices, different storage devices being deployed in different regions in terms of geographical location; and a control device connected with the storage devices and the server clusters, configured to distribute a model file of a to-be-deployed model to the storage devices, and when deploying an inference service of the to-be-deployed model, mount the storage devices to a target server cluster in a plurality of server clusters that are deployed in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0009] According to another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory storing an executable program; and a processor configured to run the program, wherein the program, when running, performs the method of any one of the above embodiments.
[0010] According to another aspect of the embodiments of the present application, a computer-readable storage medium is also provided, including a stored executable program, wherein the executable program, when running, controls a device where the computer-readable storage medium is located to perform the method of any one of the above embodiments.
[0011] In the embodiments of the present application, a model file of a to-be-deployed model can be obtained; the model file can be distributed to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; and the storage devices can be mounted to a target server cluster in a plurality of server clusters that are deployed in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters, thereby improving the processing efficiency of the model in different regions; it is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to utilize the network between the storage devices and the server clusters in the same region to realize high-speed interconnection between the storage devices and the server clusters, eliminate the transmission of cross-regional files, thereby shortening the loading time of the model file, improving the processing efficiency of the model in different regions, and further solving the technical problem of low efficiency of model cross-regional deployment in the related art.
[0012] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute the limitation of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0013] The drawings described herein are used to provide further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute the improper limitation of the present application. In the drawings:
[0014] Figure 1 is a schematic diagram of a hardware environment of a virtual reality device according to a model deployment method of an embodiment of the present application;
[0015] Figure 2 is a flowchart of a model deployment method according to Embodiment 1 of the present application;
[0016] Figure 3 is a structural diagram of a model deployment method according to an embodiment of the present application;
[0017] Figure 4 is a flowchart of another model deployment method according to an embodiment of the present application;
[0018] Figure 5 is a flowchart of another model deployment method according to an embodiment of the present application;
[0019] Figure 6 is a flowchart of a model deployment method according to Embodiment 2 of the present application;
[0020] Figure 7 is a flowchart of a model deployment method according to Embodiment 3 of the present application;
[0021] Figure 8 is a schematic diagram of a model deployment device according to Embodiment 4 of the present application;
[0022] Figure 9 is a schematic diagram of a model deployment device according to Embodiment 5 of the present application;
[0023] Figure 10 is a schematic diagram of a model deployment device according to Embodiment 6 of the present application;
[0024] Figure 11 is a schematic diagram of a model deployment system according to Embodiment 7 of the present application;
[0025] Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] The technical solutions provided by the present application are mainly implemented by using large model technology. Here, the large model refers to a deep learning model with large-scale model parameters, which can typically include hundreds of millions, billions, tens of billions, hundreds of billions or even tens of billions of model parameters. The large model can also be referred to as a foundation model. Through large-scale unlabeled corpus pre-training, a pre-trained model with hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLM) and multi-modal pre-training models.
[0029] It should be noted that, in practical applications, large models can be fine-tuned using a small number of samples to adapt them to different tasks. For example, large models can be widely used in Natural Language Processing (NLP), computer vision, and other fields. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios for large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. In this embodiment, the deployment of a large model in a model deployment scenario is used as an example for explanation.
[0030] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0031] Network Attached Storage (NAS): Provides file sharing services over a network and has other additional functions, offering a convenient and reliable way to store and share data;
[0032] Cloud Object Storage Service (OSS): Provides cloud storage services that can be used to store and manage massive amounts of unstructured data;
[0033] Virtual Private Cloud (VPC): This is a virtual network environment created in a cloud computing environment, which can provide an isolated and secure way to deploy and manage cloud resources;
[0034] Elastic Algorithm Service (EAS): This is a platform for deploying and managing machine learning models. EAS provides users with a convenient way to deploy, run, and manage their own machine learning algorithms and models, and offers scalable computing resources and high-performance computing capabilities.
[0035] Container orchestration platform (Kubernetes, or K8S for short): can organize and manage containerized applications, helping users simplify the deployment, management and scaling of applications.
[0036] To solve the above problems, the application provides a method for multi-region storage of storage devices, which deploys storage devices and services in the same virtual private network in the same region to achieve high-speed interconnection, eliminates cross-region file transmission, thereby shortens the loading time of model files, and further improves the service elasticity.
[0037] Embodiment 1
[0038] According to the embodiments of the application, a model deployment method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0039] Considering that the model parameter of the large model is large, and the operation resource of the mobile terminal is limited, the above-mentioned model deployment method provided by the embodiments of the application can be applied to the application scenario as shown in Figure 1 , but is not limited thereto. Figure 1 is a schematic diagram of the hardware environment of a virtual reality device according to the model deployment method of the embodiments of the application. In the application scenario as shown in Figure 1 , the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of networks. The client device 20 here can include but is not limited to: a smartphone, a tablet computer, a notebook computer, a palm computer, a personal computer, a smart home device, a vehicle-mounted device, etc. The client device 20 can interact with the user through a graphical user interface to realize the calling of the large model, and thus realize the method provided by the embodiments of the application.
[0040] In the embodiments of the application, the system composed of the client device and the server can execute the following steps: the client device executes the model file of the to-be-deployed model; the server executes the model file of the to-be-deployed model, and a plurality of server clusters; distributes the model file to a plurality of storage devices; and mounts the storage devices to the target server cluster in the same region as the storage devices are deployed, so as to deploy the to-be-deployed model to the plurality of server clusters. It should be noted that in the case that the running resource of the client device can meet the deployment and running conditions of the large model, the embodiments of the application can be performed in the client device.
[0041] Under the above running environment, the application provides a model deployment method as shown in Figure 2 . Figure 2 is a flowchart of the model deployment method according to Embodiment 1 of the application. As shown in Figure 2 , the method can include the following steps:
[0042] Step S202, obtaining a model file of a to-be-deployed model.
[0043] The to-be-deployed model described above can be any model to be deployed. The to-be-deployed model can be a large model, a neural network model, a language processing model, or an image processing model, where the large model can refer to a machine learning or deep learning model with a large number of parameters and strong computing power. In the embodiments of the present application, the to-be-deployed model is taken as an example of a large model, and the to-be-deployed model can be determined according to actual needs.
[0044] The model file described above can be a model file of a trained to-be-deployed model, where the model file includes but is not limited to parameters and structures of the model. Here, only an example of the model file is given, and the specific model file can be defined according to actual use.
[0045] In an optional embodiment, the model file of the to-be-deployed model can be obtained, and a plurality of server clusters can be determined, where different server clusters are deployed in different regions in terms of geographical location.
[0046] The plurality of server clusters described above can be used to provide services for users. The different regions described above can be different countries, different provinces, different cities, etc., and the different regions are not limited specifically herein.
[0047] In an optional embodiment, after the to-be-deployed model is trained and the model file is generated, the model file can be stored in a cloud server. When the to-be-deployed model is deployed, the model file can be called from the cloud server and distributed to storage devices in a plurality of different regions.
[0048] Step S204, distributing the model file to a plurality of storage devices.
[0049] Wherein, different storage devices are deployed in different regions in terms of geographical location.
[0050] The storage device described above can be a hardware storage device, for example, it can be a NAS, but is not limited thereto.
[0051] In an optional embodiment, the model file can be distributed to a plurality of storage devices according to a distribution request initiated by a user. Optionally, the model management platform can call the model file from the cloud server according to the distribution request initiated by the user, and distribute the model file to the plurality of storage devices.
[0052] Step S206, mounting the storage device to a target server cluster in the plurality of server clusters that is deployed in the same region as the storage device, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0053] The mounting described above refers to connecting the storage device to the servers of the target server cluster, so that the model files in the storage device can be visible and accessible on the servers of the target server cluster.
[0054] In an optional embodiment, the storage device can be mounted on the target server cluster in the same region as the storage device is deployed, so that the user can quickly access the model files on the target server cluster through the mounted storage device, thereby more efficiently using the services provided by the model.
[0055] Through the above steps, the model file of the to-be-deployed model can be obtained; the model file is distributed to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; the storage device is mounted on a target server cluster in a plurality of server clusters in the same region as the storage device is deployed, so that the to-be-deployed model is deployed to the plurality of server clusters, thereby improving the processing efficiency of the model in different regions; it is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to realize high-speed interconnection between the storage devices and the server clusters in the same region by using the network between the storage devices and the server clusters, eliminate the transmission of cross-regional files, thereby shorten the loading time of the model file, thereby improve the processing efficiency of the model in different regions, and further solve the technical problem of low efficiency of model cross-regional deployment in the related art.
[0056] In the above embodiments of the present application, the storage device comprises a network attached storage, wherein distributing the model file to a plurality of storage devices comprises sending the model file to a plurality of network attached storages through a public network and storing the model file in the plurality of network attached storages.
[0057] The network attached storage described above can be a storage device that provides file sharing services through a network, and has other additional functions, wherein the network attached storage can also provide a convenient and reliable way for users to store and share data.
[0058] The public network described above can be an open, shared and widely covered network.
[0059] In an optional embodiment, since the coverage range of the public network is wide, the model file can be sent to a plurality of network attached storages through the public network, so that the plurality of network attached storages can all receive the model file, and the plurality of network attached storages can store the model file after receiving the model file.
[0060] In another optional embodiment, in order to improve the security of the model file storage process, the model file can be encrypted to obtain an encrypted model file before being sent to the plurality of network attached storages through the public network, and then the encrypted model file is sent to the plurality of network attached storages through the public network. After receiving the encrypted model file, the plurality of network attached storages can decrypt the encrypted model file to obtain the model file and store the model file.
[0061] In the above embodiments of the present application, when deploying the inference service of the to-be-deployed model, the storage device is mounted to the target server cluster in the same region as the storage device in the plurality of server clusters, including: determining the target virtual private network corresponding to the storage device based on the target region where the storage device is deployed; obtaining the server cluster corresponding to the target virtual private network to obtain the target server cluster, wherein different server clusters correspond to different virtual private networks; and mounting the storage device to the server in the target server cluster when deploying the inference service of the to-be-deployed model.
[0062] The above target region can be determined according to the geographical position of the storage device, or can be determined according to the country, province, etc. of the storage device deployment. The determination method of the target region is not limited here, and the method of determining the target region of the storage device deployment can be set according to the actual situation.
[0063] The above target virtual private network (Virtual Private Cloud, VPC for short) can provide an isolated and secure way to deploy and manage resources, and can also improve the transmission speed between files. Optionally, the target virtual private network can be a virtual network environment created in a cloud computing environment.
[0064] In an optional embodiment, the target virtual private network corresponding to the storage device can be determined based on the target region where the storage device is deployed, and the server cluster connected to the target virtual private network can be obtained. One or more server clusters connected to the target virtual private network are selected as the target server cluster. Optionally, the server cluster closest to the storage device in geographical position can be selected from the server clusters connected to the target virtual private network as the target server cluster. The method of determining the target server cluster is only illustrative and is not specifically limited, and can be set according to the actual situation.
[0065] In another optional embodiment, after obtaining the target region where the storage device is deployed, the server cluster in the target region can be searched according to the target virtual private network, and one or more server clusters in the target region can be selected as the target server cluster. Alternatively, the server cluster closest to the storage device can be selected as the target server cluster. The specific method of determining the target server cluster is not limited here.
[0066] In yet another optional embodiment, when deploying the inference service of the to-be-deployed model, the storage device can be connected to a server in the target server cluster, so that the model file in the storage device can be read through the intranet on the server.
[0067] In the above embodiments of the present application, after mounting the storage device to the target server cluster in the same region as the storage device in the plurality of server clusters, the method further comprises: constructing an elastic scheduling cluster of the target server cluster; and determining the computing resources required for deploying the to-be-deployed model based on the elastic scheduling cluster.
[0068] The elastic scheduling cluster described above is a cluster management method based on cloud computing and virtualization technology, which can automatically adjust the computing resources in the cluster according to the changes in system load to meet different needs. The elastic scheduling cluster can adjust the cluster size according to the load, including increasing or decreasing computing nodes to cope with high load or low load, which can improve the flexibility and availability of the system, and also save resources and costs.
[0069] The elastic scaling capability described above can automatically adjust the resources and capacity of the target service cluster according to real-time needs and load conditions to meet the needs of users; the elastic scaling capability can automatically increase or decrease the number of server instances in the target server cluster according to different conditions to achieve load balancing and high availability.
[0070] In an optional embodiment, the elastic scheduling cluster can be constructed according to an online service platform (Elastic Algorithm Service, EAS for short) or a container orchestration platform (Kubernetes, K8S for short), and the elastic scaling capability can be provided for the target service cluster according to the elastic scheduling cluster, so that the target service cluster can automatically adjust the system resources and capacity to meet the needs of users according to the actual needs and load conditions.
[0071] In another optional embodiment, when the load of the target server cluster is low, the number of servers can be reduced to save costs, and when the load of the target server cluster is high, the number of servers can be increased to improve performance and capacity; optionally, the number of servers in the target server cluster can be managed according to pre-set rules, or the number of servers in the target server cluster can be dynamically adjusted according to real-time monitoring and analysis.
[0072] In yet another optional embodiment, after building the elastic scheduling cluster of the target server cluster, the computing resources required for deploying the to-be-deployed model, such as memory, storage, etc., need to be determined, and the computing resources are configured and allocated according to the computing requirements of the model. By determining the computing resources required for deployment, it can be ensured that the to-be-deployed model can normally run on the elastic scheduling cluster and meet the required performance and availability requirements, and sufficient computing resources can be provided by building the elastic scheduling cluster to support the deployment and running of the to-be-deployed model.
[0073] In the above embodiments of the present application, the method further comprises: distributing the preset resources to the plurality of server clusters; and distributing the model file to the plurality of storage devices, including: distributing the model file to the plurality of storage devices in the case that the distribution of the preset resources is completed.
[0074] The above-mentioned preset resources can be various resources shared and allocated in the server cluster, wherein the preset resources can be computing resources, network resources, storage resources, software resources, and load balancing resources, which are only described as examples here, and the specific preset resources can be set according to actual conditions.
[0075] The above-mentioned preset resources can be pre-set resource types and resource amounts, and the preset resources are not limited here.
[0076] The above-mentioned preset resources can also be determined according to the computing resources required for deploying the to-be-deployed model, and the preset resources are distributed to the plurality of server clusters, so as to facilitate subsequent deployment of the to-be-deployed model by the plurality of server clusters.
[0077] In an optional embodiment, the model management platform can be used to provide preset resources, which include but are not limited to management cluster metadata, service deployment, automatic / periodic scaler, monitoring data, model distribution, monitoring resource usage, and gray / rollback.
[0078] In another optional embodiment, the preset resources can be distributed to the plurality of server clusters, and the model file can be distributed to the plurality of storage devices in the case that the distribution of the preset resources is completed, so that the plurality of storage devices can use the preset resources of the target servers in the server cluster to deploy the to-be-deployed model according to the model file.
[0079] In the above embodiments of the present application, the model file of the to-be-deployed model is obtained from the central warehouse, wherein the model file is pre-uploaded to the central warehouse.
[0080] The central warehouse described above can be an object storage service (OSS), but is not limited thereto, which is only an example. The central warehouse can be used to provide cloud storage services and can be used to store and manage massive unstructured data.
[0081] The central warehouse described above can communicate with the server cluster and can also perform file transmission with the storage devices on the server cluster. According to the model deployment instruction of the user, the central warehouse can automatically distribute the model file of the to-be-deployed model to the plurality of storage devices mounted on the plurality of server clusters.
[0082] In an optional embodiment, since the central warehouse has a large memory, after the training of the to-be-deployed model is completed, the model file of the to-be-deployed model can be pre-uploaded to the central warehouse. When the to-be-deployed model needs to be deployed, the model file corresponding to the to-be-deployed model can be retrieved from the central warehouse according to the model deployment instruction, and the model file can be distributed to the plurality of storage devices.
[0083] In the above embodiments of the present application, the to-be-deployed model is a large model.
[0084] The large model described above can be a model with a large storage capacity and strong computing power.
[0085] Figure 3 FIG. 1 is a structural diagram of a model deployment method according to an embodiment of the present application. As shown in FIG. 1, a model management and control platform can store a model file of a to-be-deployed model in a central warehouse, wherein the model management and control platform can be used to provide a preset resource, such as cluster metadata management, service deployment, automatic / periodic scaler, monitoring data, model distribution, monitoring resource usage, gray / rollback, and multiple account management. The preset resource can be distributed to a plurality of server clusters, and after the distribution of the preset resource is completed, the model file can be distributed to a plurality of storage devices. Figure 3 Figure 3 As shown in FIG. 1, the system can include region A, region B, and region C, wherein region A includes server cluster A, server cluster B, and network-attached storage, region B includes server cluster A, server cluster B, and network-attached storage, and region C includes server cluster A, server cluster B, and network-attached storage. The network-attached storage in different regions can be mounted on the servers in the server cluster, so that the agents in the server cluster can store the model files in the network-attached storage, and the servers in the server cluster can load the model files in the network-attached storage. The elastic scheduling cluster can be constructed according to the online service platform or the container orchestration platform, and the elastic scheduling cluster can provide the target service cluster with the elastic scaling capability, so that the target service cluster can automatically adjust the system resources and capacity according to the actual demand and load condition to meet the demand of the user.
[0086] The present application can include two model deployment scenarios. The first model deployment scenario can be a model, program, data, and image separate deployment scenario, Figure 4 is a flowchart of another model deployment method according to an embodiment of the present application, as shown in Figure 4 The method includes the following steps:
[0087] In step S401, the preset resources are distributed to the plurality of server clusters.
[0088] Figure 3 The model management platform in the system can be used to provide the above-mentioned preset resources.
[0089] In step S402, the model files are distributed to the plurality of storage devices after the distribution of the preset resources is completed.
[0090] Optionally, if the model files are distributed to the plurality of storage devices before the distribution of the preset resources is completed, an error prompt is returned, and the step of distributing the model files to the plurality of storage devices is performed again after the distribution of the preset resources is completed.
[0091] In step S403, the target virtual private network corresponding to the storage device is determined based on the target region where the storage device is deployed.
[0092] In step S404, the server cluster corresponding to the target virtual private network is obtained, and the target server cluster is obtained.
[0093] In step S405, the storage device is mounted on the server in the target server cluster when deploying the inference service of the to-be-deployed model.
[0094] The second model deployment scenario can be an image deployment scenario, and in this process, the resource distribution process does not need to be performed, and the service deployment can be directly performed, Figure 5is a flowchart of another model deployment method according to an embodiment of the present application, as shown in the figure, the method comprises: Figure 5
[0095] In step S501, a model file of a to-be-deployed model is acquired.
[0096] In step S502, the model file is distributed to a plurality of storage devices.
[0097] In step S503, a target virtual private network corresponding to the storage device is determined based on a target region where the storage device is deployed.
[0098] In step S504, a target server cluster corresponding to the target virtual private network is acquired, to obtain the target server cluster.
[0099] In step S505, when deploying an inference service of the to-be-deployed model, the storage device is mounted to a server in the target server cluster.
[0100] The following are the experimental results of single machine testing. The present application mainly uses NAS to mount the model file. A 13B parameter GPT model (25GB size) is used to compare OSS mounting of the model file and NAS mounting of the model file:
[0101] 1. Test using the dd command, time dd if= / home / workspace / model / pytorch_model.bin of= / dev / null bs=100K count=100000000;
[0102] a. Use OSS to mount the model file. 25716611305 bytes (26GB, 24GiB) have been copied, 115.811s, 222MB / s;
[0103] b. Use NAS to mount the model file. 25716611305 bytes (26GB, 24GiB) have been copied, 4.23957s, 6.1GB / s;
[0104] 2. Large model inference service startup speed test
[0105] a. Use OSS to mount the model file. The total service startup time is 5m28.259s;
[0106] b. Use NAS to mount the model file. The total service startup time is 4m2.845s;
[0107] Through comparison, it can be seen that the NAS storage startup time is better than OSS.
[0108] The application aims to provide a system for improving the multi-region and large-scale elastic capability of a large model service. The system is based on NAS storage, realizes distribution of model files in multiple regions, and automatically mounts the NAS of the local region to load the model file when deploying the inference service. Since the NAS and the inference service are in the same regional VPC, high-speed interconnection can be achieved, thereby shortening the model loading time, accelerating the service startup speed, and improving the elastic capability.
[0109] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0110] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0111] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, and of course it can also be realized by hardware. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method of each embodiment of the present application.
[0112] Embodiment 2
[0113] According to the embodiments of the present application, a model deployment method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0114] Figure 6 is a flowchart of a model deployment method according to Embodiment 2 of the present application, as Figure 6As shown, the method comprises the following steps:
[0115] In step S602, in response to receiving a model file of a model to be deployed, the model file is stored to a central repository.
[0116] In step S604, in response to receiving a model distribution request, the model file is distributed to a plurality of storage devices.
[0117] The different storage devices are deployed in different regions in terms of geographical location, and different server clusters are also deployed in the different regions.
[0118] The model distribution request can be a request for pre-distribution of the model to the storage devices when the model to be deployed is deployed in different regions. Optionally, the model distribution request can be generated by touching the interactive interface, and the specific manner of generating the model distribution request is not limited herein.
[0119] In step S606, in response to receiving a server deployment request, the storage device is mounted to a target server cluster in the plurality of server clusters that is deployed in the same region as the storage device based on the server deployment request, so that the model to be deployed is deployed to the plurality of server clusters.
[0120] The server deployment request can be a request generated by a user when the user needs to deploy the model to be deployed. Optionally, the server deployment request can be generated by touching the interactive interface, and the specific manner of generating the server deployment request is not limited herein.
[0121] Through the above steps, in response to receiving a model file of a model to be deployed, the model file is stored to a central repository; in response to receiving a model distribution request, the model file is distributed to a plurality of storage devices, wherein the different storage devices are deployed in different regions in terms of geographical location, and different server clusters are also deployed in the different regions; and in response to receiving a server deployment request, the storage device is mounted to a target server cluster in the plurality of server clusters that is deployed in the same region as the storage device based on the server deployment request, so that the model to be deployed is deployed to the plurality of server clusters, thereby improving the processing efficiency of the model in different regions. It is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to realize high-speed interconnection between the storage devices and the server clusters in the same region, eliminate cross-regional file transmission, thereby shortening the loading time of the model file, improving the processing efficiency of the model in different regions, and further solving the technical problem of low efficiency of model cross-regional deployment in the related art.
[0122] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same scheme, application scenario and implementation process as provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0123] Embodiment 3
[0124] According to the embodiments of the present application, a model deployment method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0125] Figure 7 is a flowchart of a model deployment method according to Embodiment 3 of the present application, as shown in Figure 7 the method comprises the following steps:
[0126] Step S702, obtaining a model file of a to-be-deployed model by calling a first interface.
[0127] The first interface includes a first parameter, and the parameter value of the first parameter is the model file.
[0128] The first interface in the above steps can be an interface for data interaction between the cloud server and the client. The client can pass the model file of the to-be-deployed model into the interface function as the first parameter of the interface function, so as to achieve the purpose of uploading the model file of the to-be-deployed model to the cloud server.
[0129] Step S704, distributing the model file to a plurality of storage devices.
[0130] Different storage devices are deployed in different regions in terms of geographical location.
[0131] Step S706, when deploying an inference service of the to-be-deployed model, mounting the storage device to a target server cluster in a plurality of server clusters which is deployed in the same region as the storage device, so as to deploy the to-be-deployed model to the plurality of server clusters, and obtain a deployment result of the to-be-deployed model.
[0132] Step S708, outputting the deployment result by calling a second interface.
[0133] The second interface includes a second parameter, and the parameter value of the second parameter is the deployment result.
[0134] The second interface described above can be an interface for data interaction between the cloud server and the client. The cloud server can pass the deployment result into the interface function as the second parameter of the interface function, so as to achieve the purpose of issuing the deployment result to the client.
[0135] By the above steps, the model file of the to-be-deployed model is obtained by calling the first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter is the model file; the model file is distributed to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; when deploying an inference service of the to-be-deployed model, the storage device is mounted on a target server cluster in a plurality of server clusters which deploys the storage device in the same region, so that the to-be-deployed model is deployed to the plurality of server clusters to obtain a deployment result of the to-be-deployed model; and the deployment result is output by calling the second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter is the deployment result, thereby improving the processing efficiency of the model in different regions. It is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to realize high-speed interconnection between the storage devices and the server clusters in the same region by using the network between the storage devices and the server clusters, eliminate the transmission of cross-regional files, thereby shortening the loading time of the model file, improving the processing efficiency of the model in different regions, and further solving the technical problem of low efficiency of model cross-regional deployment in the related art.
[0136] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios and implementation processes as the schemes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0137] Embodiment 4
[0138] According to the embodiments of the present application, a model deployment device for implementing the above model deployment method is further provided, Figure 8 is a schematic diagram of a model deployment device according to Embodiment 4 of the present application, as Figure 8 shown, the device 800 includes an acquisition module 802, a distribution module 804, and a mounting module 806.
[0139] The acquisition module is configured to acquire a model file of a to-be-deployed model; the distribution module is configured to distribute the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; and the mounting module is configured to mount the storage device on a target server cluster in a plurality of server clusters which deploys the storage device in the same region, so that the to-be-deployed model is deployed to the plurality of server clusters.
[0140] It should be noted that the above obtaining module 802, distribution module 804, and mounting module 806 correspond to steps S202 to S206 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the above disclosed Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in the memory and processed by one or more processors, and the above modules can also be run in the server 10 provided in Embodiment 1 as part of the device.
[0141] In the above embodiments of the present application, the storage device includes a network attached storage, wherein the distribution module is further configured to send the model file to the plurality of network attached storages through the public network and store the model file in the plurality of network attached storages.
[0142] In the above embodiments of the present application, the mounting module is further configured to determine a target virtual private network corresponding to the storage device based on a target region where the storage device is deployed, obtain a target server cluster corresponding to the target virtual private network, and obtain the target server cluster, wherein different server clusters correspond to different virtual private networks, and the storage device is mounted to a server in the target server cluster when deploying an inference service of the to-be-deployed model.
[0143] In the above embodiments of the present application, the mounting module is further configured to determine a target virtual private network corresponding to the storage device based on a target region where the storage device is deployed, obtain a target server cluster corresponding to the target virtual private network, and obtain the target server cluster, wherein different server clusters correspond to different virtual private networks.
[0144] In the above embodiments of the present application, the device further includes a construction module and a providing module.
[0145] The construction module is configured to construct an elastic scheduling cluster, and the providing module is configured to provide the target server cluster with an elastic scaling capability based on the elastic scheduling cluster.
[0146] In the above embodiments of the present application, the distribution module is further configured to distribute a preset resource to the plurality of server clusters, and the distribution module is further configured to distribute the model file to the plurality of storage devices in a case where the distribution of the preset resource is completed.
[0147] In the above embodiments of the present application, the obtaining module is further configured to obtain a model file of the to-be-deployed model from a central warehouse, wherein the model file is pre-uploaded to the central warehouse.
[0148] In the above embodiments of the present application, the to-be-deployed model is a large model.
[0149] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same scheme, application scenario and implementation process as provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0150] Embodiment 5
[0151] According to the embodiments of the present application, a model deployment device for implementing the above model deployment method is further provided, Figure 9 is a schematic diagram of a model deployment device according to Embodiment 5 of the present application, as Figure 9 shown, the device 900 includes a storage module 902, a distribution module 904, and a mounting module 906.
[0152] The storage module is configured to store the model file into the central repository in response to receiving the model file of the model to be deployed; the distribution module is configured to distribute the model file stored in the central repository to a plurality of storage devices in response to receiving the model distribution request, wherein different storage devices are deployed in different regions in geographical position, and different server clusters are also deployed in different regions; and the mounting module is configured to mount the storage device to a target server cluster in the same region as the storage device in the plurality of server clusters based on the server deployment request, so as to deploy the model to be deployed to the plurality of server clusters.
[0153] It should be noted that the storage module 902, the distribution module 904, and the mounting module 906 correspond to steps S602 to S606 in Embodiment 2, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in the memory and processed by one or more processors, and the above modules can also be a part of the device and can run in the server 10 provided in Embodiment 1.
[0154] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same scheme, application scenario and implementation process as provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0155] Embodiment 6
[0156] According to the embodiments of the present application, a model deployment device for implementing the above model deployment method is further provided, Figure 10 is a schematic diagram of a model deployment device according to Embodiment 6 of the present application, as Figure 10 shown, the device 1000 includes an acquisition module 1002, a distribution module 1004, a mounting module 1006, and an output module 1008.
[0157] The obtaining module is configured to obtain a model file of a to-be-deployed model by calling a first interface, wherein the first interface comprises a first parameter, and a parameter value of the first parameter is the model file.
[0158] It should be noted that the obtaining module 1002, the distribution module 1004, the mounting module 1006, and the output module 1008 correspond to steps S702 to S708 in Embodiment 3, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be run in the server 10 provided in Embodiment 1 as part of the device.
[0159] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios, implementation processes as the scheme provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0160] Embodiment 7
[0161] According to the embodiments of the present application, a model deployment system for implementing the above model deployment method is further provided, Figure 11 is a schematic diagram of a model deployment system according to Embodiment 7 of the present application, as Figure 11 shown, the system 1100 comprises:
[0162] a plurality of server clusters 1102, different server clusters are deployed in different regions in terms of geographical location;
[0163] a plurality of storage devices 1104, different storage devices are deployed in different regions in terms of geographical location;
[0164] a control device 1106 connected with the storage device and the server cluster, configured to distribute a model file of a to-be-deployed model to the storage device, and mount the storage device to a target server cluster deployed in the same region as the storage device, so that the to-be-deployed model is deployed to a plurality of server clusters.
[0165] In the above embodiments of the present application, the storage device 1104 comprises a network-attached storage connected to the control device through a public network.
[0166] In the above embodiments of the present application, the target server cluster is connected to the storage device through a target virtual private network.
[0167] In the above embodiments of the present application, the system further comprises an elastic scheduling cluster for providing elastic scaling capability for the target server cluster.
[0168] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same scheme, application scenario and implementation process as provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0169] Embodiment 8
[0170] The embodiments of the present application can provide an electronic device, which can be any one of the electronic devices in the electronic device group. Alternatively, in the present embodiment, the above electronic device can also be replaced by a terminal device such as a mobile terminal.
[0171] Alternatively, in the present embodiment, the above electronic device can be located in at least one network device of a plurality of network devices of a computer network.
[0172] In the present embodiment, the above electronic device can execute program codes of the following steps in the model deployment method: obtaining a model file of a to-be-deployed model and a plurality of server clusters, wherein different server clusters are deployed in different regions in terms of geographical location; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; and mounting the storage devices to target server clusters deployed in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0173] Alternatively, Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 12 the computer terminal A can include one or more (only one is shown in the figure) processors 102, a memory 104, a storage controller, and a peripheral interface, wherein the peripheral interface is connected with a radio frequency module, an audio module, and a display.
[0174] The memory can be configured to store software programs and modules, such as program instructions / modules corresponding to the model deployment method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned model deployment method. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0175] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a model file of a to-be-deployed model; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; mounting the storage devices to a target server cluster in a plurality of server clusters in which the storage devices are deployed in the same region, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0176] Optionally, the above-mentioned processor can further execute program codes of the following steps: sending the model file to a plurality of network-attached storages through a public network and storing the model file in the plurality of network-attached storages.
[0177] Optionally, the above-mentioned processor can further execute program codes of the following steps: determining a target server cluster deployed in a target region based on the target region in which the storage device is deployed; and mounting the storage device to a server in the target server cluster.
[0178] Optionally, the above-mentioned processor can further execute program codes of the following steps: determining a target virtual private network corresponding to the storage device based on the target region in which the storage device is deployed; obtaining a server cluster corresponding to the target virtual private network to obtain a target server cluster, wherein different server clusters correspond to different virtual private networks; and mounting the storage device to a server in the target server cluster when deploying an inference service of the to-be-deployed model.
[0179] Optionally, the above-mentioned processor can further execute program codes of the following steps: constructing an elastic scheduling cluster; and providing an elastic scaling capability for the target server cluster based on the elastic scheduling cluster.
[0180] Optionally, the above-mentioned processor can further execute program codes of the following steps: distributing a preset resource to a plurality of server clusters; and distributing the model file to a plurality of storage devices in a case where the distribution of the preset resource is completed.
[0181] Optionally, the processor can further execute program codes of the following steps: obtaining a model file of the model to be deployed from the central warehouse, wherein the model file is pre-uploaded to the central warehouse.
[0182] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: in response to receiving the model file of the model to be deployed, storing the model file to the central warehouse; in response to receiving the model distribution request, distributing the model file to the plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position, and different server clusters are also deployed in different regions; and in response to receiving the server deployment request, based on the server deployment request, mounting the storage device to a target server cluster in the plurality of server clusters in the same region as the storage device is deployed, so as to deploy the model to be deployed to the plurality of server clusters.
[0183] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: obtaining a model file of the model to be deployed by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter is the model file; distributing the model file to the plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; when deploying the inference service of the model to be deployed, mounting the storage device to a target server cluster in the plurality of server clusters in the same region as the storage device is deployed, so as to deploy the model to be deployed to the plurality of server clusters to obtain a deployment result of the model to be deployed; and outputting the deployment result by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter is the deployment result.
[0184] By adopting the embodiment of the present application, a model deployment method is provided, which includes: obtaining a model file of a model to be deployed; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; and mounting the storage device to a target server cluster in a plurality of server clusters in the same region as the storage device is deployed, so as to deploy the model to be deployed to the plurality of server clusters, thereby improving the processing efficiency of the model in different regions. It is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to realize high-speed interconnection between the storage devices and the server clusters in the same region by using the network between the storage devices and the server clusters, eliminate the transmission of cross-regional files, thereby shortening the loading time of the model file, improving the processing efficiency of the model in different regions, and further solving the technical problem of low efficiency of model cross-regional deployment in the related art.
[0185] Those skilled in the art can understand that, Figure 12The structure shown is only schematic. The computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or other terminal device. Figure 12 The above electronic device is not limited in structure. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1. Figure 12 Figure 12 The above electronic device is not limited in structure. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1.
[0186] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the hardware related to the terminal device by a program, which can be stored in a computer readable storage medium, and the storage medium can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.
[0187] Embodiment 6
[0188] The embodiments of the present application also provide a computer readable storage medium. Optionally, in the embodiment, the above-mentioned storage medium can be used to save the program code executed by the model deployment method provided in Embodiment 1.
[0189] Optionally, in the embodiment, the above-mentioned storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0190] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a model file of a to-be-deployed model, wherein different server clusters are deployed in different regions in geographical position; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in geographical position; and mounting the storage devices to target server clusters in the plurality of server clusters which are deployed in the same region as the storage devices, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0191] Optionally, the above-mentioned storage medium is further configured to store program code for performing the following steps: sending the model file to a plurality of network-attached storages through a public network and storing the model file in the plurality of network-attached storages.
[0192] Optionally, the storage medium is further configured to store program code for performing the following steps: determining a target virtual private network corresponding to the storage device based on a target region where the storage device is deployed; obtaining a target server cluster corresponding to the target virtual private network to obtain a target server cluster, wherein different server clusters correspond to different virtual private networks; and mounting the storage device to a server in the target server cluster when deploying an inference service of the to-be-deployed model.
[0193] Optionally, the storage medium is further configured to store program code for performing the following steps: constructing an elastic scheduling cluster; and providing an elastic scaling capability for the target server cluster based on the elastic scheduling cluster.
[0194] Optionally, the storage medium is further configured to store program code for performing the following steps: distributing a preset resource to a plurality of server clusters; and distributing a model file to a plurality of storage devices in a case where the distribution of the preset resource is completed.
[0195] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a model file of a to-be-deployed model from a central warehouse, wherein the model file is pre-uploaded to the central warehouse.
[0196] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: storing the model file to the central warehouse in response to receiving the model file of the to-be-deployed model; distributing the model file to a plurality of storage devices in response to receiving a model distribution request, wherein different storage devices are deployed in different regions in terms of geographical location, and different server clusters are also deployed in different regions; and mounting the storage device to a target server cluster in a plurality of server clusters that is deployed in the same region as the storage device based on a server deployment request, so as to deploy the to-be-deployed model to the plurality of server clusters.
[0197] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a model file of a to-be-deployed model by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter is the model file; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; mounting the storage device to a target server cluster in a plurality of server clusters that is deployed in the same region as the storage device when deploying an inference service of the to-be-deployed model, so as to deploy the to-be-deployed model to the plurality of server clusters to obtain a deployment result of the to-be-deployed model; and outputting the deployment result by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter is the deployment result.
[0198] By adopting the embodiment of the present application, a model deployment method is provided, which comprises: obtaining a model file of a to-be-deployed model; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical position; and mounting the storage devices to a target server cluster in a plurality of server clusters which is in the same region as the storage devices are deployed in, so as to deploy the to-be-deployed model to the plurality of server clusters, thereby improving the processing efficiency of the model in different regions. It is easy to note that the model file can be distributed to the storage devices deployed in different regions, so as to realize high-speed interconnection between the storage devices and the server clusters in the same region by using the network between the storage devices and the server clusters in the same region, eliminate the transmission of cross-regional files, thereby shortening the loading time of the model file, improving the processing efficiency of the model in different regions, and further solving the technical problem of low efficiency of model cross-regional deployment in the related art.
[0199] The above sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0200] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0201] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the embodiment described above is only a schematic and illustrative, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.
[0202] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0203] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0204] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0205] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.
Claims
1. A model deployment method, characterized by, The method comprises: obtaining a model file of a to-be-deployed model; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; when deploying an inference service of the to-be-deployed model, mounting the storage devices to a target server cluster in a plurality of server clusters in which the storage devices are deployed in the same region, so that the to-be-deployed model is deployed to the plurality of server clusters, wherein the mounting is connecting the storage devices to servers in the target server cluster, and the storage devices are devices providing file sharing services through a network.
2. The method of claim 1, wherein, The storage devices comprise network-attached storages, wherein distributing the model file to a plurality of storage devices comprises: sending the model file to a plurality of network-attached storages through a public network and storing the model file in the plurality of network-attached storages.
3. The method of claim 1, wherein, When deploying the inference service of the to-be-deployed model, mounting the storage devices to the target server cluster in the plurality of server clusters in which the storage devices are deployed in the same region comprises: determining a target virtual private network corresponding to the storage devices based on a target region in which the storage devices are deployed; obtaining a server cluster corresponding to the target virtual private network to obtain the target server cluster, wherein different server clusters correspond to different virtual private networks; mounting the storage devices to the servers in the target server cluster.
4. The method of claim 1, wherein, After mounting the storage devices to the target server cluster in the plurality of server clusters in which the storage devices are deployed in the same region, the method further comprises: constructing an elastic scheduling cluster of the target server cluster; determining computing resources required for deploying the to-be-deployed model based on the elastic scheduling cluster.
5. The method of claim 1, wherein, The method further comprises: distributing a preset resource to a plurality of server clusters; distributing the model file to a plurality of storage devices comprises: distributing the model file to a plurality of storage devices after the distribution of the preset resource is completed.
6. The method of claim 1, wherein, Obtaining a model file of a to-be-deployed model comprises: obtaining the model file of the to-be-deployed model from a central repository, wherein the model file is pre-uploaded to the central repository.
7. The method of claim 1, wherein, The to-be-deployed model is a large language model.
8. A model deployment method characterized by comprising: The method comprises: in response to receiving a model file of a to-be-deployed model, storing the model file in a central repository; in response to receiving a model distribution request, distributing the model file stored in the central repository to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location, and different server clusters are also deployed in the different regions; In response to receiving the server deployment request, based on the server deployment request, mounting the storage device to a target server cluster in a plurality of server clusters in a same region as the storage device is deployed in, so as to deploy the to-be-deployed model to the plurality of server clusters, wherein the mounting is connecting the storage device to a server in the target server cluster, and the storage device is a device providing a file sharing service through a network.
9. A model deployment method characterized by comprising: The method comprises: obtaining a model file of a to-be-deployed model by calling a first interface, wherein the first interface comprises a first parameter, and a parameter value of the first parameter is the model file; distributing the model file to a plurality of storage devices, wherein different storage devices are deployed in different regions in terms of geographical location; when deploying an inference service of the to-be-deployed model, mounting the storage device to a target server cluster in a plurality of server clusters in a same region as the storage device is deployed in, so as to deploy the to-be-deployed model to the plurality of server clusters, to obtain a deployment result of the to-be-deployed model; outputting the deployment result by calling a second interface, wherein the second interface comprises a second parameter, and a parameter value of the second parameter is the deployment result, wherein the mounting is connecting the storage device to a server in the target server cluster, and the storage device is a device providing a file sharing service through a network.
10. A model deployment system, comprising: The method comprises: a plurality of server clusters, different server clusters are deployed in different regions in terms of geographical location; a plurality of storage devices, different storage devices are deployed in the different regions in terms of geographical location; a control device connected to the plurality of storage devices and the plurality of server clusters, for distributing a model file of a to-be-deployed model to a plurality of the storage devices, and when deploying an inference service of the to-be-deployed model, mounting the storage device to a target server cluster in a same region as the storage device is deployed in, so as to deploy the to-be-deployed model to the plurality of server clusters, wherein the mounting is connecting the storage device to a server in the target server cluster, and the storage device is a device providing a file sharing service through a network.
11. The system of claim 10, wherein, The storage device comprises: a network-attached storage connected to the control device through a public network.
12. The system of claim 10, wherein, The server cluster and the storage device in the same region are connected through a virtual private network.
13. An electronic device, comprising: The method comprises: a memory storing an executable program; a processor for running the program, wherein the program performs the method of any one of claims 1 to 9 when running.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a stored executable program, wherein the computer-readable storage medium controls the device where the computer-readable storage medium is located to perform the method of any one of claims 1 to 9 when the executable program runs.
Citation Information
Patent Citations
Method for accessing three-dimensional model through distributed architecture
CN115794424A