Deployment method and device of machine learning model, electronic equipment and storage medium
By obtaining the device resource threshold and model requirements, selecting multiple devices and determining the target model deployment method, the problem of single device resource limitations is solved, and efficient machine learning model deployment and seamless container management are achieved.
Patent Information
- Application Number
- CN202410637653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-21
AI Technical Summary
Due to the limited memory space of a single device, it is impossible to support the deployment of some machine learning models. The deployment efficiency of machine learning models in the existing technology is low, and the business side needs to handle the relationship between multiple containers on its own, resulting in an unfriendly deployment.
By obtaining the device resource threshold and the resource requirements of the model, multiple devices are identified, and a model deployment method that matches the target business is selected. The container relationship between devices is indicated, containers are created, and the model is deployed.
It improves the deployment efficiency of machine learning models, eliminates the need for business-side handling of container relationships, and supports customized, seamless usage.
Smart Images

Figure CN120994355A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a machine learning model deployment method and device, electronic equipment and storage medium. BACKGROUND
[0002] Model deployment is a key task in the field of machine learning and artificial intelligence, which involves converting a trained model into a format that can be directly used by an application or service, and deploying it to a target environment. Due to the limitation of video memory space resources of a single device for deploying a machine learning model, the capacity cannot support the deployment of part of the machine learning models (such as a model with a parameter quantity greater than a parameter quantity threshold), which needs multiple devices to cooperate to complete the deployment of the machine learning model. It has become a prominent problem to deploy the machine learning model for online inference to support the application of the scene.
[0003] In the related art, the business side usually needs to handle the relationship between multiple containers for deploying a machine learning model by itself, which is not friendly to the use of the business side, and the deployment efficiency of the machine learning model is low. SUMMARY
[0004] The embodiments of the present application provide a machine learning model deployment method and device, electronic equipment and storage medium, which can improve the deployment efficiency of the machine learning model.
[0005] The technical solutions of the embodiments of the present application are as follows:
[0006] The embodiments of the present application provide a machine learning model deployment method, which comprises:
[0007] Obtaining a device resource quantity threshold and a device resource quantity required by a machine learning model of a deployment target service;
[0008] When the device resource quantity exceeds the device resource quantity threshold, determining multiple devices for deploying the machine learning model;
[0009] From multiple candidate model deployment modes, determining a target model deployment mode matched with the target service, the target model deployment mode indicating a relationship between containers on each of the devices in the multiple devices;
[0010] According to the target model deployment mode, creating at least one container on each of the devices, and deploying the machine learning model into the at least one container on each of the devices.
[0011] The embodiments of the present application provide a machine learning model deployment device, which comprises:
[0012] The model obtaining module is configured to obtain a device resource amount threshold and a device resource amount required by a machine learning model to be deployed for a target service;
[0013] The device determining module is configured to determine a plurality of devices for deploying the machine learning model when the device resource amount exceeds the device resource amount threshold.
[0014] The mode determining module is configured to determine a target model deployment mode matching the target service from a plurality of candidate model deployment modes, the target model deployment mode indicating a relationship between containers on each of the devices in the plurality of devices.
[0015] The model deploying module is configured to create at least one container on each of the devices according to the target model deployment mode, and deploy the machine learning model into the at least one container on each of the devices.
[0016] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:
[0017] The memory is configured to store computer executable instructions or computer programs.
[0018] The processor is configured to execute the computer executable instructions or computer programs stored in the memory to implement the method for deploying a machine learning model provided in the embodiments of the present application.
[0019] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores computer executable instructions or computer programs, and is configured to be executed by a processor to implement the method for deploying a machine learning model provided in the embodiments of the present application.
[0020] A computer program product is provided in an embodiment of the present application, and the computer program product comprises computer executable instructions or computer programs, and the computer executable instructions or computer programs are executed by a processor to implement the method for deploying a machine learning model provided in the embodiments of the present application.
[0021] The embodiments of the present application have the following beneficial effects:
[0022] According to the embodiment of the present application, the device resource threshold and the device resource required by the machine learning model of the target service are obtained, and then when the device resource exceeds the device resource threshold, the devices for deploying the machine learning model are determined, and then the target model deployment mode matched with the target service is determined from the plurality of candidate model deployment modes, the target model deployment mode indicates the relationship between the containers on each device in the plurality of devices, at least one container is created on each device according to the target model deployment mode, and the machine learning model is deployed into the at least one container on each device. In this way, the target model deployment mode matched with the target service can be selected from the plurality of pre-set candidate model deployment modes to realize the deployment of the machine learning model, and the target model deployment mode indicates the relationship between the containers on each device in the plurality of devices, so that the business side does not need to process the relationship between the plurality of containers for deploying the machine learning model by itself, and the deployment efficiency of the machine learning model is improved. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 is an application mode schematic diagram of the machine learning model deployment method provided by the embodiment of the present application;
[0024] Figure 2 is a structural schematic diagram of an electronic device provided by the embodiment of the present application;
[0025] Figure 3A is a first flow schematic diagram of the machine learning model deployment method provided by the embodiment of the present application;
[0026] Figure 3B is a second flow schematic diagram of the machine learning model deployment method provided by the embodiment of the present application;
[0027] Figure 3C is a third flow schematic diagram of the machine learning model deployment method provided by the embodiment of the present application;
[0028] Figure 4 is a structural schematic diagram of an application program module supporting a heterogeneous model inference service provided by the embodiment of the present application;
[0029] Figure 5 is a hybrid module construction process schematic diagram of the machine learning model inference provided by the embodiment of the present application;
[0030] Figure 6 is a logic flow schematic diagram of a small model module processing an inference service request provided by the embodiment of the present application;
[0031] Figure 7 is a logic design flow schematic diagram of a machine learning model deployment unit provided by the embodiment of the present application;
[0032] Figure 8is a flowchart of a modular deployment model provided by an embodiment of the present application.
[0033] It should be noted that the above-mentioned "first", "second" are only used to distinguish different schemes, and do not represent the advantages or disadvantages of the schemes or the priority in the implementation process. DETAILED DESCRIPTION
[0034] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without making creative labor fall within the scope of protection of the present application.
[0035] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0036] In the following description, the terms "first", "second", "third" are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0037] The related data collection processing in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and within the scope of authorization of laws and regulations and the personal information subject, carry out subsequent data use and processing.
[0038] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as processing circuitry or memory) or a combination thereof. Similarly, one processor (or multiple processors or memory) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0040] Before further detailing the embodiments of the present application, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.
[0041] 1) Isomorphic model: refers to a model composed of different types of components or elements, which may differ in nature, structure, function or form, but they work together in the model to complete a specific task or express a specific phenomenon.
[0042] 2) Idempotent: is a concept used to describe the effect of an operation or function when applied multiple times, which is the same as a single application.
[0043] 3) Disaster recovery: is the ability of an organization or enterprise to quickly and effectively recover business operations in the event of a major natural disaster, man-made destruction, or other unexpected events.
[0044] 4) Computing power container: a containerization technology used to deploy and manage computing resources, focusing on providing a high-performance computing and data processing environment. Users can package computing tasks and applications into a container, deploy it to a computing cluster or cloud environment, and use the computing resources in the cluster for efficient parallel computing.
[0045] 5) Workload: an application running on a container orchestration engine, represented by a container group, which is an application, also known as a workload.
[0046] 6) Cloud-native workload: a mode of building and running modern applications on cloud platforms, which can fully utilize the elasticity, scalability and distributed characteristics of cloud computing.
[0047] The embodiments of the present application provide a machine learning model deployment method, device, electronic device and storage medium, which can improve the deployment efficiency of the machine learning model and support business side customization and non-invasive use.
[0048] The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a notebook computer, tablet computer, desktop computer, set-top box, smart phone, smart speaker, smart watch, smart television, vehicle terminal, etc. Various types of terminals, and can also be implemented as a server. The following will illustrate an exemplary application when the device is implemented as a server.
[0049] Referring to Figure 1 , Figure 1 is an application mode schematic diagram of the machine learning model deployment method provided by the embodiments of the present application, and the example Figure 1The application relates to a server 200, a network 300 and a terminal 400. The terminal 400 connects nodes through the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0050] In the deployment process of the machine learning model, the terminal 400 sends a model deployment request to the server 200 in response to a model deployment instruction of a user, the server 200 acquires a device resource amount threshold and a device resource amount required by a machine learning model of a deployment target service in response to the model deployment request, determines a plurality of devices for deploying the machine learning model when the device resource amount exceeds the device resource amount threshold, determines a target model deployment mode matched with the target service from a plurality of candidate model deployment modes, creates at least one container on each device according to the target model deployment mode, deploys the machine learning model into the at least one container on each device, obtains a model deployment result, and returns the model deployment result to the terminal 400. The terminal 400 receives and displays the model deployment result. In this way, the user can learn the deployment mode of the machine learning model through the terminal 400.
[0051] The application embodiment can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology and the like applied based on a cloud computing business model, can form a resource pool, and is used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background service of a technical network system needs a large amount of computing and storage resources, such as a video website, a picture website and more portal websites. With the high development and application of the Internet industry, and the promotion of the demand for search services, social networks, mobile commerce and open collaboration, in the future, each item can have its own hash code identification mark and needs to be transmitted to a background system for logical processing. Different levels of data will be processed separately, and various types of industry data need strong system support, which can only be realized through cloud computing.
[0052] The application embodiment can also be implemented through machine learning. Machine learning (ML) is a multi-field interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and the like. Machine learning is a subject that studies how a computer simulates or implements a human learning behavior to obtain new knowledge or skills, reorganizes an existing knowledge structure to continuously improve the performance of the computer. Machine learning is the core of artificial intelligence and is the fundamental approach to enabling a computer to have intelligence, and is applied in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and inductive learning. A pre-trained model is the latest development achievement of deep learning, which integrates the above technologies.
[0053] The embodiment of the application can also be implemented through a pre-training model. The pre-training model (PTM) is also called a cornerstone model, a large model, and refers to a deep neural network (DNN) with a large number of parameters. The PTM is trained on a large amount of unlabeled data, and the function approximation capability of the large parameter DNN enables the PTM to extract common features from the data. Through fine tuning, parameter efficient fine tuning (PEFT), prompt tuning and other technologies, the PTM is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in a few-shot or zero-shot scenario. The PTM can be divided into a language model (ELMO, BERT, GPT), a visual model (swin-transformer, ViT, V-MOE), a speech model (VALL-E), a multi-modal model (ViBERT, CLIP, Flamingo, Gato) and the like according to the data modality processed. The multi-modal model refers to a model that establishes a feature representation of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface connecting multiple specific task models. The machine learning model in the embodiment of the application can be a pre-training model. By deploying the pre-training model on multiple devices through the embodiment of the application, the efficiency of deploying the pre-training model online can be improved.
[0054] In some embodiments, the server (for example, the server 200) can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, and the like, but is not limited thereto. The terminal and the server can be connected directly or indirectly through wired or wireless communication, and the embodiment of the application is not limited in this regard.
[0055] Referring to Figure 2 , Figure 2 is a structural schematic diagram of an electronic device provided by the embodiment of the application. The electronic device can be a terminal or a server, Figure 2The electronic device 500 shown includes at least one processor 410, a memory 450, at least one network interface 420. The various components in the electronic device 500 are coupled together by a bus system 440. It is understood that the bus system 440 is used for communicating data between the components. The bus system 440 includes a data bus, a power bus, a control bus, and a state signal bus. However, for clarity, only the data bus is shown in Figure 2
[0056] The processor 410 can be an integrated circuit chip that has processing capability, such as a general purpose processor, a Digital Signal Processor (DSP), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc.
[0057] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical drives, etc. The memory 450 optionally includes one or more storage devices remotely located from the processor 410.
[0058] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. Non-volatile memory can be read only memory (ROM), and volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0059] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or superset thereof, which are described below.
[0060] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks.
[0061] The network communication module 452 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 420, examples of which include Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.
[0062] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, Figure 2 A deployment apparatus 455 of the machine learning model stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a model acquisition module 4551, a device determination module 4552, a mode determination module 4553, and a model deployment module 4554. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of the modules will be described below.
[0063] In some embodiments, the terminal or server can implement the data processing method provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; can be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a cloud computing APP or an instant messaging APP; or can be a small program that can be embedded into any APP, i.e., a program that only needs to be downloaded into a browser environment to run. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins.
[0064] The deployment method of the machine learning model provided by the embodiments of the present application will be described in conjunction with an exemplary application and implementation of the server device provided by the embodiments of the present application.
[0065] Next, the deployment method of the machine learning model provided by the embodiments of the present application will be described. As described above, the electronic device implementing the deployment method of the machine learning model provided by the embodiments of the present application can be a terminal, a server, or a combination of the two. Next, the deployment method of the machine learning model provided by the embodiments of the present application will be described taking the electronic device as a server as an example. Referring to Figure 3A , Figure 3A is a first flowchart of the deployment method of the machine learning model provided by the embodiments of the present application, which will be described in conjunction with the steps shown in Figure 3A .
[0066] In step 301, the device resource threshold and the device resource required by the machine learning model of the deployment target service are acquired.
[0067] Here, the device resource amount threshold value can be pre-set; the device resource amount can be the number of devices required for deploying the machine learning model, the amount of video memory space required for deploying the machine learning model, etc. When performing target service deployment access, the machine learning model corresponding to the target service is obtained, and the device resource amount required by the machine learning model is the resource amount size required for the machine learning model to run. The size of the machine learning model is usually measured by indicators such as the number of parameters, the number of network structure layers, and the number of neurons. The machine learning model deployed on the device can process the business inference request of the target service and output the corresponding business inference result.
[0068] In step 302, when the device resource amount exceeds the device resource amount threshold value, a plurality of devices for deploying the machine learning model are determined.
[0069] Here, when the device resource amount required for deploying the machine learning model exceeds the device resource amount threshold value, it is determined that the machine learning model needs to be deployed through a plurality of devices, so the plurality of devices for deploying the machine learning model can be determined, for example, the device identifiers of the plurality of devices for deploying the machine learning model can be obtained to determine which devices to deploy the machine learning model on. When the device resource amount required for deploying the machine learning model exceeds the device resource amount threshold value, the machine learning model is identified as a large model; when the device resource amount required for deploying the machine learning model is lower than the device resource amount threshold value, the machine learning model is identified as a small model.
[0070] For example, when the whole machine device body has 4 graphics cards, each graphics card has a video memory resource amount of 24G, and the video memory space resource of the whole machine device is 96G, the device resource amount threshold value is set to 96G, when the device resource amount required for the machine learning model is greater than 96G, the machine learning model is marked as a large model, and the machine learning model is deployed on a plurality of devices running in cooperation.
[0071] In step 303, from a plurality of candidate model deployment modes, a target model deployment mode matched with the target service is determined.
[0072] Here, the target model deployment mode indicates the relationship between containers on each device in the plurality of devices.
[0073] It should be noted that the plurality of candidate model deployment manners can be pre-set, and when a machine learning model of a target service is deployed, a target model deployment manner matched with the target service can be determined from the plurality of candidate model deployment manners of deploying machine learning models. For example, a first correspondence relationship between a service identifier of the target service and a model deployment manner can be pre-set, and then a target model deployment manner corresponding to the target service is determined according to the service identifier of the target service and the first correspondence relationship; a second correspondence relationship between a service type of the target service and a model deployment manner can be pre-set, and then a target model deployment manner corresponding to the target service is determined according to the service type of the target service and the second correspondence relationship; and a model deployment manner adapted to a service request amount of the target service can also be determined according to the size of the service request amount, such as setting a third correspondence relationship between a request amount interval and a model deployment manner, and then a model deployment manner corresponding to a target interval is determined as the target model deployment manner according to the target interval in which the request amount of the target service is located and the third correspondence relationship. The relationship between the containers on the device is a combined connection relationship between the containers, and the combined connection relationship can be a master-slave connection relationship or an idempotent relationship. According to the target model deployment manner, the relationship between the containers on each device for deploying the machine learning model is configured.
[0074] In step 304, at least one container is created on each device according to the target model deployment manner, and the machine learning model is deployed into the at least one container on each device.
[0075] In some embodiments, when the service request amount of the target service is higher than or equal to a service request amount threshold, a first model deployment manner in the plurality of candidate model deployment manners is determined as the target model deployment manner. For each device, a master container is created on the device according to the first model deployment manner, and at least one slave container connected with the master container is created.
[0076] Here, the first model deployment manner refers to a relationship between containers on each device in the plurality of devices being a master-slave connection relationship. The service request amount threshold can be pre-set according to actual needs, and when the service request amount of the target service currently processed by the machine learning model is higher than or equal to the service request amount threshold, the first model deployment manner is determined as the target model deployment manner, and a master container and at least one slave container are created on each device for deploying the machine learning model. The service request amount can be an average service request amount of the target service within a target time period (such as 24 hours, 3 days, 7 days, etc.), or a maximum service request amount of the target service within the target time period.
[0077] In some embodiments, for the first model deployment manner, the machine learning model can be deployed into at least one container on each device by performing the following steps: for each device, respectively performing the following process: determining the model unit of the machine learning model deployed on the device; and respectively deploying the model unit into each slave container.
[0078] Here, each slave container is respectively used to process a service inference request of a target service to obtain a sub-unit inference result; and the master container is used to determine a unit inference result of the model unit for the service inference request based on each sub-unit inference result.
[0079] The model unit is a part of the machine learning model, and multiple model units constitute a machine learning model. The model unit of the machine learning model is deployed in each slave container on the device, and each slave container deploys the same model unit. The master container on the device is used to receive a service inference request of a target service, and the master container sends the service inference request to each slave container. Each slave container processes the received service inference request to obtain a sub-unit inference result output by each slave container. The master container aggregates each sub-unit inference result to output a unit inference result of the model unit for the service inference request.
[0080] For example, referring to Figure 7 , Figure 7 is a deployment unit logical design flowchart of a machine learning model provided by an embodiment of the present application. Mode 1 in region 705 is a master-slave collaborative mode in a container group. The model unit of the machine learning model is deployed in multiple slave containers. A service inference request is sent to the master container. Multiple slave containers establish connections with the master container. The master container forwards the service inference request to each slave container. Each slave container processes the service inference request to obtain multiple sub-unit inference results. The master container aggregates the multiple sub-unit inference results to output a unit inference result.
[0081] In the embodiment of the present application, when the service request quantity of the target service is higher than or equal to the service request quantity threshold, the first model deployment manner is determined as the target model deployment manner, so that the deployment method of the machine learning model meets the demand of the target service, and the processing efficiency of the service request is improved.
[0082] In some embodiments, when the service request quantity of the target service is lower than the service request quantity threshold, the second model deployment manner in the multiple candidate model deployment manners is determined as the target model deployment manner. For each device, a container management process and a cascaded multiple containers managed by the container management process are created on the device according to the second model deployment manner.
[0083] Here, the second model deployment manner indicates that the relationship between the containers on each device in the plurality of devices is an idempotent connection relationship, i.e., there is no role-based priority relationship between the containers. When the service request quantity of the target business currently processed by the machine learning model is lower than the service request quantity threshold, the second model deployment manner is determined as the target model deployment manner, and one container management process and a plurality of cascaded containers are created on each device for deploying the machine learning model.
[0084] In some embodiments, for the second model deployment manner, the machine learning model can be deployed into at least one container on each device by performing the following steps: for each device, respectively performing the following processing: determining a model unit in the machine learning model that is deployed on the device; deploying the model unit into a plurality of cascaded containers, the first container in the plurality of containers being a first container, and the containers in the plurality of containers other than the first container being second containers.
[0085] Here, the first container is used to process the business inference request of the target business to obtain the output of the first container, and the second container is used to process based on the output of the previous container of the second container to obtain the output of the second container. The container management process is used to pass the business inference request to the first container, and output the output of the last second container as the unit inference result of the model unit for the business inference request.
[0086] The model unit is a part of the machine learning model, and a plurality of model units constitute a machine learning model. The model unit of the machine learning model is deployed in a plurality of cascaded containers on a device, each container deploying a different sub-model unit in the same model unit, and each sub-model unit is a processing step in the unit processing steps of the model unit. The container management process on the device is used to receive a business inference request of a target business, the container management process sends the business inference request to the first container in the plurality of cascaded containers, the first container processes the received business inference request to obtain the inference result of the first container output, the second container processes the inference result of the first container output to obtain the inference result of the second container output, the inference result of the second container output is taken as the unit inference result of the model unit for the business inference request, and the container management process outputs the unit inference result.
[0087] In some embodiments, for each device, a container management process and a container managed by the container management process are created on the device in the second model deployment manner. When deploying the machine learning model, a model unit of the machine learning model deployed on the device is determined, the model unit of the machine learning model is deployed in the container, the container management process on the device is configured to receive a service inference request of a target service, the container management process sends the service inference request to the container, the container processes the received service inference request, the inference result output by the container is taken as a unit inference result of the model unit for the service inference request, and the container management process outputs the unit inference result.
[0088] For example, referring to Figure 7 , Figure 7 is a logical design flow diagram of a deployment unit of a machine learning model provided by an embodiment of the present application. The manner 2 in the area 705 is a container management process-worker (container) cooperative manner in a container group, the model unit of the machine learning model is deployed in a plurality of cascaded worker containers, there is no role-prioritized relationship between each worker container, a service inference request is sent to the container management process, the container management process forwards the service inference request to the first worker container, the first worker container processes the service inference request, sends the output inference result to the second worker container, the second worker container processes the inference result output by the first worker container, sends the output inference result to the third worker container, the third worker container processes the inference result output by the second worker container, obtains the inference result output by the third worker container, takes the inference result output by the third worker container as a unit inference result, and the container management process outputs the unit inference result, so that the service inference request is processed by the plurality of worker containers in succession.
[0089] In the embodiment of the present application, when the service request amount of the target service is lower than the service request amount threshold, the second model deployment manner is determined as the target model deployment manner, so that the deployment method of the machine learning model meets the demand of the target service, improves the computing power inference service of the machine learning model, and thus improves the accuracy of processing the service request.
[0090] In some embodiments, referring to Figure 3B , Figure 3B is a second flow diagram of a deployment method of a machine learning model provided by an embodiment of the present application. Figure 3A The step 304 shown can be implemented by Figure 3B The steps 3041 to 3043, which are specifically described as follows.
[0091] In step 3041, a template description file of the target model deployment mode is generated.
[0092] Here, the template description file is used to describe the relationship between containers on each device indicated by the target model deployment mode, and the template description file is a file that describes the target model deployment mode in a string format.
[0093] For example, when the target model deployment mode is a first model deployment mode, i.e., the relationship between containers on each device is a master-slave connection relationship, the template description file of the first model deployment mode is described by the string ms. When the target model deployment mode is a second model deployment mode, i.e., the relationship between containers on each device is an idempotent connection relationship, i.e., there is no role-based priority relationship between containers (container management-worker), the template description file of the second model deployment mode is described by the string listen.
[0094] In step 3042, for each device, at least one container is created on the device, and the relationship between containers on the device is configured based on the template description file, to obtain a target device that is configured.
[0095] At least one container is created on each device for deploying a machine learning model, and a combined connection relationship between containers on the device is configured according to the content described in the template description file in a string format, to obtain a target device with a specified container combination relationship.
[0096] For example, referring to Figure 5 , Figure 5 is a schematic diagram of a hybrid module construction process of machine learning model inference provided by an embodiment of the present application. According to the template description file of the first model deployment mode, the containers included in the device (container group) of the cloud-native workload 507 of the large model can be configured in an M-W collaborative logic relationship, i.e., a master-worker connection mode. The region 508 is a target device (container group) after the logic relationship configuration, and the master container and the worker container included in the target device belong to a master-slave logic relationship.
[0097] In step 3043, the machine learning model is deployed into at least one container on each target device.
[0098] The machine learning model is deployed in at least one container on the target device with a specified container combination relationship, so that the machine learning model is containerized and deployed.
[0099] Based on steps 3041 to 3043, a template description file for the deployment method of the target model is generated, and the relationship between containers on the device is configured based on the template description file to obtain the configured target device. The machine learning model is then deployed to at least one container on each target device, realizing the containerized deployment of the machine learning model on multiple devices. This improves the deployment efficiency of the machine learning model, and the business side does not need to handle the multi-dimensional combination connection relationship between containers, supporting seamless use by the business side.
[0100] In some embodiments, multiple devices used to deploy machine learning models constitute a first deployment unit, see [link to relevant documentation]. Figure 3C , Figure 3C This is a schematic diagram of the third process of the data processing method provided in the embodiments of this application. Figure 3A After step 304 shown, the following can also be executed: Figure 3C Step 305 will be explained in detail below.
[0101] In step 305, in response to the scaling instruction, the machine learning model is deployed to at least one second deployment unit based on the target model deployment method.
[0102] Here, the first deployment unit is a group module that integrates multiple devices. At least one container is created on each device, and these at least one container forms a container group. Therefore, a container group also refers to a device. The scaling instruction is an instruction to increase the number of the first deployment units. The target model deployment method is the relationship between containers on each device as described in the template description file. Based on the target model deployment method, the number of the first deployment units is increased to obtain a second deployment unit with the same combination relationship between containers on each device in the first deployment unit. At the same time, the machine learning model is deployed to at least one second deployment unit.
[0103] In some embodiments, the business inference request of the target business is processed through a target deployment unit among multiple deployment units, which includes a first deployment unit and at least one second deployment unit. After deploying the machine learning model to at least one second deployment unit based on the target model deployment method, one of the following steps can be performed: in response to a scaling down instruction, a third deployment unit that is different from the target deployment unit is deleted from the multiple deployment units; in response to a failure of the target deployment unit, the target deployment unit is replaced with a fourth deployment unit that is different from the target deployment unit from the multiple deployment units.
[0104] Here, the target deployment unit is a deployment unit currently processing the business inference request, i.e., a group module currently processing the business inference request, the plurality of deployment units are a plurality of group modules, the business inference request of the target business is processed through a target group module in the plurality of group modules, and the plurality of group modules include a first group module and at least one second group module.
[0105] The scaling-down instruction is an instruction for reducing the number of deployment units, and deleting a third deployment unit different from the target deployment unit from the plurality of deployment units, i.e., deleting a third group module different from the target group module from the plurality of group modules. The third deployment unit can be a deployment unit specified by a user or a randomly determined deployment unit.
[0106] When the target deployment unit fails, the target deployment unit is replaced by a fourth deployment unit different from the target deployment unit in the plurality of deployment units, and the fourth deployment unit can be a deployment unit specified by a user or a deployment unit with a load amount lower than a load amount threshold determined according to the load situation of the deployment unit, i.e., the target group module is replaced by a fourth group module different from the target group module in the plurality of group modules.
[0107] For example, referring to Figure 5 , Figure 5 is a schematic diagram of a hybrid module construction process for machine learning model inference provided by an embodiment of the present application. For a large model workload 506 processing a business inference request, a large model inference cloud-native workload 507 is first translated, and the deployment of the large model is completed through the form of a deployment unit (group module), i.e., a plurality of devices (container groups) need to be integrated into the same deployment unit, and the devices (container groups) in the deployment unit can include one container or a plurality of containers. The deployment unit (group module) in the region 508 is configured after the container relationship, and the master container and the worker container included in the device (container group) in the deployment unit belong to a master-slave connection relationship, and the number of deployment units can be increased, reduced, and replaced in operation through a unit deployment unit.
[0108] In an embodiment of the present application, a first deployment unit is composed of a plurality of devices for deploying a machine learning model, and in response to a scaling-up instruction, the machine learning model is deployed to at least one second deployment unit based on a target model deployment manner, so as to implement the deployment of the machine learning model in a plurality of deployment units and improve the inference computing power service of the machine learning model.
[0109] In some embodiments, when the device resource amount is lower than the device resource amount threshold, a target device for deploying the machine learning model is determined; a plurality of containers are created on the target device, and the machine learning model is deployed into each container on the target device, and the relationship between each container is independent of each other.
[0110] When the device resource amount required for running the machine learning model is lower than the device resource amount threshold, the machine learning model is marked as a small model, and the machine learning model is deployed on a device by obtaining the device identifier for deploying the machine learning model. On the device for deploying the machine learning model, a plurality of independent containers are created, so that the machine learning model is deployed in each container on the device.
[0111] In some embodiments, after the machine learning model is deployed into each container on the target device, the following steps can be performed: in response to a scaling instruction, at least one new container is created on the target device, and the machine learning model is deployed into each new container.
[0112] The machine learning model marked as a small model is deployed into each container on a device, the scaling instruction is an instruction to increase the number of containers on the device, in response to the scaling instruction, the number of containers on the device is increased to obtain a plurality of new containers, and the machine learning model is deployed into the plurality of new containers.
[0113] In some embodiments, the business inference request of the target business is processed by the target container on the target device; after the machine learning model is deployed into each container on the target device, the following steps can be performed: one of the following processing is performed: in response to a scaling instruction, a third container different from the target container on the target device is deleted; in response to a failure of the target container, the target container is replaced by a fourth container different from the target container on the target device.
[0114] Here, the target container is the container currently processing the business inference request, and the business inference request of the target business is processed by the target container on the device. If the target container has an exception and cannot process the business inference request, other containers different from the target container can continue to process the business inference request.
[0115] For example, referring to Figure 6 , Figure 6 is a logical flow diagram of a module processing an inference service request provided by the small model of the embodiments of the present application. The module 601 includes a plurality of independent containers, each of which can independently receive and process an inference service request and output an inference service result obtained by processing the inference service request. Among them, the container 602 and the container 603 are abnormal containers, and the remaining containers can seamlessly process the inference service request and output the inference service result.
[0116] The shrinkage instruction is an instruction for reducing the number of containers on a device. In response to the shrinkage instruction, other containers different from the target container are deleted on the device. When the target container fails, the target container is replaced by other containers different from the target container on the device.
[0117] For example, with reference to Figure 5 , Figure 5 is a schematic diagram of a hybrid module construction process for machine learning model inference provided by an embodiment of the present application. For small model workloads 502 processing business inference requests, they are first translated into cloud-native workloads 503 for small model inference, and the deployment of small models is completed through a single-container mode. The containers in the cloud-native workloads 503 for small models can be configured with idempotent logical relationships. The target device after container relationship configuration is in region 504. The containers in the target device have idempotent relationships and no logical coordination relationship between them. The increase, decrease, and failure replacement operations of the number of containers can be performed in the target device.
[0118] In an embodiment of the present application, when the device resource amount is less than the device resource amount threshold, a target device for deploying a machine learning model is determined. In response to a scaling instruction, at least one new container is created on the target device, and the machine learning model is deployed into each new container. By performing the increase operation of the number of containers on a device, the inference computing power service of the machine learning model is improved.
[0119] In some embodiments, the machine learning model deployment method provided by the embodiments of the present application can be applied in the field of cloud technology. A cloud server obtains a device resource amount threshold and a device resource amount required for deploying a machine learning model of a target business. When the device resource amount exceeds the device resource amount threshold, a plurality of devices for deploying the machine learning model are determined. From a plurality of candidate model deployment modes, a target model deployment mode matching the target business is determined. The target model deployment mode indicates the relationship between the containers on each device in the plurality of devices. At least one container is created on each device according to the target model deployment mode, and the machine learning model is deployed into at least one container on each device. This enables the deployment of the machine learning model on a cloud computing platform. In this way, when deploying the machine learning model on the cloud platform, the business side does not need to handle the combination and connection relationship between the containers, improving the deployment online efficiency of the machine learning model and supporting the customized and non-intrusive use of the business side.
[0120] By using the above embodiments of the application, the device resource threshold and the device resource required by the machine learning model to be deployed are obtained, and then when the device resource exceeds the device resource threshold, the plurality of devices for deploying the machine learning model are determined, and then the target model deployment mode matched with the target service is determined from the plurality of candidate model deployment modes, the target model deployment mode indicates the relationship between the containers on each device in the plurality of devices, at least one container is created on each device according to the target model deployment mode, and the machine learning model is deployed into the at least one container on each device. In this way, the target model deployment mode matched with the target service can be selected from the plurality of pre-set candidate model deployment modes to implement the deployment of the machine learning model, and the target model deployment mode indicates the relationship between the containers on each device in the plurality of devices, so that the service side does not need to process the relationship between the plurality of containers for deploying the machine learning model by itself, and the deployment efficiency of the machine learning model is improved.
[0121] In the following, an example application of the deployment method of the machine learning model provided by the embodiments of the application in a heterogeneous model deployment scenario corresponding to an actual inference service will be described.
[0122] With the increasing development of AI large models, large models with hundreds of millions of parameters are gradually trained, and how to deploy the large models for online inference to support scenario-based applications has become a prominent problem. Due to the limitation of the video memory space resource of a single device for deploying a heterogeneous model, the capacity cannot complete the packing of a heterogeneous large model, i.e., a plurality of devices need to cooperate to complete the deployment of the heterogeneous large model.
[0123] In the related art, a heterogeneous model is containerized and deployed on a device to fully utilize the video memory resource of the device to process an inference service request. This use method is suitable for the deployment of a heterogeneous small model, supports the horizontal expansion of the heterogeneous small model inference service, and supports the automatic disaster recovery logic after the failure of the container on the device. In the deployment process of a heterogeneous large model, the following problems exist: the heterogeneous large model needs to cooperate with a plurality of devices to normally run the inference service, and the service side usually needs to process the combination logic relationship between the plurality of containers for deploying the heterogeneous large model by itself, the horizontal expansion and disaster recovery capability of the container are lacking, and the use of the service side is not friendly.
[0124] The embodiments of the application propose a deployment method of a machine learning model to solve the problems in the related art, and the following improvements are included compared with the related art:
[0125] (1) On the basis of the cloud native workload, a hybrid module for heterogeneous model inference is constructed to perform the containerization deployment of a heterogeneous small model and a heterogeneous large model in the module, and the deployment online efficiency of the heterogeneous model is improved.
[0126] (2) According to the device memory resources and the size of the heterogeneous model, the heterogeneous model is marked as a heterogeneous large model or a heterogeneous small model. For the scenario of deploying a heterogeneous large model in multiple devices, a group form module is constructed, the containerization deployment of the heterogeneous large model is performed in a unit group module (deployment unit), and the logical relationship of multi-dimensional combination of containers is supported in the group module. The horizontal scaling and fault replacement of containers are performed in the unit group module granularity.
[0127] The mixed module for heterogeneous model reasoning provided by the embodiments of the present application can cover the single-machine deployment scenario of a heterogeneous small model and the multi-machine deployment scenario of a heterogeneous large model. For the scenario of deploying a heterogeneous small model in a single device, a single-container mode module is constructed, the containerization deployment of the heterogeneous small model is performed, and the horizontal scaling and vertical fault replacement of containers in the module are supported. For the scenario of deploying a heterogeneous large model in multiple devices, a group form module is designed, the containerization deployment of the heterogeneous large model is performed in a unit group module, and a specified container combination logical relationship is supported in the group module, such as master-slave association, idempotent association, etc. The horizontal scaling and fault replacement of containers are performed in the unit group module granularity.
[0128] For example, referring to Figure 4 , Figure 4 is a structural schematic diagram of an application module supporting a heterogeneous model reasoning service provided by the embodiments of the present application. In the computing power container technology support layer, different application modules are combined to provide computing power resources for the reasoning service of the heterogeneous model. The application modules include an application module 401, an application module 402, and an application module 403. The product layer includes a 7B model reasoning service 404, a 13B model reasoning service 405, and a large model reasoning service 406. The 7B model is a model with a parameter quantity of about 7 billion large neural network models, and the 13B model is a model with a parameter quantity of about 13 billion large neural network models. The application module 401 includes multiple computing power container groups. The computing power containers included in the computing power container groups can be in an idempotent relationship. The number of computing power containers can be expanded. The application module 401 is suitable for supporting the reasoning service of a small model, such as the 7B model reasoning service 404 in Figure 4 The application module 402 includes multiple group modules. Each group module includes multiple computing power container groups. The computing power containers included in the computing power container groups can be in an idempotent relationship or a master-slave relationship. The application module 402 is suitable for supporting the reasoning service of a large model, such as Figure 4The group module in the application module 403 includes two types of computing power container groups: computing power container group-1 and computing power container group-2, which are suitable for supporting the 13B model inference service 405 and the large model inference service 406.
[0129] For example, referring to Figure 5 , Figure 5 is a schematic diagram of a hybrid module construction process of machine learning model inference provided by an embodiment of the present application. When deploying the inference service 501, the user judges the size of the heterogeneous model corresponding to the inference service. If the heterogeneous model can complete the deployment operation of the model through a single device, the heterogeneous model is marked as a small model; if the heterogeneous model cannot complete the deployment operation of the model through a single device, the heterogeneous model is marked as a large model.
[0130] For the small model workload 502 processing the inference service request, it is first translated into a cloud-native workload 503 of small model inference, and the deployment of the small model is completed through a single container. Since the computing power service of the small model inference needs to meet the horizontal scaling of the workload module replica number and the automatic replacement of the failed container in the module, and these requirements can be supported by the self-closing loop mode of the cloud-native workload, it is necessary to convert the small model workload into a cloud-native workload of the small model. The containers in the cloud-native workload 503 of the small model can be configured with idempotent logical relationships. The workload module after logical relationship configuration is in region 504, and the containers in the module have an idempotent relationship and no logical coordination relationship. The single-container expansion, contraction and fault replacement operations can be performed in the module. The logical flow of the containers in the module processing the inference service request is in region 505, and the single container can receive and process the request.
[0131] For example, referring to Figure 6 , Figure 6 is a schematic diagram of a logical flow of a module processing an inference service request of a small model provided by an embodiment of the present application. The module 601 includes a plurality of independent containers, each of which can independently receive an inference service request and output an inference service result obtained by processing the inference service request. Among them, the container 602 and the container 603 are abnormal containers, and the remaining containers can seamlessly process the inference service request and output the inference service result to the business side.
[0132] For example, referring to Figure 5For the large model workload 506 processing the inference service request, it is first translated into the cloud-native workload 507 of large model inference, and the deployment of the large model is completed through the mode of the group module, that is, multiple container groups need to be integrated into the same group module, and the container groups in the group module can include one container or multiple containers, and each container group can be a device. The containers included in the container groups in the cloud-native workload 507 of the large model can be configured in an M-W collaborative logical relationship, that is, a Master-Worker master-slave logical association. The group module in region 508 is after logical relationship configuration, and the master container and the worker container included in the container groups in the group module belong to the master-slave logical relationship, and the expansion, contraction and fault replacement operations of the containers can be performed through the group module granularity of the unit. The logical flow of the containers included in the container groups in the group module in region 509 processing the inference service request, and the M container (master container) can receive the inference service request and forward the inference service request to the W container (worker container).
[0133] For the large model workload 510 processing the inference service request, it is first translated into the cloud-native workload 511 of large model inference, and the deployment of the large model is completed through the mode of the group module. The containers included in the container groups in the cloud-native workload 511 of the large model can be configured in an idempotent collaborative logical relationship, that is, a worker-worker idempotent logical association. The group module in region 512 is after logical relationship configuration, and the two worker containers included in the container groups in the group module belong to the idempotent logical relationship, and the expansion, contraction and fault replacement operations of the containers can be performed through the group module granularity of the unit. The logical flow of the containers included in the container groups in the group module in region 513 processing the inference service request, and the first W container (worker container) can receive and process the inference service request, and forward the inference service result to the second W container (worker container) to enable the second W container to process the inference service request in turn.
[0134] For example, referring to Figure 7 , Figure 7 is a logical design flow diagram of a deployment unit of a machine learning model provided by an embodiment of the present application. In combination with the steps of Figure 7 , the logical design flow of the group module provided by the embodiment of the present application is explained.
[0135] In step 701, the inference deployment is accessed.
[0136] The heterogeneous model corresponding to the inference service is accessed, and when the heterogeneous model is a large model, the inference workload module of the large model is logically designed.
[0137] In step 702, the group logic strategy is configured.
[0138] According to the evaluation of the business, the configuration strategy of the group is selected, that is, the cooperative logic relationship between the containers included in the container group in the group module.
[0139] In step 703, the cloud native workload is generated.
[0140] After determining the configuration strategy of the group module as described above, the generated inference cloud native workload is generated to realize the generation of the target workload 704.
[0141] In the inference computing power service of the large model, the unit computing power service granularity is the granularity of the group module. According to the strategy configured when the business is accessed, the cooperative logic relationship between the containers included in the container group in each group module is executed. The cooperative mode in area 705 is mode 1, mode 2 and other modes, which corresponds to the group module configuration strategy. Mode 1 is a master-slave cooperative mode. The inference service request is sent to the master container. Other slave containers respectively establish connections with the master container. The master container forwards the inference service request to each slave container. Each slave container executes the response of the inference service request to obtain multiple inference results. The master container completes the aggregation of the inference results and outputs the content of the request response. Mode 2 is a container management-worker cooperative mode. The containers have no priority. The container management process completes the reception of the inference service request. The container management process forwards the inference service request to the first worker container. The first worker container processes the inference service request and forwards the output inference result to the second worker container. The second worker container processes the inference result output by the first worker container and sends the output inference result to the third worker container. The third worker container processes the inference result output by the second worker container to obtain the inference result output by the third worker container. In this way, the inference service request is executed by multiple worker containers in a progressive manner. Finally, the container management process outputs the content of the request response. Other cooperative modes can be other strategies configured by the group module, such as: multiple worker containers are combined into a computing group of the inference service request. The inference service request is executed by the computing group through multiple rounds of interactive calculation to output the inference result.
[0142] For example, referring to Figure 8 , Figure 8 is a flowchart of the modular deployment model provided by the embodiments of the present application. In combination with Figure 8The model deployment design process provided by the embodiments of the present application is explained.
[0143] In step 801, the inference business performs containerization access.
[0144] Access the inference business to obtain a heterogeneous model corresponding to the inference business.
[0145] In step 802, the model size is judged.
[0146] The operation and maintenance personnel can evaluate and judge the threshold configuration value of the model size according to the video memory resource size of the device used to deploy the heterogeneous model, such as: a single device has 4 graphics cards, and each graphics card has a video memory resource of 24G, and the video memory resource of a single device is 96G. The threshold configuration value is set to 96G. The heterogeneous model whose required device resource is less than 96G is executed as a small model containerization, and the heterogeneous model whose required device resource is greater than 96G is executed as a large model containerization, that is, the large model needs to be containerized by multiple devices.
[0147] In step 803, it is judged whether the model size exceeds the configuration threshold.
[0148] When the model size does not exceed the configuration threshold, go to step 804 to execute small model containerization. When the model size exceeds the configuration threshold, go to step 808 to configure the group module strategy.
[0149] In step 805, a small model cloud-native workload is created.
[0150] A small model inference workload is generated on k8s (a container orchestration engine), such as a deployment workload.
[0151] In step 806, horizontal replica scaling is performed.
[0152] At least one container is created in the small model inference workload module, and the replicas of each container are scaled to obtain a plurality of newly added containers. Go to step 807 to complete the containerization operation and realize the deployment of the small model in each container on the device.
[0153] In step 808, the group module strategy is configured.
[0154] According to the evaluation of the business, the user configures the combination connection mode between the containers in the group module according to the actual demand, and the group module strategy includes master-slave connection mode, container management-worker mode.
[0155] In step 809, a template description file is generated according to the strategy.
[0156] The configured group module strategy can generate a template description file. According to the content of the configured group module strategy, a file described in a string format is generated, for example: when the configured group module strategy is a master-slave connection mode, a file described by the string ms is generated; when the configured group module strategy is a container management-worker mode, a file described by the string listen is generated.
[0157] In step 810, a cloud-native workload of a large model is created.
[0158] A large model inference workload is generated on k8s (a container orchestration engine), for example: a deployment workload.
[0159] In step 811, a unit granularity group module is generated based on the template description file.
[0160] The template description file includes a logical combination mode of multiple containers in the group module. According to the content of the template description file, a unit group module with a specified logical relationship is generated.
[0161] In step 812, horizontal replica expansion of the module is performed.
[0162] The workload of the large model performs replica expansion of the group module based on the template description file, and obtains multiple newly added group modules. Go to step 813 to complete the packing operation, and realize the modular deployment of the large model on multiple devices.
[0163] In the above application scenario of heterogeneous model deployment, by constructing a hybrid module for heterogeneous model inference, the single-machine deployment scenario of a heterogeneous small model and the multi-machine deployment scenario of a heterogeneous large model are covered, and the deployment online efficiency of the heterogeneous model is improved. For the scenario of deploying a heterogeneous large model on multiple devices, by designing a group form module, containerized deployment of the heterogeneous large model is performed in a unit group module, and a specified container combination logical relationship is supported in the group module. The horizontal scaling and fault replacement of the container are performed at the unit group module granularity, and the inference computing power service of the heterogeneous large model is improved. The container combination logic is loosely coupled in architecture, and the business side does not need to handle the multi-dimensional combination logical relationship between containers, thereby reducing the computing cost.
[0164] The following continues to illustrate an exemplary structure of the implementation of the machine learning model deployment apparatus 455 provided in the embodiments of the present application as a software module. In some embodiments, as shown in FIG. 6, the machine learning model deployment apparatus 455 includes a model description module 601, a model training module 602, a model storage module 603, a model inference module 604, a model management module 605, and a model service module 606. Figure 2As shown, the software modules stored in the deployment device 455 of the machine learning model in the memory 450 can include: a model acquisition module 4551 configured to acquire a device resource amount threshold and a device resource amount required by the machine learning model of the deployment target service; a device determination module 4552 configured to determine a plurality of devices for deploying the machine learning model when the device resource amount exceeds the device resource amount threshold; a manner determination module 4553 configured to determine, from a plurality of candidate model deployment manners, a target model deployment manner matched with the target service, the target model deployment manner indicating a relationship between containers on each device in the plurality of devices; and a model deployment module 4554 configured to create at least one container on each device according to the target model deployment manner, and deploy the machine learning model into the at least one container on each device.
[0165] In some embodiments, the manner determination module 4553 is further configured to, when the service request amount of the target service is higher than or equal to the service request amount threshold, determine a first model deployment manner in the plurality of candidate model deployment manners as the target model deployment manner.
[0166] In some embodiments, the model deployment module 4554 is further configured to, for each device, create one master container on the device according to the first model deployment manner, and create at least one slave container connected to the master container respectively.
[0167] In some embodiments, the model deployment module 4554 is further configured to, for each device, respectively perform the following processing: determine a model unit in the machine learning model deployed on the device; and deploy the model unit into each slave container respectively; wherein each slave container is configured to process a service inference request of the target service to obtain a sub-unit inference result; and the master container is configured to determine a unit inference result of the model unit for the service inference request based on each sub-unit inference result.
[0168] In some embodiments, the manner determination module 4553 is further configured to, when the service request amount of the target service is lower than the service request amount threshold, determine a second model deployment manner in the plurality of candidate model deployment manners as the target model deployment manner.
[0169] In some embodiments, the model deployment module 4554 is further configured to, for each device, create one container management process and a plurality of cascaded containers managed by the container management process on the device according to the second model deployment manner.
[0170] In some embodiments, the model deployment module 4554 is further configured to determine, for each device, a model unit in the machine learning model to be deployed on the device; deploy the model unit in a plurality of containers in a cascade, a first container in the plurality of containers being the first container, and a container other than the first container in the plurality of containers being the second container; wherein the first container is configured to process a service inference request of a target service to obtain an output of the first container, and the second container is configured to process based on an output of a previous container of the second container to obtain an output of the second container; and the container management process is configured to pass the service inference request to the first container, and output the output of the last second container as a unit inference result of the model unit for the service inference request.
[0171] In some embodiments, the model deployment module 4554 is further configured to generate a template description file of the target model deployment mode, the template description file being configured to describe relationships between containers on each device indicated by the target model deployment mode; for each device, create at least one container on the device, and configure relationships between containers on the device based on the template description file to obtain a configured target device; and deploy the machine learning model to at least one container on each target device.
[0172] In some embodiments, a plurality of devices for deploying the machine learning model form a first deployment unit; and the model deployment module 4554 is further configured to, after creating at least one container on each device according to the target model deployment mode and deploying the machine learning model to at least one container on each device, in response to an expansion instruction, deploy the machine learning model to at least one second deployment unit based on the target model deployment mode.
[0173] In some embodiments, a service inference request of a target service is processed through a target deployment unit in a plurality of deployment units, the plurality of deployment units including the first deployment unit and at least one second deployment unit; and the model deployment module 4554 is further configured to, after deploying the machine learning model to at least one second deployment unit based on the target model deployment mode, perform one of the following: in response to a contraction instruction, delete a third deployment unit in the plurality of deployment units different from the target deployment unit; and in response to a failure of the target deployment unit, replace the target deployment unit with a fourth deployment unit in the plurality of deployment units different from the target deployment unit.
[0174] In some embodiments, the model deployment module 4554 is further configured to determine a target device for deploying the machine learning model when an amount of device resources is less than a device resource threshold; create a plurality of containers on the target device, and deploy the machine learning model to each container on the target device, the relationships between the containers being independent of each other.
[0175] In some embodiments, the model deployment module 4554 is further configured to, after deploying the machine learning model into the containers on the target device, in response to a scaling-out instruction, create at least one new container on the target device and deploy the machine learning model into the new containers.
[0176] In some embodiments, the service inference request of the target service is processed by a target container on the target device; and the model deployment module 4554 is further configured to, after deploying the machine learning model into the containers on the target device, perform one of the following: in response to a scaling-in instruction, delete a third container on the target device that is different from the target container; and in response to a failure of the target container, replace the target container with a fourth container on the target device that is different from the target container.
[0177] Embodiments of the present application provide a computer program product including computer executable instructions or a computer program stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions or the computer program from the computer readable storage medium, and the processor executes the computer executable instructions or the computer program, so that the electronic device performs the machine learning model deployment method provided in the embodiments of the present application.
[0178] Embodiments of the present application provide a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions or the computer program stored therein, when executed by a processor, cause the processor to perform the machine learning model deployment method provided in the embodiments of the present application, for example, as shown in the machine learning model deployment method. Figure 3A
[0179] In some embodiments, the computer readable storage medium can be a FRAM, a ROM, a PROM, an EPROM, an EEPROM, a flash memory, a magnetic surface memory, an optical disc, or a CD-ROM, etc. memory; or can be various devices including one or any combination of the above memories.
[0180] In some embodiments, the computer executable instructions can be in the form of a program, software, software module, script or code, written in any form of programming language (including a compiled or interpreted language, or a declarative or procedural language), and can be deployed in any form, including being deployed as a standalone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0181] By way of example, computer readable media can include computer- storage media and communication media. Computer-storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer-storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.
[0182] By way of example, computer readable media can include computer- storage media and communication media. Computer-storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer-storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.
[0183] The foregoing is merely illustrative of the principles of this application and various modifications can be made by those skilled in the art. Other implementations are within the scope of the following claims.
Claims
1. A method for deploying a machine learning model, characterized in that, The method includes: Obtain the threshold for device resources and the amount of device resources required to deploy the machine learning model for the target business; When the amount of device resources exceeds the device resource threshold, multiple devices are identified for deploying the machine learning model. From multiple candidate model deployment methods, a target model deployment method that matches the target business is determined, wherein the target model deployment method indicates the relationship between containers on each of the multiple devices; According to the target model deployment method, at least one container is created on each of the devices, and the machine learning model is deployed to at least one container on each of the devices.
2. The method according to claim 1, characterized in that, The step of determining the target model deployment method that matches the target business from multiple candidate model deployment methods includes: When the number of service requests for the target service is higher than or equal to the service request volume threshold, the first model deployment method among the multiple candidate model deployment methods is determined as the target model deployment method. The step of creating at least one container on each of the devices according to the target model deployment method includes: for each of the devices, according to the first model deployment method, creating a main container on the device, and creating at least one slave container that is connected to the main container respectively.
3. The method according to claim 2, characterized in that, Deploying the machine learning model to at least one container on each of the said devices includes: The following processes are performed on each of the aforementioned devices: Identify the model units deployed on the device in the machine learning model; The model units are deployed in each of the slave containers; Each of the sub-containers is used to process the business reasoning request of the target business to obtain the sub-unit reasoning result; the main container is used to determine the unit reasoning result of the model unit for the business reasoning request based on the reasoning results of each sub-unit.
4. The method according to claim 1, characterized in that, The step of determining the target model deployment method that matches the target business from multiple candidate model deployment methods includes: When the number of service requests for the target service is lower than the service request volume threshold, the second model deployment method among the multiple candidate model deployment methods is determined as the target model deployment method; The step of creating at least one container on each of the devices according to the target model deployment method includes: for each of the devices, according to the second model deployment method, creating a container management process and multiple cascaded containers managed by the container management process on the device.
5. The method according to claim 4, characterized in that, Deploying the machine learning model to at least one container on each of the said devices includes: The following processes are performed on each of the aforementioned devices: Identify the model units deployed on the device in the machine learning model; The model unit is deployed in the cascaded multiple containers, the first container of which is the first container, and the containers other than the first container of which are the second containers; Wherein, the first container is used to process the business reasoning request of the target business and obtain the output of the first container; the second container is used to process the output of the previous container of the second container and obtain the output of the second container. The container management process is used to pass the business inference request to the first container, and to output the output of the last second container as the unit inference result of the model unit for the business inference request.
6. The method according to claim 1, characterized in that, The step of creating at least one container on each of the devices according to the target model deployment method, and deploying the machine learning model to at least one container on each of the devices, includes: Generate a template description file for the target model deployment method, wherein the template description file is used to describe the relationship between containers on each of the devices indicated by the target model deployment method; For each of the aforementioned devices, at least one container is created on the device, and based on the template description file, the relationship between the containers on the device is configured to obtain the configured target device; The machine learning model is deployed to at least one container on each of the target devices.
7. The method according to claim 1, characterized in that, The plurality of devices used to deploy the machine learning model constitute a first deployment unit; after creating at least one container on each of the devices according to the target model deployment method, and deploying the machine learning model to at least one container on each of the devices, the method further includes: In response to the scaling instruction, the machine learning model is deployed to at least one second deployment unit based on the target model deployment method.
8. The method according to claim 7, characterized in that, The business reasoning request of the target business is processed through a target deployment unit among multiple deployment units, the multiple deployment units including the first deployment unit and the at least one second deployment unit; After deploying the machine learning model to at least one second deployment unit based on the target model deployment method, the method further includes: Perform one of the following processes: In response to the scaling-down command, a third deployment unit that is different from the target deployment unit is deleted from the plurality of deployment units; In response to a fault in the target deployment unit, the target deployment unit is replaced with a fourth deployment unit that is different from the target deployment unit among the plurality of deployment units.
9. The method according to claim 1, characterized in that, The method further includes: When the device resource quantity is lower than the device resource quantity threshold, a target device is determined for deploying the machine learning model; Multiple containers are created on the target device, and the machine learning model is deployed to each of the containers on the target device, with the relationships between the containers being independent of each other.
10. The method according to claim 9, characterized in that, After deploying the machine learning model to each of the containers on the target device, the method further includes: In response to the scaling instruction, at least one new container is created on the target device, and the machine learning model is deployed to each of the new containers.
11. The method according to claim 10, characterized in that, The business reasoning request of the target service is processed through the target container on the target device; After deploying the machine learning model to each of the containers on the target device, the method further includes: Perform one of the following processes: In response to a reduction command, a third container on the target device that is different from the target container is deleted; In response to a fault in the target container, the target container is replaced with a fourth container on the target device that is different from the target container.
12. A device for deploying a machine learning model, characterized in that, The device includes: The model acquisition module is used to obtain the device resource threshold and the amount of device resources required to deploy the machine learning model for the target business. The device determination module is used to determine multiple devices for deploying the machine learning model when the device resource quantity exceeds the device resource quantity threshold. The method determination module is used to determine the target model deployment method that matches the target business from multiple candidate model deployment methods, wherein the target model deployment method indicates the relationship between containers on each of the multiple devices; The model deployment module is used to create at least one container on each of the devices according to the target model deployment method, and deploy the machine learning model to at least one container on each of the devices.
13. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the deployment method of the machine learning model according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method for deploying the machine learning model according to any one of claims 1 to 11.
15. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method for deploying the machine learning model according to any one of claims 1 to 11.