A model deployment method, device, equipment and storage medium
By obtaining and deploying model configuration information and metadata in Kubernetes clusters and generating yaml configuration files, the problems of missing metadata, lack of computing resource monitoring and flexible expansion in model management are solved, and the model deployment is simplified and efficient monitoring is achieved.
Patent Information
- Application Number
- CN202210448435.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-26
AI Technical Summary
In the prior art, model management and deployment methods have problems such as missing metadata information, lack of model evaluation results, lack of computing resource monitoring and lack of elastic expansion, resulting in high difficulty in deploying models, low efficiency and ineffective monitoring of computing resources.
By obtaining model configuration information and metadata, uploading it to the test Kubernetes cluster and generating yaml configuration files, monitoring the data occupied by computing resources, supporting dynamic and elastic expansion, and deploying models in a production environment to realize real-time monitoring and evaluation of model computing resources.
It simplifies the model deployment process, improves the efficiency of developers, supports dynamic and elastic expansion, can increase the total model computing resources at any time, and better monitor the model's computing resources.
Smart Images

Figure CN114721674B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a model deployment method, apparatus, device, and storage medium. Background Art
[0002] With the continuous penetration of artificial intelligence technology based on big data, machine learning, and deep learning into all walks of life, the management and deployment of machine learning models built through big data are important links in the application of artificial intelligence in the industry.
[0003] Although the methods and systems for model management and deployment proposed in the market today can uniformly manage the metadata of models and the running status of model services, and to a certain extent solve the problems of model management and deployment, these systems still have some disadvantages.
[0004] 1. Manually deploy on a single server, and downtime is required for model service updates in the later stage, with high deployment difficulty.
[0005] 2. Unable to monitor model computing resources and lacking unified management of model post-evaluation results.
[0006] 3. Do not support elastic expansion of model computing resources and intelligent configuration of model computing resources. Summary of the Invention
[0007] Embodiments of the present invention provide a model deployment method, apparatus, device, and storage medium, which solve the problems of missing metadata information in model management, lack of model evaluation results, missing monitoring of computing resources in model deployment, and lack of elastic expansion. It can simplify the model deployment process, improve the efficiency of developers, support dynamic elastic expansion, can increase the total model computing resources at any time, and can better monitor the computing resources of the model.
[0008] In a first aspect, an embodiment of the present invention provides a model deployment method, including:
[0009] Obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information;
[0010] Upload the model file and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database;
[0011] Generate a first yaml configuration file according to the model configuration information;
[0012] Receive a test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file.
[0013] Further, receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file includes:
[0014] Receive a test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file. The first target pod obtains a model file according to a local directory and performs a test according to the model file.
[0015] Further, after receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, the first target pod obtains a model file according to a local directory, and performs a test according to the model file, it further includes:
[0016] Obtain the computing resource occupancy data sent by the test Kubernetes cluster;
[0017] Determine computing resource parameters according to the computing resource occupancy data and write the computing resource parameters into a background database.
[0018] Further, after determining computing resource parameters according to the computing resource occupancy data and writing the computing resource parameters into a background database, it further includes:
[0019] Generate a second yaml configuration file according to the computing resource parameters and the model configuration information;
[0020] Receive a release instruction and send the release instruction to the production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file;
[0021] Obtain the model running log and the model evaluation result by accessing the GPRC port, and display the model running log and the model evaluation result.
[0022] In a second aspect, an embodiment of the present invention further provides a model deployment device, and the device includes:
[0023] An acquisition module, configured to acquire model configuration information, metadata, and model files, wherein the model configuration information includes: environment image information;
[0024] An upload module, configured to upload the model files and the model configuration information to the local of a test Kubernetes cluster, and write the model configuration information into a background database;
[0025] A generation module, configured to generate a first yaml configuration file according to the model configuration information;
[0026] A receiving module, configured to receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys a model according to the first yaml configuration file.
[0027] Further, the receiving module is specifically configured to:
[0028] Receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, the first target pod obtains a model file according to a local directory, and performs a test according to the model file.
[0029] Further, the receiving module is further configured to:
[0030] After receiving a test instruction, and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, the first target pod obtains a model file according to a local directory, and performs a test according to the model file, acquire computing resource occupancy data sent by the test Kubernetes cluster;
[0031] Determine computing resource parameters according to the computing resource occupancy data, and write the computing resource parameters into the background database.
[0032] Further, the receiving module is further configured to:
[0033] After determining computing resource parameters according to the computing resource occupancy data, and writing the computing resource parameters into the background database, generate a second yaml configuration file according to the computing resource parameters and the model configuration information;
[0034] Receive a release instruction, and send the release instruction to a production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file.
[0035] Obtain the model running log and the model evaluation result by accessing the GRPC port, and display the model running log and the model evaluation result.
[0036] In a third aspect, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the model deployment method as described in any one of the embodiments of the present invention.
[0037] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the model deployment method as described in any one of the embodiments of the present invention.
[0038] In the embodiment of the present invention, by obtaining model configuration information, metadata, and model files, where the model configuration information includes: environment image information; uploading the model file and the model configuration information to the local of the test Kubernetes cluster, and writing the model configuration information into the background database; generating a first yaml configuration file according to the model configuration information; receiving a test instruction, and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file, the problems of missing metadata information in model management, lack of model evaluation results, missing monitoring of computing resources in model deployment, and lack of elastic scaling are solved. It can simplify the process of model deployment, improve the efficiency of developers, support dynamic elastic scaling, can increase the total computing resources of the model at any time, and can better monitor the computing resources of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 is a flowchart of a model deployment method in an embodiment of the present invention;
[0041] Figure 2 is a schematic structural diagram of a model deployment device in an embodiment of the present invention;
[0042] Figure 3 is a schematic structural diagram of an electronic device in an embodiment of the present invention;
[0043] Figure 4 It is a schematic structural diagram of a computer-readable storage medium including a computer program in an embodiment of the present invention. Specific embodiments
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings. Furthermore, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.
[0045] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc. Furthermore, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.
[0046] The term "including" and its variations used in the present invention are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment".
[0047] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0048] Figure 1 It is a flowchart of a model deployment method provided for an embodiment of the present invention. This embodiment is applicable to the situation of model deployment. This method can be executed by the model deployment device in the embodiment of the present invention. The model deployment device can be implemented in a software and / or hardware manner, such as Figure 1 As shown, the model deployment method specifically includes the following steps:
[0049] S110, obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information.
[0050] Among them, the way to obtain the environment image information can be to select the environment image information for model training in the image repository.
[0051] Specifically, the ways to obtain the model configuration information, metadata, and model file can be: fill in other model configuration information and metadata except the environment image information in the model factory, select the environment image information for model training in the image repository, and upload the local model file to the platform.
[0052] S120, upload the model file and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database.
[0053] Specifically, uploading the model file and the model configuration information to the local of the test Kubernetes cluster and writing the model configuration information into the background database can be, for example, filling in other model configuration information and metadata except the environment image information in the model factory, selecting the environment image information for model training in the image repository, uploading the local model file to the platform, and the platform packs and uploads the model file and configuration information uploaded by the model developer to the local directory of the test Kubernetes cluster, and writes the configuration information into the background database.
[0054] S130, generate a first yaml configuration file according to the model configuration information.
[0055] Specifically, the way to generate a first yaml configuration file according to the model configuration information can be: automatically generate the yaml configuration file of the pod used for model training in the test Kubernetes cluster, and obtain the model file by mounting the local directory through the pod.
[0056] S140, receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file.
[0057] Specifically, the way to receive a test instruction can be: the model developer applies to deploy the model in the test and trial operation environment.
[0058] Specifically, receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file can be, for example, the model developer applies to deploy the model in the test and trial operation environment, and the platform deploys the model in the test Kubernetes cluster after resource review.
[0059] Optionally, receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys a model according to the first yaml configuration file includes:
[0060] Receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, and the first target pod obtains a model file according to a local directory and performs a test according to the model file.
[0061] Optionally, after receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, and the first target pod obtains a model file according to a local directory and performs a test according to the model file, it further includes:
[0062] Obtaining computing resource occupancy data sent by the test Kubernetes cluster;
[0063] Determining computing resource parameters according to the computing resource occupancy data and writing the computing resource parameters into a background database.
[0064] Specifically, the method for obtaining the computing resource occupancy data sent by the test Kubernetes cluster may be: monitoring the computing resource occupancy data during model operation through the Kubernetes cluster monitoring component Prometheus.
[0065] Optionally, after determining computing resource parameters according to the computing resource occupancy data and writing the computing resource parameters into a background database, it further includes:
[0066] Generating a second yaml configuration file according to the computing resource parameters and the model configuration information;
[0067] Receiving a release instruction and sending the release instruction to the production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file;
[0068] Obtaining model operation logs and model evaluation results by accessing a GPRC port and displaying the model operation logs and the model evaluation results.
[0069] Specifically, the platform analyzes the data of the resources occupied by the model, automatically configures the most suitable computing resource parameters for the model, and writes them into the platform background database. When the developer publishes the model with one click on the platform, the platform uploads the packaged model file and configuration file to the production environment Kubernetes cluster. The platform automatically generates the yaml configuration file of the pod for model training in the production environment Kubernetes cluster based on the computing resource configuration information, image information, and model configuration information in the database, and exposes the GRPC ports for model operation logs and model result evaluation. The platform accesses the ports external to the model pod, and the front end displays the model operation logs and the model result evaluation page.
[0070] In a specific example, the embodiment of the present invention proposes a multi-language type model management and deployment system based on a Kubernetes cluster, including: a model factory management module, an image library management module, a model resource quota management module, a model service management module, and a model result monitoring module; the model factory management module is used to uniformly manage the model and store the metadata of the model, and the model is uploaded through a local model file; the image library management module is used to provide basic environment image support for model training and publishing, and users can freely select basic environment images that support different modeling languages; the resource quota management module makes the best allocation of model computing resources based on the test and trial operation conditions of the model; the service management module is used to deploy the model and manage the model life cycle according to the model management method and resource quota management method output by the model factory management module and the resource quota management module; the model result monitoring module is used to monitor the model operation status and model evaluation results in real time. The relevant functional modules of the system are as follows:
[0071] 1. The functions of the model factory management module also include: management of model parameters, model version management, and one-click service publishing.
[0072] 2. The basic environment images include built-in images and user-defined images. The built-in images include the dependency environments of fixed frameworks, and the user-defined images include the environment images defined by users except for the dependency environments of fixed frameworks.
[0073] 3. The functions of the resource quota management module include resource application, resource approval, resource monitoring, and elastic scaling.
[0074] 4. The deployment of the model in the service management module supports the deployment of built-in fixed frameworks and user-defined frameworks.
[0075] 5. The service monitoring module provides a visual interface to monitor the performance of each model service after deployment and the operation status of each model service in real time by the service management module.
[0076] In another specific example, the machine learning model management platform described in the multi-language type model management and deployment system based on the Kubernetes cluster is implemented according to the following steps:
[0077] Step 1: The model developer fills in the model configuration information and metadata except for the environment image information in the model factory and uploads the local model file to the platform.
[0078] Step 2: The model developer selects the environment image for model training in the image library.
[0079] Step 3: The platform packages and uploads the model file and configuration information uploaded by the model developer to the local directory of the test Kubernetes cluster.
[0080] Step 4: The platform writes the configuration information input by the model developer into the background database, automatically generates the yaml configuration file of the pod for model training in the test Kubernetes cluster, and obtains the model file by mounting the local directory through the pod.
[0081] Step 5: The model developer applies to deploy the model in the test and trial operation environments. After resource review, the platform deploys the model in the test Kubernetes cluster and monitors the computing resource occupancy data during model operation through the Kubernetes cluster monitoring component Prometheus.
[0082] Step 6: The platform analyzes the data of the resources occupied by the model and automatically configures the most suitable computing resource parameters for the model and writes them into the platform background database.
[0083] Step 7: The developer publishes the model with one click on the platform, and the platform uploads the packaged model file and configuration file to the production environment Kubernetes cluster.
[0084] Step 8: The platform automatically generates the yaml configuration file of the pod for model training in the production environment Kubernetes cluster based on the computing resource parameters and model configuration information in the database, and exposes the GRPC ports for model operation logs and model result evaluation.
[0085] Step 9: The platform accesses the port external to the model pod, and the front end displays the model operation log and model result evaluation page.
[0086] In the embodiment of the present invention, by analyzing the data of the resources occupied by the model in the test and trial operation environments, the most suitable computing resource parameters are automatically configured for the model. The Kubernetes cluster is used for model training, and the model training tasks run in the form of pods, which simplifies the model deployment process and supports dynamic expansion of computing resources.
[0087] The multi-language type model management and deployment system based on the Kubernetes cluster in the embodiments of the present invention simplifies the model deployment process and improves the efficiency of developers. The multi-language type model management and deployment system based on the Kubernetes cluster supports dynamic elastic scaling, and the total model computing resources can be increased at any time. The multi-language type model management and deployment system based on the Kubernetes cluster has better monitoring of the computing resources of the model.
[0088] In the technical solution of this embodiment, by obtaining model configuration information, metadata, and model files, where the model configuration information includes: environment image information; uploading the model file and the model configuration information to the local of the test Kubernetes cluster, and writing the model configuration information into the background database; generating a first yaml configuration file according to the model configuration information; receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file, the problems of missing metadata information in model management, lack of model evaluation results, missing monitoring of computing resources in model deployment, and lack of elastic scaling are solved. It can simplify the model deployment process, improve the efficiency of developers, support dynamic elastic scaling, can increase the total model computing resources at any time, and can better monitor the computing resources of the model.
[0089] Figure 2 It is a schematic structural diagram of a model deployment device provided by an embodiment of the present invention. This embodiment is applicable to the situation of model deployment. The device can be implemented in software and / or hardware, and the device can be integrated in any device that provides model deployment functions, such as Figure 2 As shown, the model deployment device specifically includes: an obtaining module 210, an uploading module 220, a generating module 230, and a receiving module 240.
[0090] Among them, the obtaining module 210 is used to obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information;
[0091] The uploading module 220 is used to upload the model file and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database;
[0092] The generating module 230 is used to generate a first yaml configuration file according to the model configuration information;
[0093] A receiving module 240, configured to receive a test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys a model according to the first yaml configuration file.
[0094] Optionally, the receiving module is specifically configured to:
[0095] Receive a test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, the first target pod obtains a model file according to a local directory, and performs a test according to the model file.
[0096] Optionally, the receiving module is further configured to:
[0097] After receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, the first target pod obtains a model file according to a local directory, and performs a test according to the model file, obtain computing resource occupancy data sent by the test Kubernetes cluster;
[0098] Determine computing resource parameters according to the computing resource occupancy data, and write the computing resource parameters into a background database.
[0099] Optionally, the receiving module is further configured to:
[0100] After determining computing resource parameters according to the computing resource occupancy data and writing the computing resource parameters into a background database, generate a second yaml configuration file according to the computing resource parameters and the model configuration information;
[0101] Receive a release instruction and send the release instruction to the production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file;
[0102] Obtain model running logs and model evaluation results by accessing the GPRC port, and display the model running logs and the model evaluation results.
[0103] The above product can execute the method provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0104] The technical solution of this embodiment obtains model configuration information, metadata, and model files. Among them, the model configuration information includes: environment image information; uploads the model files and the model configuration information to the local of the test Kubernetes cluster, and writes the model configuration information into the background database; generates a first yaml configuration file according to the model configuration information; receives a test instruction, and sends the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file, solving the problems of missing metadata information in model management, lack of model evaluation results, missing calculation resource monitoring in model deployment, and lack of elastic scaling. It can simplify the process of model deployment, improve the efficiency of developers, support dynamic elastic scaling, can increase the total model computing resources at any time, and can better monitor the computing resources of the model.
[0105] Figure 3 It is a schematic structural diagram of an electronic device in an embodiment of the present invention. Figure 3 It shows a block diagram of an exemplary electronic device 12 suitable for implementing the embodiments of the present invention. Figure 3 The shown electronic device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0106] As Figure 3 As shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0107] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0108] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0109] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 3 not shown, typically referred to as a "hard disk drive"). Although Figure 3 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (compact disc-read only memory (CD-ROM), digital video disc-read only memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data media interfaces. The system memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0110] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.
[0111] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. In addition, in the electronic device 12 of this embodiment, the display 24 does not exist as an independent entity, but is embedded in the mirror. When the display surface of the display 24 is not displaying, the display surface of the display 24 visually merges with the mirror surface. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN)) and / or a public network, such as the Internet, through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) systems, tape drives, and data backup storage systems, etc.
[0112] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, for example, implementing the model deployment method provided by the embodiments of the present invention:
[0113] Obtain model configuration information, metadata, and model files, wherein the model configuration information includes: environment image information;
[0114] Upload the model file and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database;
[0115] Generate a first yaml configuration file according to the model configuration information;
[0116] Receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file.
[0117] Figure 4A schematic structural diagram of a computer-readable storage medium containing a computer program in an embodiment of the present invention. An embodiment of the present invention provides a computer-readable storage medium 61, on which a computer program 610 is stored. When the program is executed by one or more processors, it implements the model deployment method provided by all embodiments of the present application:
[0118] Obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information;
[0119] Upload the model file and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database;
[0120] Generate a first yaml configuration file according to the model configuration information;
[0121] Receive a test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file.
[0122] Any combination of one or more computer-readable media can be adopted. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0123] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0124] The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0125] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0126] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.
[0127] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the “C” language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0129] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not, in some cases, constitute a limitation on the unit itself.
[0130] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and the like.
[0131] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A model deployment method, characterized in that, Including: Obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information; Upload the model files and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database; Generate a first yaml configuration file according to the model configuration information; Receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file; Obtain the computing resource occupancy data sent by the test Kubernetes cluster; Determine computing resource parameters according to the computing resource occupancy data, and write the computing resource parameters into the background database; Generate a second yaml configuration file according to the computing resource parameters and the model configuration information; Receive a release instruction, and send the release instruction to the production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file; Obtain the model running log and the model evaluation result by accessing the GPRC port, and display the model running log and the model evaluation result.
2. The method according to claim 1, characterized in that Receiving a test instruction and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file includes: Receiving a test instruction, and sending the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file, and the first target pod obtains the model file according to the local directory and performs tests according to the model file.
3. A model deployment device, characterized in that, Including: An obtaining module, configured to obtain model configuration information, metadata, and model files, where the model configuration information includes: environment image information; An uploading module, configured to upload the model files and the model configuration information to the local of the test Kubernetes cluster, and write the model configuration information into the background database; A generating module, configured to generate a first yaml configuration file according to the model configuration information; A receiving module, configured to receive a test instruction, and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster deploys the model according to the first yaml configuration file; The receiving module is further configured to: obtain the computing resource occupancy data sent by the test Kubernetes cluster; Determine computing resource parameters according to the computing resource occupancy data, and write the computing resource parameters into the background database; Generate a second yaml configuration file according to the computing resource parameters and the model configuration information; Receive the release instruction and send the release instruction to the production environment Kubernetes cluster, so that the production environment Kubernetes cluster performs production according to the second yaml configuration file; Obtain the model operation log and the model evaluation result by accessing the GPRC port, and display the model operation log and the model evaluation result.
4. The device according to claim 3, characterized in that, The receiving module is specifically used for: Receive the test instruction and send the test instruction to the test Kubernetes cluster, so that the test Kubernetes cluster generates a first target pod according to the first yaml configuration file. The first target pod obtains the model file according to the local directory and performs tests according to the model file.
5. An electronic device, characterized in that, Includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the processors implement the method according to any one of claims 1-2.
6. A computer-readable storage medium containing a computer program, having the computer program stored thereon, characterized in that, When the program is executed by one or more processors, it implements the method according to any one of claims 1-2.
Citation Information
Patent Citations
Model online deployment method and device
CN112015519A
Container-based machine learning process training task execution method and system
CN112418438A