A Distributed Container Management Method and System for AI Applications Based on K8s
Through the distributed container management method of AI applications based on K8s, the AI applications are automatically deployed and managed, and the problems of edge-side deployment and resource management are solved, automated deployment and resource scheduling are realized, and the diversity needs of edge services are met.
Patent Information
- Application Number
- CN202410951779.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-07-16
AI Technical Summary
When deploying and managing AI applications, the prior art faces problems such as model format conversion, resource management, automatic updates and version fallbacks. Especially in edge-side application scenarios, multiple data streams need to be processed and near-real-time services are provided, and resources cannot be expanded arbitrarily.
Adopting the distributed container management method of AI applications based on K8s, the application code is obtained through the development container of K8s and uploaded to the data warehouse. The automated deployment unit creates resource requests based on the code push request, sets up workflows, and allocates work nodes through the node scheduling model to realize automatic construction, testing, mirroring and deployment.
It realizes automated deployment and management, solves the problems of self-recovery and redundancy of node failures, and meets the diversity needs of edge services by optimizing resource scheduling, reduces redundant data transmission, and realizes automated deployment integration and reuse.
Smart Images

Figure CN118860651B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Kubernetes cluster deployment, and particularly relates to a method and system for distributed container management of AI applications based on K8s. Background Art
[0002] As an intelligent service, AI algorithm models are being widely used in various data analysis and understanding tasks. Since the computing and storage resources of end devices are limited and intelligent services need to be provided to users at the edge side, it is necessary to deploy AI applications to a specific platform for execution. However, due to the different model formats and representation methods developed by different frameworks, tasks such as model format conversion and correctness verification need to be carried out before model deployment, and after model deployment, development of other data processing processes, assembly of applications, and performance optimization for key processes also need to be carried out; the edge-side application scenarios are open, and different applications need to be supported and multiple data streams need to be processed simultaneously, and near-real-time services need to be provided. At the same time, the resources on the edge server cannot be arbitrarily scaled; in the traditional application processing method, users need to download, install, and run the application on the local machine and use the physical resources of the local machine. If the application is updated, a new program needs to be downloaded and installed again or an update package needs to be downloaded. It is difficult to achieve automatic update and version rollback, and a version that conforms to the configuration of one's own machine needs to be downloaded and installed, which also increases the development workload for software providers. The prior art usually deploys applications using a cloud management platform, and when deploying AI applications in the cloud, it is necessary to balance the application service quality and resource utilization in the face of dynamically changing loads. Summary of the Invention
[0003] To solve the above problems existing in the prior art, the present invention provides a method for distributed container management of AI applications based on K8s.
[0004] The object of the present invention can be achieved by the following technical solutions:
[0005] S1: Obtain application code through the development container of K8s, and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger.
[0006] S2: The automated deployment unit creates a resource request according to the code push request, sets up a workflow, and allocates corresponding working nodes to the resource management component according to the resource request through the node scheduling model; the working nodes pull the application code corresponding to the code push request, build a working space, and execute and test the pulled application code.
[0007] S3: Mirror the files that have passed the tests in the working space to obtain mirror files, add mirror tags, push the mirror files to the mirror repository, and end the workflow; after the workflow ends, the automated deployment unit sends an API request to the front-end edge server;
[0008] S4: The front-end edge server automatically creates a scheduling unit according to the API request and runs the mirror file, conducts environment tests on the running mirror file, and publishes the mirror file that has passed the tests to the worker node network.
[0009] Specifically, the data warehouse includes a code version storage repository and a mirror repository; the code version storage repository is used to store the historical records of application code versions, and the mirror repository is used to store the images of container running components.
[0010] Specifically, the trigger is a callback mechanism implemented by the HTTP protocol and is used to connect to the service endpoint that receives requests in the automated deployment unit.
[0011] Specifically, the automated deployment unit automatically builds the node environment and initializes the master node and worker nodes, constructs a distributed cluster based on the master node and worker nodes; divides the deep neural network structure of the AI application into a segmented set, and creates resource requests according to the segmented set and the computable resources of the worker nodes.
[0012] Specifically, the node environment includes a built master node, worker nodes, a node automatic cleaning module, a container network intercommunication module, and a front-end automatic scaling module; the master node is used to manage and control the worker nodes, the worker nodes are used to run containerized application programs, the node automatic cleaning module monitors the health status and resource utilization rate of the nodes, and performs automatic cleaning and recovery when the nodes fail or the resource utilization rate is too high to ensure the stable operation of the cluster; the container network intercommunication module provides network intercommunication capabilities for the resource management components in the K8s cluster, and the front-end automatic scaling module is used to automatically scale the front end of the application according to the node load.
[0013] Specifically, the calculation method of the node scheduling model is as follows:
[0014] Calculate the priority value of the resource management component in combination with the actual resource request, and the calculation formula is:
[0015]
[0016] Among them, V(P j ) is the priority value of the resource management component, C is the unit constant coefficient, node cpu is the total CPU of the available nodes, M cpuLet \(CPU_{req}\) be the CPU resource request amounts of the scheduled resource management component and the to-be-scheduled resource management component on available nodes, node mem Let \(M\) be the total memory resources of available nodes mem Let \(Mem_{req}\) be the memory resource request amounts of the scheduled resource management component and the to-be-scheduled resource management component on available nodes, num(P j ) be the number of replicas that the resource management component needs to run, images(P j ) be the size of the image used by the resource management component to run;
[0017] Arrange the resource management components in descending order according to the priority values to obtain a resource priority queue. Calculate the load values of all worker nodes in the cluster based on the load balancing of worker nodes and the distributed cluster. The calculation formula is:
[0018]
[0019] where \(LS(N i ) is the load value of node N i , \(L cpu (N i ) is the CPU usage rate of node N i , \(L mem (N i ) is the memory usage rate of node N i , \(L disk (N i ) is the disk usage rate of node N i , \(L network (N i ) is the network bandwidth usage rate of node N i , \(L gpu (N i ) is the GPU usage rate of node N i , \(a\) is the number of CPU cores, \(S cpu (N i ) is the CPU frequency of node N i , \(S mem (N i ) is the memory capacity of node N i , \(S disk (N i ) is the disk capacity of node N i , \(S network (N i ) is the network bandwidth of node N i , \(b\) is the number of GPU stream processors, \(S gpu (N i ) is the GPU frequency of node N i on.
[0020] Arrange the working nodes in ascending order according to the load value to obtain a node queue;
[0021] Select the resource management component at the head of the resource priority queue, traverse the node queue. If the set of nodes whose remaining resources in the node queue are available for the resource management component at the head is not empty, extract the feasible nodes to form a queue of nodes to be optimized. Select the optimal node from the queue of nodes to be optimized through an optimization strategy, and bind the optimal node to the resource management component at the head; If the set of nodes whose remaining resources in the node queue are available for the resource management component at the head is empty, move the resource management component at the head to the end of the resource priority queue and wait for the next scheduling.
[0022] Specifically, the optimization strategy is: calculate the matching degree by performing similarity matching between the resource requests of the resource management component and the remaining resources of the working nodes, and use the working node with the maximum matching degree as the optimal node of the resource management component.
[0023] A distributed container management system for AI applications based on K8s, including: an application development module, a scheduling optimization module, an integrated deployment module, and a test release module;
[0024] The application development module is used to obtain application code through the development container of K8s and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger;
[0025] The scheduling optimization module is used for the automated deployment unit to create a resource request according to the code push request, set up a workflow, and allocate corresponding working nodes to the resource management component according to the resource request through a node scheduling model; the working node pulls the application code corresponding to the code push request, constructs a working space to execute and test the pulled application code;
[0026] The integrated deployment module is used to mirror the files that have passed the test in the working space to obtain a mirror file, add a mirror tag, push the mirror file to the mirror warehouse and end the workflow; the automated deployment unit sends an API request to the front-end edge server after the workflow ends;
[0027] The test release module is used to automatically create a scheduling unit and run the mirror file by the front-end edge server according to the API request, perform an environment test on the running mirror file, and release the mirror file that has passed the test to the working node network.
[0028] The beneficial effects of the present invention are:
[0029] By automating the construction of the node environment, offline packaging and cross-node communication between any two containers are achieved, and the self-recovery and redundancy problems of node failures are solved. New applications can be quickly published through the front-end page. By setting up a node scheduling model, the resource management component and node resources are optimized and scheduled to meet the diverse resource requirements faced by edge service requests and the complexity of analysis services. The scheduling priority is set according to the actual resource requests of the Pods of the components where the AI applications are located. In the optimization process, the matching degree between the resource requirements of the application and the available resources of the node is combined to reduce the transmission of redundant data, and automated deployment integration and reuse are realized. Brief Description of the Drawings
[0030] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings.
[0031] Figure 1 It is a schematic flowchart of a method for managing distributed containers of AI applications based on K8s according to the present invention. Detailed Embodiments
[0032] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific embodiments, structures, features and their effects of the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0033] Please refer to Figure 1 , a method for managing distributed containers of AI applications based on K8s;
[0034] S1: Obtain the application code through the development container of K8s and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger;
[0035] S2: The automated deployment unit creates a resource request according to the code push request, sets up a workflow, and allocates corresponding working nodes to the resource management component according to the resource request through the node scheduling model; the working nodes pull the application code corresponding to the code push request, build a working space and execute and test the pulled application code;
[0036] S3: Mirror the files that pass the test in the working space to obtain a mirror file, add a mirror tag, push the mirror file to the mirror warehouse and end the workflow; the automated deployment unit sends an API request to the front-end edge server after the workflow ends;
[0037] S4: The front-end edge server automatically creates a scheduling unit and runs the mirror file according to the API request, performs an environment test on the running mirror file, and publishes the mirror file that passes the test to the working node network.
[0038] Specifically, the data warehouse includes a code version storage warehouse and an image warehouse; the code version storage warehouse is used to store the historical records of application code versions, and the image warehouse is used to store the images of container running components.
[0039] Specifically, the trigger is a callback mechanism implemented by the HTTP protocol and is used to connect to the service endpoint that receives requests in the automated deployment unit.
[0040] Specifically, the automated deployment unit automatically builds a node environment and initializes the master node and worker nodes, and constructs a distributed cluster based on the master node and worker nodes; divides the deep neural network structure of the AI application into a segmented set, and creates a resource request according to the segmented set and the computable resources of the worker nodes.
[0041] Specifically, the node environment includes a master node builder, worker nodes, a node automatic cleaning module, a container network intercommunication module, and a front-end automatic scaling module; the master node is used to manage and control the worker nodes, the worker nodes are used to run containerized application programs, the node automatic cleaning module monitors the health status and resource utilization rate of the nodes, and performs automatic cleaning and recovery when the nodes fail or the resource utilization rate is too high to ensure the stable operation of the cluster; the container network intercommunication module provides network intercommunication capabilities for the resource management components in the K8s cluster, and the front-end automatic scaling module is used to automatically scale the front-end of the application program according to the node load.
[0042] In this embodiment, the master node is constructed by passing the configuration parameters required by K8S into the system through the applyKubeadm() method of the kubeadmTempST structure; the environment of the node is configured and the image is loaded by executing the automation script through the Apply() method of the defineKubeadm structure; the Cgroups driver type of Docker is detected first through the genKubeadmConfigFiles() method of the generateKubeadm structure, and then the Kubelet configuration file is automatically generated according to the type and the initialization starts. The worker node is constructed by executing the automation script through the Apply() method of the defineKubeadm structure to configure the node environment and load the image; the Cgroups driver type of Docker is detected first through the genKubeadmConfigFiles() method of the generateKubeadm structure, and then the configuration file of the Node node is generated through the genKubeAdmConfigFile(); finally, the Node node is automatically added to the Master node to join the load balancing by executing commands remotely through the joinmaster.sh script via SSH. The continuous integration and deployment of the system is to automatically load the offline Jenkins-master and Jenkins-slave images and rebuild the Jenkins-slave image through the initJenkins.sh script; then load the Pipeline and Blueocean components and start the images simultaneously through the composeJenkins.sh script, and then distribute the Jenkinsfile to the root directory of each project file by scp; manually add a Webhook to the code repository of each project and execute the script pushCode.sh by the developer to automatically push the code to detect whether the hook is configured successfully; finally, after manual testing by the tester, confirm in the Blueocean interface to select and push the new application image to the image repository.
[0043] Specifically, the node scheduling model calculation method is as follows:
[0044] Calculate the priority value of the resource management component in combination with the actual resource request, and the calculation formula is:
[0045]
[0046] Among them, V(P j ) is the priority value of the resource management component, C is the unit constant coefficient, node cpu is the total CPU of the available nodes, M cpu is the CPU resource request amount of the scheduled resource management component and the to-be-scheduled resource management component on the available nodes, node memis the total available memory resource of the node, in M mem is the memory resource request amount of the scheduled resource management component and the to-be-scheduled resource management component on the available nodes, num(P j ) is the number of replicas that the resource management component needs to run, images(P j ) is the size of the image used by the resource management component to run;
[0047] Arrange the resource management components in descending order according to the priority value to obtain a resource priority queue. Calculate the load value of all worker nodes in the cluster according to the load balance of the worker nodes and the distributed cluster. The calculation formula is:
[0048]
[0049] Among them, LS(N i ) is the load value of node N i , L cpu (N i ) is the CPU usage rate of node N i , L mem (N i ) is the memory usage rate of node N i , L disk (N i ) is the disk usage rate of node N i , L network (N i ) is the network bandwidth usage rate of node N i , L gpu (N i ) is the GPU usage rate of node N i , a is the number of CPU cores, S cpu (N i ) is the CPU frequency of node N i , S mem (N i ) is the memory capacity of node N i , S disk (N i ) is the disk capacity of node N i , S network (N i ) is the network bandwidth of node N i , b is the number of GPU stream processors, S gpu (N i ) is the GPU frequency on node N i .
[0050] Arrange the worker nodes in ascending order according to the load value to obtain a node queue;
[0051] Select the resource management component at the head of the resource priority queue, traverse the node queue. If the set of nodes with remaining resources in the node queue available for the resource management component at the head is not empty, extract the feasible nodes to form a queue of nodes to be optimized, select the optimal node from the queue of nodes to be optimized through an optimization strategy, and bind the optimal node to the resource management component at the head. If the set of nodes with remaining resources in the node queue available for the resource management component at the head is empty, move the resource management component at the head to the end of the resource priority queue and wait for the next scheduling.
[0052] In this embodiment, external expansion is performed on the default scheduler. The default scheduler can be called according to requirements. The scheduling extension program is written in the Golang language. It is necessary to first add a declaration of the optimization strategy in the default policy and add a configuration file of the extended scheduling program in the yaml file of the default scheduling component. In this way, there is no infringement on the source code. It is possible to select the default scheduler and also call the external extension program to meet business requirements, avoiding the problem of lacking the built-in basic scheduling algorithm of the default scheduler due to completely using a custom scheduler.
[0053] Specifically, the optimization strategy is as follows: Calculate the matching degree by performing similarity matching between the resource requests of the resource management component and the remaining resources of the working nodes, and use the working node with the maximum matching degree as the optimal node of the resource management component.
[0054] A distributed container management system for AI applications based on K8s, including: an application development module, a scheduling optimization module, an integrated deployment module, and a test release module;
[0055] The application development module is used to obtain application code through the development container of K8s and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger;
[0056] The scheduling optimization module is used for the automated deployment unit to create a resource request according to the code push request, set up a workflow, and allocate corresponding working nodes to the resource management component according to the resource request through a node scheduling model; the working node pulls the application code corresponding to the code push request, constructs a working space to execute and test the pulled application code;
[0057] The integrated deployment module is used to mirror the files that have passed the test in the working space to obtain a mirror file, add a mirror tag, push the mirror file to the mirror warehouse and end the workflow; the automated deployment unit sends an API request to the front-end edge server after the workflow ends;
[0058] The test release module is used to automatically create a scheduling unit according to the API request through the front-end edge server and run the image file, perform an environment test on the running image file, and release the image file that passes the test to the worker node network.
[0059] The computer storage medium according to the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0060] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0061] The program code contained on the computer-readable medium can be transmitted with any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above. The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0062] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation of the present invention in any form. Although the present invention has been disclosed as above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the technical content disclosed above within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A distributed container management method for AI applications based on K8s, characterized in that: include: S1: Obtain application code through the K8s development container and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger; S2: the automated deployment unit creates a resource request according to the code push request, sets a workflow, and allocates corresponding working nodes to the resource management component according to the resource request through a node scheduling model; The node scheduling model calculation method is: The priority value of the resource management component is calculated in combination with the actual resource request, and the calculation formula is: , in, V ( P j ) is the priority value of the resource management component, C is the unit constant coefficient, node cpu is the total amount of CPU available on the node, M cpu is the CPU resource request of the scheduled resource management components and the resource management components to be scheduled on the available nodes. node mem is the total amount of available node memory resources, M mem is the memory resource request of the scheduled resource management components and the resource management components to be scheduled on the available nodes, num ( P j ) is the number of copies of the resource management component to run, images ( P j ) is the size of the image used by the resource management component to run; The resource management components are arranged from large to small according to the priority values to obtain a resource priority queue, and the load values of all working nodes in the cluster are calculated according to the load balancing of the working nodes and the distributed cluster. The calculation formula is: , in, L S( N i ) is a node N i The load value, L cpu ( N i ) is a node N i CPU usage, L mem ( N i ) is a node N i Memory usage, L disk ( N i ) is a node N i Disk usage of L network ( N i ) is a node N i The network bandwidth usage, L gpu ( N i ) is a node N i GPU usage, a is the number of CPU cores, S cpu ( N i ) is a node N i CPU frequency, S mem ( N i ) is a node N i Memory capacity, S disk ( N i ) is a node N i Disk capacity, S network ( N i ) is a node N i The network bandwidth b is the number of GPU stream processors, S gpu ( N i ) is a node N i The frequency of the GPU; Arrange the working nodes from small to large according to the load values to obtain a node queue; Select the head resource management component of the resource priority queue, traverse the node queue, if the node set of remaining resources of the nodes in the node queue for use by the head resource management component is not empty, extract feasible nodes to form a queue to be optimized, select the best node in the queue to be optimized through the optimization strategy, and bind the best node to the head resource management component; if the node set of remaining resources of the nodes in the node queue for use by the head resource management component is empty, put the head resource management component at the end of the resource priority queue, waiting for the next scheduling; The preferred strategy is: perform similarity matching calculation on the resource request of the resource management component and the resource surplus of the working node to obtain a matching degree, and use the working node with the maximum matching degree as the optimal node of the resource management component; The working node pulls the application code corresponding to the code push request, builds a workspace to execute and test the pulled application code; S3: mirroring the files that have passed the test in the workspace to obtain a mirror file, adding a mirror tag, pushing the mirror file to the mirror warehouse and ending the workflow; the automatic deployment unit sends an API request to the front-end edge server after the workflow ends; S4: The front-end edge server automatically creates a scheduling unit according to the API request and runs the image file, performs an environmental test on the running image file, and publishes the image file that passes the test to the working node network.
2. The method according to claim 1, characterized in that The data warehouse includes a code version storage warehouse and an image warehouse; the code version storage warehouse is used to store the historical records of application code versions, and the image warehouse is used to store the images of container running components.
3. The method according to claim 1, characterized in that The trigger is a callback mechanism implemented by the HTTP protocol, and is used to connect to the service endpoint receiving the request in the automated deployment unit.
4. The method according to claim 1, characterized in that: The automated deployment unit automatically builds a node environment and initializes a master node and a working node, and builds a distributed cluster based on the master node and the working node; divides the deep neural network structure of the AI application into a segment set, and creates a resource request based on the segment set and the computable resources of the working node.
5. The method according to claim 4, characterized in that The node environment includes building a master node, a working node, a node automatic cleaning module, a container network intercommunication module, and a front-end automatic expansion module; the master node is used to manage and control the working node, the working node is used to run containerized applications, the node automatic cleaning module monitors the health status and resource utilization of the node, and automatically cleans and recovers when the node fails or the resource utilization is too high to ensure the stable operation of the cluster; the container network intercommunication module provides network intercommunication capabilities for the resource management components in the K8s cluster, and the front-end automatic expansion module is used to automatically expand the front end of the application according to the node load.
6. A K8s-based AI application distributed container management system, used to execute the method as claimed in any one of claims 1 to 5, characterized in that: include: Application development module, scheduling optimization module, integrated deployment module, test release module; The application development module is used to obtain application code through the K8s development container and upload the application code to the data warehouse; the data warehouse sends a code push request to the automated deployment unit through a trigger; The scheduling optimization module is used for the automatic deployment unit to create a resource request according to the code push request, set a workflow, and allocate corresponding working nodes to the resource management component according to the resource request through the node scheduling model; The node scheduling model calculation method is: The priority value of the resource management component is calculated in combination with the actual resource request, and the calculation formula is: , in, V ( P j ) is the priority value of the resource management component, C is the unit constant coefficient, node cpu is the total amount of CPU available on the node, M cpu is the CPU resource request of the scheduled resource management components and the resource management components to be scheduled on the available nodes. node mem is the total amount of available node memory resources, M mem is the memory resource request of the scheduled resource management components and the resource management components to be scheduled on the available nodes, num ( P j ) is the number of copies of the resource management component to run, images ( P j ) is the size of the image used by the resource management component to run; The resource management components are arranged from large to small according to the priority values to obtain a resource priority queue, and the load values of all working nodes in the cluster are calculated according to the load balancing of the working nodes and the distributed cluster. The calculation formula is: , in, L S( N i ) is a node N i The load value, L cpu ( N i ) is a node N i CPU usage, L mem ( N i ) is a node N i Memory usage, L disk ( N i ) is a node N i Disk usage of L network ( N i ) is a node N i The network bandwidth usage, L gpu ( N i ) is a node N i GPU usage, a is the number of CPU cores, S cpu ( N i ) is a node N i CPU frequency, S mem ( N i ) is a node N i Memory capacity, S disk ( N i ) is a node N i Disk capacity, S network ( N i ) is a node N i The network bandwidth b is the number of GPU stream processors, S gpu ( N i ) is a node N i The frequency of the GPU; Arrange the working nodes from small to large according to the load values to obtain a node queue; Select the head resource management component of the resource priority queue, traverse the node queue, if the node set of remaining resources of the nodes in the node queue for use by the head resource management component is not empty, extract feasible nodes to form a queue to be optimized, select the best node in the queue to be optimized through the optimization strategy, and bind the best node to the head resource management component; if the node set of remaining resources of the nodes in the node queue for use by the head resource management component is empty, put the head resource management component at the end of the resource priority queue, waiting for the next scheduling; The preferred strategy is: perform similarity matching calculation on the resource request of the resource management component and the resource surplus of the working node to obtain a matching degree, and use the working node with the maximum matching degree as the optimal node of the resource management component; The working node pulls the application code corresponding to the code push request, builds a workspace to execute and test the pulled application code; The integrated deployment module is used to mirror the files that have passed the test in the workspace to obtain a mirror file, add a mirror tag, push the mirror file to the mirror warehouse and end the workflow; the automated deployment unit sends an API request to the front-end edge server after the workflow ends; The test publishing module is used to automatically create a scheduling unit and run the image file according to the API request through the front-end edge server, perform environmental testing on the running image file, and publish the image file that passes the test to the working node network.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the K8s-based AI application distributed container management method as described in any one of claims 1-5.
8. A storage medium containing computer executable instructions, characterized in that: When executed by a computer processor, the computer executable instructions are used to execute the K8s-based AI application distributed container management method as described in any one of claims 1-5.
Citation Information
Patent Citations
Container automatic arrangement method for distributed training of deep learning model
CN115794385A
Application program automatic deployment method and system based on cloud native
CN116755794A