Container resource recommendation method and device, electronic equipment, storage medium and product

By analyzing the historical operation data and target utilization rate of the container, the amount of resources and number of replicas required by the container are calculated, which solves the problem that the existing technology cannot recommend the container specifications and number of replicas at the same time, and realizes efficient resource utilization and maximizes system performance.

CN120066687AInactive Publication Date: 2025-05-30GUANGZHOU YAXIN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510553455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot simultaneously implement reasonable recommendations for container specifications and replica counts, which makes it difficult to maximize system performance and optimize resource utilization.

Method used

By determining the Pod instance and historical operation data of the target application running, the required total operation resources and target utilization rate are calculated, and the first resource amount and number of replicas are determined. Based on the proportional relationship of these data, the target resource quantity and target replica number are calculated to ensure the reasonable allocation and utilization of resources.

Benefits of technology

It realizes reasonable recommendations for container resources and replicas, improves system performance and resource utilization, and ensures the normal operation of the application and the maximum utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066687A_ABST
    Figure CN120066687A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a container resource recommendation method and device, electronic equipment, a storage medium and a product, and relates to the field of cloud computing. The method comprises the steps of determining at least one Pod instance for running a target application program, and obtaining historical running data of at least one type of running resources of each Pod instance; for each type of operation resources, determining the total resource quantity of the operation resources required for operating the target application program according to the historical operation data, and determining the first resource quantity of the operation resources based on the target utilization rate and the total resource quantity; if it is determined that the reference resource quantity is smaller than the first resource quantity, determining the number of copies of the operation resources based on the proportional relation between the reference resource quantity and the first resource quantity; selecting the maximum copy number as a target copy number; and determining a target resource quantity based on a proportional relationship between the target copy number and the first resource quantity, and taking the target resource quantity as the resource quantity of operation resources required for operating the target application program. The resource quantity and the copy number are reasonably recommended at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology. Specifically, this application relates to a method, device, electronic device, storage medium, and product for recommending container resources. Background Art

[0002] The application proportion of cloud container technology in the field of cloud computing is increasing compared with traditional virtualization technology. Especially, the rise of AIGC technology and the natural fit with container virtualization also promote the development of cloud container technology. For scenarios such as large-scale AI model training and online services, container virtualization technology can provide higher flexibility, scalability, and resource utilization. In this case, reasonable scheduling of container resource specifications and replicas is particularly crucial to ensure the maximization of system performance and the optimization of resource utilization.

[0003] However, current related container resource recommendation methods cannot reasonably recommend both container specifications and replicas at the same time. Summary of the Invention

[0004] Embodiments of this application provide a method, device, electronic device, storage medium, and product for recommending container resources, which are used to solve the technical problem that it is impossible to reasonably recommend both container specifications and replicas at the same time.

[0005] According to the first aspect of the embodiments of this application, a method for recommending container resources is provided. The method includes: determining at least one Pod instance running a target application program, and obtaining historical running data of at least one type of running resource of each of the at least one Pod instance; For each type of running resource, determining the total resource amount of the running resource required to run the target application program according to the historical running data of the running resource, and determining the first resource amount of the running resource based on the target utilization rate of the running resource determined in advance and the total resource amount, where the target utilization rate of the running resource is used to represent the distribution of the utilization rates of all Pod instances for the running resource; For each type of running resource, if it is determined that the reference resource amount is less than the first resource amount, then based on the proportional relationship between the reference resource amount and the first resource amount, determining the number of replicas of the running resource, where the reference resource amount is the maximum value of the resource amount allocated when the running resource runs on a node; the node is used to run the Pod instance; Selecting the largest number of replicas among the numbers of replicas of various types of running resources as the target number of replicas; For each type of running resource, based on the proportional relationship between the target number of replicas and the first resource amount, determining the target resource amount of the running resource, and using the target resource amount as the resource amount of the corresponding running resource required for the Pod instance to run the target application program.

[0006] In a possible implementation, if the amount of reference resources is not less than the first amount of resources, determine that the number of replicas of the running resources is 1; select the largest number of replicas from the numbers of replicas of various running resources as the target number of replicas.

[0007] In another possible implementation, the historical running data includes the peak value of the amount of resources used by the running resources within a preset time period; determine the peak value of the amount of resources used by the running resources within a preset time period and the historical number of replicas of the Pod instances running the target application; determine the product between the peak value and the historical number of replicas to obtain the total amount of resources of the running resources required to run the target application.

[0008] In yet another possible implementation, determine a first ratio between the total amount of resources and the target utilization rate, and use the first ratio as the first amount of resources.

[0009] In yet another possible implementation, determine a second ratio between the first amount of resources and the amount of reference resources, and perform a ceiling operation on the second ratio to obtain the number of replicas.

[0010] In yet another possible implementation, determine a third ratio between the first amount of resources and the target number of replicas, and use the third ratio as the target amount of resources.

[0011] In yet another possible implementation, determine the running requirements of the target application; the running requirements are used to characterize the requirements for the resource utilization rate of the Pod instances running the target application; Determine the target utilization rate corresponding to the running requirements according to the pre-stored mapping relationship; wherein, the target utilization rate is: The average value of the utilization rates of all Pod instances; The median of the utilization rates of all Pod instances, or The product of the utilization rate peak of the utilization rates of all Pod instances and a preset peak coefficient; In the mapping relationship, each type of running requirement corresponds to a target utilization rate.

[0012] In yet another possible implementation, the running resources include at least two of the following: CPU resources; Memory resources; GPU resources; NPU resources; Bandwidth resources.

[0013] According to the second aspect of the embodiments of the present application, there is provided a container resource recommendation device, and the device includes: The first determination module is configured to determine at least one Pod instance running the target application and obtain historical running data of at least one type of running resources of each of the at least one Pod instance; The second determination module is configured to, for each type of running resources, determine the total resource amount of the running resources required to run the target application according to the historical running data of the running resources, and determine the first resource amount of the running resources based on the target utilization rate and the total resource amount of the running resources determined in advance, where the target utilization rate of the running resources is used to represent the distribution of the utilization rates of the running resources by all Pod instances; The third determination module is configured to, for each type of running resources, if it is determined that the reference resource amount is less than the first resource amount, determine the number of replicas of the running resources based on the proportional relationship between the reference resource amount and the first resource amount, where the reference resource amount is the maximum value of the resource amount allocated when the running resources run on the node; the node is used to run the Pod instance; The selection module is configured to select the largest number of replicas from the numbers of replicas of various types of running resources as the target number of replicas; The fourth determination module is configured to, for each type of running resources, determine the target resource amount of the running resources based on the proportional relationship between the target number of replicas and the first resource amount, and use the target resource amount as the resource amount of the corresponding running resources required for the Pod instance to run the target application.

[0014] According to the third aspect of the embodiments of the present application, there is provided an electronic device, which includes a memory, a processor, and a computer program stored on the memory. When the processor executes the program, the steps of the method provided in the first aspect are implemented.

[0015] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.

[0016] According to the fifth aspect of the embodiments of the present application, there is provided a computer program product, which includes computer instructions stored in a computer-readable storage medium. When the processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, the computer device is caused to execute the steps of the method provided in the first aspect.

[0017] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are: The container resource recommendation method provided by the embodiments of this application determines at least one Pod instance running the target application and the historical operation data of at least one type of running resource of each Pod instance. For each type of running resource obtained, the total resource amount of the running resource required to run the target application is determined according to the historical operation data of the running resource. Based on the target utilization rate and the total resource amount of the running resource determined in advance, the first resource amount of the running resource is determined. If the reference value representing the maximum value of the resource amount allocated for the running resource to run on the node is less than the first resource amount, it means that the current node cannot provide the first resource amount required to run the target application. Therefore, based on the proportional relationship between the reference resource amount and the first resource amount, the number of replicas of the running resource is determined. After determining the number of replicas of each type of running resource, in order to ensure that the determined number of replicas can provide the resource amount required by all running resources, the largest number of replicas is selected as the target number of replicas. Since the first resource amount to be provided is fixed, based on the proportional relationship between the target number of replicas and the first resource amount, the target resource amount of the running resource is determined, and the target resource amount is used as the resource amount of the corresponding running resource required for the Pod instance to run the target application. By collecting the historical operation conditions of each actual Pod instance running the target application and the global resource utilization rate, the target number of replicas of the Pod instances required to run the target application and the target resource amount provided by the Pod instances for each type of running resource are calculated, realizing reasonable recommendations for both the resource amount and the number of replicas at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description in the embodiments of this application.

[0019] Figure 1 It is a schematic diagram of the system architecture for implementing the container resource recommendation method provided by the embodiments of this application; Figure 2 It is a schematic flowchart of a container resource recommendation method provided by the embodiments of this application; Figure 3 It is a schematic flowchart of a container resource recommendation method provided by the embodiments of this application; Figure 4 It is a schematic diagram of the system architecture for implementing the container resource recommendation method provided by the embodiments of this application; Figure 5 It is a schematic diagram of the structure of a container resource recommendation device provided by the embodiments of this application; Figure 6 It is a schematic diagram of the structure of an electronic device provided by the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0021] Those skilled in the art of this technology can understand that unless specifically stated, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by this technology field. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0022] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0023] The related technologies will be described below: Currently, for the container resource recommendation method commonly used in the industry, the scheduling algorithm is relatively simple, and the implementation method completely depends on the vertical and horizontal scaling technologies of containers. Therefore, the container specification recommendation strategy and the container replica number recommendation strategy are divided into two mutually exclusive recommendation strategies. Customers can only be forced to choose one of them, and service providers cannot provide a selection recommendation for the two scheduling schemes either. There is a lack of a unified scheduling strategy that combines the two.

[0024] In view of the above at least one technical problem or area for improvement in the related art, the present application proposes a container resource recommendation method. By determining at least one Pod instance running the target application program and the historical operation data of at least one type of operation resource for each Pod instance; for each type of operation resource obtained, determine the total resource amount of the operation resource required to run the target application program according to the historical operation data of the operation resource, and determine the first resource amount of the operation resource based on the pre-determined target utilization rate and total resource amount of the operation resource. If the reference value representing the maximum value of the resource amount allocated for the operation resource running on the node is less than the first resource amount, it indicates that the current node cannot provide the first resource amount required to run the target application program. Therefore, based on the proportional relationship between the reference resource amount and the first resource amount, determine the number of replicas of the operation resource. After determining the number of replicas of each type of operation resource, in order to ensure that the determined number of replicas can provide the resource amount required for all operation resources, select the largest number of replicas as the target number of replicas. Since the first resource amount to be provided is fixed, based on the proportional relationship between the target number of replicas and the first resource amount, determine the target resource amount of the operation resource, and use the target resource amount as the resource amount of the corresponding operation resource required for the Pod instance to run the target application program. By collecting the historical operation conditions of each actual Pod instance running the target application program and the global resource utilization rate, calculate the target number of replicas of the Pod instances required to run the target application program and the target resource amount provided by the Pod instances for each type of operation resource, thus realizing a reasonable recommendation for both the resource amount and the number of replicas.

[0025] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0026] Figure 1 It is a schematic diagram of the system architecture for implementing the container resource recommendation method provided by the embodiment of the present application, where the system architecture includes: a terminal 120 and a server 140.

[0027] The terminal 120 installs and runs an application program for the container resource recommendation method. The terminal 120 is used to collect the historical operation data and target utilization rate of at least one type of operation resource for each Pod instance, and determine the target resource amount and the target number of replicas based on the collected data.

[0028] The terminal 120 is connected to the server 140 through a wireless network or a wired network.

[0029] The server 140 includes at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. Schematically, the server 140 includes a first processor 144 and a first memory 142. The first memory 142 includes a display module 1421, a control module 1422, and a receiving module 1423. The server 140 is used to provide background services for the application of the container resource recommendation method. Optionally, the server 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, a distributed computing architecture is adopted between the server 140 and the terminal 120 for collaborative computing.

[0030] Optionally, the device types of the terminal include at least one of a smart phone, a tablet computer, an e - book reader, a laptop computer, and a desktop computer.

[0031] Those skilled in the art can know that the number of the above - mentioned terminals can be more or less. For example, the above - mentioned terminal can be only one, or the above - mentioned terminals can be dozens or hundreds, or a larger number. The embodiments of the present application do not limit the number and device types of the terminals.

[0032] An embodiment of the present application provides a container resource recommendation method, as Figure 2 shown, the method includes: S101, determine at least one Pod instance running the target application, and obtain the historical running data of at least one type of running resources of each of the at least one Pod instance.

[0033] In the embodiment of the present application, a Pod instance is the smallest deployment unit in Kubernetes. A Pod instance encapsulates all the resources required to run an application. A Pod instance can contain one or more containers. After a Pod instance is created and assigned to a certain node, it starts to run the application.

[0034] In the embodiment of the present application, the target application is the target application for which container resources are to be recommended. Container resource recommendation includes: Pod instance resource specification recommendation and Pod instance replica number recommendation. The Pod instance resource specification refers to the resources required to run the target application, and the Pod instance replica number refers to the number of Pod instances deployed for the target application.

[0035] In the embodiment of the present application, the target application is containerized and deployed in a Kubernetes cluster. Containerization refers to the process of packaging an application and all its dependencies into a lightweight, portable container. Kubernetes is an open - source container orchestration platform used for automating the deployment, scaling, and management of containerized applications.

[0036] In the embodiments of the present application, determining at least one Pod instance running the target application is to determine the usage of various types of running resources for each Pod instance when running the target application within a historical time period, that is, to obtain the historical running data of at least one type of running resource.

[0037] In the embodiments of the present application, the running resources provided by the Pod instance when running the target application are used to jointly support the normal operation and efficient management of the application in the Pod instance. There are multiple categories of running resources provided by the Pod instance, and the running resources are the resources required to run the target application, such as CPU resources and memory resources in computing resources, etc.

[0038] In the embodiments of the present application, various types of running resources of the Pod instance are collected through Prometheus, so as to obtain at least one type of running resource of each of the at least one Pod instance corresponding to the target application.

[0039] In the embodiments of the present application, by obtaining the historical running conditions of various types of running resources of each Pod instance when running the target application within the historical event segment, the resource consumption of the target application during operation can be effectively and accurately analyzed.

[0040] S102. For each type of running resource, determine the total resource amount of the running resources required to run the target application according to the historical running data of the running resources, and determine the first resource amount of the running resources based on the target utilization rate and the total resource amount of the running resources determined in advance.

[0041] In the implementation of the present application, the historical running data of the running resources is data reflecting the consumption of the running resources by the Pod instance when running the target application, such as the peak value of the running resources within a preset period, the average consumption of the running resources within a preset period, and other data.

[0042] In the embodiments of the present application, for each type of running resource, after determining the historical running data of each Pod instance, according to the historical running data, the total consumption of the running resources when running the target application can be determined, that is, the total resource amount of the running resources required to run the target application within the historical time period to which the obtained historical running data belongs can be obtained.

[0043] In the embodiments of the present application, the target utilization rate of the running resources is used to represent the distribution of the utilization rates of the running resources by all Pod instances respectively.

[0044] In the embodiment of the present application, since the user may have optimized and improved the operation of the resource cluster through an optimization solution and expects to improve the utilization rate of the running resources, therefore, by collecting the utilization rate of each Pod instance in the entire resource cluster for the running resources, the distribution of the utilization rate of the running resources of all Pod instances in the entire resource cluster is determined, and the target utilization rate of the resource is obtained.

[0045] In the embodiment of the present application, the collected utilization rate of the running resources can be the utilization rate of the running resources of all Pod instances in a certain container cluster, or the utilization rate of the running resources of all Pod instances at the region level, or the utilization rate of the running resources of all Pod instances in multiple remote container clusters, or the utilization rate of the running resources of all Pod instances at the server level. The specific selection is determined based on the granularity of the optimization required by the user for the resource pool.

[0046] In one example, for memory resources, the utilization rate of the memory resources of all Pod instances at the single-cluster level is collected. According to the utilization rate of the memory resources of all Pod instances, the target utilization rate of the memory can be determined, that is, the distribution of the utilization rate of the memory resources of all Pod instances at the entire single-cluster level is determined. For example, by sorting the utilization rate of the above-mentioned memory resources, the target utilization rates such as peak utilization rate, average utilization rate, and median utilization rate that can characterize the distribution of the utilization rate of the memory resources can be determined.

[0047] In the embodiment of the present application, for each type of running resources, based on the total amount of the running resources required for running the target application program and the target utilization rate of this type of resources, a first resource amount can be determined, and the first resource amount is used for subsequent calculation of the resource amount and the number of replicas to be recommended.

[0048] S103, for each type of running resources, if it is determined that the reference resource amount is less than the first resource amount, then based on the proportional relationship between the reference resource amount and the first resource amount, the number of replicas of the running resources is determined.

[0049] In the embodiment of the present application, the reference resource amount is the maximum value of the resource amount allocated when the running resources are running on the node; the node is used to run the Pod instance.

[0050] In the embodiment of the present application, the node is used to run the Pod instance and provides running resources for the Pod instance to run the target application program. The node is a working machine in the Kubernetes cluster, and the node can be a physical machine or a virtual machine.

[0051] In the embodiments of the present application, since the resources that a node can allocate to a Pod instance for running an application are limited, for each type of running resource, the maximum value of the amount of resources that can be allocated from the current node to a Pod instance is used as the reference resource amount, such as the amount of the maximum memory resource that can be allocated by the node.

[0052] In the embodiments of the present application, for each type of running resource, if the reference resource amount is less than the first resource amount, it indicates that currently, corresponding running resources need to be provided for multiple Pod instances to run the target application. Since the first resource amount is the total resource amount required to run the target application, based on the proportional relationship between the reference resource amount and the first resource amount, that is, in the case of determining the total resource amount required to be consumed and the resource amount that a single Pod instance can provide, the number of replicas representing the number of Pod instances can be obtained. Through the Pod instances with the number of replicas, the node can provide the resource amount required for the corresponding running resources needed to run the target application.

[0053] S104. Select the largest number of replicas among the numbers of replicas of various types of running resources as the target number of replicas.

[0054] In the embodiments of the present application, since for different running resources, the first resource amount required to run the target application is different, and the reference resource amounts of various types of running resources that a node can provide when running the target application for a single Pod instance are also different, the numbers of replicas of various types of running resources are not unified. To meet the demand for various types of running resources when running the target application, the largest number of replicas is selected as the target number of replicas, which is the number of Pod instances for finally running the target application.

[0055] S105. For each type of running resource, based on the proportional relationship between the target number of replicas and the first resource amount, determine the target resource amount of the running resource, and use the target resource amount as the resource amount of the corresponding running resource required for the Pod instance to run the target application.

[0056] In the embodiments of the present application, since there may be a difference between the finally determined target number of replicas and the numbers of replicas originally calculated for various types of running resources, to ensure the maximum utilization of running resources, for each type of running resource, based on the proportional relationship between the determined target number of replicas and the first resource amount, re-determine the resource amount of the running resource that a single Pod instance needs to provide, that is, the target resource amount. Finally, use the target resource amount as the resource amount of the running resource required for each Pod instance to run the target application.

[0057] In the embodiment of the present application, after determining the target number of replicas and the target resource amounts corresponding to various types of running resources, the container orchestration system is used to specify Pod instances with the target number of replicas to run the target application, and the various types of running resources required by the Pod instances when running the target application are configured based on the target resource amounts.

[0058] In the embodiment of the present application, the FinOps technology is a cloud cost management and optimization solution that provides a systematic methodology for organizations, enterprises, and teams. The container resource recommendation method provided by the embodiment of the present application can, based on the FinOps technology, focus on how to improve resource utilization and reduce cloud costs through refined resource confirmation, management, and optimization strategies.

[0059] The container resource recommendation method provided by the embodiment of the present application determines the historical running data of at least one Pod instance running the target application and at least one type of running resource for each Pod instance. For each type of running resource obtained, the total resource amount of the running resource required to run the target application is determined according to the historical running data of the running resource. The first resource amount of the running resource is determined based on the pre-determined target utilization rate and total resource amount of the running resource. If the reference value representing the maximum value of the resource amount allocated for the running resource to run on the node is less than the first resource amount, it indicates that the current node cannot provide the first resource amount required to run the target application. Therefore, based on the proportional relationship between the reference resource amount and the first resource amount, the number of replicas of the running resource is determined. After determining the number of replicas of various types of running resources, in order to ensure that the determined number of replicas can provide the resource amounts required by all running resources, the largest number of replicas is selected as the target number of replicas. Since the first resource amount to be provided is fixed, the target resource amount of the running resource is determined based on the proportional relationship between the target number of replicas and the first resource amount, and the target resource amount is used as the resource amount of the corresponding running resource required by the Pod instance to run the target application. By collecting the historical running conditions of each actual Pod instance running the target application and the global resource utilization rate, the target number of replicas of the Pod instances required to run the target application and the target resource amounts provided by the Pod instances for various types of running resources are calculated, realizing reasonable recommendations for both the resource amount and the number of replicas at the same time.

[0060] Based on the above embodiments, as an optional embodiment, if the reference resource amount is not less than the first resource amount, the number of replicas of the running resource is determined to be 1; The largest number of replicas is selected from the number of replicas of various types of running resources as the target number of replicas.

[0061] In the embodiments of the present application, for each type of operating resource, if the reference resource amount is not less than the first resource amount, it indicates that the reference resource amount of this operating resource provided by the node for a single Pod instance can meet the resource amount required when running the target application. Therefore, it is determined that the number of replicas of this operating resource is 1, that is, only one Pod instance is used.

[0062] In the embodiments of the present application, after determining the number of replicas of each type of operating resource according to the reference resource amount and the first resource amount, the largest number of replicas is selected from the numbers of replicas of each type of operating resource as the target number of replicas, ensuring that the number of selected Pod instances can meet the demand for each type of operating resource.

[0063] By determining the number of replicas of the required Pod instances based on the maximum allocable reference resource amount provided by the current node for the Pod instance and the first resource amount, the number of replicas can be determined according to the actual performance of the current cluster, ensuring the maximized utilization of resources.

[0064] Based on the above embodiments, as an optional embodiment, the historical operation data includes the peak value of the used resource amount of the operating resource within a preset time period; determine the peak value of the used resource amount of the operating resource within a preset time period and the historical number of replicas of the Pod instance running the target application; determine the product between the peak value and the historical number of replicas to obtain the total resource amount of the operating resource required for running the target application.

[0065] In the embodiments of the present application, the historical operation data includes the peak value of the used resource amount of the operating resource within a preset time period. Since the target application is run by at least one Pod instance, for each type of operating resource, each Pod instance has a peak value of the used resource amount within a preset time period for this operating resource. The peak value with the largest numerical value is selected from the above multiple peak values to participate in the calculation of the total resource amount of the operating resource.

[0066] In the embodiments of the present application, since the target application is run by at least one Pod instance, for each type of operating resource, it is necessary to determine the historical number of replicas of the Pod instance running the target application within a preset time period, multiply the largest peak value by the historical number of replicas to obtain the product between the peak value and the historical number of replicas, and use the above product as the total resource amount of the operating resource required for running the target application.

[0067] In the embodiments of the present application, for each type of operating resource, when calculating the total resource amount of the operating resource, the peak value is adjusted by setting a peak coefficient. The peak coefficient can be a value such as 0.98, 0.95, or 0.92. Multiply the obtained peak value by the peak coefficient to obtain the adjusted peak value, and then multiply the adjusted peak value by the historical copy number to obtain the total resource amount of the operating resource. By setting the peak coefficient, some abnormal peak values can be eliminated, avoiding waste of operating resources.

[0068] In the embodiments of the present application, for each type of operating resource, when calculating the total resource amount of the operating resource, an adjustment coefficient can also be set to adjust the total resource amount of the operating resource. Multiply the product between the calculated peak value and the historical copy number by the adjustment coefficient to obtain the total resource amount of the operating resources required for the running target application.

[0069] It should be noted that during the process of calculating the total resource amount, the collection period (preset time period), peak coefficient, and adjustment coefficient of the historical operation data can be specifically set in combination with the actual business requirements and big data analysis technology.

[0070] In the above solution, for each type of operating resource, obtain the peak value of the operating resource within the preset time period, the historical copy number of the Pod instances of the running target application, and the peak coefficient, and determine the total resource amount of the operating resource required for the running target application. Determining the total resource amount through historical data makes the target resource amount allocated for the operating resource more accurate and reasonable subsequently.

[0071] Based on the above embodiments, as an alternative embodiment, determine the first ratio between the total resource amount and the target utilization rate, and use the first ratio as the first resource amount.

[0072] In the embodiments of the present application, since various operating resources cannot be fully utilized, it is necessary to determine the resource amount that actually needs to be provided for each operating resource, that is, the first resource amount, according to the pre-determined target utilization rate.

[0073] In the embodiments of the present application, for each type of operating resource, by calculating the first ratio between the total resource amount and the target utilization rate, and using the first ratio as the resource amount, the first resource amount that actually needs to be provided is determined.

[0074] In the above solution, by considering the utilization rate of each operating resource during the process of the Pod instance running the target application, the actual demand for the resource amount of each type of operating resource during the process of running the target application is determined more accurately and pertinently.

[0075] Based on the above embodiments, as an alternative embodiment, determine a second ratio between the first resource amount and the reference resource amount, and perform a ceiling operation on the second ratio to obtain the number of replicas.

[0076] In the embodiments of the present application, since the reference resource amount that can be allocated to a Pod instance by a single node is insufficient to meet the requirements for running the target application, therefore, through the second ratio between the first resource amount and the reference resource amount, it can be determined how, with the minimum number of replicas, the running resources provided by the Pod instances can meet the requirements for running the target application.

[0077] In the embodiments of the present application, for each type of running resource, since the obtained second ratio may not be an integer when calculating the second ratio between the first resource amount and the reference resource amount, and the number of replicas must be an integer, therefore, after determining the second ratio, perform a ceiling operation on the second ratio to obtain the number of replicas of this type of running resource.

[0078] In the above solution, since the reference resources that can be allocated by a single node are insufficient to meet the first resource amount for running the target application, therefore, according to the ratio between the actually required first resource amount and the reference resource amount, determine the number of replicas of the required Pod instances, that is, select the most suitable number of replicas according to the allocable situation of the actual nodes, ensuring the maximization of performance.

[0079] Based on the above embodiments, as an alternative embodiment, determine a third ratio between the first resource amount and the target number of replicas, and use the third ratio as the target resource amount.

[0080] In the embodiments of the present application, since the maximum number of replicas is selected from the number of replicas of all running resources as the target number of replicas for each running resource, and the first resource amounts corresponding to each type of running resource are different, therefore, for each type of running resource, calculate the third ratio between the first resource amount and the target number of replicas, and use the third ratio as the target resource amount, that is, for each type of running resource, after determining the first resource amount actually required for running the target application and the target number of replicas of the Pod instances providing the running resources, the required resource amount of each Pod instance during the running of the target application can be determined based on the third ratio.

[0081] In the above solution, based on the third ratio of the first resource amount and the target number of replicas, determine the resource amount of the corresponding running resources required by each Pod instance when finally running the target application, ensuring the normal operation of the application while improving the overall resource utilization rate of cloud resources.

[0082] Based on the above embodiments, as an alternative embodiment, determine the running requirements of the target application; the running requirements are used to characterize the requirements for the resource utilization rate of the Pod instances running the target application. Determine the target utilization rate corresponding to the running requirements according to the pre-stored mapping relationship. Wherein, the target utilization rate is: The average value of the utilization rates of all Pod instances; The median of the utilization rates of all Pod instances, or The product of the utilization rate peak among the utilization rates of all Pod instances and a preset peak coefficient; In the mapping relationship, each type of running requirement corresponds to a target utilization rate.

[0083] In the embodiments of the present application, for each type of running resource, when the target utilization rate is the average value of the utilization rates of all Pod instances, it can be obtained by collecting the utilization rates of all Pod instances for this running resource; when the target utilization rate is the median of the utilization rates of all Pod instances, the utilization rates of all Pod instances for this running resource can be sorted, and then the utilization rate ranked in the middle position is selected as the target utilization rate; when the target utilization rate is the product of the utilization rate peak among the utilization rates of all Pod instances and a preset peak coefficient, the utilization rates of all Pod instances for this running resource are sorted in descending order, and then the utilization rate ranked first is selected as the target utilization rate.

[0084] In the embodiments of the present application, different application programs have different running requirements, that is, different application programs have different requirements for the resource utilization rate of the Pod instances running the application program. Some have extremely high requirements for the resource utilization rate, or some have a more dispersed resource utilization rate distribution. Therefore, the corresponding target utilization rate is selected according to the running requirements for different target application programs.

[0085] In the embodiments of the present application, a mapping relationship is pre-created, and the running requirements are divided into three categories. Each category of running requirements has its corresponding target utilization rate. After determining the running requirements of the target application program, the target utilization rate corresponding to the target application program is determined according to the running requirement mapping relationship, so as to perform subsequent calculation of the first resource quantity.

[0086] In the embodiments of the present application, when the running scenario of the target application program is that the utilization rates of each Pod instance are relatively concentrated and the global resource pool is optimized well, it is more representative to select the average value of the utilization rates of all Pod instances as the target utilization rate.

[0087] In an embodiment of the present application, when the running scenario of the target application is a scenario where the utilization rates of each Pod instance are dispersed and the global resources are not fully optimized, the median of the utilization rates of all Pod instances is selected as the target utilization rate.

[0088] In an embodiment of the present application, when the target application mainly belongs to offline jobs such as big data analysis, video rendering, and machine learning, or in other scenarios with extremely high requirements for resource utilization, the product of the peak utilization rate among the utilization rates of all Pod instances and a preset peak coefficient is selected as the target utilization rate.

[0089] In an embodiment of the present application, after selecting the target utilization rate corresponding to the target application, a utilization rate adjustment coefficient can be set to adjust the target utilization rate, and the adjustment coefficient is continuously optimized through data analysis or prediction techniques.

[0090] In the above solution, according to the running requirements of the target application, the corresponding target utilization rate is determined to achieve targeted processing of target applications with different running requirements, improving the accuracy of the target replica number and target resource amount.

[0091] Based on the above embodiments, as an alternative embodiment, the running resources include at least two of the following: CPU resources; Memory resources; GPU resources; NPU resources; Bandwidth resources.

[0092] In an embodiment of the present application, CPU resources are the computing capabilities of the central processing unit used by a Pod instance when running the target application; memory resources are the random access storage required during container runtime to support the execution of applications within the container; GPU resources (graphics processing units) are used to accelerate compute-intensive workloads, especially in scenarios such as machine learning, deep learning, data analysis, and image processing; in a cloud computing environment, NPUs can be deployed in servers, virtual machines, or containers to accelerate the execution of AI applications. With the development of containerization technology, NPU resources are also integrated into containers and can be efficiently scheduled and used in a cloud environment; in cloud container technology, bandwidth resources are an important factor. In application scenarios highly dependent on network communication, as more and more applications migrate to the cloud platform, the management and optimization of bandwidth become particularly important because a large amount of data exchange and communication may be involved in a containerized environment. The effective management of bandwidth resources can improve the overall performance of the container cluster, reduce latency, and ensure efficient network traffic allocation.

[0093] Reference Figure 3As shown, it exemplarily shows a schematic flow diagram of a container resource recommendation method, and the specific content is as follows: S201. Determine all Pod instances running the target application; S202. For each Pod instance, collect the peak values of CPU resources and memory resources within a preset time period respectively; S203. Select the maximum peak value from the obtained multiple peak values of CPU resources, multiply it by the peak coefficient to obtain the adjusted peak value, and multiply the adjusted peak value by the historical replica number of the Pod instance running the target application to obtain the total resource amount of CPU resources; S204. Select the maximum peak value from the obtained multiple peak values of memory resources, multiply it by the peak coefficient to obtain the adjusted peak value, and multiply the adjusted peak value by the historical replica number of the Pod instance running the target application to obtain the total resource amount of memory resources; S205. Collect the utilization rate of CPU resources of all Pod instances in the cluster, and calculate the target utilization rate corresponding to the CPU resources according to the running requirements of the target application and the preset mapping relationship; S206. Collect the utilization rate of memory resources of all Pod instances in the cluster, and calculate the target utilization rate corresponding to the memory resources according to the running requirements of the target application and the preset mapping relationship; S207. Calculate the first ratio between the total resource amount of CPU resources and the target utilization rate corresponding to the CPU resources, and use the first ratio as the first resource amount of CPU resources; S208. Calculate the first ratio between the total resource amount of memory resources and the target utilization rate corresponding to the memory resources, and use the first ratio as the first resource amount of memory resources; S209. Judge whether the reference resource amount of CPU resources is less than the first resource amount of CPU resources. If it is less, execute step S210; otherwise, execute step S211; S210. Calculate the second ratio between the first resource amount of CPU resources and the reference resource amount of CPU resources, and perform ceiling processing on the second ratio to obtain the replica number of CPU resources.

[0094] S211. Determine that the replica number of CPU resources is 1; S212. Judge whether the reference resource amount of memory resources is less than the first resource amount of memory resources. If it is less, execute step S213; otherwise, execute step S214; S213. Calculate the second ratio between the first resource amount of memory resources and the reference resource amount of memory resources, and perform ceiling processing on the second ratio to obtain the replica number of memory resources.

[0095] S214, determine that the number of copies of memory resources is 1; S215, select the larger number of copies from the number of copies of CPU resources and the number of copies of memory resources as the target number of copies of CPU resources and memory resources; S216, calculate the third ratio between the first resource amount of CPU resources and the target number of copies, and use the third ratio as the target resource amount of CPU resources; S217, calculate the third ratio between the first resource amount of memory resources and the target number of copies, and use the third ratio as the target resource amount of memory resources.

[0096] It should be noted that the above steps S203, S205, S207, S209, S210 or S211 and steps S204, S206, S208, S212, S213 or S214 can be executed simultaneously or successively.

[0097] The embodiment of the present application provides a container resource recommendation method, which can improve the overall resource utilization rate of cloud resources while ensuring the normal operation of the application, realize the effective management of the resource usage process, and ultimately reduce the customer's cloud usage cost.

[0098] Reference Figure 4 As shown, it exemplarily shows a schematic diagram of a system architecture for implementing the container resource recommendation method, and the specific content is as follows: The system architecture provided by the embodiment of the present application includes an acquisition module 401, a resource management module 402, and a resource recommendation module 403.

[0099] The acquisition module 401 interacts with Prometheus to obtain the historical operation data of various running resources in each Pod instance running the target application program and the utilization rate of all Pod instances in the cluster for various running resources, and performs a secondary cache on the obtained data for internal provision to the resource management module 402 and the resource recommendation module 403.

[0100] The resource management module 402 is used to determine the total resource amount of various running resources required to run the target application program according to the historical operation data of various running resources in each Pod instance running the target application program uploaded by the acquisition module, and is also used to store the relevant information of the target application program that needs resource recommendation.

[0101] The resource recommendation module 403 interacts with Kube-scheduler, and is used to perform multiple rounds of calculations according to the data obtained from the resource management module 402 and the acquisition module 401 according to a preset algorithm, and finally obtain the target number of copies and the target resource amount of running the target application program, and generate a recommendation plan.

[0102] An embodiment of the present application provides a container resource recommendation device, such as Figure 5 shown. The container resource recommendation device 50 may include: a first determination module 501, a second determination module 502, a third determination module 503, a selection module 504, and a fourth determination module 505.

[0103] Specifically, the first determination module 501 is configured to determine at least one Pod instance running a target application program, and obtain historical running data of at least one type of running resource of each of the at least one Pod instance; The second determination module 502 is configured to, for each type of running resource, determine the total resource amount of the running resources required to run the target application program according to the historical running data of the running resources, and determine the first resource amount of the running resources based on the target utilization rate of the running resources determined in advance and the total resource amount, where the target utilization rate of the running resources is used to represent the distribution of the utilization rates of the running resources by all Pod instances; The third determination module 503 is configured to, for each type of running resource, if it is determined that the reference resource amount is less than the first resource amount, determine the number of replicas of the running resources based on the proportional relationship between the reference resource amount and the first resource amount, where the reference resource amount is the maximum value of the resource amount allocated when the running resources are running on a node; the node is used to run the Pod instance; The selection module 504 is configured to select the largest number of replicas from the numbers of replicas of various types of running resources as the target number of replicas; The fourth determination module 505 is configured to, for each type of running resource, determine the target resource amount of the running resources based on the proportional relationship between the target number of replicas and the first resource amount, and use the target resource amount as the resource amount of the corresponding running resource required for the Pod instance to run the target application program.

[0104] The container resource recommendation device provided by the embodiment of the present application determines at least one Pod instance running the target application program and the historical operation data of at least one type of operation resource of each Pod instance. For each type of operation resource obtained, the total resource amount of the operation resource required to run the target application program is determined according to the historical operation data of the operation resource. Based on the target utilization rate and the total resource amount of the operation resource determined in advance, the first resource amount of the operation resource is determined. If the reference value representing the maximum value of the resource amount allocated for the operation resource running on the node is less than the first resource amount, it indicates that the current node cannot provide the first resource amount required to run the target application program. Therefore, based on the proportional relationship between the reference resource amount and the first resource amount, the number of replicas of the operation resource is determined. After determining the number of replicas of each type of operation resource, in order to ensure that the determined number of replicas can provide the resource amount required for all operation resources, the largest number of replicas is selected as the target number of replicas. Since the first resource amount to be provided is fixed, based on the proportional relationship between the target number of replicas and the first resource amount, the target resource amount of the operation resource is determined, and the target resource amount is used as the resource amount of the corresponding operation resource required for the Pod instance to run the target application program. By collecting the historical operation conditions of each actual Pod instance running the target application program and the global resource utilization rate, the target number of replicas of the Pod instances required to run the target application program and the target resource amount provided by the Pod instances for each type of operation resource are calculated, realizing reasonable recommendations for both the resource amount and the number of replicas at the same time.

[0105] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, which will not be elaborated here.

[0106] Further, in a possible implementation manner, if the reference resource amount is not less than the first resource amount, the number of replicas of the operation resource is determined to be 1; from the number of replicas of each type of operation resource, the largest number of replicas is selected as the target number of replicas.

[0107] In another possible implementation manner, the historical operation data includes the peak value of the resource amount used by the operation resource within a preset time period; the peak value of the resource amount used by the operation resource within a preset time period and the historical number of replicas of the Pod instance of the target application program are determined; the product between the peak value and the historical number of replicas is obtained to obtain the total resource amount of the operation resource required for the target application program.

[0108] In still another possible implementation manner, the first ratio between the total resource amount and the target utilization rate is determined, and the first ratio is used as the first resource amount.

[0109] In another possible implementation, a second ratio between the first resource amount and the reference resource amount is determined, and the ceiling operation is performed on the second ratio to obtain the number of replicas.

[0110] In another possible implementation, a third ratio between the first resource amount and the target number of replicas is determined, and the third ratio is used as the target resource amount.

[0111] In another possible implementation, the running requirements of the target application are determined; the running requirements are used to characterize the requirements for the resource utilization rate of the Pod instances running the target application; The target utilization rate corresponding to the running requirements is determined according to the pre-stored mapping relationship; wherein, the target utilization rate is: The average value of the utilization rates of all Pod instances; The median of the utilization rates of all Pod instances, or The product of the peak utilization rate among the utilization rates of all Pod instances and a preset peak coefficient; In the mapping relationship, each type of running requirement corresponds to a target utilization rate.

[0112] In another possible implementation, the running resources include at least two of the following: CPU resources; Memory resources; GPU resources; NPU resources; Bandwidth resources.

[0113] An electronic device (computer device / equipment / system) is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to perform the steps of the container resource recommendation method. Compared with the related art, it can achieve: the container resource recommendation method.

[0114] In an optional embodiment, an electronic device is provided, as Figure 6 shown, Figure 6 The electronic device 4000 shown includes: a second processor 4001 and a second memory 4003. Among them, the second processor 4001 and the second memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0115] The second processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The second processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0116] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0117] The second memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0118] The second memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the second processor 4001 to execute. The second processor 4001 is used to execute the computer program stored in the second memory 4003 to implement the steps shown in the foregoing method embodiments.

[0119] Among them, the electronic device package may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0120] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding content shown in the foregoing method embodiments can be implemented. Compared with the prior art, it can be realized that by collecting the historical running conditions of each actual Pod instance running the target application program and the global resource utilization rate, the target number of replicas of the Pod instances required to run the target application program, and the target resource amount provided by the Pod instances for various running resources are calculated, and reasonable recommendations for both the resource amount and the number of replicas are realized.

[0121] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0122] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor can implement the steps and corresponding contents of the foregoing method embodiments. Compared with the prior art, it can be achieved that by collecting the historical running conditions of each actual Pod instance running the target application program and the global resource utilization rate, the target number of replicas of the Pod instances required to run the target application program, as well as the target resource amounts provided by the Pod instances for various types of running resources, are calculated, realizing reasonable recommendations for both the resource amount and the number of replicas simultaneously.

[0123] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims, and drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0124] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0125] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A container resource recommendation method, characterized in that: include: Determine at least one Pod instance running the target application, and obtain historical running data of at least one type of running resources of each of the at least one Pod instance; For each type of operating resource, determining a total resource amount of the operating resource required to run the target application according to historical operating data of the operating resource, and determining a first resource amount of the operating resource based on a predetermined target utilization rate of the operating resource and the total resource amount; The target utilization of the running resources is used to represent the distribution of the utilization of the running resources by all Pod instances; For each type of running resource, if it is determined that the reference resource amount is less than the first resource amount, the number of copies of the running resource is determined based on the proportional relationship between the reference resource amount and the first resource amount, and the reference resource amount is the maximum value of the resource amount allocated when the running resource is running on the node; the node is used to run the Pod instance; From the number of copies of various running resources, select the largest number of copies as the target number of copies; For each type of operating resources, the target resource amount of the operating resources is determined based on the proportional relationship between the target number of copies and the first resource amount, and the target resource amount is used as the resource amount of the corresponding operating resources required by the Pod instance to run the target application.

2. The method according to claim 1, characterized in that The method further comprises: If the reference resource amount is not less than the first resource amount, determining the number of copies of the running resource to be 1; From the number of replicas of various running resources, select the largest number of replicas as the target number of replicas.

3. The method according to claim 1, characterized in that The historical operation data includes a peak value of the amount of resources used by the operation resources within a preset time period; The determining the total amount of the running resources required to run the target application according to the historical running data of the running resources includes: Determine the peak value of the amount of resources used by the running resources within a preset time period and the number of historical copies of the Pod instance running the target application; The product of the peak value and the number of historical copies is determined to obtain the total resource amount of the running resources required to run the target application.

4. The method according to claim 1, characterized in that: The determining the first resource amount of the operating resource based on the predetermined target utilization rate of the operating resource and the total resource amount comprises: A first ratio between the total resource amount and the target utilization rate is determined, and the first ratio is used as the first resource amount.

5. The method according to claim 1, characterized in that The determining the number of copies of the running resource based on the proportional relationship between the reference resource amount and the first resource amount includes: A second ratio between the first resource amount and the reference resource amount is determined, and the second ratio is rounded up to obtain the number of replicas.

6. The method according to claim 1, characterized in that: The determining the target resource amount of the running resource based on the proportional relationship between the target number of replicas and the first resource amount includes: A third ratio between the first resource amount and the target number of replicas is determined, and the third ratio is used as the target resource amount.

7. The method according to claim 1, characterized in that The method for determining the target utilization rate includes: Determine the operation requirements of the target application; the operation requirements are used to characterize the requirements for resource utilization of the Pod instance running the target application; The target utilization rate corresponding to the operation demand is determined according to the pre-stored mapping relationship; wherein the target utilization rate is: The average utilization of all Pod instances; The median utilization of all Pod instances, or The product of the utilization peak of all Pod instances and the preset peak factor; In the mapping relationship, each type of operation demand corresponds to a target utilization rate.

8. The method according to claim 1, characterized in that: The operating resources include at least two of the following: CPU resources; Memory resources; GPU resources; NPU resources; Bandwidth resources.

9. A container resource recommendation device, characterized in that: include: A first determination module is used to determine at least one Pod instance running a target application, and obtain historical operation data of at least one type of operation resource of each of the at least one Pod instance; A second determination module is used to determine, for each type of operating resource, a total resource amount of the operating resource required to run the target application according to historical operating data of the operating resource, and determine a first resource amount of the operating resource based on a predetermined target utilization rate of the operating resource and the total resource amount, wherein the target utilization rate of the operating resource is used to represent the distribution of utilization rates of the operating resource by all Pod instances; A third determination module is used to determine, for each type of running resource, if it is determined that the reference resource amount is less than the first resource amount, the number of copies of the running resource based on the proportional relationship between the reference resource amount and the first resource amount, the reference resource amount being the maximum value of the resource amount allocated when the running resource is running on the node; the node is used to run the Pod instance; A selection module is used to select the largest number of copies as the target number of copies from the number of copies of various running resources; The fourth determination module is used to determine the target resource amount of the running resources for each type of running resources based on the proportional relationship between the target number of copies and the first resource amount, and use the target resource amount as the resource amount of the corresponding running resources required for the Pod instance to run the target application.

10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • System capacity estimation method and device

    CN114356577A

  • Intelligent analysis and evaluation method for surveying and mapping data

    CN117271683A

  • Resource configuration method and device, electronic equipment and storage medium

    CN118467101A

  • Container processing method and device and electronic equipment

    CN119292772A

  • Real time statistical computation in embedded systems

    US6785632B1