Resource management method and device, computer equipment and storage medium

By filtering and releasing resources of running services with low priority or low activity on the target server, the service startup failure caused by insufficient resources in the computer system is solved, and efficient resource utilization and normal service operation are achieved.

CN120066794APending Publication Date: 2025-05-30BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237865.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a computer system, when the server's running resources are insufficient, the relevant services cannot be started or run normally, resulting in waste of resources and unavailable services.

Method used

By obtaining the remaining resource information of the target server and the running information of the running services, filter out the running services with low priority or low activity as the target services, and release their occupied resources to support the operation of the services to be run.

Benefits of technology

It effectively solves the problem of service startup failure caused by insufficient resources, improves resource utilization, reduces resource waste, and ensures the normal operation of services to be run.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066794A_ABST
    Figure CN120066794A_ABST
Patent Text Reader

Abstract

The invention provides a resource management method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The resource management method comprises the steps of obtaining a to-be-operated service and a target server used for operating the to-be-operated service; the target server supports operation of at least one operated service; under the condition that the residual resources of the target server cannot meet the operation of the to-be-operated service, screening out a target service from the at least one operated service at least according to the operation information of each operated service in the target server; and releasing occupied resources of the target service, so that the target server supports operation of the to-be-operated service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of computer technology, and particularly relates to a resource management method, an apparatus thereof, a computer device, and a storage medium. Background Art

[0002] The operation of related services deployed in a computer system requires certain operating resources for support, such as scheduling Central Processing Unit (CPU) resources and memory resources, etc. However, during the process of starting related services, if the operating resources of the server are insufficient, the related service will remain in a pending start state until the resources meet the operating requirements and then it will operate normally; or if the operating resources of the server are insufficient, the operation will fail directly. Summary of the Invention

[0003] This disclosure aims to solve at least one of the technical problems existing in the prior art, and provides a resource management method, an apparatus thereof, a computer device, and a storage medium.

[0004] In a first aspect, the technical solution adopted to solve the technical problems of this disclosure is a resource management method, including:

[0005] Obtain a service to be run and a target server for running the service to be run; at least one already-running service can run on the target server;

[0006] In a case where the remaining resources of the target server cannot meet the running of the service to be run, screen out a target service from at least one of the already-running services at least according to the running information of each of the already-running services in the target server;

[0007] Release the occupied resources of the target service for the target server to support the running of the service to be run.

[0008] In some embodiments, screening out a target service from at least one of the already-running services at least according to the running information of each of the already-running services in the target server includes:

[0009] Screen out a target service from at least one of the already-running services according to the running information of each of the already-running services in the target server and the running information of the service to be run.

[0010] In some embodiments, the running information includes a priority;

[0011] The screening out a target service from at least one of the already-running services according to the running information of each of the already-running services in the target server and the running information of the service to be run includes:

[0012] Filter out a first service from at least one of the running services according to the priorities corresponding to each of the running services and the priority corresponding to the to-be-run service;

[0013] Determine the target service from the first service.

[0014] In some embodiments, the running information includes a priority;

[0015] The filtering out a target service from at least one of the running services according to the running information of each of the running services in the target server and the running information of the to-be-run service includes:

[0016] Filter out a first service from at least one of the running services according to the priorities corresponding to each of the running services and the priority corresponding to the to-be-run service;

[0017] Determine the target service from the filtered first service.

[0018] In some embodiments, the filtering out a first service from at least one of the running services according to the priorities corresponding to each of the running services and the priority corresponding to the to-be-run service includes:

[0019] Filter out the running services whose priorities are lower than the priority of the to-be-run service from the priorities comparison results of each of the running services and the to-be-run service in the target server, and denote them as the first service.

[0020] In some embodiments, the running information further includes activity; the determining the target service from the filtered first service includes:

[0021] When the activity of the first service continuously is lower than its own set value for a first duration reaching a preset shutdown time threshold, determine the first service as the target service.

[0022] In some embodiments, for any of the first services, the process of determining the activity includes:

[0023] Obtain the resource utilization rate of the resources occupied by the first service and the first preset period of the first service;

[0024] Determine the average resource utilization rate of the resources occupied by the first service within the first preset period according to the resource utilization rate of the resources occupied by the first service within the first preset period, and use it as the activity of the first service.

[0025] In some embodiments, the running information includes activity;

[0026] Filter out a target service from at least one of the running services according to at least the running information of each of the running services in the target server, including:

[0027] Judge whether a first duration during which the activity of the running service continuously is lower than its own set value reaches a preset shutdown time threshold, and determine the running service for which the first duration reaches the preset shutdown time threshold;

[0028] Filter out the target service from the running services for which the first duration reaches the preset shutdown time threshold.

[0029] In some embodiments, there are multiple target services;

[0030] The resource management method further includes:

[0031] For any of the target services, determine a second duration during which the first duration exceeds the preset shutdown time threshold;

[0032] Arrange the multiple target services in order from longest to shortest according to the second duration, and determine the arrangement order of the multiple target services.

[0033] In some embodiments, releasing the occupied resources of the target service for the target server to support the running of the to-be-run service includes:

[0034] Release the occupied resources of the target service in sequence according to the arrangement order of the multiple target services until the target server supports the running of the to-be-run service.

[0035] In some embodiments, for any of the first services, the process of determining the activity includes:

[0036] Obtain the resource utilization rate of the occupied resources of the first service and a first preset period of the first service;

[0037] Determine the average utilization rate of the occupied resources of the first service within the first preset period according to the resource utilization rate of the occupied resources of the first service within the first preset period, and use it as the activity of the first service.

[0038] In some embodiments, the resource management method further includes:

[0039] For any of the running services in the target server, obtain the activity of the running service and its own set value;

[0040] Within a second preset period, if the activity of the running service continuously remains lower than its own set value, downgrade the current priority of the running service by a preset adjustment amount.

[0041] Within a second preset period, if the activity of the running service is higher than or equal to its own set value, upgrade the current priority of the running service by a preset adjustment amount.

[0042] In some embodiments, before upgrading the current priority of the running service according to the second preset adjustment amount, it further includes:

[0043] If the current priority has reached the initial set value of the priority of the running service, stop the upgrade operation of the running service.

[0044] In some embodiments, the resource management method further includes:

[0045] For the running service that has been downgraded, increase the second preset period corresponding to the next downgrade.

[0046] In some embodiments, obtaining the target server for running the to-be-run service includes:

[0047] Determine the target server for running the to-be-run service according to the type of the to-be-run service.

[0048] In some embodiments, the running information includes occupied resources; there are multiple target servers of the same type as the to-be-run service;

[0049] The screening of the target service from at least one running service according to the running information of each running service in the target server at least includes:

[0050] Determine the remaining resources of each target server according to the total resources and occupied resources of each target server.

[0051] According to the remaining resources of each target server, the occupied resources of each running service in each target server, and the occupied resources required by the to-be-run service, screen out the to-be-transferred server and the to-be-received server from multiple target servers, and determine the target service in the to-be-transferred server;

[0052] The releasing of the occupied resources of the target service for the target server to support the running of the to-be-run service includes:

[0053] Transfer the target service in the server to be transferred to the server to be received for running, so as to release the occupied resources of the target service in the server to be transferred and support the running of the service to be run.

[0054] In some embodiments, the resource management method further includes:

[0055] When the remaining resources after the server to be transferred releases the occupied resources do not meet the running requirements of the service to be run, according to the priorities of the remaining running services corresponding to the target server and the priority of the service to be run, screen out the second service from the remaining running services;

[0056] Determine whether the second service is the target service according to the activity corresponding to the second service.

[0057] In some embodiments, the resource management method further includes:

[0058] Detect the occupied resources of each running service in each server in the cluster according to a third preset period;

[0059] According to the occupied resources of each running service in each server, screen out the third service to be transferred and the server to be received from multiple running services;

[0060] Transfer the third service to the server to be received to release the resources of the original server where the third service is located.

[0061] In some embodiments, the resource management method further includes:

[0062] Deploy the service to be run on the target server by using the application software Kubernetes.

[0063] In a second aspect, an embodiment of the present disclosure further provides a resource management device, including: a data acquisition module, a service screening module, and a resource release module;

[0064] The data acquisition module is configured to acquire a service to be run and a target server for running the service to be run; the target server currently supports the running of multiple running services;

[0065] The service screening module is configured to screen out a target service from multiple running services at least according to the running information of each running service in the target server when the remaining resources of the target server cannot meet the running requirements of the service to be run;

[0066] The resource release module is configured to release the occupied resources of the target service for the target server to support the operation of the to-be-run service.

[0067] In a third aspect, an embodiment of the present disclosure further provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the resource management method according to any one of the first aspects are executed.

[0068] In a fourth aspect, an embodiment of the present disclosure further provides a computer non-transitory readable storage medium. A computer program is stored on the computer non-transitory readable storage medium. When the computer program is run by a processor, the steps of the resource management method according to any one of the first aspects are executed. Description of the Drawings

[0069] Figure 1 It is a flowchart of the resource management method provided by the embodiment of the present disclosure.

[0070] Figure 2 It is a schematic diagram of relevant information of each running service in the target server provided by the embodiment of the present disclosure.

[0071] Figure 3 It is a specific flowchart of screening a target service provided by the embodiment of the present disclosure.

[0072] Figure 4 It is a specific flowchart of sorting multiple target services and releasing resources in sequence provided by the embodiment of the present disclosure.

[0073] Figure 5 It is a flowchart of the calculation process of activity provided by the embodiment of the present disclosure.

[0074] Figure 6 It is a flowchart of upgrading and downgrading a running service provided by the embodiment of the present disclosure.

[0075] Figure 7 It is a schematic diagram of downgrading and upgrading a service provided by the embodiment of the present disclosure.

[0076] Figure 8 It is another specific flowchart of screening a target service provided by the embodiment of the present disclosure.

[0077] Figure 9 It is a schematic diagram of a service transfer process provided by the embodiment of the present disclosure.

[0078] Figure 10Another specific flowchart for screening target services provided by the embodiments of the present disclosure.

[0079] Figure 11 Another schematic diagram of the service transfer process provided by the embodiments of the present disclosure.

[0080] Figure 12 Schematic diagram of the resource management device provided by the embodiments of the present disclosure.

[0081] Figure 13 Schematic diagram of the structure of a computer device provided by the embodiments of the present disclosure. Detailed implementation manners

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0083] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "including" or "comprising" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0084] As used in this disclosure, "multiple" or "several" means two or more. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0085] In the related art, there are some problems in computer systems: First, the operating resources of clusters or individual servers are not concentrated enough, and some services are run too dispersedly, resulting in the inability to start and run related services normally. Second, some secondary or ordinary services always occupy a certain amount of operating resources, but the operation of this service fails to generate a value that matches the resources it occupies, or play a practical application role that matches the resources it occupies. Third, the resource allocation control of the computer system is not precise enough, which is not conducive to resource scheduling.

[0086] Before formally introducing the resource management method provided by the embodiments of this disclosure, special terms and a specific application scenario involved in the resource management method provided by the embodiments of this disclosure are described to facilitate the understanding of the embodiments of this disclosure.

[0087] Explanation of special terms: 1. Kubernetes, abbreviated as K8s, is an open-source one used to manage service-oriented applications on multiple hosts in a cloud platform. The goal of Kubernetes is to make the deployment of service-oriented applications simple and efficient (powerful). Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.

[0088] 2. The function of LABEL is to define the type of variables or labels, and the segment attributes and offset attributes of variables or labels are determined by the position where this statement is located.

[0089] 3. As the operation and control core of a computer system, the CPU is the final execution unit for information processing and program running.

[0090] 4. Memory is an important component of a computer, also known as internal memory and main memory. It is used to temporarily store the operation data in the CPU and the data exchanged with external memories such as hard disks. It is a bridge for communication between the external memory and the CPU. All programs in the computer run in the memory, and the performance of the memory affects the overall performance of the computer. As long as the computer starts running, the operating system will transfer the data to be operated from the memory to the CPU for operation. When the operation is completed, the CPU will send out the result.

[0091] 5. A container group (pod) is a set of one or more services (or multiple Docker containers) with shared storage or network resources, and is the smallest deployable computing unit that can be created and managed in a platform environment. In the present disclosure, the container group (pod) technology is a lightweight and portable software packaging and delivery method. It allows applications and their dependencies to be packaged into one or more container groups (pods), which can run in any environment as long as the corresponding container group (pod) runtime environment is installed. This enables developers to maintain consistent application behavior in different environments (such as development, testing, and production).

[0092] When the entire service system to be managed is relatively large, K8s is a very good choice. In addition to having the functions of container technology, it also has the functions of automatically deploying, scaling, and managing containerized applications, and secondary development is both mature and convenient. A resource management method provided by an embodiment of the present disclosure may be a method for dynamically managing resources based on K8s. The execution subject of this resource management method may be a computer device with a certain computing ability.

[0093] It should be noted that the "running service" mentioned in the present disclosure, such as the service container group (pod) to be run obtained and the running service supported on the target server, can be understood as the smallest deployable computing unit created and managed in the platform environment, that is, the container group (pod). Of course, if the resource management method is no longer implemented based on K8s, which is an application with container technology, then correspondingly the "running service" can also be the "smallest deployable computing unit" involved in other applications. For the sake of convenience of understanding, the embodiment of the present disclosure takes the "running service" as the "container group (pod)" as an example for description, and other application scenarios are not listed one by one.

[0094] Figure 1 The flowchart of the resource management method provided by the embodiment of the present disclosure is as Figure 1 shown, including steps S11 to S13.

[0095] S11. Obtain the service to be run and the target server for running the service to be run.

[0096] In this step, the service to be run refers to the service container group (pod) to be started or the service container group (pod) to be deployed.

[0097] Running the to-be-run service requires the support of a server. That is to say, a target server needs to be found to run the to-be-run service. Specifically, according to the type (label) of the to-be-run service, a server that matches the type can be screened out from the cluster as the target server. It should be noted that the server only supports the running of relevant services that match the type. Multiple servers are deployed in the cluster, and multiple container groups (pods) are deployed in the servers. Each server is configured with a server type to support the running of the same type of container groups (pods). Among them, the target server supports the running of at least one already-running service.

[0098] S12. When the remaining resources of the target server cannot meet the running requirements of the to-be-run service, at least based on the running information of each already-running service in the target server, a target service is screened out from at least one already-running service.

[0099] Specifically, based on the total resources of the target server and the resources occupied by each already-running service, the remaining resources of the target server can be determined. According to the remaining resources of the target server and the running resources required by the to-be-run service, it is judged whether the current target server meets the running requirements of the to-be-run service. If not, at least based on the running information of each already-running service in the target server, a target service is screened out from at least one already-running service.

[0100] Given the already-running services running on the target server, the resources occupied by each already-running service are known, and the total resources of the target server are known. Therefore, based on the total resources of the target server and the resources occupied by each already-running service, the remaining resources of the target server can be determined. Exemplarily, three already-running services are running on the target server, denoted as pod_1, pod_2, and pod_3 respectively. Among them, the resources occupied by pod_1 include: 4 cores of CPU and 512Mi of memory; the resources occupied by pod_2 include: 3 cores of CPU and 512Mi of memory; the resources occupied by pod_3 include: 2 cores of CPU and 512Mi of memory. The total resources of the target server include: 12 cores of CPU and 2048Mi of memory. Finally, it is determined that the remaining resources of the target server include: 3 cores of CPU and 512Mi of memory. Given that the running resources required by the to-be-run service include 5 cores of CPU and 512Mi of memory. Therefore, it can be determined that the remaining resources of the target server, 3 cores of CPU and 512Mi of memory, cannot meet the running requirements of the to-be-run service.

[0101] The running information of the service may include occupied resources and / or running parameters, where the occupied resources refer to the running resources required for the service to run on the target server. The running parameters include priority and / or activity; among them, the priority is used to characterize the importance of the service, and the numerical representation is 0 to 10. The smaller the number, the higher the priority and the higher the importance. The activity is used to characterize the active situation of the running service in the system, such as the traffic usage or the CPU usage, etc. The priority is a running parameter pre-configured by the user for the relevant services (including the to-be-run service and the running service), and the activity is a running parameter calculated for the running service within a certain period of time.

[0102] In the case where the remaining resources of the target server cannot meet the running requirements of the to-be-run service, at least the target service can be screened out from at least one running service according to the occupied resources and / or running parameters of each running service in the target server.

[0103] S13. Release the occupied resources of the target service for the target server to support the running of the to-be-run service.

[0104] Based on step S12, it can be known that the target service may be a service with a lower importance than the to-be-run service. Therefore, the target service can be appropriately shut down to release the occupied resources of the target service on the target server for the target server to support the running of the to-be-run service. Or, the target service can be transferred to other servers for running to release its occupied resources for the target server to support the running of the to-be-run service.

[0105] The embodiments of the present disclosure realize the operation management operation of the service by means of the running information of the service, can flexibly screen out the target service from the running services of the target server, and release the occupied resources of the target service in the target server, which can effectively reduce the problem of resource waste. Especially in the case of resource tension, the resources can be reasonably allocated to the greatest extent, and the maximum value can be generated with the least resource waste.

[0106] In some embodiments, the present disclosure can realize the deployment, management, and running of relevant services (or container groups pod) by means of the K8s technology. Specifically, after the target server releases the occupied resources of the target service, the to-be-run service can be deployed on the target server by using the application software (K8s) to support the running of the to-be-run service by the target server. Hereinafter, the present disclosure will be described by taking the automatic deployment and management of the container group (pod) by the application software (K8s) as an example.

[0107] In some embodiments, management parameters are created in advance for each service (including running services and services to be run), including priority, activity level, self-set value corresponding to the activity level, first preset period, second preset period, and preset shutdown time threshold. The self-set value corresponding to the activity level refers to the set value that is preset for the activity level and conforms to the current service activity critical situation. The first preset period is mainly used to calculate the activity level. The second preset period is mainly used for raising or lowering the priority corresponding to the service. The preset shutdown time threshold is mainly used to define whether the service can be shut down.

[0108] As Figure 2 shown, it is a schematic diagram of the relevant information of each running service in the target server. Among them, the target server 01 includes multiple running services, denoted as pod_1, pod_2, pod_3, pod_4, pod_5, and pod_6 respectively. The relevant information of each running service includes the priority in the management parameters, the self-set value corresponding to the activity level, the first preset period, the second preset period, and the preset shutdown time threshold, as well as the CPU resource occupancy and memory resource occupancy in the occupied resources. Of course, the occupied resources can also include parameters such as GPU, TPU, and NPU. However, the control logic of parameters such as GPU, TPU, and NPU is the same as that of CPU and memory parameters. Therefore, this disclosure takes CPU and memory parameters as examples for illustration, and the same applies to other resource parameters. This disclosure will not elaborate further in the embodiments.

[0109] Among them, 02 represents the default management parameters globally configured in the target server 01. That is, if the parameter priority, the set value corresponding to the activity level, the first preset period, and the second preset period are not set in a single service (container group pod), the single service (container group pod) shall be subject to the default management parameters of the target server. For example, if pod_3 does not set the second preset period and the preset shutdown time threshold, then it shall be subject to the default management parameters of the target server, with the second preset period being 5 days and the preset shutdown time threshold being 30 days.

[0110] In some embodiments, for S12, the target service can be screened out from at least one running service according to the running information of each running service and the running information of the service to be run in the target server. Specifically, taking the running information including priority as an example, it specifically includes S12-1-1 to S12-1-2, as Figure 3 shown.

[0111] S12-1-1. Screen out the first service from at least one running service according to the priority corresponding to each running service and the priority corresponding to the service to be run.

[0112] As Figure 2As shown, given the priorities corresponding to each running service in the target server, based on the comparison results of the priorities of each running service and the to-be-run service in the target server, the running services with priorities lower than that of the to-be-run service are filtered out and denoted as the first services. Exemplarily, given that the priority of the to-be-run service is 5, it should be noted that the higher the number corresponding to the priority, the lower the priority. Therefore, the running services in the target server with priorities lower than 5 are Pod_4 and Pod_6 respectively. Thus, Pod_4 and Pod_6 can be denoted as the first services.

[0113] S12-1-2. Determine the target service from the filtered first services.

[0114] In a possible implementation manner, the running information further includes activity. When the activity of the first service continuously remains lower than its own set value for a first duration and reaches a preset shutdown time threshold, the first service is determined as the target service. The activity can be the utilization rate of specified resources (such as resource parameters like flux, CPU, GPU, TPU, NPU, or memory, etc.) within a first preset period. Here, the first duration can be understood as the continuous time during which the activity of the first service has been lower than its own set value.

[0115] For example, if the first duration (such as 40 days) during which the activity of Pod_4 continuously remains lower than its own set value of 0.4 exceeds the preset shutdown time threshold of 30 days, then Pod_4 is determined as the target service. Another example, if the first duration (such as 5 days) during which the activity of Pod_6 continuously remains lower than its own set value of 0.6 does not reach the preset shutdown time threshold of 30 days, then it is determined that Pod_6 does not belong to the target service. On the contrary, if the first duration (such as 30 days) during which the activity of Pod_6 continuously remains lower than its own set value of 0.6 reaches the preset shutdown time threshold of 30 days, then Pod_6 is determined as the target service.

[0116] In another possible implementation manner, if there are multiple filtered first services, then further compare the priorities of each first service, and take the first service with the lowest priority as the target service. For example, the priority of Pod_6 is 9, which is lower than the priority of Pod_4 which is 7. Therefore, Pod_6 can be determined as the target service.

[0117] In another possible implementation manner, if there are multiple filtered first services, all the first services are taken as the target services. For example, both the filtered Pod_4 and Pod_6 are taken as the target services.

[0118] In another possible implementation manner, the target service can be determined manually from the filtered first services according to actual application requirements. Manual filtering can set filtering rules according to the current application scenario and experience, which is not limited in this disclosure.

[0119] In some embodiments, for S12, the target service can be directly screened out from at least one running service according to the running information of each running service in the target server. Specifically, taking the running information including activity as an example; it can be determined whether the first duration during which the activity of the running service continuously remains lower than its own set value reaches a preset shutdown time threshold, and the running service whose first duration reaches the preset shutdown time threshold can be determined. Here, the preset shutdown time threshold can be set to a relatively high number of days, such as 100 days, which means that the running service whose first duration reaches the preset shutdown time threshold is most likely a service with extremely low activity, and at this time, it is allowed to shut it down. Therefore, the target service can be determined from the running services whose first duration reaches the preset shutdown time threshold.

[0120] Exemplarily, the first duration (such as 110 days) during which the activity of Pod_4 continuously remains lower than its own set value of 0.4 exceeds the preset shutdown time threshold of 100 days, then it can be directly determined that this Pod_4 is the target service.

[0121] In some embodiments, for the above-mentioned step S12-1-2, when the running information includes activity, if there are multiple target services, such as Pod_4 and Pod_6, it is necessary to further judge their activity conditions, specifically including S21 to S23, as Figure 4 shown.

[0122] S21. For any target service, determine the second duration during which the first duration exceeds the preset shutdown time threshold.

[0123] The second duration = the first duration - the preset shutdown time threshold. For example, the first duration corresponding to Pod_4 is 40 days, and the preset shutdown time threshold is 30 days, then the second duration corresponding to Pod_4 is 10 days; the first duration corresponding to Pod_6 is 60 days, and the preset shutdown time threshold is 30 days, then the second duration corresponding to Pod_6 is 30 days. It can be seen that Pod_6 has lower activity than Pod_4, that is, the longer the second duration, the lower the activity of the corresponding target service.

[0124] S22. Arrange the multiple target services in the order from the longest to the shortest second duration, and determine the arrangement order of the multiple target services.

[0125] Continuing the example in S21, since Pod_6 has lower activity than Pod_4, the arrangement order of the multiple target services is Pod_6 first and Pod_4 second. If the target service is to be shut down, the resources can be released in this arrangement order, that is, the resources of the target service are released in the order from the longest to the shortest second duration.

[0126] Further, continuing from S22, continue asFigure 4 As shown, release the occupied resources of the target service, specifically including S23.

[0127] S23. Release the occupied resources of the target service in the arranged order of multiple target services until the target server supports the operation of the service to be run.

[0128] Continuing with the example of S22, "release in sequence" means releasing in the order of the second durations corresponding to multiple target services from long to short. First, close the target service Pod_6 with relatively lower activity. If the running resources released by closing Pod_6 are not enough for the service to be run, then further close the target service Pod_4 with relatively lower activity. If the running resources released by closing Pod_6 are enough for the service to be run, then do not traverse Pod_4 backward, that is, ensure the continuous operation of Pod_4. Therefore, in this embodiment, strict control is performed on the closing operation of the service according to the screening conditions of the multiple activities of the running service, and the reliability of the running service (such as Pod_4) is mainly protected, and dynamic scaling management is performed on each service.

[0129] In some embodiments, the specific process of calculating the activity includes S121 to S122, as Figure 5 shown.

[0130] S121. Obtain the resource utilization rate of the occupied resources of the first service and the first preset period of the first service.

[0131] The system knows the resource occupancy of the first service at any time, such as the resource utilization rate of the occupied resources, which can specifically be the flux utilization rate, CPU utilization rate, GPU utilization rate, TPU utilization rate, NPU utilization rate, or memory utilization rate, etc. at each moment. The first preset period is a preset resource parameter, as Figure 2 shown. The first preset period of pod_4 is 30s, and the first preset period of pod_6 is 45s.

[0132] S122. Determine the average utilization rate of the occupied resources of the first service in the first preset period according to the resource utilization rate of the occupied resources of the first service in the first preset period, and use it as the activity of the first service.

[0133] This step defines the activity as the average utilization rate of specified resources within the first preset period. Here, the "specified resources" can be, for example, flux resources, CPU resources, GPU resources, TPU resources, NPU resources, or memory resources, etc. Specifically, it can be flexibly set according to the actual application situation, and the embodiments of the present disclosure do not make specific limitations. Calculate the average utilization rate of the specified resources within 30s for pod_4 as the activity of pod_4. Calculate the average utilization rate of the specified resources within 45s for pod_6 as the activity of pod_4.

[0134] For the above-mentioned various embodiments and their combinations, it can be seen that on the basis of K8s technology, the present disclosure adds precise control parameters such as priority, activity, preset shutdown time threshold, etc., which is more conducive to service management and avoids waste of cluster resources. Among them, the success rate of service startup can be improved through priority. Not only based on the conditions set by k8s itself, but also when resources are insufficient, the shutdown of unimportant services can be achieved according to the priority, activity, and preset shutdown time threshold of the currently running services to meet the resources required by the services to be run during operation. At the same time, the present disclosure can also dynamically adjust the management parameters in the cluster according to time, place, and people, set flexible management methods for various complex scenarios, and have generality while expanding personalized adjustment functions.

[0135] In some embodiments, the resource management method provided by the present disclosure further includes monitoring the running services, and the level of the priority of the running services can be flexibly adjusted according to the current state of the running services to achieve the rise and fall of the priority, so as to ensure that the operation of the cluster is more conducive to production needs. Specifically, as Figure 6 shown, the resource management method further includes S31 to S33.

[0136] S31: For any running service in the target server, obtain the activity and its own set value of the running service.

[0137] Regarding the calculation of the activity of the running service in this step, the calculation process of the activity in S121 to S122 above can be referred to, and the repeated parts will not be elaborated. As Figure 2 shown, for any running service, the set value corresponding to the activity can be directly obtained.

[0138] S32: Within the second preset period, if the activity of the running service continuously remains lower than its own set value, downgrade the current priority of the running service according to the preset adjustment amount.

[0139] As Figure 2 shown, the second preset period is a management parameter set in advance. The preset adjustment amount can be flexibly set according to the actual application scenario. In the present disclosure, the preset adjustment amount is taken as 1 level for illustration.

[0140] As Figure 7 shown, taking Pod_1 as an example, within the second preset period of 5 days, if the activity of Pod_1 remains lower than its own set value of 0.7, the current priority 5 of Pod_1 can be downgraded according to the preset adjustment amount of level 1. The numerical representation is an increase of 1 digit, that is, the numerical representation of the priority of Pod_1 is adjusted from 5 to 6, which means a downgrade of 1 level.

[0141] S33. Within the second preset period, if the activity of the running service is higher than or equal to its own set value, upgrade the current priority of the running service according to the preset adjustment amount.

[0142] Similarly, continuing as Figure 7 shown, taking Pod_1 as an example, within the second preset period of 5 days, if the activity of Pod_1 is at least once higher than or equal to its own set value of 0.7, upgrade the current priority 6 of Pod_1 according to the preset adjustment amount of level 1. The numerical representation is a decrease of 1 digit, that is, the numerical representation of the priority of Pod_1 is adjusted from 6 to 5, which means an upgrade of 1 level.

[0143] It should be noted that closing the currently running service is a relatively dangerous operation. According to the principle of service operation priority, within the second preset period, as long as the activity is higher than its own set value once, it means that the currently running service is not a service that has not been used for a long time. Therefore, the upgrade condition can be met and the upgrade can be allowed. However, if the current priority has reached the initial set value of the priority of the running service, the upgrade operation of the running service can be stopped. Exemplarily, continuing the example of S33, within the second preset period of 5 days, if the activity of Pod_1 is at least once higher than or equal to its own set value of 0.7, and the current priority of Pod_1 has been upgraded to 5, then the upgrade operation of the running service can be stopped at this time, and the current priority of Pod_1 is kept unchanged at 5. That is to say, the upgrade operation is not an endless operation, and the upgrade can only reach the initial set value set at the time of service creation at most.

[0144] In some embodiments, for the running service that has been downgraded, the second preset period corresponding to the next downgrade can be increased. Exemplarily, continuing the example of S32, within the second preset period of 5 days, the priority of Pod_1 is downgraded from 5 to 6. To ensure the reliability of the running service and prevent rapid consecutive downgrades, the second preset period corresponding to the next downgrade can be increased, for example, it can be increased in integer multiples (2 times, 3 times, 4 times, etc.), such as from 5 days to 10 days. Thus, when the priority downgrade operation is performed next time, the second preset period in S32 is updated to the latest increased second preset period.

[0145] In this embodiment, the degradation period of the already degraded running service is increased, so as to ensure the running reliability of the running service.

[0146] The embodiment of the present disclosure adds a function for monitoring the running process of services on the basis of K8s, which can optimize the running mode of the entire cluster and make the running of the cluster more conducive to production needs. On the basis of the original K8s, the running status of the entire cluster services can be monitored more carefully, and specific metrics can be specified for monitoring (such as traffic, memory, CPU, GPU, TPU, NPU, etc.). Different metrics can be monitored for different services to perform service upgrade and downgrade operations, providing favorable operation guarantees for the deployment and running of subsequent services to be run.

[0147] In some embodiments, according to the type of the service to be run, servers with matching types are screened out from the cluster. In one case, there are multiple target servers of the same type as the service to be run. Taking the running information including occupied resources as an example, for S12, target services are screened out from at least one running service, specifically including S12-2-1 to S12-2-2, as Figure 8 shown.

[0148] S12-2-1. Determine the remaining resources of each target server according to the total resources and occupied resources of each target server.

[0149] Remaining resources = total resources - occupied resources. Among them, the occupied resources refer to the sum of the occupied resources of each running service in the target server. As Figure 9As shown in the figure, the multiple target servers are Node-1, Node-2 and Node-3. The total resources of Node-1 include: CPU has 12 cores and memory has 2048Mi. The occupied resources of Node-1 include the resources occupied by pod_1 (CPU has 4 cores and memory has 512Mi), the resources occupied by pod_2 (CPU has 3 cores and memory has 512Mi), and the resources occupied by pod_3 (CPU has 2 cores and memory has 512Mi). The remaining resources of Node-1 include: CPU has 3 cores and memory has 512Mi. The total resources of Node-2 include: CPU has 12 cores and memory has 2048Mi. The occupied resources of Node-2 include the resources occupied by pod_4 (CPU has 2 cores and memory has 512Mi) and the resources occupied by pod_7 (CPU has 5 cores and memory has 1024Mi). The remaining resources of Node-2 include: CPU has 5 cores and memory has 512Mi. The total resources of Node-3 include: CPU has 12 cores and memory has 2048Mi. The resources occupied by Node-3 include the resources occupied by pod_5 (CPU has 5 cores and memory is 1024Mi), and the resources occupied by pod_6 (CPU has 5 cores and memory is 1024Mi). The remaining resources of Node-3 include: CPU has 2 cores and memory is 0Mi.

[0150] S12-2-2. Based on the remaining resources of each target server, the occupied resources of each running service in each target server, and the occupied resources required for the services to be run, select the servers to be transferred and the servers to be received from multiple target servers, and determine the target services in the servers to be transferred.

[0151] Given the remaining resources of each target server, the occupied resources of each running service (pod_1 to pod_7) in each target server, and the occupied resources required by the services to be run (for example, 3 CPU cores and 1024 Mi of memory), K8s can be used to manage and schedule the running services in the target server, and centralize the services of the same type but running in a distributed manner in the cluster. This can release the remaining resources of a single target server to a greater extent, which is conducive to supporting the operation of the services to be run. Figure 9As shown, it can be seen that the remaining resources of the entire cluster can meet the running requirements of the service to be run. Therefore, the service transfer operation can be executed. Specifically, based on the remaining resources of each target server, the occupied resources of each running service in each target server, and the occupied resources required by the service to be run, Node-2 is selected from Node-1 to Node-3 as the server to be transferred, Node-1 is selected as the server to be received, and pod_4 is determined as the target service in the server to be transferred, Node-2. After that, the target service in the server to be transferred can be transferred to the server to be received for running, so as to release the occupied resources of the target service in the server to be transferred and support the running of the service to be run. As Figure 9 shown, at this time, the remaining resources of Node-1 include: 1 core of CPU and 0 Mi of memory; the remaining resources of Node-2 include: 7 cores of CPU and 1024 Mi of memory; the remaining resources of Node-3 include: 2 cores of CPU and 0 Mi of memory. Here, Node-2 can be used to support the running of the service to be run. Therefore, Node-2 can be used to deploy and run the service to be run.

[0152] Or, Node-1 is selected from Node-1 to Node-3 as the server to be transferred, Node-2 is selected as the server to be received, and pod_3 is determined as the target service in the server to be transferred, Node-1. After that, the target service in the server to be transferred can be transferred to the server to be received for running, so as to release the occupied resources of the target service in the server to be transferred and support the running of the service to be run. At this time, the remaining resources of Node-1 include: 5 cores of CPU and 1024 Mi of memory; the remaining resources of Node-2 include: 3 cores of CPU and 0 Mi of memory; the remaining resources of Node-3 include: 2 cores of CPU and 0 Mi of memory. Here, Node-1 can be used to support the running of the service to be run. Therefore, Node-1 can be used to deploy and run the service to be run.

[0153] In this embodiment, based on the running status and resource usage of the entire cluster, on the basis of k8s service management, the dynamic transfer of services is realized, the usage of the resources of the entire cluster is evenly distributed, and the uneven distribution of the work of the servers in the cluster is avoided.

[0154] In some embodiments, if the remaining resources of the server to be transferred after releasing the occupied resources do not meet the running requirements of the service to be run, the operation of shutting down the service can be continued. As Figure 10 shown, it specifically includes S41 to S42.

[0155] S41. When the remaining resources after the occupied resources are released by the server to be transferred do not meet the running requirements of the service to be run, according to the priorities corresponding to the remaining running services in the target server and the priority corresponding to the service to be run, a second service is screened out from the remaining running services.

[0156] Here, the screening process of the second service is the same as the principle of the screening process of the first service in S12-1-1, and the repeated part will not be elaborated.

[0157] S42. Determine a target service from the screened second services.

[0158] Here, the process of determining whether the second service is the target service is the same as the principle of the process of determining whether the first service is the target service in S12-1-2.

[0159] In a possible implementation manner, when the third duration during which the activity of the second service continues to be lower than its own set value reaches a preset shutdown time threshold, the second service is determined as the target service. Here, the second duration can be understood as the duration during which the activity of the second service has been lower than its own set value.

[0160] In another possible implementation manner, if multiple second services are screened out, the priorities of each second service are further compared, and the second service with the lowest priority is used as the target service.

[0161] In another possible implementation manner, if multiple second services are screened out, all the first services are used as the target services.

[0162] In another possible implementation manner, the target service can be determined manually from the screened second services according to the actual application requirements. Manual screening can set screening rules according to the application scenario and experience at that time, and the present disclosure does not limit it.

[0163] Continuing from S42, if the second service is determined to be the target service, the occupied resources of the target service can be further released according to S21 to S23 above for the server to be transferred to support the running of the service to be run, and the repeated part will not be elaborated.

[0164] The embodiments of the present disclosure first perform a service transfer operation. When the service transfer operation cannot meet the running requirements of the service to be run, the service shutdown operation is continued, and finally the normal running of the service to be run is ensured.

[0165] In some embodiments, when the remaining resources after the occupied resources are released by the server to be transferred do not meet the running requirements of the service to be run, the target service can be determined according to the activity corresponding to the remaining running services in the server to be transferred.

[0166] Specifically, it can be determined whether the fourth duration during which the activity of the remaining running services in the server to be transferred continues to be lower than its own set value reaches a preset shutdown time threshold, and the running services whose fourth duration reaches the preset shutdown time threshold are determined. Here, the preset shutdown time threshold can be set to a relatively high number of days, which means that the running services whose fourth duration reaches the preset shutdown time threshold are likely to be services with extremely low activity. At this time, they are allowed to be directly shut down. Therefore, the target services can be determined from the running services whose fourth duration reaches the preset shutdown time threshold.

[0167] In some embodiments, in addition to performing relevant transfer operations when there is a demand for services to be run as described above, service transfer can also be performed within a fixed period (i.e., periodic transfer). Specifically, as Figure 11 shown, it specifically includes S51 to S53.

[0168] S51. Detect the occupied resources of each running service in each server in the cluster according to the third preset period.

[0169] S52. According to the occupied resources of each running service in each server, screen out the third service to be transferred and the server to receive from multiple running services.

[0170] S53. Transfer the third service to the server to receive, so as to release the resources of the original server where the third service is located.

[0171] Among them, the third preset period is a preset fixed period. Every time a third preset period passes, the service transfer operation is started to run the scattered services centrally, and as many resources as possible are left to provide guarantee for the operation of new services, avoiding unnecessary resource waste.

[0172] The screening conditions here can be to release the occupied resources of a single server (the server to be transferred) as much as possible, and to occupy the resources of a single server (the server to receive) as much as possible. In this way, the services (pods) in the server to receive can be concentrated, thereby releasing the resources of the server to be transferred and providing operation guarantee for the subsequent deployment of new services.

[0173] The embodiments of the present disclosure are secondary development based on the original K8s to achieve dynamic resource management in a more accurate and flexible manner, make full use of existing resources, and maximize the utilization rate, thereby solving the technical problems that the existing service management method is too rigid and prone to resource waste.

[0174] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0175] In addition, an apparatus for resource management corresponding to the resource management method is provided in the embodiments of the present disclosure. Since the principle of problem-solving of the apparatus in the embodiments of the present disclosure is similar to that of the above-mentioned resource management method in the embodiments of the present disclosure, the implementation of the resource management apparatus can refer to the implementation of the method, and the repeated parts will not be described again.

[0176] Figure 12 is a schematic diagram of the resource management apparatus provided in the embodiments of the present disclosure, as Figure 12 shown, the resource management apparatus includes a data acquisition module 61, a service screening module 62, and a resource release module 63.

[0177] The data acquisition module 61 is configured to acquire a service to be run and a target server for running the service to be run; the target server currently supports the running of multiple running services.

[0178] It should be noted that the data acquisition module 61 in the embodiments of the present disclosure is configured to execute step S11 in the above-mentioned resource management method, and the repeated parts will not be described again.

[0179] The service screening module 62 is configured to screen out a target service from multiple running services at least according to the running information of each running service in the target server when the remaining resources of the target server cannot meet the running of the service to be run;

[0180] It should be noted that the service screening module 62 in the embodiments of the present disclosure is configured to execute step S12 in the above-mentioned resource management method, and the repeated parts will not be described again.

[0181] The resource release module 63 is configured to release the occupied resources of the target service for the target server to support the running of the service to be run.

[0182] It should be noted that the resource release module 63 in the embodiments of the present disclosure is configured to execute step S13 in the above-mentioned resource management method, and the repeated parts will not be described again.

[0183] In some embodiments, the service screening module 62 is specifically configured to screen out a target service from at least one running service according to the running information of each running service in the target server and the running information of the service to be run.

[0184] In some embodiments, the running information includes a priority; the service screening module 62 is specifically configured to screen out a first service from at least one running service according to the priorities corresponding to each running service and the priority corresponding to the service to be run; and determine the target service from the screened first service.

[0185] In some embodiments, the service screening module 62 screens out a first service from at least one running service according to the priorities corresponding to the running services and the priority corresponding to the service to be run, specifically including: screening out the running services whose priorities are lower than the priority of the service to be run according to the comparison results of the priorities of the running services and the service to be run in the target server, and denoting them as the first services.

[0186] In some embodiments, the running information further includes activity; the service screening module 62 determines a target service from the screened first services, specifically including: when the first duration during which the activity of the first service continuously remains lower than its own set value reaches a preset shutdown time threshold, determining the first service as the target service.

[0187] In some embodiments, the service screening module 62 determines the activity specifically including: obtaining the resource utilization rate of the resources occupied by the first service and the first preset period of the first service; determining the average utilization rate of the resources occupied by the first service within the first preset period according to the resource utilization rate of the resources occupied by the first service within the first preset period, and using it as the activity of the first service.

[0188] In some embodiments, the running information includes activity; the service screening module 62 is specifically configured to determine whether the first duration during which the activity of the running service continuously remains lower than its own set value reaches a preset shutdown time threshold, and determine the running services whose first duration reaches the preset shutdown time threshold; screening out the target service from the running services whose first duration reaches the preset shutdown time threshold.

[0189] In some embodiments, there are multiple target services; the service screening module 62 is further configured to, for any target service, determine a second duration during which the first duration exceeds the preset shutdown time threshold; arranging the multiple target services in the order from long to short according to the second duration, and determining the arrangement order of the multiple target services.

[0190] In some embodiments, the resource release module 63 is further configured to sequentially release the occupied resources of the target services according to the arrangement order of the multiple target services until the target server supports the operation of the service to be run.

[0191] In some embodiments, the resource management device further includes a promotion and demotion module 64; the promotion and demotion module 64 is configured to, for any running service in the target server, obtain the activity of the running service and its own set value; within the second preset period, if the activity of the running service continuously remains lower than its own set value, demote the current priority of the running service according to a preset adjustment amount; within the second preset period, if the activity of the running service is higher than or equal to its own set value, promote the current priority of the running service according to a preset adjustment amount.

[0192] In some embodiments, the upgrade / downgrade module 64 is further configured to stop the upgrade operation of the running service if the current priority has reached the initial set value of the priority of the running service before upgrading the current priority of the running service by the second preset adjustment amount.

[0193] In some embodiments, the upgrade / downgrade module 64 is further configured to increase the second preset period corresponding to the next downgrade for the running service that has been downgraded.

[0194] In some embodiments, the data acquisition module 61 acquires the target server for running the to-be-run service, specifically including: determining the target server for running the to-be-run service according to the type of the to-be-run service.

[0195] In some embodiments, the running information includes occupied resources; there are multiple target servers of the same type as the to-be-run service; the service screening module 62 is specifically configured to determine the remaining resources of each target server according to the total resources and occupied resources of each target server; according to the remaining resources of each target server, the occupied resources of each running service in each target server, and the occupied resources required by the to-be-run service, screen out the to-be-transferred server and the to-be-received server from multiple target servers, and determine the target service in the to-be-transferred server; the resource release module 63 is specifically configured to transfer the target service in the to-be-transferred server to the to-be-received server for running, so as to release the occupied resources of the target service in the to-be-transferred server to support the running of the to-be-run service.

[0196] In some embodiments, the service screening module 62 is further configured to, when the remaining resources after the to-be-transferred server releases the occupied resources do not meet the running requirements of the to-be-run service, screen out the second service from the remaining running services according to the priorities corresponding to the remaining running services in the to-be-transferred server and the priority corresponding to the to-be-run service; determine whether the second service is the target service according to the activity corresponding to the second service.

[0197] In some embodiments, the service screening module 62 is further configured to, when the remaining resources after the to-be-transferred server releases the occupied resources do not meet the running requirements of the to-be-run service, determine the target service according to the activity corresponding to the remaining running services in the to-be-transferred server.

[0198] In some embodiments, the resource management device further includes a service transfer module 65; the service transfer module 65 is configured to detect the occupied resources of each running service in each server in the cluster according to the third preset period; screen out the third service to be transferred and the server to be received from multiple running services according to the occupied resources of each running service in each server; transfer the third service to the server to be received to release the resources of the original server where the third service is located.

[0199] In some embodiments, the resource management device further includes a service deployment module 66; the service deployment module 66 is configured to use the application software Kubernetes to deploy the service to be run on the target server.

[0200] Figure 13 Schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure. As Figure 13 shown, an embodiment of the present disclosure provides a computer device including: one or more processors 701, a memory 702, and one or more I / O interfaces 703. One or more programs are stored on the memory 702. When the one or more programs are executed by the one or more processors, the one or more processors implement the resource management method as described in any of the above embodiments; one or more I / O interfaces 703 are connected between the processor and the memory and are configured to implement information interaction between the processor and the memory.

[0201] Among them, the processor 701 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 702 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) 703 is connected between the processor 701 and the memory 702 and can implement information interaction between the processor 701 and the memory 702, including but not limited to a data bus (Bus), etc.

[0202] In some embodiments, the processor 701, the memory 702, and the I / O interface 703 are interconnected through a bus 704 and are further connected to other components of the computing device.

[0203] According to an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium. A computer program is stored on the non-transitory computer-readable storage medium. When the program is executed by a processor, the steps in the resource management method as described in any of the above embodiments are implemented.

[0204] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a machine-readable medium. The computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part and / or installed from a removable medium. When the computer program is executed by a central processing unit CPU, the above functions defined in the system of the present disclosure are executed.

[0205] It should be noted that the computer non-transitory readable medium shown in this disclosure can be a computer readable signal medium, a computer readable storage medium, or any combination of the two. The computer readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in this disclosure, the computer readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium can also be any computer non-transitory readable storage medium other than the computer readable storage medium, and this computer non-transitory readable storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer non-transitory readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the foregoing module, segment of a program, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually represent an execution in substantially parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0207] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present disclosure, and the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also considered within the scope of protection of the present disclosure.

Claims

1. A resource management method, wherein: include: Acquire a service to be run and a target server for running the service to be run; The target server supports the operation of at least one running service; In the case that the remaining resources of the target server cannot meet the needs of the service to be run, filtering out the target service from at least one of the already run services according to at least the running information of each of the already run services in the target server; The occupied resources of the target service are released so as to be used by the target server to support the operation of the service to be operated.

2. The resource management method according to claim 1, wherein: Filtering a target service from at least one of the running services according to at least the running information of each of the running services in the target server includes: A target service is selected from at least one of the already running services according to the running information of each of the already running services and the running information of the services to be run in the target server.

3. The resource management method according to claim 2, wherein: The operation information includes priority; The step of filtering out a target service from at least one of the already running services according to the running information of each of the already running services and the running information of the services to be run in the target server comprises: Filtering a first service from at least one of the already running services according to the priorities corresponding to the services already running and the priorities corresponding to the services to be run; The target service is determined from the filtered first services.

4. The resource management method according to claim 3, wherein: The step of selecting a first service from at least one of the already running services according to the priorities corresponding to the already running services and the priorities corresponding to the services to be run includes: According to the comparison result between the priorities of the already running services and the services to be run in the target server, the already running services whose priorities are lower than the priorities of the services to be run are screened out and recorded as the first services.

5. The resource management method according to claim 3, wherein: The operation information also includes activity; The determining the target service from the filtered first services includes: When the activity of the first service continues to be lower than its own set value for a first period of time and reaches a preset shutdown time threshold, the first service is determined as the target service.

6. The resource management method according to claim 5, wherein: For any of the first services, the process of determining the activity level includes: Obtaining resource utilization of resources occupied by the first service and a first preset period of the first service; According to the resource utilization rate of the resources occupied by the first service in the first preset period, an average utilization rate of the resources occupied by the first service in the first preset period is determined and used as the activity level of the first service.

7. The resource management method according to claim 1, wherein: The operation information includes activity; Filtering a target service from at least one of the running services according to at least the running information of each of the running services in the target server includes: Determine whether a first duration during which the activity of the running service is continuously lower than a value set by the running service reaches a preset shutdown time threshold, and determine the running service for which the first duration reaches the preset shutdown time threshold; The target service is selected from the running services whose first duration reaches the preset shutdown time threshold.

8. The resource management method according to claim 5 or 7, wherein: The target service includes multiple; The resource management method further includes: For any of the target services, determining a second time duration during which the first time duration exceeds the preset shutdown time threshold; Arrange the plurality of target services in descending order of the second durations to determine an arrangement order of the plurality of target services.

9. The resource management method according to claim 8, wherein: The releasing the occupied resources of the target service so as to be used by the target server to support the operation of the service to be operated includes: According to the arrangement order of the plurality of target services, the occupied resources of the target services are released in sequence until the target server supports the operation of the service to be operated.

10. The resource management method according to any one of claims 1 to 7, wherein: The resource management method further includes: For any of the running services in the target server, obtain the activity and the self-set value of the running service; In a second preset period, if the activity of the running service is continuously lower than the self-set value, the current priority of the running service is downgraded according to a preset adjustment amount; In the second preset period, if the activity of the running service is higher than or equal to the self-set value, the current priority of the running service is upgraded according to the preset adjustment amount.

11. The resource management method according to claim 10, wherein: Before the current priority of the running service is upgraded according to the second preset adjustment amount, the method further includes: If the current priority has reached the initially set value of the priority of the running service, the upgrade operation on the running service is stopped.

12. The resource management method according to claim 10, wherein: The resource management method further includes: For the running service that has been downgraded, the second preset period corresponding to the next downgrade is increased.

13. The resource management method according to claim 1, wherein: Obtaining a target server for running the service to be run, including: According to the type of the service to be run, a target server for running the service to be run is determined.

14. The resource management method according to any one of claims 1 to 6, wherein: The operation information includes occupied resources; the target servers of the same type as the service to be run include multiple; The step of filtering out a target service from at least one of the already running services according to at least the running information of each of the already running services in the target server comprises: Determining the remaining resources of each of the target servers according to the total resources and occupied resources of each of the target servers; According to the remaining resources of each of the target servers, the occupied resources of each of the running services in each of the target servers, and the occupied resources required by the service to be run, the server to be transferred and the server to be received are screened out from the multiple target servers, and the target service in the server to be transferred is determined; The releasing the occupied resources of the target service so as to be used by the target server to support the operation of the service to be operated includes: The target service in the server to be transferred is transferred to the server to be received for operation, so as to release the occupied resources of the target service in the server to be transferred, so as to support the operation of the service to be operated.

15. The resource management method according to claim 14, wherein: The resource management method further includes: When the remaining resources after the server to be transferred releases the occupied resources do not meet the running requirements of the service to be run, selecting a second service from the remaining running services according to the priorities corresponding to the remaining running services in the server to be transferred and the priorities corresponding to the service to be run; Determine whether the second service is a target service according to the activity corresponding to the second service.

16. The resource management method according to claim 14, wherein: The resource management method further includes: When the remaining resources after the server to be transferred releases the occupied resources do not meet the running requirements of the service to be run, a target service is determined according to the activities corresponding to the remaining running services in the server to be transferred.

17. The resource management method according to claim 1, wherein: The resource management method further includes: According to a third preset period, detecting the occupied resources of each of the running services in each server in the cluster; Filtering the third service to be transferred and the server to be received from the plurality of the running services according to the occupied resources of each of the running services in each of the servers; The third service is transferred to the server to be received to release resources of the original server where the third service is located.

18. The resource management method according to claim 1, wherein: The resource management method further includes: The service to be run is deployed on the target server using the application software Kubernetes.

19. A resource management device, wherein: include: Data acquisition module, service screening module and resource release module; The data acquisition module is configured to acquire the service to be run and the target server for running the service to be run; The target server currently supports the operation of multiple running services; The service screening module is configured to screen out a target service from a plurality of already running services at least according to the running information of each of the already running services in the target server when the remaining resources of the target server cannot meet the running of the service to be run; The resource release module is configured to release the occupied resources of the target service so that the target server can support the operation of the service to be run.

20. A computer device, wherein: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the resource management method as described in any one of claims 1 to 18 are performed.

21. A computer non-transitory readable storage medium, wherein: The computer non-transitory readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the resource management method according to any one of claims 1 to 18 are executed.