Method and device for supporting elastic expansion capability by computing power resource of intelligent computing center
By monitoring the service call frequency in the intelligent computing center and dynamically adjusting the computing resource allocation according to the service level, the startup delay problem caused by the loading of large model files and service images under the Serverless architecture is solved, and more efficient computing resource utilization and user experience improvement are achieved.
Patent Information
- Application Number
- CN202411345361.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-06-03
AI Technical Summary
When applying Serverless architecture in the intelligent computing center, we face significant startup delay problems caused by loading large model files and service images, especially during cold startup, which affects the user experience and limits the effective utilization of high-performance computing resources.
By monitoring the call frequency of the target service, automatically judge its service level, and dynamically adjust the allocation of computing resources according to the service level. Specific measures include: for high-frequency services, deploying multiple container instances; for medium-frequency services, using a single container instance; for low-frequency services, using pre-cache resources to quickly deploy container instances; for uncache services, directly pulling resources from the data source for instantiation.
It effectively avoids the service startup delay, improves the user experience, and improves the computing resource utilization rate and performance of high-performance computing resources in the intelligent computing center.
Smart Images

Figure CN120086003A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of computing power infrastructure, and in particular, to a method and device for enabling elastic scaling of computing power resources in an intelligent computing center. Background Technique
[0002] With the development of artificial intelligence technology and computing power technology, the concept of an intelligent computing center has emerged. An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly to provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). An intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0003] Serverless is an execution architecture of cloud computing, where the cloud provider automatically manages the allocation of machine resources. This model allows users to write and deploy code without having to worry about the operation and maintenance of the underlying servers. Users only pay according to the actual computing consumption, without the need to purchase and maintain a fixed number of servers or virtual machines.
[0004] Applying the Serverless architecture to an intelligent computing center, the intelligent computing center can utilize its computing power resources more effectively. The Serverless architecture ensures that computing power resources are only used when actually needed, which helps to optimize the use of computing power resources and reduce waste.
[0005] However, there are also the following problems: In the application of an intelligent computing center, especially in scenarios involving deep learning, model files are often very large, usually ranging from dozens of GB to hundreds of GB. The loading and initialization processes of such large model files consume a large amount of time and computing power resources. And the container images of services usually contain all the dependencies and libraries required to run the services. For complex applications, the size of these images may be around 1 to 3 GB, and loading these large images also takes a long time, especially when starting the service for the first time. And if the Serverless architecture is applied in an intelligent computing center, the service usually starts when a request arrives, which means that each new request may trigger a "cold start". If the default scheduler is used and the startup process is not optimized, then loading large model files and service images may cause startup delays, sometimes even taking more than ten minutes (such as thirty minutes), and such delays are unacceptable for applications that require quick responses.
[0006] In summary, when applying the Serverless architecture in an intelligent computing center, the on-demand startup feature of the Serverless architecture brings the problem of initial startup latency. Especially when it comes to loading large model files and service images, the startup latency of the service is particularly obvious, which not only affects the user experience but also limits the effective utilization of the high-performance computing power resources in the intelligent computing center and restricts the performance improvement of the high-performance computing power resources in the intelligent computing center. Summary of the Invention
[0007] Embodiments of the present application provide a method and device for enabling elastic scaling of computing power resources in an intelligent computing center to solve the technical problem that in an environment combining an intelligent computing center and the Serverless architecture, the task requirements of processing large model files and service images lead to significant startup latency problems (especially during cold startups), which not only affect the user experience but also limit the effective utilization of the high-performance computing power resources in the intelligent computing center and restrict the performance improvement of the high-performance computing power resources in the intelligent computing center.
[0008] To solve the above technical problems, the present application is implemented as follows:
[0009] In a first aspect, embodiments of the present application provide a method for enabling elastic scaling of computing power resources in an intelligent computing center, the method comprising:
[0010] Step S1: Monitor the invocation frequency of a target service;
[0011] Step S2: Based on the invocation frequency, determine the current service level to which the target service belongs, where the service levels include: high-frequency service, medium-frequency service, low-frequency service, and cacheless service;
[0012] Step S3: Run the target service based on the current service level to which the target service belongs;
[0013] The step S3 includes:
[0014] Step S31: When the current service level to which the target service belongs is the high-frequency service, run the target service based on at least two deployed container instances and the computing power resources relied on by the at least two container instances;
[0015] Step S32: When the current service level to which the target service belongs is the medium-frequency service, run the target service based on one deployed container instance and the computing power resources relied on by the one container instance;
[0016] Step S33: When the current service level to which the target service belongs is the low-frequency service, run the target service based on the resources of the target service cached in the intelligent computing center;
[0017] Step S34: When the service level to which the target service currently belongs is the cacheless service, pull the resources of the target service from the data source of the target service, deploy a container instance based on the pulled resources, and run the target service based on the container instance and the computing power resources on which the container instance depends.
[0018] Optionally, step S1 includes:
[0019] Step S11: Monitor the queries per second (QPS) of the target service;
[0020] Step S12: Determine the invocation frequency based on the QPS.
[0021] Optionally, step S2 includes:
[0022] Step S21: When the invocation frequency is greater than a first preset threshold, determine that the service level to which the target service currently belongs is the high-frequency service;
[0023] Step S22: When the invocation frequency is greater than a second preset threshold and less than or equal to the first preset threshold, determine that the service level to which the target service currently belongs is the medium-frequency service;
[0024] Step S23: When the invocation frequency is greater than a third preset threshold and less than or equal to the second preset threshold, determine that the service level to which the target service currently belongs is the low-frequency service;
[0025] Step S24: When the invocation frequency is less than or equal to the third preset threshold and greater than or equal to zero, determine that the service level to which the target service currently belongs is the cacheless service; where the third preset threshold is less than the second preset threshold and the second preset threshold is less than the first preset threshold.
[0026] Optionally, step S31 further includes:
[0027] Step S311: When the service level to which the target service currently belongs is the high-frequency service, for every preset number of increases in the invocation frequency of the target service, the number of container instances among the at least two container instances increases by one.
[0028] Optionally, step S3 further includes:
[0029] Step S35: Allocate a service queue for the target service according to the service level to which the target service currently belongs, and run the target service, where the service queue includes at least one of the following: high-frequency service queue, medium-frequency service queue, low-frequency service queue. If the service level to which the target service currently belongs is the cacheless service, the service queue of the target service is the low-frequency service queue.
[0030] Optionally, step S35 includes:
[0031] Step S351: Determine whether the target service has a previous belonging service queue;
[0032] Step S352: If the target service has a previous belonging service queue, re-allocate a service queue for the target service according to the previous belonging service queue and the currently belonging service level, and run the target service.
[0033] Optionally, step S352 includes:
[0034] Step S3521: If the target service has a previous belonging service queue, perform at least one of the following operations according to the previous belonging service queue and the currently belonging service level: service queue upgrade operation, service queue downgrade operation, and service queue level unchanged operation, and run the target service;
[0035] The service upgrade operation includes at least one of the following: upgrading the target service from the low-frequency service queue to the medium-frequency service queue, upgrading the target service from the medium-frequency service queue to the high-frequency service queue;
[0036] The service downgrade operation includes at least one of the following: downgrading the target service from the high-frequency service queue to the medium-frequency service queue, downgrading the target service from the medium-frequency service queue to the low-frequency service queue;
[0037] Wherein, when the disk capacity of the node where the target service is located reaches the preset capacity threshold, the target service is removed from the low-frequency service queue.
[0038] In a second aspect, an embodiment of the present application provides a device for the elastic scaling ability of computing power resources in an intelligent computing center. The device includes:
[0039] A monitoring module, configured to monitor the invocation frequency of the target service;
[0040] An execution module, configured to determine the service level to which the target service currently belongs according to the invocation frequency, where the service level includes: high-frequency service, medium-frequency service, low-frequency service, and cacheless service;
[0041] Run the target service according to the service level to which the target service currently belongs;
[0042] The execution module is further configured to, when the service level to which the target service currently belongs is the high-frequency service, run the target service based on at least two deployed container instances and the computing power resources relied on by the at least two container instances;
[0043] When the service level to which the target service currently belongs is the medium-frequency service, run the target service based on one deployed container instance and the computing power resources relied on by the one container instance;
[0044] When the service level to which the target service currently belongs is the low-frequency service, run the target service based on the resources of the target service cached in the intelligent computing center;
[0045] When the service level to which the target service currently belongs is the non-cached service, pull the resources of the target service from the data source of the target service, deploy one container instance based on the pulled resources, and run the target service based on the container instance and the computing power resources relied on by the container instance.
[0046] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of a method for an intelligent computing center computing power resource to support elastic scaling ability as described in the first aspect are implemented.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the steps of a method for an intelligent computing center computing power resource to support elastic scaling ability as described in the first aspect are implemented.
[0048] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of a method for an intelligent computing center computing power resource to support elastic scaling ability as described in the first aspect are implemented.
[0049] In the embodiments of the present application, first, the service call frequency is monitored to ensure that the intelligent computing center can monitor the usage of services and provide data support for subsequent resource allocation and service level adjustment. Subsequently, the service level is automatically determined according to the call frequency, including high-frequency, medium-frequency, low-frequency, and cacheless services. This level classification enables the intelligent computing center to adopt the most suitable operation mode for different services. Specifically, for high-frequency services, the intelligent computing center deploys multiple container instances and utilizes sufficient computing power resources to ensure high availability and low-latency response of the services; for medium-frequency services, they are processed through a single container instance to ensure the economy of computing power resource usage while meeting service requirements; for low-frequency services, the pre-cached resources in the intelligent computing center are used to quickly deploy container instances. Although instantiation is required during response, the startup time is reduced by using cached resources; for cacheless services, the latest resources are directly pulled from the data source for instantiation to ensure the timeliness and accuracy of the data. Generally, the intelligent computing center can dynamically adjust the allocation of computing power resources according to the actual needs of services, realize the elastic scaling ability of computing power resources, avoid service startup latency, improve the user experience, and avoid waste of computing power resources, thereby improving the utilization rate of computing power resources in the intelligent computing center and the performance of high-performance computing power resources in the intelligent computing center. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0051] Figure 1 is a flowchart of a method for an intelligent computing center to support the elastic scaling ability of computing power resources provided by an embodiment of the present application;
[0052] Figure 2 is a schematic diagram of task queue division in an intelligent computing center provided by an embodiment of the present application;
[0053] Figure 3 is a system architecture diagram of an intelligent computing center to support the elastic scaling ability of computing power resources provided by an embodiment of the present application;
[0054] Figure 4 is a structural block diagram of a device for an intelligent computing center to support the elastic scaling ability of computing power resources provided by an embodiment of the present application;
[0055] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0057] The following briefly explains the technical terms involved in the present application.
[0058] The "computing power" described in the present application refers to the ability of a computer device or a computing / data center to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to process information data and output a target result. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0059] The "Computational Power (CP)" described in the present application is an ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (Floating Point Operations Per Second, FLOPS, where 1 EFLOPS = 10^18 FLOPS). The larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs (Central Processing Unit) or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 . In the embodiments of the present application, the computing power uses the half-precision floating-point computing power number (FP16) of the graphics card.
[0060] The "carrying capacity" (Network Power, NP) described in the present application is the manifestation of the data transmission ability of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. The carrying capacity involves network transmission inside and between data centers and is a comprehensive indicator to measure the network transmission scheduling ability. In the embodiments of the present application, the carrying capacity uses the video memory bandwidth. In the embodiments of the present application, the storage capacity uses the video memory bandwidth.
[0061] The "Storage Power (SP)" described in this application is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator to measure the data storage ability of a data center, including external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability. In the embodiments of this application, the storage power uses the video memory capacity.
[0062] The "computing power infrastructure" described in this application is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power. It can realize the centralized computing, storage, transmission, and application of information, presenting characteristics such as multi-dimensional ubiquity, intelligent agility, security and reliability, and green and low-carbon. It is of great significance for promoting industrial transformation and upgrading, empowering China's scientific and technological innovation, meeting people's beautiful life, and realizing high-efficiency social governance.
[0063] The "computing power" described in this application includes general computing power CP 通用 , intelligent computing power CP 智能 , and super computing power CP 超级 .
[0064] The "general computing power" described in this application refers to the computing ability provided by servers based on CPU chips, which is used to support basic general computing such as cloud computing and edge computing.
[0065] The "intelligent computing power" described in this application refers to the large-scale deployment of intelligent computing centers for various artificial intelligence innovation applications based on special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0066] The "super computing power" described in this application is mainly the computing ability provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0067] The "intelligent computing center" described in this application refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0068] The "computing power resources" described in this application refer to the technologies and facilities required for the development of the digital society, with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0069] The "elastic scaling ability" described in this application can also be referred to as automatic scaling, that is, it can automatically update the workload resources to meet the demand. It allows the intelligent computing center to dynamically increase or decrease computing power resources according to the actual workload demand. This ability is a core feature of modern cloud computing services and large-scale distributed systems, and is particularly suitable for application environments with large demand fluctuations.
[0070] The Serverless architecture described in this application is a software design pattern, not referring to the absence of physical servers or virtual machines, but rather "service without awareness". In this architecture, developers do not need to manage servers (such as configuring, maintaining, and updating the operating system or software), and the cloud service provider is responsible for all underlying infrastructure management. Developers only need to focus on writing and deploying code.
[0071] The Serverless architecture is usually billed according to the number of times a function is executed or the execution time, and is very suitable for processing intermittent or irregular loads, such as API (Application Programming Interface) requests or event responses. Serverless solves a series of platform-type problems such as resource hosting, scheduling, and operation and maintenance management, enabling developers to focus more on business requirements and application logic without having to consider the application, creation, management, and maintenance of service resources.
[0072] Specifically, Serverless provides convenience for developers by solving the following problems:
[0073] 1. Resource Management and Maintenance: In the traditional server deployment model, developers need to manage the server lifecycle themselves, including operating system installation, updates, hardware maintenance, etc. In the Serverless architecture, these tasks are handled by the service provider, and developers don't need to worry about the physical or virtual details of the server, thus greatly reducing the complexity and cost of operation and maintenance.
[0074] 2. Elastic Scaling: In the traditional server deployment model, developers need to pre-configure server resources according to predicted business demands. However, business traffic is often unpredictable, which may lead to over-provisioning or under-provisioning of resources. The Serverless architecture can automatically adjust computing resources according to actual needs, ensuring high availability and response speed of the service and avoiding waste of resources.
[0075] 3. Cold Start Problem: In the traditional server model, if there are no requests for a long time, the service may encounter the cold start problem, that is, when a request arrives, it takes extra time and resources to start the service. The Serverless architecture reduces the latency caused by cold start through preloading or caching technology, improving the response speed of the service.
[0076] 4. High Availability and Disaster Recovery: Serverless services usually have a high availability design, which can automatically perform failover and recovery to ensure the continuous operation of the service. In addition, since the service provider is responsible for maintaining and updating the infrastructure, developers can focus more on the development and optimization of business logic.
[0077] 5. Cost Optimization: In the Serverless mode, users only pay for the resources actually used. This pay-as-you-go model helps reduce the operating costs of enterprises, especially in the case of unstable business volumes, and can avoid over-provisioning and waste of resources.
[0078] Please refer to Figure 1 , Figure 1 which shows a method for supporting elastic scaling ability of computing power resources in an intelligent computing center according to an embodiment of the present application. As Figure 1 shown, the method includes:
[0079] Step S1: Monitor the call frequency of the target service;
[0080] Step S2: Determine the service level to which the target service currently belongs according to the call frequency;
[0081] wherein the service levels include: high-frequency service, medium-frequency service, low-frequency service, and non-cached service;
[0082] Step S3: Run the target service according to the service level to which the target service currently belongs.
[0083] In step S1, the intelligent computing center first needs to monitor the call frequency of the target service, that is, the number of times the service is requested per second or per minute. The call frequency is basic data, which is crucial for subsequent judgment of the computing power resources required by the service and the level to which the service belongs. And the monitoring can be real-time monitoring or time-segmented monitoring, which can be set according to the specific needs of users.
[0084] In a possible implementation, step S1 includes: step S11: monitoring the QPS (Queries Per Second) of the target service; step S12: determining the call frequency based on the QPS. That is, the call frequency can be determined by monitoring the QPS of the target service. QPS refers to the number of queries processed per second, which is an important indicator for measuring system performance and load. By monitoring the QPS and determining the call frequency of the service based on this data, system administrators and developers can better understand the performance status of the service and make reasonable decisions on the allocation of computing power resources (specifically described later). This can not only help optimize the utilization efficiency of computing power resources, but also ensure that the service can provide stable responses according to actual needs, thereby improving user satisfaction and the reliability of the intelligent computing center.
[0085] In step S2, based on the data collected in step S1, the target service can be evaluated and classified into corresponding service levels: high-frequency, medium-frequency, low-frequency, or cacheless service. This step is the core of the decision-making process, and different levels of services will correspond to different resource configurations and service operation modes.
[0086] Specifically, the judgment of the service level usually depends on a preset threshold. In a possible implementation, step S2 includes:
[0087] Step S21: When the call frequency is greater than the first preset threshold, it is determined that the current service level to which the target service belongs is a high-frequency service;
[0088] Step S22: When the call frequency is greater than the second preset threshold and less than or equal to the first preset threshold, it is determined that the current service level to which the target service belongs is a medium-frequency service;
[0089] Step S23: When the call frequency is greater than the third preset threshold and less than or equal to the second preset threshold, it is determined that the current service level to which the target service belongs is a low-frequency service;
[0090] Step S24: When the call frequency is less than or equal to the third preset threshold and greater than or equal to zero, it is determined that the current service level to which the target service belongs is a cacheless service; where the third preset threshold is less than the second preset threshold and the second preset threshold is less than the first preset threshold.
[0091] That is to say, if the invocation frequency of a service exceeds the first preset threshold, it indicates that the service receives a very high frequency of requests. This service is classified as a "high-frequency service", meaning that it requires more resources and stronger scaling capabilities to meet the continuous high demand. If the invocation frequency of the service is greater than the second preset threshold but does not exceed the first preset threshold, it indicates that the request volume of the service is moderate. This service is classified as a "medium-frequency service", and such a service requires moderate computing power resources and does not need to scale as frequently as high-frequency services. If the invocation frequency of the service is greater than the third preset threshold but less than or equal to the second preset threshold, it indicates that the service occasionally receives requests. This service is classified as a "low-frequency service". Generally, such services do not need to continuously occupy a large amount of resources and can be started on demand. If the invocation frequency of the service is less than or equal to the third preset threshold, it indicates that the requests for the service are very few, almost close to none or only requested under specific conditions, or it is the first request. This service is classified as a "cacheless service", and such services do not need to pre-cache data or resources, and the necessary data can be loaded from the source for each request.
[0092] And the setting of the thresholds can be obtained based on the analysis of historical data, the performance requirements of the service, and cost-effectiveness, etc. The first preset threshold (the highest) is used to define the most active service, while the third preset threshold (the lowest) is used to define the service with the lowest activity. These thresholds can help the intelligent computing center automatically identify and adapt to the operating requirements of different services and optimize the allocation of computing power resources.
[0093] In a specific application scenario, the response of a high-frequency service is within 100ms, and the average invocation is more than 5 / s. The main characteristics are high invocation frequency, low latency, and multiple container instances for horizontal scaling support at the same time;
[0094] The response of a medium-frequency service is within 100ms, and the average invocation is 5 / s to 1 / min. The main characteristics are medium invocation frequency, low latency, small concurrency, and only one container supporting the service;
[0095] The first response of a low-frequency service is within 10s, and the average invocation is greater than 1 / min. The main characteristics are low invocation frequency, large first latency, and the background server has a corresponding mapping, and it is necessary to ensure that the model files and image files affecting the deployment speed are cached on the server.
[0096] Among them, the mapping of the background server refers to a data structure that records the correspondence between services and the server nodes on which they are deployed, as well as the location information of the available model files and image files on each node. This mapping relationship is the basis for the scheduler to quickly deploy and manage services.
[0097] In a cacheless service, when the service needs to be scheduled, the corresponding resources need to be directly pulled from the data source of the service.
[0098] In step S3, the target service can be run according to the service level to which the target service currently belongs. Step S3 can be further divided into steps S31 to S34.
[0099] Step S31: When the service level to which the target service currently belongs is a high-frequency service, run the target service based on at least two deployed container instances and the computing power resources relied on by the at least two container instances;
[0100] Step S32: When the service level to which the target service currently belongs is a medium-frequency service, run the target service based on one deployed container instance and the computing power resources relied on by the one container instance;
[0101] Step S33: When the service level to which the target service currently belongs is a low-frequency service, run the target service based on the resources of the target service cached in the intelligent computing center;
[0102] Step S34: When the service level to which the target service currently belongs is a non-cached service, pull the resources of the target service from the data source of the target service, deploy one container instance based on the pulled resources, and run the target service based on the container instance and the computing power resources relied on by the container instance.
[0103] It should be noted that for high-frequency services, the intelligent computing center will deploy at least two container instances to handle high request volumes, ensuring that the service does not experience performance bottlenecks due to the overload of a single instance. And these container instances will be configured with sufficient computing power resources, such as multiple CPU (Central Processing Unit) cores and sufficient memory, and even including GPU (Graphics Processing Unit) resources, to handle concurrent requests. For medium-frequency services, their call frequency is relatively low, and usually one container instance can meet the requirements. The deployed container instance can be configured with appropriate computing power resources to balance performance and cost efficiency. For low-frequency services, the intelligent computing center will not continuously run an active container instance, and in terms of computing power resource configuration, the resources cached in the intelligent computing center can be used to run the target service, reducing the startup time to save resources without affecting the response speed. For non-cached services, the latest resources need to be pulled from the data source each time to deploy the service, which is applicable to scenarios with extremely high requirements for data real-time performance. And the deployed container instance will be configured immediately according to the pulled data and resources to ensure the latest state of the service and the accuracy of the data. On the other hand, non-cached services are usually services requested for the first time. Since the service is requested for the first time and its call frequency is not yet clear, there is no need and no way to cache resources in advance. Then, the corresponding resources can be pulled from the data source of the service, and after the service is requested for the first time, it can be determined as a low-frequency service.
[0104] In a possible implementation, step S31 further includes:
[0105] Step S311: When the current service level of the target service is a high-frequency service, for every preset number of increases in the invocation frequency of the target service, among at least two container instances, the number of container instances increases by one.
[0106] It should be noted that when the current service level of the target service is a high-frequency service, if it is monitored that the invocation frequency increases to reach the preset number, the intelligent computing center can automatically perform a container expansion operation, that is, every time the invocation frequency increases to reach this preset threshold, the intelligent computing center will add a new instance to the deployed container instance cluster. For example: If two instances are initially configured, then it will increase to three when the first expansion is triggered, and so on. Thus, increasing the container instances can help disperse the load of processing requests, avoid overloading of a single or a few instances, thereby maintaining the response speed and reliability of the service. And multi-instance deployment can improve the overall availability of the service. Even if one instance fails, other instances can still continue to provide services. Through this automatic expansion strategy, the intelligent computing center can dynamically manage its computing power resources, ensure sufficient processing capacity when the task demand increases, and not waste computing power resources when the task demand decreases.
[0107] In a possible implementation, step S3 further includes:
[0108] Step S35: According to the current service level to which the target service belongs, allocate a service queue for the target service and run the target service, where the service queue includes at least one of the following: high-frequency service queue, medium-frequency service queue, low-frequency service queue. Among them, if the current service level to which the target service belongs is a non-caching service, the service queue of the target service is the low-frequency service queue.
[0109] Among them, the service queue is different categories divided according to the invocation frequency of the service, and the purpose is to optimize resource allocation and management according to the load and performance requirements of each service. As Figure 2 shown, the high-frequency queue is a data structure used to store high-frequency services, and one service ID corresponds to multiple hosts (referring to multiple nodes where the current model service is located); the medium-frequency queue is a data structure used to store medium-frequency services, and one service ID corresponds to 1 host (the node where the current model service is located); the low-frequency queue is a data structure used to store low-frequency services, and one service ID corresponds to 1 host (referring to the node where the current model service cache, model, and image are located). And it should be noted that Figure 2 For Figure 3 a part of the screenshot, reference can be made to Figure 3 for understanding Figure 2 .
[0110] In a possible implementation, step S35 includes: Step S351: Determine whether the target service has a previous service queue to which it belongs; Step S352: If the target service has a previous service queue to which it belongs, reassign a service queue for the target service according to the previous service queue to which it belongs and the current service level to which it belongs, and run the target service. Among them, step S352 includes: Step S3521: If the target service has a previous service queue to which it belongs, perform at least one of the following operations according to the previous service queue to which it belongs and the current service level to which it belongs: a service queue upgrade operation, a service queue downgrade operation, and a service queue level unchanged operation, and run the target service; The service upgrade operation includes at least one of the following: upgrade the target service from a low-frequency service queue to a medium-frequency service queue, upgrade the target service from a medium-frequency service queue to a high-frequency service queue; The service downgrade operation includes at least one of the following: downgrade the target service from a high-frequency service queue to a medium-frequency service queue, downgrade the target service from a medium-frequency service queue to a low-frequency service queue; Among them, when the disk capacity of the node where the target service is located reaches a preset capacity threshold, the target service is removed from the low-frequency service queue. That is to say, it is also possible to continuously monitor the storage resource utilization rate of the node (server) where the target service is located. When the storage resource utilization rate reaches a certain ratio, the historical cache of the server will be cleared, and the corresponding service will be downgraded, that is, the target service will be removed from the low-frequency service queue.
[0111] It should be noted that the intelligent computing center first checks whether each service has a "previous service queue to which it belongs", that is, the queue previously assigned, to understand the historical status and change trend of the service. If the service has historical queue information, the intelligent computing center will compare the current service call frequency with the previous status to determine whether queue changes are needed. Such changes include: service upgrade (such as upgrading from low frequency to medium frequency, or from medium frequency to high frequency), service downgrade (such as downgrading from high frequency to medium frequency, or from medium frequency to low frequency), or remaining at the current queue level unchanged. According to the current demand and historical performance of the service, corresponding queue adjustment operations are performed, which includes dynamically moving the service from one queue to another to optimize resource usage and response time. In addition, the adjustment operation also needs to consider the actual usage of computing power resources. For example, if the disk capacity of the node where a certain service is located reaches a preset threshold, it will trigger the operation of removing the service from the low-frequency queue to prevent resource overload and system performance degradation. In summary, through this flexible queue management strategy, the intelligent computing center can adapt to various operating conditions and load changes, thereby improving the overall service quality and user satisfaction.
[0112] In summary, in the embodiments of the present application, first, the service call frequency is monitored to ensure that the intelligent computing center can monitor the usage of services, providing data support for subsequent resource allocation and service level adjustment. Subsequently, according to the call frequency, the level of the service is automatically determined, including high-frequency, medium-frequency, low-frequency, and non-cached services. This level classification enables the intelligent computing center to adopt the most suitable operation mode for different services. Specifically, for high-frequency services, the intelligent computing center deploys multiple container instances, using sufficient computing power resources to ensure high availability and low-latency response of the services. For medium-frequency services, they are processed by a single container instance to ensure the economy of computing power resource usage while meeting service requirements. For low-frequency services, the pre-cached resources in the intelligent computing center are used to quickly deploy container instances. Although instantiation may be required during response, the startup time is reduced by using cached resources. For non-cached services, the latest resources are directly pulled from the data source for instantiation to ensure the timeliness and accuracy of the data. Generally, the intelligent computing center can dynamically adjust the allocation of computing power resources according to the actual needs of the services, realizing the elastic scaling ability of computing power resources, avoiding service startup delays, improving the user experience, and avoiding waste of computing power resources, thus improving the utilization rate of computing power resources in the intelligent computing center and the performance of high-performance computing power resources in the intelligent computing center.
[0113] The method shown in the embodiments of the present application (as Figure 1 shown) can be applied to the system architecture of the intelligent computing center. As Figure 3 shown, the system consists of the following main parts: a router, a scheduler, and a monitor.
[0114] The monitor is mainly responsible for collecting the QPS sent by each service container, summarizing it, and sending it to the scheduler. The monitor is also used to collect the disk cache usage of each service node, and when it is determined that the disk cache usage reaches the preset capacity threshold, it notifies the scheduler to perform a service degradation operation for clearing the cache.
[0115] The router mainly performs the following two parts of work:
[0116] 1. Forwarding routing. Specifically, it uniformly accepts user requests and checks whether there is a service ID corresponding to the service ID carried in the user request in the memory mapping table. If it already exists, indicating that the service has been deployed (the service has been deployed and run on a certain server), then according to the mapping information in the memory mapping table, the corresponding server is found, and the user request is forwarded to this server to reduce the response time and improve service efficiency. If it does not exist, the scheduler is notified for scheduling, and the request is suspended, waiting for the notification of the scheduling result. After receiving the notification of successful scheduling, the router will associate the service ID and the server and store them in the mapping table.
[0117] 2. Update the mapping information. Specifically, when the monitor detects that the traffic of a certain service has increased to a certain threshold (for example, the call frequency of a certain service reaches a certain threshold or the QPS of a certain service reaches a certain threshold), it will trigger a service upgrade operation. The monitor will notify the scheduler, and after calculating the appropriate scheduling node through a preset scheduling algorithm, the scheduler will notify the router, and the router can append the corresponding relationship between the service ID and the server in the mapping table. Similarly, when the monitor detects that the traffic of a certain service has decreased to a certain threshold (for example, the call frequency of a certain service has decreased to a certain threshold or the QPS of a certain service has decreased to a certain threshold), it will trigger a service downgrade operation. The monitor will notify the scheduler, and after calculating the appropriate scheduling node through a preset scheduling algorithm, the scheduler will notify the router, and the router can delete the corresponding relationship between the service ID and the server in the mapping table.
[0118] The scheduler mainly performs the following tasks:
[0119] Receive scheduling information, calculate the appropriate scheduling node for service scheduling according to a preset scheduling algorithm; execute service upgrade operations and service downgrade operations; clear the cache (in the service downgrade operation, if the disk capacity of the node where the service is located reaches the preset threshold, the scheduler will perform the operation of clearing the cache).
[0120] The scheduler first needs to master the resource information of all nodes in the intelligent computing center and manage services in 4 categories according to the different call frequencies of the services. They are high-frequency services, medium-frequency services, low-frequency services, and cacheless services.
[0121] In specific application scenarios, the characteristics of the 4 service categories are as follows:
[0122] High-frequency services: The response is within 100ms, and the average call is more than 5 / s. The main characteristics are high call frequency, low latency, and horizontal scaling support by multiple container instances at the same time;
[0123] Medium-frequency services: The response is within 100ms, and the average call is 5 / s - 1 / min. The main characteristics are medium call frequency, low latency, small concurrency, and only one container supporting the service;
[0124] Low-frequency services: The first response is within 10s, and the average call is more than 1 / min. The main characteristics are low call frequency, large first latency, and the background server has a corresponding mapping, and it is necessary to ensure that the model files and image files affecting the deployment speed are cached on the server.
[0125] Among them, the mapping of the background server refers to a data structure that records the correspondence between services and the server nodes on which they are deployed, as well as the location information of the model files and image files available on each node. This mapping relationship is the basis for the scheduler to quickly deploy and manage services.
[0126] The difference between low-frequency services and medium-frequency services is as follows: A medium-frequency service has one container supporting the service. A low-frequency service has caches of model files and image files related to the service, but the corresponding containers are not deployed, that is, the container instances do not run continuously, but necessary model files and image files are cached in advance on appropriate server nodes. When the service is called, the container instances can be quickly deployed and started, leveraging the cached resources to ensure response time and reduce startup latency.
[0127] Cacheless service: When a service needs to be scheduled, the corresponding resources need to be directly pulled from the data source of the service.
[0128] Corresponding to high-frequency services, medium-frequency services, and low-frequency services, there are also the following types of queues:
[0129] High-frequency queue: A data structure used to store high-frequency services, where one service ID corresponds to multiple hosts (referring to multiple nodes where the current model service is located);
[0130] Medium-frequency queue: A data structure used to store medium-frequency services, where one service ID corresponds to 1 host (the node where the current model service is located);
[0131] Low-frequency queue: A data structure used to store low-frequency services, where one service ID corresponds to 1 host (referring to the node where the current model service cache, model, and image are located).
[0132] When the scheduler selects the server nodes for deploying services, the first thing to consider is whether these nodes have the resources to meet the service requirements; it will check the requirements of each service in the above queues against the resource configurations of each node, and preferentially select the nodes that can best meet these service requirements for deployment to accelerate scheduling; it will execute scheduling changes based on the scheduling change requests from the monitor. Scheduling changes are divided into two types: service upgrade and service downgrade.
[0133] Service upgrade: For the model service that is loaded for the first time (without cache), it will be added to the low-frequency queue, and the router will be notified to update the route. For the model service in the low-frequency queue, when the call frequency exceeds a threshold (for example, greater than or equal to 1 time / minute), the service will be removed from the low-frequency queue, added to the medium-frequency queue, and the router will be notified to update the route. For the model service in the medium-frequency queue, when the call frequency exceeds a threshold (for example, greater than or equal to 5 times / second), the service will be removed from the medium-frequency queue, added to the high-frequency queue, and multiple instances will be started for load balancing, and the router will be notified to update the route. For the model service in the high-frequency queue, when the call frequency exceeds a certain threshold (for example, the call frequency increases by 10 times / second each time), a container instance will be started for load balancing, and the router will be notified to update the route.
[0134] Service downgrade: For the model service in the high-frequency queue, when the call frequency decreases by a certain number of times on average over a period of time (for example, 10 minutes) (for example, from 20 times / second to 10 times / second), the destruction of an instance will be triggered, and the router will be notified to update the route. For the model service in the high-frequency queue, when the call frequency drops to a certain threshold (for example, less than 5 times / second) on average over a period of time (for example, 10 minutes), it will be removed from the high-frequency queue and added to the medium-frequency queue, and the router will be notified to update the route. For the model service in the medium-frequency queue, when the call frequency drops to a certain threshold (for example, 1 time / minute) on average over a period of time (for example, 10 minutes), the service will be removed from the medium-frequency queue, added to the low-frequency queue, and the single container will be destroyed to release resources, and wait for the next scheduling, and the router will be notified to update the route. When the cache (disk) usage of the scheduled node reaches a certain threshold (for example, 65%), the low-frequency service will be downgraded to a cacheless service, the service will be removed from the low-frequency queue, and the disk cache (the model file and image file corresponding to the service) of the service will be deleted.
[0135] It should be noted that the scheduled computing power resources (such as the GPU cluster in the intelligent computing center) are responsible for starting containers, pulling images, pulling models, starting services, executing specific prediction tasks, and clearing caches.
[0136] Generally, the intelligent computing center can dynamically adjust the allocation of computing power resources according to the actual needs of the service, realize the elastic scaling ability of computing power resources, avoid the startup delay of the service, improve the user experience, and avoid the waste of computing power resources, improve the utilization rate of computing power resources in the intelligent computing center, and improve the performance of high-performance computing power resources in the intelligent computing center.
[0137] Figure 4 Shows a device for supporting the elastic scaling ability of computing power resources in an intelligent computing center according to an embodiment of the present application, as Figure 4 shown, the device 40 includes:
[0138] A monitoring module 41, configured to monitor the call frequency of the target service;
[0139] An execution module 42, configured to determine the current service level of the target service according to the call frequency, where the service levels include: high-frequency service, medium-frequency service, low-frequency service, and non-cached service;
[0140] Run the target service according to the current service level of the target service;
[0141] The execution module 42 is further configured to, when the current service level of the target service is a high-frequency service, run the target service based on at least two deployed container instances and the computing power resources relied on by the at least two container instances;
[0142] When the current service level of the target service is a medium-frequency service, run the target service based on one deployed container instance and the computing power resources relied on by the one container instance;
[0143] When the current service level of the target service is a low-frequency service, run the target service based on the resources of the target service cached in the intelligent computing center;
[0144] When the current service level of the target service is a non-cached service, pull the resources of the target service from the data source of the target service, deploy one container instance based on the pulled resources, and run the target service based on the container instance and the computing power resources relied on by the container instance.
[0145] In a possible implementation, the monitoring module 41 is further configured to monitor the queries per second QPS of the target service; determine the call frequency based on the QPS.
[0146] In a possible implementation, the execution module 42 is further configured to, when the call frequency is greater than a first preset threshold, determine that the current service level of the target service is a high-frequency service;
[0147] When the call frequency is greater than a second preset threshold and less than or equal to the first preset threshold, determine that the current service level of the target service is a medium-frequency service;
[0148] When the call frequency is greater than a third preset threshold and less than or equal to the second preset threshold, determine that the current service level of the target service is a low-frequency service;
[0149] When the call frequency is less than or equal to the third preset threshold and greater than or equal to zero, determine that the current service level of the target service is a non-cached service; where the third preset threshold is less than the second preset threshold and the second preset threshold is less than the first preset threshold.
[0150] In a possible implementation, the execution module 42 is further configured to, when the service level to which the target service currently belongs is a high-frequency service, increase the number of container instances by one among at least two container instances every time the invocation frequency of the target service increases by a preset number of times.
[0151] In a possible implementation, the execution module 42 is further configured to allocate a service queue for the target service according to the service level to which the target service currently belongs, and run the target service, where the service queue includes at least one of the following: a high-frequency service queue, a medium-frequency service queue, and a low-frequency service queue. If the service level to which the target service currently belongs is a non-buffered service, the service queue of the target service is a low-frequency service queue.
[0152] In a possible implementation, the execution module 42 is further configured to determine whether the target service has a previous service queue to which it belonged; if the target service has a previous service queue to which it belonged, re-allocate a service queue for the target service according to the previous service queue to which it belonged and the current service level to which it belongs, and run the target service.
[0153] In a possible implementation, the execution module 42 is further configured to, if the target service has a previous service queue to which it belonged, perform at least one of the following operations according to the previous service queue to which it belonged and the current service level to which it belongs: a service queue upgrade operation, a service queue downgrade operation, and a service queue level unchanged operation, and run the target service;
[0154] The service upgrade operation includes at least one of the following: upgrading the target service from a low-frequency service queue to a medium-frequency service queue, and upgrading the target service from a medium-frequency service queue to a high-frequency service queue;
[0155] The service downgrade operation includes at least one of the following: downgrading the target service from a high-frequency service queue to a medium-frequency service queue, and downgrading the target service from a medium-frequency service queue to a low-frequency service queue;
[0156] Among them, when the disk capacity of the node where the target service is located reaches a preset capacity threshold, the target service is removed from the low-frequency service queue.
[0157] In summary, in the embodiments of the present application, first, the service call frequency is monitored to ensure that the intelligent computing center can monitor the usage of services, providing data support for subsequent resource allocation and service level adjustment. Subsequently, based on the call frequency, the service level is automatically determined, including high-frequency, medium-frequency, low-frequency, and cacheless services. This level classification enables the intelligent computing center to adopt the most suitable operation mode for different services. Specifically, for high-frequency services, the intelligent computing center deploys multiple container instances to ensure high availability and low-latency response of the service by utilizing sufficient computing power resources. For medium-frequency services, a single container instance is used to process, ensuring the economy of computing power resource usage while meeting service requirements. For low-frequency services, the pre-cached resources in the intelligent computing center are used to quickly deploy container instances. Although instantiation may be required during response, the startup time is reduced by using cached resources. For cacheless services, the latest resources are directly pulled from the data source for instantiation to ensure the timeliness and accuracy of data. Generally, the intelligent computing center can dynamically adjust the allocation of computing power resources according to the actual needs of services, realizing the elastic scaling ability of computing power resources, avoiding service startup latency, improving the user experience, and avoiding waste of computing power resources, thereby improving the utilization rate of computing power resources in the intelligent computing center and the performance of high-performance computing power resources in the intelligent computing center.
[0158] Please refer to Figure 5 , the embodiments of the present application also provide an electronic device 50, including a processor 51, a memory 52, and a computer program stored on the memory 52 and executable on the processor 51. When the computer program is executed by the processor 51, it implements each process of the above-mentioned method embodiment for the elastic scaling ability of the computing power resources in the intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0159] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned method embodiment for the elastic scaling ability of the computing power resources in the intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0160] The embodiments of the present application also provide a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned method embodiment for the elastic scaling ability of the computing power resources in the intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0161] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device that includes a series of elements not only includes those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device that includes such element.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.
[0163] The embodiments of this application have been described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them fall within the protection scope of this application.
Claims
1. A method for supporting elastic scaling of computing resources in an intelligent computing center, characterized in that: The method comprises: Step S1: monitor the calling frequency of the target service; Step S2: judging the service level to which the target service currently belongs according to the calling frequency, wherein the service levels include: high-frequency service, medium-frequency service, low-frequency service and non-cached service; Step S3: running the target service according to the service level to which the target service currently belongs; The step S3 comprises: Step S31: When the service level to which the target service currently belongs is the high-frequency service, the target service is run based on at least two deployed container instances and computing resources on which the at least two container instances rely; Step S32: When the service level to which the target service currently belongs is the medium frequency service, the target service is run based on a deployed container instance and computing resources on which the container instance relies; Step S33: when the service level to which the target service currently belongs is the low-frequency service, running the target service based on the resources of the target service cached in the intelligent computing center; Step S34: When the service level of the target service is the cache-free service, pull the resources of the target service from the data source of the target service, deploy a container instance based on the pulled resources, and run the target service based on the container instance and the computing resources on which the container instance relies.
2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: monitoring the number of queries per second (QPS) of the target service; Step S12: Determine the calling frequency based on the QPS.
3. The method according to claim 1, characterized in that The step S2 comprises: Step S21: when the calling frequency is greater than a first preset threshold, determining that the service level to which the target service currently belongs is the high-frequency service; Step S22: when the calling frequency is greater than the second preset threshold and less than or equal to the first preset threshold, determining that the service level to which the target service currently belongs is the medium frequency service; Step S23: when the calling frequency is greater than the third preset threshold and less than or equal to the second preset threshold, determining that the service level to which the target service currently belongs is the low-frequency service; Step S24: When the calling frequency is less than or equal to the third preset threshold and greater than or equal to zero, determine that the service level to which the target service currently belongs is the cache-free service; wherein the third preset threshold is less than the second preset threshold and less than the first preset threshold.
4. The method according to claim 1, characterized in that: The step S31 further includes: Step S311: When the service level to which the target service currently belongs is the high-frequency service, every time the calling frequency of the target service increases by a preset number of times, the number of container instances in the at least two container instances increases by one.
5. The method according to claim 3, characterized in that: The step S3 further comprises: Step S35: According to the service level to which the target service currently belongs, a service queue is allocated for the target service, and the target service is run, wherein the service queue includes at least one of the following: a high-frequency service queue, a medium-frequency service queue, and a low-frequency service queue. If the service level to which the target service currently belongs is the cacheless service, the service queue of the target service is the low-frequency service queue.
6. The method according to claim 5, characterized in that The step S35 comprises: Step S351: Determine whether the target service has a previous service queue; Step S352: If the target service has a previous service queue, a service queue is reallocated for the target service according to the previous service queue and the current service level, and the target service is run.
7. The method according to claim 6, characterized in that The step S352 includes: Step S3521: If the target service has a previous service queue, then according to the previous service queue and the current service level, perform at least one of the following operations: service queue upgrade operation, service queue downgrade operation, and service queue level unchanged operation, and run the target service; The service upgrade operation includes at least one of the following: upgrading the target service from the low-frequency service queue to the medium-frequency service queue, upgrading the target service from the medium-frequency service queue to the high-frequency service queue; The service downgrade operation includes at least one of the following: downgrading the target service from the high-frequency service queue to the medium-frequency service queue, downgrading the target service from the medium-frequency service queue to the low-frequency service queue; When the disk capacity of the node where the target service is located reaches a preset capacity threshold, the target service is removed from the low-frequency service queue.
8. A device for supporting elastic scalability of computing resources in an intelligent computing center, characterized in that: The device comprises: Monitoring module, used to monitor the calling frequency of the target service; An execution module, configured to determine the service level to which the target service currently belongs according to the calling frequency, wherein the service level includes: high-frequency service, medium-frequency service, low-frequency service and non-cached service; Running the target service according to the service level to which the target service currently belongs; The execution module is further configured to run the target service based on at least two deployed container instances and computing resources supported by the at least two container instances when the service level to which the target service currently belongs is the high-frequency service; When the service level to which the target service currently belongs is the medium frequency service, running the target service based on a deployed container instance and computing resources on which the container instance relies; When the service level to which the target service currently belongs is the low-frequency service, running the target service based on the resources of the target service cached in the intelligent computing center; When the service level to which the target service currently belongs is the cache-free service, resources of the target service are pulled from the data source of the target service, a container instance is deployed based on the pulled resources, and the target service is run based on the container instance and the computing resources on which the container instance relies.
9. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a method for supporting elastic scaling capabilities of computing resources in an intelligent computing center as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for supporting elastic scaling capabilities of computing resources in an intelligent computing center as described in any one of claims 1 to 7.
11. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a method for supporting elastic scaling capabilities of computing resources in an intelligent computing center as described in any one of claims 1-7.
Citation Information
Cited By
Method and device for supporting elastic scalability of computing power resources of intelligent computing center
CN120560858A
Method and device for supporting elastic scaling capability of intelligent computing center computing power resources
CN120560858B