Memory resource scheduling method and system, computing equipment and electronic equipment

By identifying and classifying the memory usage characteristics of database instances and formulating targeted resource scheduling strategies, the problem of low memory resource scheduling efficiency in the existing technology is solved, and the precise, dynamic and differentiated management of resources is realized, and the stability and performance of services are improved.

CN119988039AActive Publication Date: 2025-05-13ALIBABA CLOUD COMPUTING CO LTD

Patent Information

Application Number
CN202510465804.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The scheduling efficiency of memory resources in the prior art leads to a decline in service availability and performance, increasing the risk of insufficient memory, and causing resource waste and cost increase.

Method used

By obtaining instances in the database, identifying the type of instances is divided according to the change of memory resources required during operation on the service platform, and determining the resource scheduling strategy corresponding to the instance type, and performing resource scheduling operations according to the policy.

Benefits of technology

Dynamic optimization of resource scheduling is realized, the scheduling efficiency and resource utilization of memory resources are improved, the risk of OOM is reduced, and the stability and performance of services are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988039A_ABST
    Figure CN119988039A_ABST
Patent Text Reader

Abstract

The invention discloses a memory resource scheduling method and system, computing equipment and electronic equipment, and relates to the field of database technologies and cloud computing. The method can comprise the steps that at least one instance in a database is obtained, the database is deployed on a service platform, and the instance is used for providing a service function of the service platform; the instance types of the instances are identified, and the instance types are divided according to the variation amplitude of memory resources needed in the running process of the instances on the service platform; resource scheduling strategies corresponding to the instance types are determined, different instance types correspond to different resource scheduling strategies, and the resource scheduling strategies are used for representing rules for scheduling memory resources from a service platform to instances; and executing resource scheduling operation on the instance according to the resource scheduling strategy. The technical problem that the scheduling efficiency of the memory resources is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of database technology and cloud computing, and more specifically, to a method, system, computing device, and electronic device for scheduling memory resources. Background Art

[0002] Currently, efficient management of memory resources is critical to ensure high performance, availability, and cost-effectiveness of services. As the tasks on databases expand and become more complex, databases face increasing resource demands, especially in terms of memory resource allocation and scheduling.

[0003] In related technologies, time series models can be used to predict future resource requirements based on historical memory usage data. In addition, in order to improve resource utilization, a fixed-ratio memory overselling strategy can be adopted, that is, the total amount of memory allocated to the instance is allowed to exceed a certain proportion of the actual physical memory. By adopting a unified resource scheduling and memory overselling strategy for instances in the database, an attempt is made to improve efficiency by simplifying the scheduling process. However, the above-mentioned one-size-fits-all strategy results in the inability to optimize resource allocation based on the memory usage characteristics of different instances. This not only affects the availability and performance of the service, increases the risk of Out of Memory (OOM), but also leads to waste of resources and rising costs. Therefore, there is still a technical problem of low scheduling efficiency of memory resources.

[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0005] Embodiments of the present application provide a method, system, computing device, and electronic device for scheduling memory resources to at least solve the technical problem of low scheduling efficiency of memory resources.

[0006] According to one aspect of an embodiment of the present application, a method for scheduling memory resources is provided. The method may include: obtaining at least one instance in a database, wherein the database is deployed on a service platform, and the instance is used to provide service functions of the service platform; identifying the instance type of the instance, wherein the instance type is divided according to the change range of the memory resources required during the operation of the instance on the service platform; determining a resource scheduling strategy corresponding to the instance type, wherein different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent the rules for scheduling memory resources from the service platform to the instance; and performing resource scheduling operations on the instance according to the resource scheduling strategy.

[0007] According to another aspect of the embodiment of the present application, a method for scheduling memory resources is provided. The method is applied to a resource management system and may include: obtaining at least one instance in a cloud database system, wherein the cloud database system is deployed on a cloud service platform, and the instance is used to provide service functions of the cloud service platform; identifying the instance type of the instance, wherein the instance type is divided according to the change range of the memory resources required during the operation of the instance on the cloud service platform; determining a resource scheduling strategy corresponding to the instance type, wherein different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent the rules for scheduling memory resources from the cloud service platform to the instance; and performing resource scheduling operations on the instance according to the resource scheduling strategy.

[0008] According to another aspect of the embodiment of the present application, another method for scheduling memory resources is provided. The method may include: obtaining the instance type of at least one instance in the database by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the instance type, the database is deployed on a service platform, and the instance is used to provide the service function of the service platform; identifying the instance type of the instance, wherein the instance type is divided according to the change range of the memory resources required during the operation of the instance on the service platform; determining the resource scheduling strategy corresponding to the instance type, wherein different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent the rules for scheduling memory resources from the service platform to the instance; performing resource scheduling operations on the instance according to the resource scheduling strategy to obtain a resource scheduling result; outputting the resource scheduling result by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the resource scheduling result.

[0009] According to another aspect of the embodiment of the present application, another memory resource scheduling system is provided. The system may include: a classifier for obtaining at least one instance in a database, wherein the database is deployed on a service platform, and the instance is used to provide service functions of the service platform; identifying the instance type of the instance, wherein the instance type is divided according to the change range of the memory resources required during the operation of the instance on the service platform; a scheduler for determining a resource scheduling strategy corresponding to the instance type, wherein different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent the rules for scheduling memory resources from the service platform to the instance; and performing resource scheduling operations on the instance according to the resource scheduling strategy.

[0010] According to another aspect of an embodiment of the present application, a computing device is also provided. The computing device may include a memory and a processor: the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions. When the above-mentioned computer-executable instructions are executed by the processor, any one of the above-mentioned methods is implemented.

[0011] According to another aspect of an embodiment of the present application, an electronic device is also provided. The electronic device may include a memory and a processor: the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions. When the above-mentioned computer-executable instructions are executed by the processor, any one of the above-mentioned methods is implemented.

[0012] According to another aspect of an embodiment of the present application, a processor is further provided, and the processor is used to run a program, wherein any one of the above methods is executed when the program is running.

[0013] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, the computer-readable storage medium including a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute any of the above methods.

[0014] According to another aspect of the embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the method of the embodiment of the present application is implemented.

[0015] In an embodiment of the present application, if memory resources need to be scheduled, an instance can be obtained from a database deployed on a service platform. The change service of the memory resources required for the process of running the acquired instance on the service platform can be analyzed to identify the instance type of the instance. Through the change range of the required memory resources, the differences between the instances can be accurately grasped, and the resource scheduling strategy can be avoided from being too unified, so as to achieve the purpose of setting scheduling strategies for instances of different instance types in a targeted manner. The resource scheduling strategy corresponding to the instance type can be determined, and the corresponding resource scheduling operation can be performed according to the resource scheduling strategy to schedule memory resources to the instance from the service platform. In this embodiment, the instance can be classified according to the change range of the memory resources required by different instances, and adaptive scheduling can be performed for different types of instances in a targeted manner. The scheduling method of memory resources in the database combined with instance classification and adaptive scheduling realizes the dynamic optimization of resource scheduling, thereby overcoming the limitations of historical data prediction, single binning algorithm, unified resource scheduling and memory overselling strategy in the related technology, realizing the precision, dynamic and differentiation of resource scheduling, significantly improving the scheduling efficiency and resource utilization of memory resources, while reducing OOM risks and ensuring the stability and performance of the service. This achieves the technical effect of improving the scheduling efficiency of memory resources and solves the technical problem of scheduling efficiency of memory resources.

[0016] It is easy to notice that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 is a schematic diagram of an application scenario of a method for scheduling memory resources according to an embodiment of the present application;

[0019] Figure 2 is a flowchart of a method for scheduling memory resources according to an embodiment of the present application;

[0020] Figure 3 is a flowchart of another method for scheduling memory resources according to an embodiment of the present application;

[0021] Figure 4 is a flowchart of another method for scheduling memory resources according to an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of a scheduling system for memory resources according to an embodiment of the present application;

[0023] Figure 6 is a schematic diagram of a cloud database memory overselling optimization system architecture based on instance classification and adaptive scheduling according to an embodiment of the present application;

[0024] Figure 7 is a flow chart of an instance classifier training method according to an embodiment of the present application;

[0025] FIG8( a ) is a schematic diagram of comparing the number of OOM errors according to an embodiment of the present application;

[0026] FIG8( b ) is a schematic diagram of a memory utilization comparison according to an embodiment of the present application;

[0027] Fig. 9 is a schematic diagram of host memory utilization according to an embodiment of the present application;

[0028] FIG10( a ) is a schematic diagram of an average distribution ratio according to an embodiment of the present application;

[0029] FIG10( b ) is a schematic diagram of an average utilization rate according to an embodiment of the present application;

[0030] FIG10( c ) is a schematic diagram of a utilization rate and the number of migration tasks according to an embodiment of the present application;

[0031] Fig.11 is a schematic diagram of a scheduling device for memory resources according to an embodiment of the present application;

[0032] Fig.12 is a schematic diagram of another memory resource scheduling device according to an embodiment of the present application;

[0033] Fig.13 is a schematic diagram of another memory resource scheduling device according to an embodiment of the present application;

[0034] Fig.14 is a structural block diagram of a computer terminal according to an embodiment of the present application;

[0035] Fig.15 is a block diagram of an electronic device according to a method for scheduling memory resources in an embodiment of the present application;

[0036] Fig.16 It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for scheduling memory resources according to an embodiment of the present application;

[0037] Fig.17 It is a structural block diagram of a computing environment of a method for scheduling memory resources according to an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0040] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0041] Memory Over-Subscription refers to a strategy in which the total amount of memory allocated to an instance exceeds the physical memory in a cloud computing system to improve resource utilization.

[0042] Service Level Objective (SLO) is a service availability and performance indicator agreed upon between the service provider and the customer, usually expressed as a percentage of service availability over a specific period of time.

[0043] Out of memory (OOM), an error that occurs when the system or application's memory requirements exceed the available physical memory, which may result in instance termination or service interruption;

[0044] Transient Instances: Instances with large and unpredictable memory usage fluctuations, which can easily lead to OOM errors;

[0045] Steady Instance: An instance with relatively stable memory usage and fluctuations within a controllable range.

[0046] Adaptive Instance Scheduler, a scheduling algorithm that combines multiple instance sizing methods to dynamically optimize resource scheduling based on instance characteristics;

[0047] Max User Reservation (MUR), a memory reservation policy that determines the total memory reservation based on the maximum potential memory demand of a single user;

[0048] Fallback mechanism is an emergency processing mechanism. When the system detects that the memory utilization is too high or the prediction is wrong, it triggers corresponding measures (for example, instance migration) to prevent OOM errors and ensure system stability.

[0049] According to an embodiment of the present application, a method for scheduling memory resources is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0050] The above-mentioned memory resource scheduling method provided in the embodiment of the present application can be applied to the following Figure 1 The application scenarios shown are not limited to these. Figure 1 is a schematic diagram of an application scenario of a method for scheduling memory resources according to an embodiment of the present application. Figure 1In the application scenario shown, the instance classification model capable of allocating instances is deployed in a database 40 associated with a server 30, and the server may be a cloud. The server 30 may be connected to one or more terminal devices 10 via a local area network connection, a wide area network connection, an Internet connection, or other types of data networks, wherein the terminal device 10 may be a client device, and the client device here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc., and the client devices together constitute a client relative to the server. An operation interface may be deployed on a graphical user interface on the client device. The terminal device 10 may interact with the user through the operation interface to implement the call of the instance classification model, thereby implementing the scheduling method of memory resources provided in the embodiment of the present application. Information exchange may be performed between the server 30 and the terminal device 10 through the network 20. An instance classification model may be deployed in the database 40 in the server, and in the embodiment of the present application, an instance classification model capable of classifying instances in the database in the present application is stored in the database 40. It can be used to call through the server 30 to identify the instance type corresponding to the acquired instance.

[0051] In an embodiment of the present application, a system consisting of a terminal device, a network, and a server can perform the following steps: Step S102, obtaining at least one instance in a database; Step S104, identifying the instance type of the instance; Step S106, determining a resource scheduling strategy corresponding to the instance type; Step S108, performing a resource scheduling operation on the instance according to the resource scheduling strategy. The above method can indirectly serve the requests and operations of the terminal device 10 (client) by memory scheduling. Memory resource scheduling can also adjust the resource classification strategy of the instance according to the frequency, type, and scale of the request of the terminal device 10. By optimizing memory resource scheduling through the above method, the database can better respond to the request of the terminal device 10, provide more stable and efficient services, and thus improve user experience and satisfaction. In short, the optimization of resource scheduling ultimately serves the terminal device 10 and improves its access quality and performance to database services.

[0052] In this embodiment, instances can be classified according to the change range of memory resources required by different instances, and adaptive scheduling can be carried out for different types of instances. Through the scheduling method of memory resources in the database that combines instance classification and adaptive scheduling, dynamic optimization of resource scheduling is achieved, thereby overcoming the limitations of related technologies based on historical data prediction, single packing algorithm, unified resource scheduling and memory overselling strategy, etc., and achieving precise, dynamic and differentiated resource scheduling, significantly improving the scheduling efficiency and resource utilization of memory resources, while reducing OOM risks and ensuring the stability and performance of the service. Thereby achieving the technical effect of improving the scheduling efficiency of memory resources and solving the technical problem of scheduling efficiency of memory resources.

[0053] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the instance classification model, the embodiments of the present application can be carried out in the client device.

[0054] Under the above operating environment, this application provides Figure 2 It should be noted that the memory resource scheduling method of this embodiment can be implemented by Figure 1 The illustrated embodiment is executed by a mobile terminal. Figure 2 is a flowchart of a method for scheduling memory resources according to an embodiment of the present application, such as Figure 2 As shown, the method may include the following steps:

[0055] Step S202, obtaining at least one instance in the database.

[0056] In the technical solution provided in the above step S202 of the present application, the database is deployed on the service platform, and the database can be a cloud database system, referred to as a cloud database. The service platform can be a server. The service platform can be a physical server, a virtualization platform, or a containerized environment. In the embodiment of the present application, it can be a cloud service platform (referred to as a cloud platform), or it can be called a cloud computing system. The instance can be used to provide the service functions of the service platform. The service functions provided by the instance can be cloud database services. In the present application, it can refer to a database instance (Database Instance, referred to as DBInstance), which is the basic unit of cloud database services. Each instance has its own unique configuration resources, including resources such as a central processing unit (CPU), memory, and disk space to support its operation and provide specific service functions. The database instance can be a representation of a single database server or a node in a database cluster. The database instance is responsible for executing core database operations such as querying, storing, and retrieving data in Structured Query Language (SQL), and is the direct provider of cloud database services to users.

[0057] Optionally, cloud database services include but are not limited to data storage, query processing, transaction management, data backup and recovery, data migration, performance optimization, security assurance and other functions. The realization of the above service functions depends on sufficient and reasonable resource allocation, especially memory, because memory is a key factor affecting database operation efficiency. This is only an example and does not specifically limit the functions covered by cloud database services.

[0058] In this embodiment, instances can be obtained from a database deployed on a service platform, that is, database instances that require resource management and optimization are identified and locked.

[0059] Optionally, before any resource scheduling is performed, the object of resource management, i.e., the database instance, can be clarified. This embodiment is the starting point of the resource allocation and optimization process, and provides a basis for the formulation of subsequent resource analysis and scheduling strategies. Each instance carries different data and task functions and is a direct responder to user requests. Ensuring the stable operation of the instance and efficient resource utilization is the key to improving the overall service quality.

[0060] Optionally, the operation of obtaining instances from the database of the service platform can proactively identify and lock database instances that require resource management and optimization for the database.

[0061] For example, a series of monitoring mechanisms can be run in the background of the service platform to monitor the operating status and resource usage of each instance in the database. For example, runtime features such as memory usage, CPU utilization, and disk input / output (I / O) can be monitored. Non-runtime features such as the role of the instance (for example, master instance, slave instance), and service level can also be monitored. The operating status of the instance can also be monitored, such as whether it is under high load, whether out of memory OOM and other events occur. The above information can be obtained in a variety of ways, such as through the management system of the service platform, monitoring components, or directly extracted from the operating environment of the instance. The above runtime features and non-runtime features can be collectively referred to as runtime metrics.

[0062] It should be noted that the above method and process of obtaining an instance and monitoring various information of the instance are only for illustration and are not specifically limited here.

[0063] Step S204: Identify the instance type of the instance.

[0064] In the technical solution provided in the above step S204 of the present application, the instance type can be obtained by dividing according to the change range of the memory resources required by the instance during the operation on the service platform, that is, the instance type can be used to indicate the fluctuation of the memory resource (referred to as memory) usage of the corresponding instance during the operation, for example, the fluctuation is large or the memory usage is relatively stable. The instance type can also be called the instance classification result, which can include transient instances and stable instances.

[0065] In this embodiment, after the instance in the database is obtained, the change range of the memory resources required by the instance during its operation on the service platform can be analyzed to determine the instance type of the instance.

[0066] Optionally, this embodiment is an important step in the database memory overselling optimization process based on instance classification and adaptive scheduling. The goal of this embodiment is to classify instances according to their memory usage characteristics. The above classification process lays the foundation for the subsequent adaptive resource scheduling strategy formulation.

[0067] Optionally, the instance type is identified mainly based on the memory usage fluctuation of the instance, which can be achieved by analyzing the runtime characteristics of the instance, such as the frequency of changes in memory usage, the magnitude of changes, the regularity of peak occurrence, etc., as well as possible non-runtime characteristics, such as the service type of the instance, the amount of pre-allocated resources, etc.

[0068] It should be noted that the characteristics of the above-mentioned analysis instance types are only for illustration and are not specifically limited here. As long as different resource scheduling strategies can be formulated for different instances to improve the efficiency of memory resource scheduling, the processes and methods are within the protection scope of the embodiments of this application.

[0069] Optionally, instances can be divided into transient instances and stable instances. Transient instances refer to instances with large and unpredictable memory usage fluctuations, while stable instances refer to instances with relatively stable memory usage patterns and fluctuations within a controllable range. The purpose of classification is to enable subsequent resource scheduling to adopt different targeted strategies to balance resource utilization and service reliability.

[0070] For example, clustering algorithms can be used to perform cluster analysis on the runtime characteristics of instances, and instances with similar memory usage patterns can be classified into the same cluster, while instances in different clusters have different usage patterns. By setting the appropriate number of clusters and cluster centers, instances can be automatically divided into transient and stable categories. With the above method, there is no need to define labels for transient and stable instances in advance, and instance usage patterns can be automatically discovered from the data and classified.

[0071] For another example, we analyze the memory usage rate or density distribution of memory usage of instances, and use statistical methods such as histogram analysis and kernel density estimation to identify abnormal points or long tails in the distribution, and mark the abnormal instances as transient instances, while the instances near the center of the distribution are considered stable instances. The above methods can effectively identify transient instances that are significantly different from stable instances without the need for complex model training.

[0072] As an optional example, anomaly detection algorithms and time series analysis can be combined to identify transient instances. The time series model is used to predict the normal memory usage pattern of the instance, and the anomaly detection algorithm is used to identify the instances that deviate from the prediction model in actual use and mark them as transient instances; while the instances that are consistent with the prediction model are considered stable instances. Through the above method, the accuracy of the prediction model and the sensitivity of anomaly detection are combined to capture the abnormal behavior of transient instances at different time scales.

[0073] It should be noted that the above-mentioned method and process for classifying instances are only for illustration and are not specifically limited here. Different methods have their own advantages and limitations, and the choice of method depends on the specific application scenario, data characteristics, and resource constraints. In actual deployment, multiple methods can be compared and evaluated to select an instance classification scheme suitable for the current environment. In this application, an instance classifier is also proposed to provide an accurate and robust instance classification method by combining feature engineering, Markov chain model, and iterative pseudo-label training.

[0074] In an embodiment of the present application, by classifying the instances, a higher priority can be given to transient instances in resource scheduling to ensure that their memory requirements are met and to avoid the occurrence of OOM errors. For stable instances, since the memory usage pattern of stable instances is more stable, a more aggressive memory overselling strategy can be adopted in resource scheduling to improve resource utilization. Through the above method, the instance type is accurately identified, laying the foundation for the subsequent formulation of resource scheduling strategies. Through the above instance classification, the utilization of memory resources is optimized, while the OOM risk is reduced to ensure the achievement of SLO. The ability to more accurately manage transient instances and stable instances in the database provides key information for adaptive resource scheduling, which helps to achieve the dual goals of maximizing resource utilization and optimizing service reliability.

[0075] Step S206: determine a resource scheduling strategy corresponding to the instance type.

[0076] In the technical solution provided in the above step S206 of the present application, different instance types correspond to different resource scheduling strategies. The resource scheduling strategy can be used to represent the rules for scheduling memory resources from the service platform to the instance.

[0077] In this embodiment, after the instance type of the instance is identified, a resource scheduling policy corresponding to the instance type may be determined.

[0078] Optionally, different resource scheduling strategies are formulated based on the identified instance types, namely transient instances and stable instances, to achieve the dual goals of efficient resource utilization and service stability.

[0079] Optionally, due to the large and unpredictable fluctuations in memory usage of transient instances, their resource scheduling strategies focus more on risk control and ensuring service stability. Therefore, the resource scheduling strategies corresponding to the transient instances may include: reserving more resources, especially memory, for transient instances to cope with the sudden increase in their demand and avoid OOM errors. Even if the memory of transient instances is allowed to be oversold, a lower oversold rate will be set to control risks and ensure that the instances can operate normally under high load. In the case of tight resources, transient instances are scheduled first to ensure that their memory requirements are met to prevent service interruptions. The memory usage of transient instances is monitored in real time. Once it is found that resource usage exceeds the preset threshold, measures are taken immediately, such as dynamically increasing resources or triggering the Fallback mechanism to avoid OOM errors. When the memory demand of transient instances increases, the instance scale is automatically expanded or the number of instances is increased to provide additional resources to meet the demand.

[0080] Optionally, compared with transient instances, the memory usage pattern of stable instances is relatively stable, and its resource scheduling strategy tends to improve resource utilization and cost-effectiveness. Therefore, the resource scheduling strategy corresponding to stable instances may include: Since the memory usage of stable instances is stable, a higher memory overselling rate can be set to improve resource utilization and maximize the economic benefits of the service platform. Use multiple methods such as the Ratio & Usage Method, the Percentile Method, the Stochastic Bin Packing (SBP) algorithm, and the Time Series Forecasting (TSF) to estimate the size of stable instances to more accurately predict their memory requirements. Use the packing algorithm, such as the First-Fit algorithm, to optimize and schedule as many stable instances as possible to the same physical machine to improve resource utilization. Under the premise of ensuring service stability, prioritize cost-effectiveness, reduce resource waste through intelligent scheduling, and reduce the cost of cloud service providers and users. Dynamically adjust the scheduling strategy according to the actual memory usage of stable instances, such as real-time adjustment of instance size, resource reservation, etc., to adapt to changing workloads.

[0081] It should be noted that the resource scheduling strategies formulated for different instance types are only examples and are not specifically limited here. The scheduling strategy for transient instances focuses more on risk control, ensuring sufficient resources to avoid OOM errors, while the strategy for stable instances focuses more on improving resource utilization, and improving the economic benefits of the service platform through methods such as memory overselling. For transient instances, a conservative resource scheduling strategy is adopted to reserve more resources; while for stable instances, a more aggressive resource scheduling strategy is adopted to try to maximize resource utilization. The resource scheduling strategy for transient instances needs to respond to changes in memory usage in real time, while the strategy for stable instances can be adjusted based on preset rules and models to reduce the need for real-time response.

[0082] In the embodiment of the present application, through the above method, differentiated management can be adopted according to the characteristics of different instance types, which not only ensures the stability of the service, but also improves the utilization of resources and realizes the optimization of resource management. In actual deployment, the above adaptive resource scheduling method can significantly reduce the occurrence rate of OOM errors, while improving the overall performance and cost-effectiveness of the service platform.

[0083] Step S208: Execute resource scheduling operations on the instance according to the resource scheduling policy.

[0084] In the technical solution provided in the above step S208 of the present application, after determining the resource scheduling strategy corresponding to the instance type, the resource scheduling operation can be performed on the instance according to the resource scheduling strategy. That is, according to the determined resource scheduling strategy of different instance types, specific resource scheduling operations are implemented to achieve optimal allocation and management of resources.

[0085] Optionally, after identifying and classifying instances, a specific scheduling policy will be applied to each type of instance. For transient instances, the policy may be more conservative to ensure sufficient memory and avoid OOM errors; while for stable instances, the scheduling policy may be more aggressive to make full use of resources and improve resource utilization.

[0086] Optionally, based on the classification results and scheduling policies, the resource allocation of the instance can be adjusted, for example, increasing or decreasing the memory quota of the instance, adjusting the CPU allocation of the instance, migrating the instance to a more suitable server, etc., to meet the resource requirements of the instance and optimize the overall resource usage.

[0087] Optionally, resource scheduling operations may include, but are not limited to: migrating instances from the current host to a host with more sufficient resources to avoid resource contention and OOM risks. Dynamically adjusting resource reservations based on instance type and current system status to ensure service stability and resource efficiency. Adjusting instance size based on predicted resource demand to optimize resource usage while ensuring service quality. When resource demand changes dramatically, instances may need to be created or destroyed quickly to meet service needs or reduce resource waste.

[0088] In the embodiment of the present application, by executing the above method, the results of previous classification and policy formulation can be converted into actual resource scheduling operations, realizing intelligent management and optimization of resources. The above adaptive resource scheduling method, combined with a deep understanding of different instance types, can more efficiently utilize the resources of the service platform, while reducing the risk of OOM errors and improving service stability and user experience.

[0089] Through the above steps S202 to S208 of the present application, if it is necessary to schedule memory resources, an instance can be obtained from the database deployed on the service platform. The change service of the memory resources required for the process of running the acquired instance on the service platform can be analyzed to identify the instance type of the instance. Through the change range of the required memory resources, the differences between the instances can be accurately grasped, and the resource scheduling strategy can be avoided from being too unified, so as to achieve the purpose of setting scheduling strategies for instances of different instance types in a targeted manner. The resource scheduling strategy corresponding to the instance type can be determined, and the corresponding resource scheduling operation can be performed according to the resource scheduling strategy to schedule memory resources to the instance from the service platform. In this embodiment, the instances can be classified according to the change range of the memory resources required by different instances, and adaptive scheduling can be performed in a targeted manner for different types of instances. Through the scheduling method of memory resources in the database that combines instance classification and adaptive scheduling, dynamic optimization of resource scheduling is achieved, thus overcoming the limitations of related technologies based on historical data prediction, single binning algorithm, unified resource scheduling and memory overselling strategy, etc., and achieving precise, dynamic and differentiated resource scheduling, significantly improving the scheduling efficiency and resource utilization of memory resources, while reducing OOM risks and ensuring the stability and performance of services. In addition, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of scheduling efficiency of memory resources is solved.

[0090] The above method of this embodiment is further introduced below.

[0091] As an optional implementation, step S204, identifying the instance type of the instance, includes: using an instance classification model to classify the instance to obtain the instance type, wherein the instance classification model is obtained by training a Markov chain model using instance type labels of instance samples.

[0092] In this embodiment, in the process of identifying the instance type of the instance, the instance classification model can be used to classify the instance to obtain the instance type of the instance. Among them, the instance classification model can be obtained by training a Markov Chain Model using the instance type label of the instance sample. The instance type label corresponds to the instance type and can include a transient instance label and a stable instance label. The Markov chain model can be a Markov chain state transition model, which can be used to improve the classification effect of the instance. The instance classification model can be an instance classifier, referred to as a classifier.

[0093] Optionally, a Markov chain model, in particular a Markov chain state transition model, is used in combination with instance type labels of instance samples for training to improve classification effect and accuracy.

[0094] Alternatively, a Markov chain model is a statistical model that can be used to describe how a system state changes over time, assuming that the future state of the system depends only on the current state and has nothing to do with the past state. In the cloud database instance classification scenario, the Markov chain model is used to model the transition of instance states, where the state can be a transient instance state or a stable instance state.

[0095] Optionally, a large number of instance samples are collected, each of which contains runtime features (e.g., memory usage, CPU utilization, etc.) and non-runtime features (e.g., instance role, customer service level, etc.) of the instance, as well as manually annotated instance type labels. The above instance type labels reflect the real state of the instance samples and are the basis for model training.

[0096] Optionally, the Markov chain model can be used to analyze the transition probability between instance states, that is, the probability of transitioning from a transient state to a stable state or vice versa. By training the Markov chain state transition model, the transition pattern between different states can be learned, which is of great value for identifying transient instances and stable instances. During the model training process, an iterative pseudo-label training method can also be used, that is, the labels predicted by the model are used as part of the training data, combined with the manually annotated labels, to continuously optimize the classification ability of the model, especially when dealing with transient instances, which can improve the sensitivity and accuracy of the model.

[0097] Optionally, the Markov chain state transition model takes into account the time dependency of instance states, enabling the classification model to classify based on the current and historical states of instances, rather than relying solely on current runtime features. This helps capture the dynamic changes in instance behavior and improves the classifier's ability to identify transient and stable instances. The iterative training process can further enhance the performance of the model, especially when dealing with class imbalance problems, reducing the misclassification of transient instances and improving classification accuracy.

[0098] Optionally, once the trained instance classification model is obtained, the real-time or newly collected instances can be classified and processed to accurately identify transient instances and stable instances. The above instance types (classification results) will be used to determine the resource scheduling strategy and subsequent resource scheduling operations.

[0099] Optionally, the input samples for training the Markov chain model can be determined by the following method: collect a certain number of labeled (i.e., known instance types) instance samples (also referred to as initial label samples) as the initial data set for model training. The above instance samples can cover various types of instances, including transient instances and stable instances, to ensure that the Markov chain model can learn the characteristics of different types of instances. In addition, a large number of unlabeled instance samples are collected, and the unlabeled instance samples can be used to generate pseudo labels, and the performance of the Markov chain model is further optimized through iterative training. The unlabeled instance samples contain the runtime features and non-runtime features of the instance, which are used for the prediction of the Markov chain model and the subsequent pseudo label generation.

[0100] Optionally, in order to apply the Markov chain model, the state of the instance needs to be defined. For example, the memory usage of the instance can be divided into three states: low, medium, and high, or a more refined state division, such as by percentage segments. The definition of the state should be based on the memory usage of the instance at different points in time. Calculate the transition probability between each state. For example, you can count the frequency of instances transitioning from one state to another, and then calculate the transition probability. For example, if an instance has a 20% probability of transitioning to a medium memory state when it is in a low memory state, then the transition probability from low to medium is 0.2. The above transition probabilities can form a state transition matrix, where each row of the matrix represents the current state of the instance, each column represents the next possible state, and the matrix elements are the corresponding state transition probabilities.

[0101] Optionally, after the Markov chain model is trained with the initial label samples, the Markov chain model can be used to predict the unlabeled samples to obtain the prediction results, and the prediction results are used as pseudo labels. Since the Markov chain model trained with the initial label samples may not be very accurate, the generated pseudo labels will have a certain error rate. Through iterative training, the accuracy of instance classification can be gradually improved. The pseudo labels are merged with the original initial label samples as a new training data set, and the above Markov chain model is trained again. The above process can be repeated many times, and after each iteration, the prediction performance of the Markov chain model will be improved because the Markov chain model is constantly learning and correcting its own errors.

[0102] Optionally, a stop criterion is needed to decide when to stop iterative training. For example, based on the convergence of the performance of the Markov chain model, if the generation of pseudo labels no longer significantly affects the performance of the Markov chain model, or when the validation set accuracy and recall rate of the Markov chain model reach a predetermined threshold, the iteration can be stopped to obtain the final trained Markov chain model as the instance classification model.

[0103] In an embodiment of the present application, accurate classification of cloud database instances is achieved by introducing a Markov chain state transition model and training in combination with instance type labels of instance samples. The above method takes into account the time dependency of instance states, improves the performance of the classification model, and has significant advantages in identifying transient instances and stable instances. The classification results will directly guide subsequent resource scheduling strategies and are a key step in achieving efficient resource utilization and service stability. Through continuous model training and optimization, the accuracy of instance classification and the overall resource management efficiency of the system can be further improved.

[0104] As an optional implementation, the instance type label includes a transient instance type label and a stable instance type label, the transient instance type label is used to indicate that the change value of the memory utilization of the instance sample exceeds the change threshold, and the stable instance type label is used to indicate that the change value of the memory utilization of the instance sample does not exceed the change threshold, and the instance classification model is used to classify the instance to obtain the instance type, including: extracting runtime features and non-runtime features from the instance, wherein the runtime features at least include the memory utilization of the instance, and the non-runtime features at least include the memory allocation of the instance; using the instance classification model, classifying the runtime features and the non-runtime features to obtain the instance type, wherein the instance type is a transient instance type or a stable instance type, the transient instance type is used to indicate that the change value of the memory utilization of the instance exceeds the change threshold, and the stable instance type is used to indicate that the change value of the memory utilization of the instance does not exceed the change threshold.

[0105] In this embodiment, in the process of classifying the instance using the instance classification model to obtain the instance type, the non-runtime features and runtime features of the instance can be extracted from the instance. The instance type can be obtained by classifying using the runtime features and non-runtime features of the instance classification model. Among them, the instance type label can include a transient instance type label and a stable instance type label. The transient instance type can be used to indicate that the change value of the memory utilization of the instance sample exceeds the change threshold. The stable instance type label can be used to indicate that the change value of the memory utilization of the instance sample does not exceed the change threshold. The stable instance indicated by the stable instance type label can also be called a steady-state instance. The instance type can include a transient instance type or a stable instance type. The transient instance type can be used to indicate that the change value of the memory utilization of the instance exceeds the change threshold. The stable instance type can be used to indicate that the change value of the memory utilization of the instance does not exceed the change threshold.

[0106] Alternatively, runtime features and non-runtime features can be used to fully describe the behavior and properties of database instances at different stages. Each represents different aspects of the instance's running process and configuration settings, and plays a vital role in instance classification and resource scheduling decisions.

[0107] Optionally, runtime features mainly reflect the dynamic behavior and resource usage of the instance at runtime. Runtime features are real-time or historical measurement data that can provide instant information on instance load and behavior patterns. Runtime features usually include but are not limited to the following aspects: memory utilization, CPU utilization, disk I / O utilization, buffer pool hit rate, transactions per second (TPS) or queries per second (QPS), network bandwidth utilization, and instance restart events. Runtime features are used to represent the resource usage and behavior patterns of instances during operation, and are a key data source for identifying transient instances and stable instances. By analyzing runtime features, you can capture the dynamic changes of instances, predict future resource requirements, and provide real-time decision-making basis for resource scheduling.

[0108] Optionally, non-runtime features focus more on the attributes that have been determined when the instance is started and configured. The above non-runtime features remain unchanged during the operation of the instance and provide static configuration information of the instance. Non-runtime features usually include but are not limited to the following aspects: CPU allocation, memory allocation, disk allocation, instance role, customer service level and instance type. Non-runtime features are used to represent the static configuration information and background attributes of the instance, and are an important part of understanding the resource requirements and classification of the instance. The above features provide the initial configuration information of the instance in the instance classification model training, especially when the instance classification and resource scheduling strategy are formulated, the non-runtime features provide a priori information on the basic resource requirements and potential behaviors of the instance.

[0109] Optionally, the instance type tag is used to distinguish the type of instance, i.e., transient instance and stable instance. The above instance type tag is determined based on whether the change value of the instance's memory utilization exceeds a preset change threshold. For the transient instance type tag, if the change value of the instance's memory utilization exceeds the change threshold within a certain time window, the instance will be marked as a transient instance. This indicates that the instance's memory demand has large fluctuations, which may cause OOM errors. For the stable instance type tag, if the change value of the instance's memory utilization does not exceed the change threshold, the instance will be marked as a stable instance. This indicates that the instance's memory demand is relatively stable and the fluctuation is within a controllable range.

[0110] In the embodiment of the present application, the runtime features and non-runtime features are used in combination to more comprehensively describe the behavior and resource requirements of the instance, improve the classification accuracy of the instance classification model and the precision of resource scheduling. The runtime features provide real-time dynamic information of the instance, and the non-runtime features provide the static configuration background of the instance. The combination of the two can better predict the future resource requirements and potential risks of the instance, and provide a richer data foundation for adaptive resource scheduling and instance classification.

[0111] As an optional implementation, runtime features and non-runtime features are classified and processed using an instance classification model to obtain instance types, including: converting runtime features and non-runtime features into target features that match the instance classification model; and classifying the target features using the instance classification model to obtain instance types.

[0112] In this embodiment, in the process of classifying the runtime features and the non-runtime features using the instance classification model, the runtime features and the non-runtime features can be converted into target features that match the instance classification model. The instance classification model can be used to classify the target features to obtain the instance type.

[0113] Optionally, feature engineering includes extracting, building and selecting features from raw data to facilitate model training and prediction. In an embodiment of the present application, feature engineering includes steps such as runtime feature extraction, non-runtime feature extraction and feature selection. Among them, for runtime feature extraction, runtime features reflect the real-time or historical resource usage of the instance during operation, including but not limited to memory utilization, CPU utilization, disk I / O utilization, etc. The above features can reveal the dynamic behavior and resource demand pattern of the instance at runtime. For non-runtime feature extraction, non-runtime features generally refer to attributes that have been determined in the startup or configuration phase of the instance, such as memory, CPU, disk allocation, instance type (master node or slave node), customer service level, etc. The above features provide static information on instance resource allocation, which helps to understand the resource demand background of the instance. For feature selection, from the extracted features, select features that have a significant impact on instance type classification, usually involving feature importance evaluation, correlation analysis and other technologies to ensure the efficiency of model training and the accuracy of classification.

[0114] Optionally, the instance classification model is used to convert and classify the runtime features and non-runtime features of the instance to obtain the instance type. The above process includes two key steps: feature conversion and model application.

[0115] Optionally, feature conversion is the process of converting raw data (runtime features and non-runtime features) into a format that can be understood and processed by the instance classification model, that is, the target feature. The purpose of the above method is to ensure that the feature data can match the instance classification model and improve the training efficiency and classification accuracy of the model. The raw data may contain noise, missing values ​​and outliers. Data preprocessing such as data cleaning, missing value filling and outlier processing can be performed to ensure the quality of feature data. Normalize or standardize numerical features to scale feature values ​​to the same range to avoid the impact caused by different dimensions between features and improve the stability of model training. Encode classification features (for example, instance type, customer service level) and convert classification features into numerical features that can be recognized by the model. Based on the original features, new features are combined or derived, for example, the correlation between memory utilization and CPU utilization is calculated, or new features are constructed based on non-runtime features and runtime features to reflect the comprehensive behavioral characteristics of the instance.

[0116] Optionally, once feature conversion is completed and target features matching the model are obtained, the instance classification model can be used for classification processing to obtain the type of instance (transient or stable). Select a machine learning model suitable for processing classification tasks, such as eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Random Forest, etc., and use a training dataset with instance type labels for model training. The goal of model training is to learn the mapping relationship between features (runtime features and non-runtime features) and instance types. After the instance classification model training is completed, the instance classification model can be used to classify and predict real-time or newly collected instance data. The instance classification model receives the target features as input and outputs the predicted instance type. The classification result (transient or stable instance) will be used for subsequent resource scheduling strategies. For transient instances, a conservative memory overselling strategy can be adopted to reduce OOM risk; for stable instances, a higher memory overselling rate can be tried to improve resource utilization.

[0117] In the embodiment of the present application, feature conversion ensures that the original feature data can be input into the model in the most appropriate form, thereby improving the training efficiency and prediction accuracy of the instance classification model. The application of the instance classification model realizes the automatic identification of the instance type, providing a scientific basis for resource scheduling. By combining runtime features and non-runtime features, the dynamic behavior and static configuration information of the instance in operation can be fully described, so that the classification model can more accurately identify transient instances and stable instances, thereby guiding the formulation and optimization of resource scheduling strategies, and achieving a balance between resource utilization and service stability. The above process embodies the idea of ​​data-driven and intelligent decision-making, and automatically identifies instance types by processing complex data features through machine learning technology, bringing intelligent and automated advantages to the resource management of cloud databases. The step of converting features to target features ensures the quality and format of the model input data, which is the basis and prerequisite for improving the performance of the classification model. Through continuous model training and optimization, as well as improvements to the feature conversion method, the accuracy of instance classification can be continuously improved, the resource scheduling strategy of the cloud database can be further optimized, and the overall performance and reliability of the system can be improved.

[0118] As an optional implementation, the method also includes: using transient instance type labels and stable instance type labels to train an initial Markov chain model; using the trained initial Markov chain model to predict the instance type of unlabeled instance samples to obtain prediction results; generating pseudo-instance type labels based on the prediction results and the state transition matrix of the instance samples, wherein the state transition matrix is ​​used to represent the probability of the instance samples performing state transitions between adjacent time periods; using the pseudo-instance type labels to train the trained initial Markov chain model to obtain a Markov chain model.

[0119] In this embodiment, the initial Markov chain model can be trained using transient instance type labels and stable instance type labels. Using the trained initial Markov chain model, the instance type of unlabeled unlabeled instance samples is predicted to obtain a prediction result. A pseudo instance type label can be generated based on the prediction model and the state transition matrix of the instance sample. The trained initial Markov chain model can be trained using the pseudo instance type label to obtain a Markov chain model. Among them, the initial Markov chain model can also be referred to as an initial model. The prediction result can be the result obtained by predicting the instance type of the unlabeled instance sample, and can also be referred to as the initial model prediction result. The pseudo instance type label can be a new pseudo label. The state transition matrix can be used to represent the probability of the instance sample performing state transition between adjacent time periods, and is used to utilize the time dependency of the instance state to improve the classification effect.

[0120] Optionally, in the cloud database memory overselling optimization method based on instance classification and adaptive scheduling in the embodiment of the present application, the Markov chain model plays an important role, especially in improving the accuracy of transient instance recognition and capturing instance state changes.

[0121] Optionally, based on existing data (e.g., runtime features and non-runtime features of instances), instances can be classified into transient instances and stable instances, and marked with corresponding type labels. For example, transient instances can be marked with transient instance type labels, and stable instances can be marked with stable instance type labels. The above type labels will be used to train the initial Markov chain model so that it can learn the transition probabilities of different instance types in different states.

[0122] Optionally, the Markov chain model is a statistical model used to describe the probability of a system state changing over time. In an embodiment of the present application, the Markov chain model learns the probability of an instance transferring from a state at one point in time (e.g., memory usage, CPU utilization) to a state at the next point in time, thereby helping to capture the time dependency of instance behavior. That is, through transient instance type labels and stable instance type labels, the initial Markov chain model can learn the changes of each instance over time, and obtain a trained initial Markov chain model.

[0123] Optionally, after the initial Markov chain model is trained, the trained initial Markov chain model can be used to predict the unlabeled instance samples, that is, to estimate the probability of the unlabeled instance samples belonging to transient instances or stable instances, and obtain corresponding prediction results. The prediction results are crucial for further optimizing the trained initial Markov chain model and generating pseudo labels.

[0124] Optionally, a pseudo instance type label can be generated using the prediction results and the state transition matrix of the instance sample. The state transition matrix records the probability of an instance transitioning between adjacent time periods. By combining it with the prediction results, it can be determined whether the state change of an unlabeled instance in different time periods is more inclined to the behavior pattern of a transient instance or a stable instance.

[0125] Optionally, the generation of pseudo labels may involve setting a threshold for the prediction probability. For example, if the probability of predicting that an instance is a transient instance exceeds a certain threshold, it is marked as a transient instance. The above method can expand the training data set and improve the generalization ability of the initial Markov chain model without relying entirely on manual labeling.

[0126] Optionally, the initial Markov chain model can be trained again using the generated pseudo instance type labels. The above process is called iterative refinement or iterative pseudo label training. By introducing more unlabeled instance samples and state transition information of unlabeled instance samples, the initial Markov chain model can learn richer and more detailed instance behavior patterns, especially in capturing complex state changes of transient instances, which can significantly improve the recognition ability of the model.

[0127] Optionally, iterative training can continue until the prediction results become stable or meet specific performance indicators. The above process helps the Markov chain model to gradually converge and improve classification accuracy, especially when dealing with class imbalance problems. By continuously strengthening the recognition of minority classes (transient instances), the recall rate can be significantly improved, thereby reducing the probability of OOM.

[0128] In an embodiment of the present application, a pseudo-label is generated by combining the dynamic behavior (state transfer matrix) of the instance and the prediction result of the initial Markov chain model, so that the model is further optimized using unlabeled data based on limited labeled data. The Markov chain model can capture the time dependency of the instance state, and the iterative optimization improves the accuracy and robustness of the model in identifying transient instances by continuously learning the behavior pattern of the instance. The above-mentioned iterative training method based on pseudo-labels is particularly effective in dealing with the problem of class imbalance, because pseudo-labels can pay more attention to a few categories (transient instances) during the training process of the Markov chain model, avoiding the problem that the positive instances (stable instances) that may occur in the traditional training method dominate the training process, resulting in the model's insufficient ability to identify transient instances. Through the above method, the Markov chain model finally obtained will more accurately reflect the behavior pattern of the instance, especially for the recognition and prediction of transient instances. It provides strong data support for subsequent resource scheduling strategies, which helps to achieve a better balance between improving resource utilization and reducing OOM risks.

[0129] As an optional implementation, determining a resource scheduling policy corresponding to an instance type includes: in response to the instance type being a stable instance type, determining a resource scheduling policy corresponding to the stable instance type according to resource demand information of the instance, wherein the resource demand information is used to represent actual memory resources required for the instance to run on the service platform, and the resource scheduling policy is used to represent rules for scheduling memory resources to the instance from the service platform according to a memory overselling policy.

[0130] In this embodiment, in the process of determining the resource scheduling policy corresponding to the instance type, if the instance type of the instance is identified as a stable instance type, the resource scheduling policy corresponding to the stable instance type can be determined according to the resource demand information of the instance. Among them, the resource demand information can be used to represent the actual memory resources required by the instance during operation on the service platform. The resource demand information is a data description of the actual memory resources required by the stable instance during operation, including memory usage statistics of the instance during normal operation, such as average memory usage, peak memory usage, and memory usage pattern. The above resource demand information is crucial for identifying the stability of the instance and predicting its future memory demand. The resource scheduling policy can be used to represent the rule of scheduling the corresponding memory resources from the service platform to the instance according to the memory overselling policy. The resource scheduling policy is a rule for dynamically allocating memory resources from the service platform to the instance based on the type of the instance (stable instance type) and resource demand information, combined with the memory overselling policy. The memory overselling policy can refer to a policy that allows the total amount of memory allocated to the instance on the service platform to exceed the available physical memory under the premise of ensuring service quality and performance, so as to improve resource utilization.

[0131] Optionally, in this embodiment, determining a resource scheduling strategy that matches the instance type is a key step to ensure efficient resource utilization and service stability. When the instance type is identified as a stable instance, formulating a resource scheduling strategy based on its resource demand information will directly affect the efficiency and risk control of memory overselling.

[0132] Optionally, when a stable instance is running, its memory usage is relatively stable and the fluctuation is within a controllable range. The resource demand information of the stable instance reflects the average or stable memory consumption of the instance over a period of time. Through instance classification models (such as XGBoost, LightGBM), combined with runtime features and non-runtime features, stable instances can be accurately identified, providing a basis for resource scheduling.

[0133] Optionally, resource demand information includes, but is not limited to, key indicators such as the instance's memory usage, CPU usage, disk I / O, and network bandwidth. For stable instances, the focus is on observing their resource consumption over a long period of time to obtain their resource requirements in a stable state. Statistical analysis may be involved, such as calculating the average value and quantile within a certain time window to more accurately reflect the instance's resource usage habits.

[0134] Optionally, for stable instances, a more aggressive memory overselling strategy can be adopted because their memory usage patterns are predictable and the risk is relatively low. Resource scheduling strategies corresponding to stable instances can include instance size estimation, dynamic resource adjustment, and optimized binning algorithms. For instance size estimation, the upper limit of memory usage of stable instances can be estimated more accurately based on resource demand information using methods such as proportion & utilization method, quantile method, and time series prediction. This helps to allocate appropriate resources to them during resource scheduling and avoid over-reservation. For dynamic resource adjustment, a dynamic resource adjustment mechanism is set for stable instances, allowing memory overselling when they are in a stable state, but when abnormal memory usage or increased service load is detected, resource allocation can be quickly adjusted to meet the actual needs of the instance. For the optimized binning algorithm, a hybrid binning algorithm can be adopted, such as an adaptive instance scheduler, which combines multiple instance size estimation methods to dynamically optimize resource allocation and scheduling strategies based on the specific needs and behavior patterns of stable instances.

[0135] Optionally, after implementing the resource scheduling strategy, the running status of the instance can be continuously monitored, including memory usage, service performance indicators, etc., to evaluate the effectiveness and potential risks of the scheduling strategy. If it is detected that the resource scheduling strategy causes service performance to degrade or OOM risk to increase, it can be optimized in the following ways: Adjust the overselling ratio, that is, dynamically adjust the memory overselling ratio based on the real-time resource demand information of the instance to ensure that the service stability is not sacrificed while improving resource utilization. Introduce the Fallback mechanism, that is, when the memory utilization is close to the warning value, the Fallback mechanism is automatically triggered, such as online instance migration, to release some resources in time, prevent the occurrence of OOM errors, and ensure the achievement of SLO. Real-time policy adjustment, that is, adjust the resource scheduling strategy in real time according to the actual operation of the instance and load changes to ensure the rationality and dynamic optimization of resource allocation.

[0136] In an embodiment of the present application, determining the resource scheduling strategy corresponding to the stable instance type is an adaptive management method based on instance characteristics. By accurately estimating the resource demand information of the stable instance, a more efficient and flexible memory overselling strategy can be adopted to improve resource utilization while ensuring service stability. The implementation of the resource scheduling strategy corresponding to the above-mentioned stable instance can significantly improve the operating efficiency of the cloud database system under high load, while reducing the OOM risk and ensuring the achievement of service-level goals. The above method embodies the core idea of ​​intelligent resource scheduling, that is, dynamically adjusting the resource allocation strategy according to the characteristics of different instances to achieve a balance between resource utilization and service reliability.

[0137] It is worth noting that the execution of resource scheduling strategies corresponding to stable instance types requires a high degree of monitoring and feedback mechanism support to ensure that when there is a deviation between the prediction and the actual operation, the resource scheduling strategy can be adjusted in time to maintain the stable operation of the system. At the same time, the optimization and iteration of resource scheduling strategies can also rely on the collection and analysis of a large amount of runtime data, as well as the continuous training and adjustment of instance classification models, to achieve a resource management effect that meets the requirements.

[0138] As an optional implementation, the method also includes: in response to the instance type being a stable instance type, determining the memory usage pattern of the instance; determining an instance size determination strategy corresponding to the memory usage pattern, wherein the instance size determination strategy is used to represent a rule for determining the instance size of the instance; determining the instance size according to the instance size determination strategy; and determining resource requirement information that meets the instance size.

[0139] In this embodiment, when the instance type is a stable instance type, the memory usage model of the instance can be determined. An instance size determination strategy corresponding to the memory usage model can be determined. According to the instance size determination strategy, the instance size of the instance is determined, and the resource requirement information that meets the instance size is determined. Among them, the memory usage model, especially the time series pattern for the stable instance, is constructed by collecting and analyzing the memory usage data during the instance operation. In an embodiment of the present application, the memory usage model of the stable instance can be expressed as stable fluctuations, trend usage and periodic patterns. The steps of constructing the memory usage model include data preprocessing (e.g., smoothing, denoising), feature extraction (e.g., extracting trends, periodic components) and model selection, for example, using models such as the AutoRegressive Integrated Moving Average model (ARIMA) for prediction.

[0140] The instance sizing policy can be used to represent the rules for determining the instance size of an instance, that is, the rules for determining the instance size after understanding the memory usage pattern of a stable instance. The instance size can be an estimated result of the instance size determined using the instance sizing policy. The instance size of a stable instance can be the maximum possible memory requirement estimated based on its memory usage pattern and the above instance sizing policy. The above process is intended to avoid over-reservation of resources while ensuring stable operation of the instance under high load. The determination of instance size is a prerequisite for the formulation of resource scheduling policies, and directly determines the resource allocation rules of the instance in a memory oversold environment.

[0141] Optionally, in the cloud database memory overselling optimization process based on instance classification and adaptive scheduling, the resource management strategy for stable instances is particularly important, because the resource management strategy for stable instances is directly related to how to maximize resource utilization while ensuring service stability and reliability.

[0142] Optionally, the memory usage pattern of a stable instance refers to the stable and predictable behavior characteristics of its memory usage during operation. By analyzing the runtime characteristics of the instance, such as time series data of memory utilization, specific memory usage patterns can be identified. For example, some instances may show a nearly linear memory usage trend, while other instances may maintain constant memory usage under a certain load.

[0143] Optionally, once the memory usage pattern of a stable instance is determined, a suitable instance size determination strategy can be selected based on the pattern. The embodiments of the present application can utilize strategies such as the ratio & usage method, the quantile method, SBP and TSF, each of which has its applicable instance pattern: the ratio & usage method is applicable to instances where the memory usage is proportional to the allocation amount, and can predict the possible peak memory usage in the future based on the ratio of the current memory usage to the allocation amount. The quantile method is applicable to instances where the memory usage follows a certain distribution, and the peak memory demand of the instance can be estimated by calculating the high quantile of the historical memory usage (such as the 99% quantile). SBP is applicable to instances where the memory usage is random but generally stable and predictable. By constructing a statistical distribution model of memory usage, such as a normal distribution, the average memory demand and potential memory usage fluctuations of the instance can be estimated. TSF is applicable to instances where the memory usage has obvious time series characteristics, such as periodic or trend changes, and time series prediction models, such as long short-term memory networks (LSTM for short), can be used to predict future memory usage.

[0144] Optionally, based on the selected instance sizing strategy, the future memory requirements of the stable instance can be estimated to determine the instance size (i.e., the amount of memory that should be allocated). The above steps are the basis for resource scheduling, ensuring that the instance will obtain resources that match the predicted memory usage and avoiding waste caused by over-allocation of resources or OOM risks caused by insufficient memory under the memory overselling strategy.

[0145] For example, for the ratio & utilization method, the memory usage can be determined by the following formula, so that the instance size can be predicted by the memory usage. That is, by calculating the memory usage, the upper limit of the potential memory demand of an instance in the future can be obtained. Then, the memory usage can be used to guide the adjustment of the instance size:

[0146]

[0147] in, Can be used to indicate the The predicted memory usage of each instance; Can be used to indicate a preset scale factor; Can be used to indicate the assignment to The total amount of memory for each instance, that is, the memory allocated to the instance; Can be used to indicate the The actual memory usage of the instance, that is, the memory used by the instance.

[0148] For another example, for the quantile method, the memory usage can be determined by the following formula, so the instance size can be predicted by the memory usage. That is, by calculating the memory usage, the upper limit of the potential memory demand of an instance in the future can be obtained. Then, the memory usage can be used to guide the adjustment of the instance size:

[0149]

[0150] in, Can be used to indicate the The predicted memory usage of each instance; Can be used to represent the quantiles of a calculated data set; Can be used to indicate the A dataset of historical memory usage for instances.

[0151] Optionally, after determining the instance size, the resource requirement information will clearly indicate the actual memory resources required by the stable instance when running on the service platform. It can include the basic configuration of the instance, statistical information on resource consumption at runtime, etc., and is the direct basis for the execution of resource scheduling policies. The determination of resource requirement information needs to consider the memory usage pattern of the instance and the predicted instance size to ensure that resource allocation not only meets the operating requirements of the stable instance, but also complies with the overall resource utilization plan of the service platform.

[0152] Optionally, the process for the Item Size Estimator can be used to predict or estimate the resource requirements, especially the memory requirements, of the cloud database instance at runtime to support resource scheduling decisions.

[0153] For example, the pseudo code of the instance size estimator flow is analyzed as follows:

[0154] :Set of n database instances ({ ,..., }).Methods: Item size estimation methods.

[0155] Ensure: :Predicted memory usage for each instance.

[0156] The above pseudo code illustrates the input and output required by the algorithm. Specifically, initialize the output set ( ) is an empty set ( ← ), which is used to store the predicted memory usage of each instance. Traverse the instance collection ( ), for each instance ( ):if( ) is a transient instance (memory usage fluctuates widely and unpredictably): set directly ( )equal( ) of the allocated memory ( ). The memory usage of transient instances is difficult to predict, so the allocated memory is used as the upper limit of the prediction to avoid OOM events caused by memory overselling. If ( ) is a stable instance (memory usage is relatively stable). ) characteristics to select the appropriate prediction method (Methodi). If (Methodi) is set to Proportional & Utilization Rate Method, the preset proportionality factor ( ) and allocate memory ( ) multiplied by the actual memory used ( ) and take the larger value as the predicted memory usage ( The above method is suitable for the case where the memory usage of the instance is proportional to the allocation amount, or the memory usage of the instance has a clear baseline level.

[0157] If (Method) is set to the quantile method, the calculation example ( )'s memory usage history data ( ) as the predicted memory usage ( ). The above strategy is applicable to the case where the memory usage of the instance shows a certain distribution pattern. Using quantiles can more robustly estimate the peak memory usage that the instance may reach. If (Methodi) is set to SBP, assuming that the memory usage follows a certain statistical distribution (for example, normal distribution), use the mean of the distribution ( ) and variance ( ) to estimate the memory usage of the instance ( ). Random binning is suitable for instances where memory usage is random but can be described by a statistical distribution model. If (Methodi) is set to TSF, a time series forecasting model is used, such as the Holt-Winters forecasting method ( )'s future memory usage peak, and take the maximum value as the predicted memory usage ( ). Time series forecasting is suitable for instances where the memory usage has obvious trends or periodic patterns and can predict future usage based on historical data.

[0158] For each estimated instance ( ), add it to ( ) collection, ensure that ( ) contains the predicted memory usage of each instance. After completing the prediction for each instance, return the collection of predicted memory usage ( ).

[0159] In summary, by classifying (transient and stable instances) and selecting an appropriate prediction method (based on instance characteristics), a predicted memory usage upper limit is provided for each instance. This helps the resource scheduler make more accurate memory overselling decisions, improves resource utilization, reduces OOM risks, and ensures service stability and resource efficiency. Through the above prediction mechanism, the scheduler can allocate and manage memory resources more intelligently in a complex and changing cloud database environment.

[0160] In an embodiment of the present application, the key to the above method is to determine an instance size determination strategy suitable for a stable instance through accurate instance classification and memory usage pattern analysis. The formulation and execution of the instance size determination strategy rely on an in-depth understanding of the runtime characteristics of the instance and accurate prediction of resource demand information, aiming to improve resource utilization while reducing OOM risks and ensuring the achievement of SLO. By adaptively adjusting the resource scheduling strategy, a balance is achieved between resource utilization efficiency and service stability. Through accurate instance size estimation, resource demand information that meets the instance size can be determined, providing a basis for dynamic resource scheduling for the service platform, and achieving efficient resource utilization and continuous and reliable services.

[0161] It is worth noting that the selection of instance size determination strategy can be based on the specific behavior characteristics and resource requirements of the instance to adapt to stable instances with different memory usage patterns. At the same time, the implementation of instance size determination strategy can also consider the overall resource status and operation goals of the service platform, such as resource utilization, service stability, etc., for comprehensive optimization. Through continuous monitoring, analysis and policy adjustment, the intelligent level of resource management and scheduling can be further improved to ensure the efficient operation of the cloud database system.

[0162] As an optional implementation, determining a resource scheduling policy corresponding to an instance type includes: in response to the instance type being a transient instance type, determining a resource scheduling policy corresponding to the transient instance type according to resource requirement information of the instance, wherein the resource requirement information is used to represent actual memory resources required for the instance during its operation on a service platform, and the resource scheduling policy is used to represent a rule for scheduling memory resources from the service platform to the instance that are greater than or equal to the actual memory resources and less than or equal to the physical memory of the service platform.

[0163] In this embodiment, in the process of determining the resource scheduling strategy corresponding to the instance type, if the instance type of the instance is a transient instance type, the resource demand information of the instance can be determined, and the resource scheduling strategy corresponding to the transient instance type can be determined. Among them, the resource demand information can be used to represent the actual memory resources required by the instance during the operation of the service platform. The resource demand information of the transient instance refers to the actual amount of memory resources that may be required during the operation of the transient instance on the service platform. Since the memory usage of the transient instance fluctuates greatly and is unpredictable, when determining its resource demand information, the potential maximum memory usage can be considered to avoid OOM errors caused by insufficient resources. The resource scheduling strategy can be used to represent the rule of scheduling memory resources greater than or equal to the actual memory resources and less than or equal to the physical memory of the service platform to the instance from the service platform. For transient instances, the resource scheduling strategy adopted by this solution is not to oversell memory to give priority to its memory requirements. In the resource scheduling strategy, not overselling memory is the core of the embodiment of the present application for processing transient instances, which aims to ensure the continuity and stability of the service and avoid OOM errors caused by insufficient resources. The execution of the above resource scheduling strategy requires a high degree of real-time monitoring and feedback mechanism support, as well as a flexible resource adjustment mechanism to ensure timely response and processing when the memory demand of transient instances suddenly increases.

[0164] Optionally, in the cloud database memory overselling optimization method based on instance classification and adaptive scheduling, the resource scheduling strategy design for transient instances is a key part to ensure service stability and reduce the occurrence of OOM events. Due to the large fluctuations and unpredictable characteristics of transient instances, if memory is blindly oversold, it is very easy to cause OOM errors, resulting in service interruption or performance degradation. Therefore, for transient instances, the embodiment of the present application proposes a more conservative and safe resource scheduling strategy.

[0165] Optionally, the resource requirement information refers to the actual memory resources required by the transient instance during operation. Due to the unpredictability of memory usage of transient instances, the determination of their resource requirement information is usually based on conservative estimates, that is, the maximum memory usage of the instance or the upper specification limit of its resources. This means that for transient instances, the system will reserve enough memory to cope with sudden peaks in memory demand and avoid OOM.

[0166] Optionally, for transient instances, the core of the resource scheduling policy is not to oversell memory, but to prioritize memory requirements. This means that the total amount of memory allocated to transient instances will be strictly limited to the physical memory range of the service platform and will not exceed the threshold of physical memory.

[0167] Optionally, the resource scheduling policy for transient instances can ensure memory upper limit constraints, dynamic detection and adjustment, and reservation of sufficient resources. For memory upper limit constraints, the amount of allocated memory for transient instances will not exceed the maximum possible memory usage indicated in its resource requirement information. For dynamic monitoring and adjustment, the memory usage of transient instances will be monitored in real time. Once it is detected that the memory usage is close to the preset upper limit, no additional memory will be allocated to it to avoid exceeding the physical memory limit. For reserving sufficient resources, the memory resources reserved for transient instances will give priority to meeting their memory needs under high load or abnormal conditions to reduce OOM risks.

[0168] Optionally, resource scheduling policies for transient instances are implemented, which means taking a more cautious approach to resource allocation. Although this may result in lower memory resource utilization for transient instances than for stable instances, the primary goal is to ensure service continuity and stability. By not overselling memory, OOM errors can be effectively prevented, which is critical for cloud database services, because any OOM error may cause serious service interruptions, affecting customer experience and task operations.

[0169] Optionally, the scheduler algorithm of the embodiment of the present application (for the Hybrid-Item BinPacking scheduler algorithm) combines the First-Fit strategy and the offline version of the resource scheduling algorithm to allocate database instances on multiple machines while taking into account the impact of memory overselling and the achievement of SLO. The input parameters may include: machine set , a collection of (m) machines ({ ,..., });Instance collection , a collection of (n) database instances ({ ,..., }); item size estimates for each instance , the set of predicted memory usage for each instance, provided by the instance size estimator algorithm; the upper bound of the bin packing constraint , provided by the module that models the impact of memory overselling on SLO. Output ( ): The instance-to-machine allocation mapping, that is, which machine each instance is assigned to.

[0170] Optionally, the specific operation process of the adaptive instance scheduler dynamically optimizes resource allocation by combining instance classification results and multiple instance size estimation methods to achieve high efficiency and flexibility in resource scheduling. The specific scheduling process can be: at the beginning of the process, each instance is considered to be in an unallocated state, that is, the value of the allocated state (allocated) is set to False (allocated ← False), which means that the instance has not been allocated to any machine. In the process of traversing instances, you can select from the instance set Select an instance from For the selected instance , will traverse the machine set Each machine in , try to assign the instance to the above machine The total memory usage of the current machine, including both deterministic and uncertain parts, can be calculated using the following formula:

[0171]

[0172] in, Can be used to indicate the total memory usage of the current machine; Can be used to represent deterministic memory usage, which can be known; Can be used to represent the mean of uncertain memory usage; Can be used to represent the standard deviation of uncertain memory usage; Can be used to indicate the number of database instances.

[0173] Optionally, the MUR strategy can be used to update the upper bound of the packing constraints , check whether the constraints are satisfied: If the above constraints are met, then the instance Divided into machines If the above constraints are not met, you can try the next machine; if all machines cannot be allocated, you can start a new machine. Repeat the above steps until all instances are allocated.

[0174] Optionally, the scheduling algorithm aims to solve the problem of memory overselling when allocating cloud database resources, with the goal of improving resource utilization while avoiding OOM errors. The above scheduler algorithm is as follows: traverse the database instance to be allocated ( ), try to find a suitable machine for each instance. , the algorithm will traverse each machine ( ), check whether the instance can be Assign to machine Calculate the current machine The remaining resource constraints , then according to the example The memory usage (deterministic and non-deterministic) and resource requirements of the instance are used to determine whether the allocation can be made. Can be safely assigned to machines If the instance is not able to find a suitable allocation location after traversing each machine, the algorithm will start a new empty machine or execute other emergency strategies to ensure that the instance can get resources.

[0175] For example, the following pseudo code can be used ∈ do means that the instance collection will be traversed Each instance in allocated ← False, when trying to allocate an instance Before, its allocation status is initialized to False, indicating that the instance has not been allocated. ∈ do, ← ,in, It is a machine that calculates The upper limit of the acceptable packing constraint is set to The smaller value of (the upper limit of the global packing constraint) and the MUR of the remaining resources of the machine. The MUR function takes into account the machine Currently running instances and new instances May result in resource adjustments to ensure resource limits are not exceeded. ← 0, ← , initialize the deterministic resource usage of the current machine ( ) and uncertain resource usage ( ). if is deterministic then, ← + , else, ← , algorithms for machines Examples on Traverse, according to the example The resource usage of the current machine is calculated based on the resource usage nature (deterministic or uncertain). For deterministic resource usage, the actual .

[0176] For uncertain resource use, the parameters of the statistical distribution (mean and variance ) to estimate potential resource usage, using Function to update. then Allocate to , allocated ← True, break, check whether the current machine can safely accommodate the new instance through the above steps By calculating the upper limit of uncertain resource usage ( ) and the upper bound on the remaining resources of the machine minus the deterministic resource usage ( ) for comparison. If the upper limit of the uncertain resource usage is less than the upper limit of the remaining resources, then the instance Can be assigned to a machine and update the instance The allocation status of is True. If notallocated then, Allocate Ii to a new empty machine. If after traversing each machine, the instance If it is still not assigned, the Fallback mechanism is enabled and the instance Allocate to a newly created empty machine to ensure that each instance gets the necessary resources.

[0177] The above adaptive instance scheduler algorithm uses fine instance classification and status evaluation, combined with the calculation of deterministic and uncertain resource usage, to dynamically check and adjust instance allocation to maximize resource utilization while reducing the risk of OOM errors. The key to the algorithm is the ability to handle different types of instance resource requirements, and to ensure that each instance can get appropriate resource allocation by starting a new machine or taking other emergency measures when resource allocation constraints arise. This adaptive and flexible scheduling strategy is the key to optimizing cloud database resource management and memory overselling.

[0178] In the embodiments of the present application, for transient instances, the core strategy of the present application is not to oversell memory, but to give priority to ensuring its memory needs. This requires accurately determining the resource demand information of the transient instance and formulating the corresponding resource scheduling strategy accordingly. Although the resource scheduling strategy of the transient instance sacrifices part of the resource utilization, it effectively reduces the OOM risk, ensures the stability of the service, and embodies the principle of security first in resource management. In an actual production environment, the implementation of the resource scheduling strategy can be closely coordinated with other modules such as instance classification, monitoring system, Fallback mechanism, etc., to jointly ensure that the cloud database service can continue to operate stably under complex and changing workloads.

[0179] It is worth noting that although the resource scheduling strategy of transient instances is relatively conservative, the physical memory resources on the service platform can still be fully utilized during the entire process of resource allocation to improve overall resource utilization. Therefore, how to effectively allocate and schedule resources between transient instances and stable instances is an important factor that needs to be considered comprehensively. Through the implementation of the embodiments of the present application, a balance between resource utilization efficiency and service stability can be achieved, providing a safer and more efficient service for the cloud database system.

[0180] As an optional implementation, the method also includes: determining current characteristic information of the instance, wherein the current characteristic information is used to represent the current characteristics of the instance running on the service platform; adjusting the resource scheduling strategy using the current characteristic information and the instance type, wherein the adjusted resource scheduling strategy matches the current characteristic information.

[0181] In this embodiment, the current characteristic information of the instance can be determined. The resource scheduling strategy is adjusted using the current characteristic information and instance type. Among them, the current characteristic information can be used to represent the current characteristics of the instance running on the service platform, that is, the current characteristic information refers to the characteristic description of the instance at the current running time, providing a snapshot of the instant resource usage and running status of the instance. In the embodiment of the present application, the memory usage of the instance is important, including but not limited to: current memory usage (the amount of physical memory currently actually used by the instance), memory usage rate (the proportion of the memory currently used by the instance to the size allocated to it), memory usage trend (the trend of changes in the memory usage of the instance over the past period of time, for example, it can be an increase, decrease or stability) and sudden memory events (detecting an abnormal increase in the memory usage of the instance in a short period of time, which may indicate transient behavior). The adjusted resource scheduling strategy matches the current characteristics.

[0182] Optionally, the embodiments of the present application take into account the dynamic changes of the instance during operation, and real-time adjustment of the resource scheduling strategy to adapt to the current characteristic information is the key to ensuring efficient resource utilization and SLO achievement. The above embodiments focus on using the current characteristic information of the instance, combined with its type, to dynamically adjust the resource scheduling strategy, thereby improving the adaptability and accuracy of the strategy.

[0183] Optionally, for instances of different instance types, their resource requirements and memory usage patterns vary significantly. Therefore, combined with the current feature information and type of the instance, the resource scheduling policy can be adjusted more accurately to adapt to the actual operating status of the instance. For transient instances, once its memory usage is detected to be close to or exceeds the preset threshold, the resource scheduling policy is adjusted immediately to increase memory allocation or trigger instance migration to avoid OOM errors. For stable instances, based on their current memory usage, the instance size estimate is dynamically adjusted to optimize resource scheduling, maximize the memory overselling rate, and avoid OOM risks. Continuously monitor the current feature information of each instance, adjust the resource scheduling policy based on real-time data, and ensure that the policy is consistent with the current state of the instance.

[0184] Optionally, by real-time monitoring of the current characteristic information of the instance, the resource scheduling strategy is dynamically adjusted to ensure that the implementation of the strategy matches the real-time needs of the instance. This not only improves the accuracy of resource scheduling, but also enhances the ability to respond to emergencies, ensuring the stability and efficiency of the service. For transient instances, the adjusted resource scheduling strategy can immediately respond to the sudden increase in memory usage and avoid OOM risks through resource adjustment. For stable instances, the adjusted resource scheduling strategy can adaptively adjust resource allocation according to the real-time resource needs of the instance to maximize resource utilization. In the face of unpredictable fluctuations in resource demand, the resource scheduling strategy can ensure the overall stable operation of the database and maintain SLO compliance even under high memory overselling rates.

[0185] In the embodiment of the present application, the above method reflects the dynamic and adaptive nature of the resource scheduling strategy. By collecting and analyzing the current characteristic information of the instance in real time, combined with the instance classification results, the resource scheduling strategy can be dynamically adjusted to better adapt to the real-time needs of the instance, ensuring efficient use of resources and stable operation of services. The implementation of the resource scheduling strategy requires a powerful real-time monitoring system, a fast decision-making mechanism and flexible resource adjustment capabilities. In the actual production environment, the above method can effectively cope with complex and changeable operating conditions, and is an important means to improve the resource utilization and service quality of the cloud database system. By real-time monitoring of the current characteristic information of the instance, changes in resource demand can be quickly discovered, and resource scheduling strategies can be adjusted in time to avoid waste or insufficiency of resources. At the same time, combined with instance classification, differentiated resource management can be implemented for different types of instances, which not only improves the dynamic allocation efficiency of resources, but also reduces the risk of service stability caused by improper resource scheduling. The implementation of the resource scheduling strategy reflects the intelligence and refinement of resource management, which helps the cloud database system to achieve efficient use of resources and continuous reliability of services under high load and complex workload scenarios.

[0186] In summary, using current feature information and instance types to adjust resource scheduling strategies is a key step in achieving dynamic resource management and optimization in the embodiments of this application. Through real-time monitoring and intelligent decision-making, the adaptability and robustness of resource scheduling are enhanced, providing cloud database services with more efficient and stable services.

[0187] As an optional implementation, the method also includes: separately determining the total resource requirement information of the instance sets under different accounts, wherein the total resource requirement information is used to represent the sum of actual memory resources required for different instances in the instance set during the operation of the service platform; determining the maximum total resource requirement information among the different total resource requirement information corresponding to different accounts; and retaining the maximum total resource requirement information.

[0188] In this embodiment, the total resource requirement information of the instance set under different accounts can be determined respectively. Among the different total resource requirement information corresponding to different accounts, the maximum total resource requirement information is determined, and the maximum total resource requirement information is retained. Among them, the account can be an account registered by the user on the service platform. The total resource requirement information can be used to represent the sum of actual memory resources required by different instances in the instance set during operation on the service platform, that is, the total memory reservation. The maximum total resource requirement information can be MUR.

[0189] Optionally, in the embodiment of the present application, introducing MUR is an important method to ensure efficient resource utilization and service stability.

[0190] Optionally, the total resource demand information refers to the sum of actual memory resources required by each instance in a set of instances under different accounts on the service platform. Determination of the total resource demand information requires analysis and statistics of the memory usage of each instance under each account. For example, by collecting and counting the actual memory usage peak of each instance during operation. Associate the resource demand of the instance with the account to which it belongs to form an account-level resource demand data set. Sum the resource demands of the instances under the account to obtain the total resource demand information of the account.

[0191] Optionally, after determining the total resource demand information of different accounts, the maximum total resource demand information of each account, that is, MUR, can be calculated. The probability that a set of instances of different accounts reaches the peak memory usage at the same time is extremely low. Therefore, when reserving resources, there is no need to reserve memory resources for the set of instances of each account to reach the maximum total resource demand information at the same time. The core of the MUR strategy is to ensure that only the set of instances of a single account can still run stably when it reaches its total resource demand information, without reserving too many memory resources for the set of instances of each account.

[0192] Optionally, retaining the maximum total resource demand information means that, according to the MUR policy, memory resources sufficient to cope with the maximum total resource demand information will be reserved for the instance set of each account, but will not exceed the above threshold. Specifically, the resource reservation strategy is as follows: Memory reservation optimization, based on the maximum total resource demand information, calculate the appropriate memory reservation amount while meeting the potential peak demand of each account. The reserved amount is dynamically adjusted, and the MUR reservation amount can be dynamically adjusted according to the real-time resource demand information and load conditions to ensure the effective use of resources. Improved resource utilization efficiency, the MUR policy avoids excessive reservation of memory resources, improves memory utilization efficiency, reduces resource waste, and ensures the stability and reliability of the service.

[0193] Optionally, the implementation of the MUR policy can also conduct in-depth analysis of the memory usage behavior of the instance, and accurately evaluate the total resource demand information of different accounts by collecting and processing the runtime data of each instance. The system must also have strong data processing capabilities and real-time decision-making mechanisms to ensure that the maximum total resource demand information can be quickly calculated and the resource reservation strategy can be adjusted in real time.

[0194] For example, the potential additional memory requirement per user can be calculated as follows:

[0195]

[0196] in, Can be used to represent users The collection of instances of ; Can be used to represent instances The memory requirement of the instance, that is, the upper limit of memory allocated to the instance (the maximum memory allowed); Can be used to represent instances The actual amount of memory used (memory utilization); Can be used to represent users Instances of The sum of potential additional memory requirements, that is, the additional memory requirement.

[0197] As another example, the total memory reserve can be determined by the following formula:

[0198]

[0199] in, Can be used to represent a collection of users; Can be used to indicate the total amount of memory that needs to be reserved (that is, the total memory reservation).

[0200] In an embodiment of the present application, the core of the MUR strategy is to determine the total resource demand information of the instance set of different accounts respectively, calculate the maximum total demand information, and then retain only the memory resource amount corresponding to the total demand information. The above strategy avoids reserving too many memory resources for each account at the same time, resulting in resource waste and low utilization. At the same time, the MUR strategy ensures that even under high memory overselling rates, the SLO of each user can be met, improving resource utilization efficiency and the ability to cope with sudden loads. By dynamically adjusting the resource reservation strategy, MUR not only improves the utilization efficiency of memory resources, but also reduces the occurrence of OOM events, ensuring the stability and continuity of services. This strategy embodies the refinement and intelligence in resource management, and is an important innovation to improve the resource utilization and service quality of cloud database systems.

[0201] As an optional implementation, the method also includes: determining the memory utilization of the instance in the service platform; in response to the memory utilization being greater than a memory utilization threshold, adjusting the memory resources scheduled to the instance, or performing a migration operation on the instance.

[0202] In this embodiment, the memory utilization of the instance in the service platform can be determined. The relationship between the memory utilization and the memory utilization threshold can be determined. If the memory utilization is greater than the memory utilization threshold, the memory resources scheduled by the instance can be adjusted, or a migration operation can be performed on the instance.

[0203] Optionally, the core of this embodiment is to monitor the memory utilization of the instance in real time, and take corresponding resource management and scheduling measures according to the comparison result between the memory utilization and a preset threshold.

[0204] Optionally, in a cloud service environment, real-time monitoring of instance memory utilization is essential for resource management and scheduling. This involves collecting and analyzing instance runtime performance data, such as memory usage, peak memory usage, average memory usage, and other indicators. In this way, you can continuously understand the resource consumption of each instance and whether the resource consumption is within the normal range.

[0205] Optionally, the memory utilization threshold is a critical point used to trigger resource adjustment or instance migration operations. The memory utilization threshold can be derived based on historical data statistics, or set according to SLO and task requirements. When the memory utilization of an instance exceeds this threshold, it can be considered that the instance's resource consumption may reach an unstable state, and action needs to be taken to prevent service interruption or performance degradation.

[0206] Optionally, when the memory utilization exceeds the threshold, the following two main measures can be taken to deal with it: adjusting memory resources and performing instance migration (Reactive Live Migration of Database Instances).

[0207] Optionally, for adjusting memory resources, you can try to reduce the memory consumption of the instance, for example, by compressing data, optimizing query plans, adding caching mechanisms, etc. In addition, you can also increase the memory allocation of the instance to meet its high load requirements and avoid OOM errors.

[0208] Optionally, for instance migration, if adjusting memory resources is not enough to solve the problem, or this is not a suitable solution, the instance migration operation can be triggered. That is, the instance can be transferred from the current resource-constrained machine to a machine with more abundant resources to release the memory resources of the current machine, reduce memory utilization, and prevent OOM errors. Instance migration needs to be performed without affecting service performance and user experience as much as possible, and can be achieved with the help of automated tools and intelligent scheduling algorithms to achieve seamless migration.

[0209] Optionally, this implementation achieves dynamic resource management and scheduling in a cloud service environment by real-time monitoring, defining thresholds, and taking measures such as resource adjustment or instance migration, so that memory allocation can be automatically adjusted according to load changes and instance behavior, keeping resource utilization in an efficient and stable state while avoiding service interruptions or performance bottlenecks.

[0210] In the embodiments of the present application, the reliability and resource utilization efficiency of cloud services can be significantly improved through the above method. Through active monitoring and intervention, the system can effectively prevent and handle the situation of tight memory resources, reduce the probability of OOM errors, and thus ensure service continuity and user satisfaction. At the same time, by optimizing resource scheduling, excessive memory retention is avoided, resource utilization is improved, and the operating costs of cloud service providers are reduced.

[0211] In summary, through real-time monitoring, threshold setting and dynamic resource adjustment mechanism, an effective memory management and instance scheduling strategy is provided for the service platform. It can optimize resource allocation and improve the overall efficiency and economy of cloud services while ensuring service stability and high availability.

[0212] As an optional implementation method, the probability of achieving the service level objective SLO of the service function is determined; the probability of achieving the target is input into a logistic regression model for analysis to obtain a memory utilization threshold, wherein the logistic regression model is used to represent the nonlinear relationship between different memory utilizations of the instance and different probabilities of achieving the target corresponding to the service function.

[0213] In this embodiment, the probability of achieving the service level objective SLO of the service function can be determined. The probability of achieving the target can be input into the logistic regression model for analysis to obtain the memory utilization threshold. The probability of achieving the target can be the probability of achieving the SLO. The logistic regression model can be used to represent the nonlinear relationship between different memory utilizations of the instance and different probability of achieving the target corresponding to the service function.

[0214] Optionally, the embodiment of the present application can apply the logistic regression model to cloud service resource management to determine the memory utilization threshold. The above implementation provides data support and decision-making basis for resource scheduling and optimization by analyzing the relationship between memory utilization and SLO compliance probability.

[0215] Optionally, the SLO compliance probability refers to the probability that the service can meet or exceed the service level indicators specified in the SLO within a specific time window. In cloud database services, SLO can involve key performance indicators such as availability, latency, and throughput. The determination of the compliance probability is usually based on historical data analysis and statistical evaluation of service quality. For example, if the statistical results show that the service can meet the SLO requirements 99% of the time in the past month, then the SLO compliance probability is 99%.

[0216] Optionally, the logistic regression model is a statistical method for analyzing the relationship between one or more independent variables and a binary dependent variable. In an embodiment of the present application, the logistic regression model is used to describe the nonlinear relationship between different memory utilizations of an instance and different attainment probabilities corresponding to a service function. Specifically, the independent variable of the logistic regression model is memory utilization, and the dependent variable is the SLO attainment probability.

[0217] Optionally, collect historical running data of the instance at different memory utilization rates, including memory usage, service response time, service availability, and other indicators. Preprocess and extract features from the collected raw data to ensure the quality and relevance of the data input to the model. Use the logistic regression algorithm and historical data to train the model to find a suitable fitting relationship between memory utilization and the probability of achieving the SLO. Evaluate the performance of the model through methods such as cross-validation to ensure its prediction accuracy and generalization ability.

[0218] Optionally, the logistic regression model can express the nonlinear relationship between memory utilization and SLO compliance probability through an S-curve (Sigmoid function). As memory utilization increases, the SLO compliance probability may first increase to a peak, and then begin to decrease due to increased resource competition and pressure. The logistic regression model can capture the above trend and provide a mathematical expression to describe it.

[0219] Optionally, based on the analysis of the logistic regression model, a memory utilization threshold can be found so that under the memory utilization threshold, the probability of the service's SLO reaching the target reaches an ideal level (for example, 95%). The above memory utilization threshold is a key parameter in resource scheduling and optimization because it balances the relationship between resource utilization and service quality. If the memory utilization threshold is exceeded, the probability of SLO reaching the target may drop rapidly, resulting in service instability; if the memory utilization threshold is lower than the memory utilization threshold, although the service stability is high, the resource utilization is low and the cost increases.

[0220] Optionally, once the memory utilization threshold is determined, the resource scheduler can use it as a decision basis to guide resource allocation. For example, if the memory utilization of the current instance approaches or exceeds the threshold, the scheduler may take measures such as adjusting resource allocation, triggering instance migration, and optimizing resource scheduling policies. For adjusting resource allocation, dynamically adjust the memory quota of the instance to ensure that the memory utilization threshold is not exceeded. For triggering instance migration, migrate memory-intensive instances to machines with more abundant memory resources to reduce the memory utilization of the original machine. For optimizing resource scheduling policies, adjust the resource scheduling policy based on instance classification and historical performance to keep the overall memory utilization below the threshold while maximizing resource utilization.

[0221] In the embodiment of the present application, the key to the above implementation is to use a logistic regression model to quantify the relationship between memory utilization and service stability, providing an objective and quantitative decision-making basis for resource management. By dynamically adjusting the memory utilization threshold, a balance can be found between improving resource utilization and ensuring service stability, thereby optimizing the overall performance and economy of cloud services.

[0222] The present application also provides another memory resource scheduling method. Figure 3 is a flowchart of another method for scheduling memory resources according to an embodiment of the present application, such as Figure 3 As shown, the method is applied to a resource management system and may include the following steps:

[0223] Step S302: obtaining at least one instance in a cloud database system, wherein the cloud database system is deployed on a cloud service platform, and the instance is used to provide service functions of the cloud service platform.

[0224] In the technical solution provided in the above step S302 of the present application, in the process of implementing the memory overselling optimization method, at least one instance in the cloud database system is obtained as an object of analysis and processing. The above steps are the basis of the entire method, ensuring that subsequent resource management and scheduling operations have specific objects to act on. Among them, the cloud database system is a key component on the cloud service platform, providing data storage, management and access services, and supporting various task applications. In a cloud environment, a database instance may run on a virtualized or containerized infrastructure, and can dynamically adjust resource allocation according to demand. Therefore, the cloud database system is closely related to the service functions of the cloud service platform and is the basis for achieving high availability and efficient resource utilization. In a cloud database system, an instance refers to a logical unit or container for running a database service. Each instance has its own independent resource quota, including CPU, memory, disk, etc. The instance provides database services for users or applications, and can dynamically expand or shrink resources according to the workload. In the memory overselling optimization method, the instance is the smallest unit of resource management and scheduling, and real-time monitoring and reasonable scheduling of the instance are the key to improving resource utilization and reducing OOM risks.

[0225] Optionally, the core operation of this embodiment is to obtain at least one instance in the cloud database system. It may involve scanning or querying the database cluster on the cloud service platform to identify the currently running instance. After obtaining the instance, the subsequent steps may perform detailed memory usage analysis, classification, resource demand prediction, etc. on it to provide data support for the formulation of resource scheduling strategies. The selection of instances has an important impact on the execution efficiency and resource management effect of subsequent steps. For example, instances with high memory usage, heavy workloads, or OOM events that have occurred can be preferentially selected for analysis because these instances are the focus of resource management. By selecting instances in a targeted manner, problems can be identified and handled more effectively, and resource scheduling can be optimized.

[0226] Step S304, identifying the instance type of the instance, wherein the instance type is divided according to the variation range of the memory resources required by the instance during its operation on the cloud service platform.

[0227] In the technical solution provided in the above step S304 of the present application, after obtaining the instance in the cloud database system, the instance type of the instance can be identified. The instance type is defined according to the change range of the memory resources required during the operation of the instance on the cloud service platform, which is crucial for resource management and scheduling. The instance type is a classification of the memory usage behavior of the instance, reflecting the stability of the memory demand of the instance during operation. In this solution, instance types are mainly divided into two categories: transient instances and stable instances. The memory demand of transient instances fluctuates greatly and is unpredictable, which can easily lead to OOM errors. The memory demand of stable instances is relatively stable, and the fluctuations are within a controllable range, which is more suitable for memory overselling.

[0228] Optionally, the purpose of identifying instance types is to provide a basis for subsequent resource management and scheduling strategies. For transient instances, a conservative resource allocation strategy needs to be adopted to avoid excessively high memory overselling rates and reduce OOM risks; while for stable instances, a higher memory overselling rate can be tried to improve resource utilization while maintaining a low OOM risk.

[0229] Optionally, instance type identification can be achieved by analyzing the runtime data and historical performance records of the instance. Specific methods may include: feature engineering, extracting features related to the magnitude of memory usage changes, such as maximum memory usage, average memory usage, peak frequency of memory usage, memory usage volatility, etc. Machine learning classification, using supervised learning algorithms, such as logistic regression, support vector machine, random forest, etc., input the extracted features into the trained model to predict whether the instance belongs to the transient or stable type. Statistical analysis, performing statistical analysis based on the historical memory usage data of the instance, for example, calculating the standard deviation or variance of memory usage fluctuations, and if the fluctuation exceeds a preset threshold, it is identified as a transient instance.

[0230] Optionally, the accuracy and efficiency of instance type identification directly affect the effectiveness of resource management. Accurate identification can ensure that resource allocation strategies match instance behavior, avoid resource waste, reduce OOM risks, and improve resource utilization. Inaccurate identification may lead to unreasonable resource allocation, such as misidentifying transient instances as stable instances, thereby increasing OOM risks; or misidentifying stable instances as transient instances, resulting in overly conservative resources and reduced utilization.

[0231] Step S306, determining a resource scheduling strategy corresponding to the instance type, wherein different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent a rule for scheduling memory resources from the cloud service platform to the instance.

[0232] In the technical solution provided in the above step S306 of the present application, after determining the instance type of the instance, a resource scheduling policy corresponding to the instance type can be determined. In the embodiment of the present application, the resource scheduling policy refers to the rules for determining how to allocate memory resources to different types of instances, which may include memory allocation priorities, thresholds, adjustment mechanisms, etc., to ensure that the instance has sufficient resource support during operation while avoiding resource waste and overload.

[0233] Optionally, in the memory overselling optimization method, corresponding resource scheduling strategies are determined for different types of instances. That is, based on the identification of instance types, refined and differentiated memory resource management rules can be formulated to ensure effective utilization of resources and stable operation of the system.

[0234] Optionally, a stable instance can be an instance with relatively stable memory usage, and a more aggressive resource scheduling strategy can be adopted. For example, a higher memory overselling rate is allowed to improve resource utilization. This can include memory size estimation based on historical data and prediction models, and adopting a split or hybrid binning algorithm for memory allocation.

[0235] Alternatively, transient instances can be used for instances with large and unpredictable fluctuations in memory demand, which requires a more conservative resource scheduling strategy. This may mean avoiding memory overselling and ensuring that instances have a fixed and sufficient memory quota to prevent OOM errors. Transient instances may require more frequent monitoring and more immediate resource adjustments to cope with sudden peaks in memory demand.

[0236] Optionally, the basis for formulating resource scheduling policies usually includes the following points: Instance classification results, determine the resource management focus of each instance based on the identified instance type. SLO, consider the availability and performance requirements of the service, ensure that the resource scheduling policy can meet the SLO, and avoid service interruption or performance degradation. Resource utilization and cost, balance the efficient use of resources and cost control, avoid excessive resource reservation, and ensure the stable operation of the cloud database system. Prediction model and real-time monitoring, combine the prediction model to estimate the future resource demand of the instance, and continuously track the resource usage through the real-time monitoring mechanism to dynamically adjust the policy.

[0237] Optionally, the resource scheduling policy can be implemented through the following mechanisms: Adaptive instance scheduler, dynamically adjusts memory allocation according to instance classification and current resource usage to optimize resource utilization efficiency. Dynamic threshold monitoring, sets dynamic thresholds for memory utilization, and triggers resource adjustment or instance migration operations when the threshold is exceeded. MUR policy, adjusts memory reservations based on the user's potential maximum memory demand to reduce resource waste. Fallback mechanism, when errors occur in prediction or scheduling, resulting in excessive memory pressure, emergency measures are taken, such as instance migration, to prevent OOM errors and ensure system stability.

[0238] Optionally, implementing differentiated resource scheduling strategies can effectively improve resource utilization, while reducing OOM risks and ensuring service stability and SLO achievement. For stable instances, adopting an aggressive scheduling strategy that allows a certain degree of memory overselling can significantly improve resource utilization; while for transient instances, adopting a conservative strategy to ensure that the instance has a fixed memory quota can effectively reduce OOM risks and avoid service interruptions.

[0239] In summary, through differentiated strategies, we ensure that resource allocation is both efficient and safe, and balance the relationship between resource utilization and service stability. The implementation of the above steps can comprehensively consider factors such as instance type, SLO requirements, resource cost, and prediction model to achieve appropriate resource management effects. Through precise resource scheduling strategies, cloud service platforms can better meet user needs, improve service quality, and reduce operating costs.

[0240] Step S308: Perform resource scheduling operations on the instance according to the resource scheduling policy.

[0241] In the technical solution provided in the above step S308 of the present application, after identifying and classifying the instances, a specific scheduling strategy will be applied to each type of instance. For transient instances, the strategy may be more conservative to ensure sufficient memory and avoid OOM errors; while for stable instances, the scheduling strategy may be more aggressive to make full use of resources and improve resource utilization.

[0242] Through the above steps S302 to S308 of the present application, at least one instance in the cloud database system is obtained; the instance type of the instance is identified; the resource scheduling strategy corresponding to the instance type is determined; and the resource scheduling operation is performed on the instance according to the resource scheduling strategy. Thus, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of low scheduling efficiency of memory resources is solved.

[0243] The above method of this embodiment is further introduced below.

[0244] As an optional implementation, step S304, identifying the instance type of the instance, includes: calling the instance classification model in the resource management system, classifying the instance, and obtaining the instance type, wherein the instance classification model is obtained by training the Markov chain model using the instance type label of the instance sample.

[0245] In this embodiment, in the process of identifying the instance type of the instance, the instance classification model in the resource management system can be called to classify the instance to obtain the instance type. The instance classification model can be obtained by training a Markov chain model using the instance type label of the instance sample.

[0246] Optionally, identifying the instance type of the instance adopts an optional implementation method, that is, using the instance classification model in the resource management system to classify the instance to determine its type. The above method focuses on using the Markov chain model for training to achieve instance type identification.

[0247] Alternatively, a Markov chain model is a statistical model used to describe the sequential changes of a system state over time, where the probability of the next state depends only on the current state and is independent of the historical state. In the classification of cloud database instances, the Markov chain model can be used to analyze the time series of instance memory usage status and predict the instance type (transient or stable).

[0248] Optionally, the training process of the instance classification model involves the following key steps: Data preparation, collecting runtime data of the instance, including memory usage, CPU load, I / O operations, etc. At the same time, marking each instance with its type (transient or stable) as a label for training. Feature selection, extracting features that are highly relevant to the instance type from the original data, such as memory usage volatility, peak memory usage, average memory usage, etc., which will be used as input to the model. Model training, using the Markov chain model to model the memory usage pattern of the instance, and through the iterative pseudo-label training method, continuously optimizing the model parameters and improving the classification accuracy. Evaluation and adjustment, during the training process, using cross-validation and other techniques to evaluate the model performance, adjusting the model parameters based on the evaluation results to ensure the generalization ability of the model on different instances.

[0249] Optionally, in the resource management system, the instance classification model can be called by other modules or components as a service or function. When the type of an instance needs to be identified, the instance classification model can be called, and the current state and historical data of the instance can be input. The instance classification model will output a probability value indicating the probability that the instance is transient or stable. Based on the output result, the instance type can be determined.

[0250] Optionally, identifying the instance type is crucial for the formulation of subsequent resource scheduling strategies. For transient instances, a more conservative resource allocation strategy may be required to reduce OOM risk; for stable instances, a higher memory overselling rate can be tried to improve resource utilization. By identifying instance types, the resource management system can manage resource allocation in the cloud database more intelligently and accurately, balancing resource utilization efficiency and service stability.

[0251] Optionally, the advantage of the Markov chain model is that it can effectively capture the temporal dependencies of instance states, and has certain advantages for the identification of transient instances. In order to ensure the long-term effectiveness of the instance classification model, instance operation data can be continuously collected, and the model can be retrained and optimized regularly. This can include: data update, regularly update the instance sample set, including new instance data and type labels to reflect changes in the current instance behavior. Model iteration, retrain the model based on the latest data, adjust the model parameters, and improve the classification accuracy. Performance monitoring, continuously monitor the classification performance of the model, such as accuracy, recall, and F1 score, and promptly detect performance degradation or deviation problems.

[0252] In the embodiment of the present application, the instance classification model trained by the Markov chain model can effectively identify the instance type and provide key information for the formulation of subsequent resource scheduling strategies. The above method combines machine learning and statistical modeling, reflecting the development trend of resource management systems in the direction of intelligence and automation. However, in order to ensure the long-term effectiveness and classification accuracy of the model, continuous data updates and model optimization can also be performed to adapt to the evolution of instance behavior patterns.

[0253] As an optional implementation, determining a resource scheduling policy corresponding to an instance type includes: in response to the instance type being a stable instance type, determining a resource scheduling policy corresponding to the stable instance type according to resource demand information of the instance, wherein the stable instance type is used to indicate that a change value of the memory utilization of the instance does not exceed a change threshold, the resource demand information is used to indicate actual memory resources required for the instance during operation on the cloud service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources from the cloud service platform to the instance according to a memory overselling policy; in response to the instance type being a transient instance type, determining a resource scheduling policy corresponding to the transient instance type according to the resource demand information of the instance, wherein the transient instance type is used to indicate that a change value of the memory utilization of the instance exceeds a change threshold, the resource demand information is used to indicate actual memory resources required for the instance during operation on the cloud service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources greater than or equal to the actual memory resources and less than or equal to the physical memory of the cloud service platform from the cloud service platform to the instance.

[0254] In this embodiment, in the process of determining the resource scheduling policy corresponding to the instance type, if the instance type is a stable instance type, the resource scheduling policy corresponding to the stable instance type can be determined according to the resource demand information of the instance. If the instance type is a transient instance type, the resource scheduling policy corresponding to the transient instance type can be determined according to the resource demand information of the instance.

[0255] Optionally, the above embodiment may be an optional implementation method for determining a resource scheduling strategy corresponding to an instance type, which embodies the core idea of ​​adopting differentiated resource management strategies for different instance types.

[0256] Optionally, for stable instance types, the resource demand information reflects the actual memory resources required by the instance during normal operation. This includes but is not limited to the historical average memory usage, peak memory usage, and immediate memory demand based on the current workload. Based on the resource demand information of the stable instance type, the resource scheduling policy can be more aggressive. Specifically, when the instance is identified as a stable type, the following points can be considered to determine the scheduling policy: Allow moderate overselling. The memory usage of stable instances is relatively stable. A certain memory overselling ratio can be set to improve resource utilization. Dynamically adjust the threshold. According to the specific resource requirements of the instance and the physical memory capacity of the cloud service platform, dynamically adjust the memory quota of the instance to ensure that resources are fully utilized while meeting resource requirements. Predictive scheduling uses historical data to predict the future resource requirements of the instance, and adjust resources in advance to avoid insufficient or excessive resources.

[0257] Optionally, in the resource management of stable instances, the resource scheduling policy will allow a certain amount of memory overselling, that is, the total amount of memory allocated to the instance can exceed the actual memory resources currently required by the instance, but will not exceed the physical memory capacity of the cloud service platform. The purpose of this policy is to maximize resource utilization efficiency while ensuring service stability.

[0258] Optionally, the resource demand information of transient instance types also reflects the actual memory resources required by the instance during operation, but the memory demand of such instances varies significantly and is difficult to predict. For transient instance types, the resource scheduling policy needs to be more conservative and flexible. Specific policies may include: In order to avoid OOM errors, the memory allocation of transient instances will not exceed their actual demand, that is, the allocated memory resources are equal to or slightly greater than the actual memory demand. Considering the possible memory peak of transient instances, the resource scheduler will reserve more memory for these instances to cope with sudden demand increases. More frequent monitoring of the memory usage of transient instances is implemented, and once resource shortages are found, immediate actions such as resource increase or instance migration are taken to prevent OOM errors. The resource scheduling policy of transient instances ensures that instances can obtain sufficient memory resources in a timely manner during operation, while not over-allocating to avoid resource waste. The memory resources allocated to transient instances will be equal to or greater than their actual demand, but not exceeding the physical memory capacity of the cloud service platform. The purpose of this policy is to reduce OOM risks and ensure service continuity and availability.

[0259] Optionally, different resource scheduling strategies can be adopted for stable instances and transient instances, which can manage resources more finely, improve resource utilization efficiency, and reduce OOM risks. By dynamically adjusting thresholds and implementing predictive scheduling, resource scheduling strategies can respond to the actual needs of instances more intelligently and reduce resource waste. For transient instances, adopting a more conservative resource allocation strategy and reserving sufficient memory can significantly reduce the occurrence of OOM errors and ensure service stability.

[0260] In the embodiments of the present application, by determining the resource scheduling strategy corresponding to the instance type, the importance of differentiated resource management for different instance types is reflected. By adopting a moderate memory overselling strategy for stable instances, and a conservative resource allocation strategy with real-time monitoring and adjustment capabilities for transient instances, the relationship between resource utilization and service stability can be effectively balanced, and the overall performance and efficiency of the cloud service platform can be improved. The implementation of the above strategies needs to rely on intelligent algorithms and modules in the resource management system, such as instance classification models, resource demand prediction models, etc., to ensure the accuracy and real-time nature of decision-making.

[0261] The embodiment of the present application also provides a method for scheduling memory resources from the Software as a Service (SAAS) side. Figure 4 is a flowchart of another method for scheduling memory resources according to an embodiment of the present application, such as Figure 4 As shown, the method may include the following steps:

[0262] Step S402, obtaining the instance type of at least one instance in the database by calling the first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the instance type, the database is deployed on the service platform, and the instance is used to provide service functions of the service platform.

[0263] Step S404, identifying the instance type of the instance, wherein the instance type is divided according to the variation range of the memory resources required by the instance during its operation on the service platform.

[0264] Step S406, determining a resource scheduling policy corresponding to the instance type, wherein different instance types correspond to different resource scheduling policies, and the resource scheduling policy is used to represent a rule for scheduling memory resources from the service platform to the instance.

[0265] Step S408: perform resource scheduling operations on the instance according to the resource scheduling strategy to obtain a resource scheduling result.

[0266] Step S410: outputting the resource scheduling result by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the resource scheduling result.

[0267] Through the above steps S402 to S410 of the present application, the instance type of at least one instance in the database is obtained by calling the first interface; the instance type of the instance is identified; the resource scheduling strategy corresponding to the instance type is determined; according to the resource scheduling strategy, the resource scheduling operation is performed on the instance to obtain the resource scheduling result; and the resource scheduling result is output by calling the second interface. Thus, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of low scheduling efficiency of memory resources is solved.

[0268] The above method of this embodiment is further introduced below.

[0269] As an optional implementation, step S404, identifying the instance type of the instance, includes: using an instance classification model to classify the instance to obtain the instance type, wherein the instance classification model is obtained by training a Markov chain model using instance type labels of instance samples.

[0270] As an optional implementation, determining a resource scheduling policy corresponding to an instance type includes: in response to the instance type being a stable instance type, determining a resource scheduling policy corresponding to the stable instance type according to resource demand information of the instance, wherein the stable instance type is used to indicate that a change value of the memory utilization of the instance does not exceed a change threshold, the resource demand information is used to indicate actual memory resources required for the instance during operation on the service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources from the service platform to the instance according to a memory overselling policy; in response to the instance type being a transient instance type, determining a resource scheduling policy corresponding to the transient instance type according to the resource demand information of the instance, wherein the transient instance type is used to indicate that a change value of the memory utilization of the instance exceeds a change threshold, the resource demand information is used to indicate actual memory resources required for the instance during operation on the service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources greater than or equal to the actual memory resources and less than or equal to the physical memory of the service platform to the instance from the service platform.

[0271] According to an embodiment of the present application, an embodiment of a scheduling system for memory resources is also provided. Figure 5 is a schematic diagram of a scheduling system for memory resources according to an embodiment of the present application, such as Figure 5 As shown, the memory resource scheduling system 500 may include: a classifier 501 and a scheduler 502 .

[0272] Classifier 501 is used to obtain at least one instance in the database, wherein the database is deployed on a service platform and the instance is used to provide service functions of the service platform; identify the instance type of the instance, wherein the instance type is divided according to the change range of memory resources required during the operation of the instance on the service platform.

[0273] In this embodiment, the classifier 501, as a core component in the service platform memory overselling optimization method, has the main function of obtaining instances deployed in the database, identifying the instance types of these instances, and describing the change range of the memory resources required during the instance operation based on the instance type. The classifier can be a pattern classifier.

[0274] Optionally, the classifier 501 obtains at least one instance from a database deployed on the service platform. The above instance can be a running instance of various types of services (such as a relational database, a data warehouse, etc.), and each instance carries certain service functions, such as data storage, query processing, etc. The purpose of obtaining the instance is to identify the subsequent instance type and determine the resource scheduling strategy.

[0275] Optionally, instance type identification is one of the key functions of classifier 501, which divides instances into transient instances and stable instances. The memory requirements of transient instances vary greatly and are difficult to predict, while the memory requirements of stable instances vary relatively less and are more stable. By identifying instance types, classifier 501 can provide differentiated resource management strategies for different instances, which is crucial for reducing OOM risks, improving resource utilization and service stability.

[0276] Optionally, the change in memory resources required by the instance during its operation on the service platform reflects the changing trend and volatility of the instance's memory usage over time. For transient instances, the change may be very large, manifested as a sudden increase or decrease in memory usage; while for stable instances, the change is smaller and the memory usage is relatively stable. The change in memory resources is one of the important bases for instance type identification.

[0277] Optionally, the classifier 501 can identify the instance type based on a Markov chain model, a machine learning algorithm (e.g., XGBoost, LightGBM, etc.) or a hybrid model. The historical operation data of the instance, including memory usage, CPU usage, I / O operation data, etc., is used to extract features related to the instance type through feature engineering, and then input into the trained classification model to predict the type of the instance. In a specific implementation, it may also be combined with an iterative pseudo-label training method to improve the model's recognition accuracy for transient instances.

[0278] Optionally, the output of classifier 501, i.e., the instance type, will directly determine the subsequent resource scheduling strategy. For instances identified as stable instances, the resource scheduling strategy can be more aggressive, allowing a certain degree of memory overselling to improve resource utilization efficiency; while for instances identified as transient instances, the resource scheduling strategy should be more conservative to ensure that the instance has sufficient memory quota to avoid OOM errors and ensure service stability.

[0279] Scheduler 502 is used to determine the resource scheduling strategy corresponding to the instance type, where different instance types correspond to different resource scheduling strategies, and the resource scheduling strategy is used to represent the rules for scheduling memory resources from the service platform to the instance; according to the resource scheduling strategy, the resource scheduling operation is performed on the instance.

[0280] In this embodiment, the scheduler 502 plays a core role in the memory overselling optimization method of the service platform. The scheduler 502 is responsible for determining the resource scheduling strategy corresponding to the instance type and performing specific resource scheduling operations.

[0281] Optionally, the scheduler 502 can determine the corresponding resource scheduling strategy based on the instance type identified by the classifier 501. This strategy defines the rules for scheduling memory resources from the service platform to the instance, and its formulation is based on the following key factors: Instance type. Transient instances and stable instances have different memory usage characteristics, so different scheduling strategies are required. SLO, ensure the availability and performance indicators of the service and avoid SLO breaches due to improper resource scheduling. Physical resource limitations. The total amount of physical memory of the service platform is a hard limit, and the resource scheduling strategy must be implemented without exceeding this total amount. Resource utilization target. While ensuring service stability, improve resource utilization and reduce resource waste.

[0282] Optionally, different instance types correspond to different resource scheduling strategies. For stable instances, that is, for instances with relatively stable memory usage, the scheduler can adopt a more aggressive strategy to allow a certain degree of memory overselling to improve resource utilization efficiency. This may include resource adjustment based on prediction models, dynamic allocation mechanisms, etc. For transient instances, that is, for transient instances with large and unpredictable memory usage fluctuations, the scheduler should adopt a more conservative strategy to ensure that the instance has sufficient memory quota, avoid OOM errors, and ensure service continuity and stability. This means not overselling memory, and even reserving additional memory to cope with sudden resource demands.

[0283] Optionally, the scheduler 502 performs specific resource scheduling operations on the instance according to the determined resource scheduling strategy. The above resource scheduling operations may include: dynamic memory allocation, adjusting the memory quota of the instance based on the instance type and current resource usage. Instance migration, when the service platform where the instance is currently located is resource-constrained or the allocation strategy changes, the scheduler can migrate the instance to other more suitable service platforms to ensure that the instance has sufficient resource support. Resource reservation and release, reserve or release memory resources according to the instance type and load to optimize the overall utilization of resources.

[0284] Optionally, the intelligence of the scheduler 502 is reflected in its ability to adaptively adjust the resource scheduling strategy according to the instance type and current resource status to achieve efficient resource utilization and stable operation of the system. Its dynamicity is reflected in its ability to respond to changes in instance resource requirements in real time and quickly adjust the scheduling strategy to adapt to the changing workload.

[0285] Optionally, the scheduler 502 can be implemented based on a variety of algorithms, such as an adaptive instance size estimation algorithm, a hybrid packing algorithm, a MUR strategy, and a Fallback mechanism for dynamic threshold monitoring and online instance migration. These algorithms and strategies together constitute the core of the scheduler, enabling it to accurately identify resource requirements and efficiently schedule resources.

[0286] In the embodiment of the present application, by implementing differentiated resource scheduling strategies, the scheduler 502 can significantly improve resource utilization while reducing OOM risks, ensuring service stability and SLO achievement. Specific advantages include: improved resource utilization efficiency, allowing moderate overselling for stable instances, and improving resource utilization efficiency. Enhanced service stability, adopting a conservative strategy for transient instances, reserving sufficient memory, reducing the occurrence of OOM errors, and ensuring service continuity and reliability. Cost-effectiveness optimization, through intelligent scheduling, reducing excessive resource reservation and reducing operating costs.

[0287] In summary, the scheduler 502 achieves effective management and efficient use of resources by intelligently determining and executing resource scheduling strategies, while ensuring the stability and reliability of services, and is an indispensable part of service platform resource management. By continuously optimizing scheduling algorithms and strategies, the scheduler 502 can better adapt to changing workloads and resource requirements, and further improve the efficiency and effectiveness of resource management.

[0288] In the embodiment of the present application, a memory resource scheduling system 500 is provided. At least one instance in the database is obtained through a classifier 501; the instance type of the instance is identified; a resource scheduling strategy corresponding to the instance type is determined through a scheduler 502; and a resource scheduling operation is performed on the instance according to the resource scheduling strategy. Thus, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of low scheduling efficiency of memory resources is solved.

[0289] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application, such as the data for verification, are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0290] The embodiments of the present application are further explained below.

[0291] Currently, in the field of cloud computing, efficient resource management is crucial to maintain performance, meet SLOs, and optimize costs for users and cloud service providers. Memory management is particularly challenging because memory often becomes a key resource bottleneck.

[0292] To improve resource utilization, cloud service providers usually adopt a memory overselling strategy, which allows more instances to be deployed on each server and allocates more memory than the actual physical memory. Although this strategy can improve resource utilization, it also brings new challenges.

[0293] Excessive memory overselling will increase host memory utilization, which in turn increases the risk of out-of-memory (OOM) errors. OOM errors can cause database instances to terminate, causing service interruptions, severely affecting service availability, and violating the SLO agreed with customers. Especially in high-utilization environments, traditional methods of predicting memory utilization based on historical data are prone to errors due to limited prediction accuracy and cannot effectively avoid OOM errors.

[0294] Taking the cloud database system in the related technology as an example, the occurrence of OOM errors observed in the production environment conforms to the Pareto principle, that is, most OOM errors originate from a small number of transient instances. The memory usage of transient instances fluctuates greatly and is unpredictable, and the prediction methods in the related technology are difficult to accurately identify and process. At the same time, the memory usage of most instances is relatively stable, but due to concerns about transient instances, the overall resource scheduling strategy is conservative and resource utilization cannot be further improved.

[0295] Therefore, a new method is urgently needed to improve memory utilization while reducing the risk of OOM errors and ensuring the achievement of SLO.

[0296] In a related technology, a prediction method based on historical data and a time series prediction model are proposed. Time series models such as ARIMA, LSTM, and Holt-Winters are used to predict the historical memory usage data of database instances. Based on the prediction results, future memory requirements are estimated to allocate and schedule resources. Statistical analysis methods use the statistical characteristics of historical data (such as mean, variance, and quantiles) to estimate future memory usage and assist in resource scheduling decisions. However, in an environment with high memory utilization, the volatility and uncertainty of memory usage increase with the above methods. Especially for transient instances, traditional prediction models are difficult to accurately capture their memory usage patterns, resulting in increased prediction errors. For instances with a sudden increase in memory usage, the prediction model may not be able to reflect the changes in a timely manner, resulting in insufficient resource allocation and increased OOM risk. To avoid the risks brought by prediction errors, a conservative resource allocation strategy is usually adopted to reserve more resources, resulting in insufficient memory resource utilization.

[0297] In another related technology, a static overselling strategy is proposed, which sets a fixed memory overselling rate (for example, the total memory allocation does not exceed 120% of the physical memory), allocates memory according to the preset ratio during resource scheduling, and allows a certain degree of overselling. However, the above method lacks dynamic adaptability. The fixed overselling rate cannot be adjusted according to the actual memory usage and load changes of the instance, and cannot adapt to variable workloads and instance behaviors. Waste of resources or increased risks: Too high an overselling rate may increase the risk of OOM and violate the SLO; too low an overselling rate cannot fully improve resource utilization, resulting in resource waste. It is impossible to treat instances differently, and the same strategy is applied to each instance, and it is impossible to optimize instances with different memory usage characteristics.

[0298] In another related technology, a single bin packing algorithm is proposed. In resource scheduling, a fixed bin packing algorithm (e.g., First-Fit) is used to allocate resources to instances without considering the memory usage pattern and risk level of the instances. However, the above method lacks flexibility. The single bin packing algorithm cannot be adjusted according to the characteristics of the instance, resulting in a rigid resource scheduling strategy that is difficult to adapt to complex and changing instance characteristics. The scheduling effect is limited, and it is impossible to effectively optimize resource utilization or reduce OOM risks. In an environment with high memory utilization, the scheduling efficiency and effect are limited. It is unable to handle special situations, such as the inability to perform special processing for transient instances, which may lead to OOM errors.

[0299] In summary, the related technologies still have problems such as insufficient prediction accuracy, lack of consideration of instance characteristics, rigid scheduling strategies, and lack of emergency mechanisms. Therefore, there are still technical problems in the scheduling efficiency of memory resources.

[0300] Furthermore, the present application provides a cloud database memory overselling optimization method based on instance classification and adaptive scheduling. By classifying instances into transient and stable states, combined with feature engineering, the memory usage pattern of instances is deeply analyzed to improve prediction accuracy. An adaptive instance scheduler is designed. According to the instance classification results, different item size estimation methods and scheduling strategies are adopted to achieve dynamic optimization of resource scheduling. The MUR memory reservation strategy is introduced to optimize memory reservation and reduce resource waste based on the user's potential maximum memory demand. The Fallback mechanism is integrated to timely trigger instance migration and other measures when the memory utilization is too high or the prediction is wrong to prevent OOM errors and ensure service availability. In instance classification, iterative pseudo-label training and Markov chain state transition model are combined to improve the performance of the classifier and enhance the recognition ability of transient instances. Through the above method, the memory overselling potential of stable instances is fully explored and the memory utilization is significantly improved. Accurately identify and specially handle transient instances to reduce the occurrence of OOM errors. Ensure the availability and reliability of services by optimizing scheduling strategies and emergency mechanisms. It has dynamic adaptability and flexibility and can cope with complex and changing workloads and emergencies. Thereby, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of scheduling efficiency of memory resources is solved.

[0301] The above method of this embodiment is further introduced below.

[0302] In the embodiments of the present application, Figure 6 is a schematic diagram of a cloud database memory overselling optimization system architecture based on instance classification and adaptive scheduling according to an embodiment of the present application, such as Figure 6 As shown in the figure, the architectural design of a cloud database memory over-subscription optimization method based on instance classification and adaptive scheduling is described. Database instances represent multiple running instances in a cloud database system, each of which has its own runtime metrics (e.g., CPU utilization, memory utilization, disk I / O, etc.) and runtime events (e.g., OOM events). The above instances are the main objects of system optimization, and their resource requirements and behavior patterns will directly affect resource scheduling strategies and the way memory over-subscription is handled. The memory over-subscription profiling module is responsible for collecting and analyzing the runtime and non-runtime characteristics of database instances in order to classify the instances. It includes two main data sources: non-runtime characteristics, such as the allocation of CPU / memory / disk, primary image, customer information, etc., which do not change with the instance running status. Runtime characteristics, such as CPU / memory / disk utilization, cache pool hit ratio, QPS, OOM events, etc., which change in real time with the instance running status.

[0303] like Figure 6As shown in the figure, the instances are classified into transient instances and stable instances by the pattern classifier. This classifier is optimized based on the Markov model and iterative pseudo-label training to improve the classification accuracy and the recognition ability of transient instances. The instance size estimator is used to evaluate the memory requirements of database instances and provide a basis for resource scheduling. It uses multiple methods to estimate instance size, including: proportion & utilization, time series prediction, random binning, and quantiles. Each method can be selected according to different instance behavior patterns to improve the accuracy of prediction and the efficiency of resource scheduling.

[0304] like Figure 6 As shown in Figure 1, the packing constraint is a constraint in the resource scheduling process, which is used to guide the decision of the instance size estimator. Constraints can include memory utilization upper limit, instance priority, etc. to ensure that resource scheduling not only meets instance requirements but also complies with the overall resource allocation rules. The adaptive instance scheduler is responsible for dynamically allocating resources using adaptive strategies based on instance classification results and size estimates. The adaptive scheduler interacts with the reservation optimizer and the core scheduler (Eigen Master) in the control plane to obtain the latest scheduling instructions and resource status information. The control plane consists of multiple sub-modules, including: the reservation optimizer, which is responsible for optimizing memory reservation policies, such as maximum user reservation, to reduce resource waste. The Eigen Master can be a central node that coordinates the scheduler and other components, responsible for global resource scheduling decisions and status monitoring.

[0305] like Figure 6 As shown, the online node (Online Node) involves the actual execution of resource scheduling, including: real-time migration of database instances, online migration of instances when memory utilization is too high or prediction errors occur, to release resources or avoid OOM errors. The local agent (Eigen Agent) is responsible for executing instance migration and other local agents for scheduling instructions. When using Kubernetes as the underlying scheduling system, the node agent (Kublet) is a key component that collaborates with the Eigen Agent to execute scheduling decisions. Modeling the Impact of Memory Over-Subscription on SLO can include a quadratic logistic regression model or other predictive models to evaluate and optimize the relationship between memory overselling and service-level objectives.

[0306] In summary, Figure 6It demonstrates the modular design concept, decomposing complex functions into multiple independent modules to facilitate independent implementation and optimization of functions. Through instance classification and size estimation, the system can intelligently identify different types of instances and adopt differentiated processing strategies to find a balance between resource utilization and service reliability. The integrated Fallback mechanism allows the system to trigger instance migration in real time when high memory load or prediction errors are detected, effectively avoiding OOM errors and ensuring service stability. The maximum user retention policy can reduce the over-reservation of resources for a single user, release more resources for other instances, and improve the overall utilization of resources. Components such as Eigen Master and Kublet ensure the coordination between global scheduling strategies and local resource adjustments, and achieve efficient and automated resource scheduling.

[0307] Alternatively, through observation of the production environment, it is found that the occurrence of OOM errors conforms to the Pareto principle: most OOM errors originate from a few transient instances. Therefore, accurate identification and special handling of these transient instances become the key to reducing OOM risks. The adoption of the Pareto principle is mainly a data-driven rule of thumb. The core idea of ​​this principle is that a few key factors often cause most problems. In actual application scenarios, about 20% of transient instances may cause 80% of OOM errors. It should be noted that the ratio of the above transient instances to stable instances is not fixed, but can be calibrated and verified based on the data in the specific production environment.

[0308] Optionally, data collection can collect runtime features and non-runtime features. Runtime features can include events such as CPU utilization, memory utilization, disk I / O utilization, buffer pool hit rate, number of queries per second (QPS), instance restart, etc. Non-runtime features can include CPU, memory, disk allocation, instance role (master node or slave node), and customer service level.

[0309] Optionally, feature engineering can include data preprocessing, feature extraction, feature selection, and feature normalization. For data preprocessing, data cleaning, missing value processing, and outlier detection can be performed. For feature extraction, useful features can be extracted from the original data, such as the volatility and growth rate of memory utilization. For feature selection, correlation analysis, principal component analysis, and other methods can be used to select features that have a significant impact on classification. For feature normalization, numerical features can be standardized to avoid the impact caused by inconsistent dimensions.

[0310] Optionally, if the change in memory utilization of an instance exceeds a certain threshold (such as 5%) within a period of time, the instance is marked as a transient instance. Otherwise, it is marked as a stable instance.

[0311] Optionally, class imbalance can be handled by: using oversampling or undersampling methods to balance the sample ratio of transient and stable instances. Using cost-sensitive learning methods to give higher weight to minority classes in model training.

[0312] Optionally, use algorithms suitable for handling class imbalance problems, such as XGBoost, LightGBM, etc. Construct a state transition matrix, that is, calculate the state transition probability of instances in adjacent time periods. Use the initial model prediction results, adjust the prediction sequence, and generate new pseudo labels. Retrain the model using pseudo labels to improve the recognition ability of transient instances.

[0313] Figure 7 is a flow chart of an instance classifier training method according to an embodiment of the present application, such as Figure 7 As shown, the method may include the following steps:

[0314] Step S702: training data.

[0315] In this embodiment, the original data set prepared for training includes runtime and non-runtime features of the database instance. These features may include, but are not limited to, CPU utilization, memory usage, disk I / O, buffer pool hit rate, queries per second (QPS), instance restart events, memory allocation, instance role (master node or slave node), and customer service level. Provide basic data for subsequent model training to ensure that the model can learn the patterns and rules of instance classification from these data.

[0316] Step S704: train the initial XGBoost model M using the training data.

[0317] In this embodiment, XGBoost is an efficient machine learning algorithm that is particularly suitable for processing large-scale data and high-dimensional features. An initial XGBoost classifier is trained using the training data prepared above. A preliminary classifier is established that can predict whether an instance belongs to a transient or stable category based on the input features.

[0318] Step S706, using the XGBoost model M to perform prediction on the original data set to obtain a prediction result.

[0319] In this embodiment, the training data set is input into the trained XGBoost model again, and the model will give a classification prediction for each instance based on the current prediction ability. The prediction performance of the model on unseen data is tested to obtain preliminary classification results.

[0320] Step S708: Use the Markov chain model to adjust the prediction result.

[0321] In this embodiment, the Markov chain model is used to capture the transition rules between instance states, and the time dependency of the instance state sequence is used to adjust the prediction results of the classifier, especially when dealing with the class imbalance problem, to enhance the recognition of transient instances. Through iterative pseudo-label training of the Markov chain model, the accuracy of the classifier's prediction of transient instances is improved and misclassification is reduced.

[0322] Step S710: Generate initial pseudo labels.

[0323] In this embodiment, based on the adjusted prediction results in step S708, pseudo labels are generated for samples that the classifier fails to accurately predict, and these pseudo labels will be used for subsequent training enhancement. By introducing pseudo labels, the model can be further trained and optimized using the labels predicted by the model itself in the absence of sufficient labeled data.

[0324] Step S712: Use the enhanced data set and pseudo labels to enhance the training data set.

[0325] In this embodiment, the pseudo labels generated in step S710 are merged with the original data set to form an enhanced data set for the next round of model training. The amount of training data for the classifier is increased, especially for samples that were misclassified in the initial model, to improve the generalization ability and classification accuracy of the model.

[0326] Step S714, train the XGBoost model M again using the enhanced training data set.

[0327] In this embodiment, the XGBoost model is retrained using the enhanced data set to iteratively optimize the prediction capability of the model. Through iterative training, the parameters of the model are adjusted so that the model can better learn the difference between transient instances and stable instances, thereby improving classification performance.

[0328] Step S716, determining whether the performance has converged.

[0329] In this embodiment, during the model training process, the performance of the model is regularly evaluated to determine whether the classification accuracy has reached a stable state, that is, whether the model has "converged". Ensure the efficiency of model training and avoid overtraining and model overfitting. When the model performance converges, the training can be stopped and the current model can be used for instance classification. If convergence is not achieved, the process can return to step S706 and continue iterating until convergence.

[0330] In an embodiment of the present application, XGBoost is used as a training model for OOM classification of database instances. During the training process, by real-time monitoring of key performance indicators (such as accuracy, precision, recall, and F1 value) on the validation set, when the improvement of the indicators is lower than the preset tolerance ε in N consecutive iterations, it can be determined that the above model has reached a convergence state, and then the training is terminated. The above method not only ensures the stability of the classifier performance, but also effectively prevents overfitting through the built-in early stopping mechanism, thereby improving the generalization ability of the model in actual scenarios.

[0331] Step S718, using the final model.

[0332] In this embodiment, when the model training process is completed, that is, when the model performance converges, the final optimized XGBoost model is obtained. The above final model can be deployed in the production environment to classify database instances in real time, guide the adaptive instance scheduler and resource management strategy, reduce OOM errors, and improve resource utilization and service stability.

[0333] In summary, the embodiment of the present application can use the XGBoost model for preliminary classification training, establish a basic instance classifier (i.e., an initial prediction model), and use the instance classifier to determine the instance classification result of the initial training sample (i.e., the initial prediction instance type), and realize Markov model integration, and perform Markov state estimation through the integrated Markov model (i.e., the instance type adjustment model). Considering that the instance state has dependence on the time series, the transfer between the stable and transient states is modeled using the Markov process. The instance classification result is input into the Markov model for modeling processing, thereby obtaining the adjusted instance classification result. The adjusted instance classification result is determined as a pseudo-label (i.e., the target sample label), and the initial training sample is determined as a pseudo-training sample (i.e., the target training sample) corresponding to the pseudo-label.

[0334] In this embodiment, the role of the pattern classifier is to select a suitable instance size estimation method for a database instance classified as a stable instance according to its time series pattern of memory usage.

[0335] Optionally, predefined rules: Proportion & Usage Method, applicable to instances where memory usage is proportional to allocation. Quantile Method, applicable to instances where memory usage presents a certain distribution. SBP, applicable to instances where memory usage is random. TSF, applicable to instances where memory usage has a clear trend or periodicity.

[0336] For example, for the ratio & utilization method, the memory usage can be determined by the following formula, so that the instance size can be predicted by the memory usage. That is, by calculating the memory usage, the upper limit of the potential memory demand of an instance in the future can be obtained. Then, the memory usage can be used to guide the adjustment of the instance size:

[0337]

[0338] in, Can be used to indicate the The predicted memory usage of each instance; Can be used to indicate a preset scale factor; Can be used to indicate the assignment to The total amount of memory for each instance, that is, the memory allocated to the instance; Can be used to indicate the The actual memory usage of the instance, that is, the memory used by the instance.

[0339] For another example, for the quantile method, the memory usage can be determined by the following formula, so the instance size can be predicted by the memory usage. That is, by calculating the memory usage, the upper limit of the potential memory demand of an instance in the future can be obtained. Then, the memory usage can be used to guide the adjustment of the instance size:

[0340]

[0341] in, Can be used to indicate the The predicted memory usage of each instance; Can be used to represent the quantiles of a calculated data set; Can be used to indicate the A dataset of historical memory usage for instances.

[0342] As an optional example, SBP can assume that memory usage follows a certain statistical distribution and use its mean and variance to estimate it, and use time series models such as Holt-Winters to predict future memory usage peaks.

[0343] For example, the pseudo code of the instance size estimator flow is analyzed as follows:

[0344] :Set of n database instances ({ ,..., }).Methods: Item size estimation methods.

[0345] Ensure: :Predicted memory usage for each instance.

[0346] The above pseudo code illustrates the input and output required by the algorithm. Specifically, initialize the output set ( ) is an empty set ( ← ), which is used to store the predicted memory usage of each instance. Traverse the instance collection ( ), for each instance ( ):if( ) is a transient instance (memory usage fluctuates widely and unpredictably): set directly ( )equal( ) of the allocated memory ( ). The memory usage of transient instances is difficult to predict, so the allocated memory is used as the upper limit of the prediction to avoid OOM events caused by memory overselling. If ( ) is a stable instance (memory usage is relatively stable). ) characteristics to select the appropriate prediction method (Methodi). If (Methodi) is set to Proportional & Utilization Rate Method, the preset proportionality factor ( ) and allocate memory ( ) multiplied by the actual memory used ( ) and take the larger value as the predicted memory usage ( The above method is suitable for the case where the memory usage of the instance is proportional to the allocation amount, or the memory usage of the instance has a clear baseline level.

[0347] If (Method) is set to the quantile method, the calculation example ( )'s memory usage history data ( ) as the predicted memory usage ( ). The above strategy is applicable to the case where the memory usage of the instance shows a certain distribution pattern. Using quantiles can more robustly estimate the peak memory usage that the instance may reach. If (Methodi) is set to SBP, assuming that the memory usage follows a certain statistical distribution (for example, normal distribution), use the mean of the distribution ( ) and variance ( ) to estimate the memory usage of the instance ( ). Random binning is suitable for instances where memory usage is random but can be described by a statistical distribution model. If (Methodi) is set to TSF, a time series forecasting model is used, such as the Holt-Winters forecasting method ( )'s future memory usage peak, and take the maximum value as the predicted memory usage ( ). Time series forecasting is suitable for instances where the memory usage has obvious trends or periodic patterns and can predict future usage based on historical data.

[0348] For each estimated instance ( ), add it to ( ) collection, ensure that ( ) contains the predicted memory usage of each instance. After completing the prediction for each instance, return the collection of predicted memory usage ( ).

[0349] In this embodiment, the adaptive instance scheduler can adopt different scheduling strategies according to the classification results of the instance, combine multiple item size estimation methods, optimize resource allocation, improve memory utilization, dynamically adjust resource scheduling, and adapt to changes in instance memory usage.

[0350] Optionally, the specific operation process of the adaptive instance scheduler dynamically optimizes resource allocation by combining instance classification results and multiple instance size estimation methods to achieve high efficiency and flexibility in resource scheduling. The specific scheduling process can be: at the beginning of the process, each instance is considered to be in an unallocated state, that is, the value of the variable allocated is set to False (allocated ← False), which means that the instance has not been allocated to any machine. In the process of traversing instances, you can select from the instance set Select an instance from For the selected instance , will traverse the machine set Each machine in , try to assign the instance to the above machine The total memory usage of the current machine, including both deterministic and uncertain parts, can be calculated using the following formula:

[0351]

[0352] in, Can be used to indicate the total memory usage of the current machine; Can be used to represent deterministic memory usage, which can be known; Can be used to represent the mean of uncertain memory usage; Can be used to represent the standard deviation of uncertain memory usage.

[0353] Optionally, the MUR strategy can be used to update the upper bound of the packing constraints , check whether the constraints are satisfied: If the above constraints are met, then the instance Divided into machines If the above constraints are not met, you can try the next machine; if all machines cannot be allocated, you can start a new machine. Repeat the above steps until all instances are allocated.

[0354] Optionally, the scheduling algorithm aims to solve the problem of memory overselling when allocating cloud database resources, with the goal of improving resource utilization while avoiding OOM errors. The above scheduler algorithm is as follows: traverse the database instance to be allocated ( ), try to find a suitable machine for each instance. , the algorithm will traverse each machine ( ), check whether the instance can be Assign to machine Calculate the current machine The remaining resource constraints , then according to the example The memory usage (deterministic and non-deterministic) and resource requirements of the instance are used to determine whether the allocation can be made. Can be safely assigned to machines If the instance is found to have a suitable allocation location, the algorithm will start a new empty machine or execute other emergency strategies to ensure that the instance can obtain resources.

[0355] For example, the following pseudo code can be used ∈ do means that the instance collection will be traversed Each instance in allocated ← False, when trying to allocate an instance Before, its allocation status is initialized to False, indicating that the instance has not been allocated. ∈ do, ← ,in, It is a machine that calculates The upper limit of the acceptable packing constraint is set to (global packing constraint upper limit) and the smaller value of MUR. The MUR function takes into account the machine Currently running instances and new instances May result in resource adjustments to ensure resource limits are not exceeded. ← 0, ← , initialize the deterministic resource usage of the current machine ( ) and uncertain resource usage ( ). if is deterministic then, ← + , else, ← , algorithms for machines Examples on Traverse, according to the example The resource usage of the current machine is calculated based on the resource usage nature (deterministic or uncertain). For deterministic resource usage, the actual .

[0356] For uncertain resource use, the parameters of the statistical distribution (mean and variance ) to estimate potential resource usage, using Function to update. then Allocate to , allocated ← True, break, check whether the current machine can safely accommodate the new instance through the above steps By calculating the upper limit of uncertain resource usage ( ) and the upper bound on the remaining resources of the machine minus the deterministic resource usage ( ) for comparison. If the upper limit of the uncertain resource usage is less than the upper limit of the remaining resources, then the instance Can be assigned to a machine and update the instance The allocation status of is True. If notallocated then, Allocate Ii to a new empty machine. If after traversing each machine, the instance If it is still not assigned, the Fallback mechanism is enabled and the instance Allocate to a newly created empty machine to ensure that each instance gets the necessary resources.

[0357] The above adaptive instance scheduler algorithm uses fine instance classification and status evaluation, combined with the calculation of deterministic and uncertain resource usage, to dynamically check and adjust instance allocation to maximize resource utilization while reducing the risk of OOM errors. The key to the algorithm is the ability to handle different types of instance resource requirements, and to ensure that each instance can get appropriate resource allocation by starting a new machine or taking other emergency measures when resource allocation constraints appear. The above adaptive and flexible scheduling strategy is the key to optimizing cloud database resource management and memory overselling.

[0358] In this embodiment, the self-adaptation of the scheduling policy is reflected in the following three aspects: For transient instances, memory is not oversold, and its memory requirements are prioritized. For stable instances, memory is oversold based on the instance size estimation results to improve resource utilization. Dynamic adjustment: According to the changes in the memory usage of the instance, the scheduling policy is updated in real time, and instance migration is triggered when necessary. That is, when the memory usage status of the instance changes, for example, from a stable type to a transient type, the scheduling policy of the instance can be selected (adjusted) accordingly to adapt to the changes in memory usage and ensure the SLA guarantee of the instance.

[0359] For example, the potential additional memory requirement per user can be calculated as follows:

[0360]

[0361] in, Can be used to represent users The collection of instances of ; Can be used to represent instances The memory requirement of the instance, that is, the upper limit of memory allocated to the instance (the maximum memory allowed); Can be used to represent instances The actual amount of memory used (memory utilization); Can be used to represent users Instances of The sum of potential additional memory requirements, that is, the additional memory requirement.

[0362] As another example, the total memory reserve can be determined by the following formula:

[0363]

[0364] in, Can be used to represent a collection of users; Can be used to indicate the total amount of memory that needs to be reserved (that is, the total memory reservation).

[0365] In an embodiment of the present application, the above-mentioned MUR memory reservation optimization can reduce the amount of memory reservation and avoid reserving memory for the maximum potential demand of each user at the same time. It can also improve resource utilization and release excess memory reservation for use by more instances. It can also meet the SLO requirements. In practice, the probability of memory bursts occurring at the same time in instances of different users is extremely low, so this strategy can meet the requirements of service reliability. When each instance scheduling calculation is performed, that is, when the MUR algorithm is running, it is triggered to determine whether the scheduling meets the MUR conditions. The MUR threshold is a hyperparameter determined based on production experience. Under the MUR threshold, the SLA can be guaranteed without wasting resources.

[0366] In this embodiment, the integration of the Fallback mechanism can take into account possible errors in the prediction and scheduling process, and an emergency mechanism is needed to prevent OOM errors and ensure service continuity. The Fallback mechanism can perform dynamic threshold monitoring, set a memory utilization threshold (such as 85%), and monitor the memory utilization of the machine in real time. When the threshold is exceeded, the Fallback mechanism is triggered. Online instance migration can also be performed to select the migration instance, that is, the appropriate instance can be selected for migration based on memory usage and service impact. Coordinated migration, that is, through the collaboration of Eigen Agent and Eigen Master, seamless migration is achieved to minimize the impact on the service. OOM prevention can also be performed. When the memory utilization is too high, measures can be taken in advance to avoid the occurrence of OOM errors.

[0367] Table 1 is a classification performance comparison result of an embodiment of the present application. As shown in Table 1, the precision and recall indicators for classifying transient instances are shown. The precision of LightGBM is 61.28% and the recall is 85.54%. XGBoost has significant improvements in both precision (79.53%) and recall (87.22%). The performance of the stacking classifier is better, with precision and recall reaching 83.47% and 87.84% respectively. The highest performance is the Stacking classifier combined with the Markov model, with a recall of 90.06%, but a slightly lower precision of 82.12%. The main focus is on the recall rate, because identifying more transient instances contributes significantly to improving memory overselling. Experimental data show that by introducing iterative optimization into the Stacking method, the recall rate is effectively improved, and the detection capability of transient instances is brought to an appropriate level.

[0368] Table 1 A comparison of classification performance

[0369]

[0370] Figure 8 (a) is a schematic diagram of a comparison of the number of OOM errors according to an embodiment of the present application. As shown in Figure 8 (a), the number of OOM errors of each machine in the cluster is shown. The number of OOM events of the Eigen+Markov method is almost as small as that of the Eigen+Opt method, and is significantly lower than the Naive Baseline method. Figure 8 (b) is a schematic diagram of a comparison of memory utilization according to an embodiment of the present application. As shown in Figure 8 (b), the cumulative distribution function of the memory utilization of each machine in the cluster is shown. The memory utilization distribution of the Eigen+Markov method almost completely overlaps with that of the Eigen+Opt method, and is significantly higher than the NaiveBaseline method. The above results further demonstrate the efficiency of the method of the embodiment of the present application in maximizing memory utilization and minimizing OOM errors.

[0371] Optionally, by comparing the number of OOM errors of different algorithms, it can be seen that no matter which algorithm is used in the scheduling process, the number of OOM errors is significantly reduced after the introduction of MUR optimization. Eigen+ achieved zero OOM errors in most clusters, which is very close to Eigen-Optimal. In contrast, even after optimization, the Naive method still had OOM errors in OLTP and OLAP clusters.

[0372] Alternatively, by comparing the memory utilization of different algorithms, we can see that Eigen+ maintains a high level of memory utilization, reaching 85.56%, which is very close to Eigen-Optimal's 86.69%. In contrast, the highest utilization of the Naive method is 70.63%, and Eigen+ achieves significant improvements while maintaining very few OOM events. These results highlight the excellent performance of the Eigen+ method, which effectively strikes a balance between high memory utilization and low OOM errors, and significantly outperforms traditional methods.

[0373] Fig. 9 is a schematic diagram of a host memory utilization according to an embodiment of the present application, such as Fig. 9 As shown in the figure, for the non-Eigen+ algorithm, in order to achieve a 95% SLO probability, the memory utilization threshold is about 72.42%. In contrast, Eigen+ significantly increases this threshold to about 88.13%, which is a huge improvement. This improvement shows that Eigen+ allows higher host memory utilization while maintaining the same SLO level as the original Eigen algorithm. Eigen+'s algorithms and optimizations more effectively manage the variability of memory resources and workloads, significantly reducing the risk of OOM events even at higher utilization levels.

[0374] FIG10(a) is a schematic diagram of an average allocation ratio according to an embodiment of the present application. As shown in FIG10(a), the effect of the Eigen+ algorithm on CPU and memory utilization and allocation within 13 weeks in a production environment is shown. The vertical dotted line marks the implementation time point of Eigen+. Before the adoption of Eigen+, the allocation of CPU and memory showed a stable growth trend, and the utilization level showed moderate fluctuations. After the adoption of Eigen+, both the allocation and utilization ratios showed a significant upward trend. Specifically, FIG10(b) is a schematic diagram of an average utilization according to an embodiment of the present application. As shown in FIG10(b), the average memory utilization continued to rise (from 41.35% to 58.25%), which corresponds to the improvement in allocation efficiency (from 75.67% to 111.88%). Importantly, no OOM event occurred during the entire monitoring period, indicating that the system operation remained stable under the action of the Eigen+ algorithm.

[0375] FIG10( c ) is a schematic diagram of utilization and the number of migration tasks according to an embodiment of the present application. As shown in FIG10( c ), the 99th percentile of the utilization ratio also continues to rise.

[0376] In the embodiment of the present application, the Pareto principle is used to find that most OOMs in cloud databases originate from a few transient instances. Based on this, complex prediction problems are converted into binary classification tasks to accurately identify and manage the above instances. Through memory overselling analysis and instance classification, resource scheduling is adaptively optimized to achieve a balance between resource utilization and service reliability. Transient and stable instances are accurately classified using feature engineering and Markov chain models. According to the instance classification results, a variety of instance size estimation methods are used to dynamically optimize resource allocation. Quantifying the impact of memory overselling on SLO, it has been verified in production that Eigen+ increases the memory allocation rate of the MySQL cluster by an average of 36.21% (from 75.67% to 111.88%) without OOM, and maintains SLO compliance. Thereby achieving the technical effect of improving the scheduling efficiency of memory resources and solving the technical problem of scheduling efficiency of memory resources.

[0377] In this embodiment, in the field of memory management and resource scheduling of cloud databases, in addition to the memory overselling optimization method based on instance classification and adaptive scheduling proposed in this solution, there are other solutions. These solutions attempt to solve the contradiction between improving resource utilization and reducing OOM risks from different angles. Several representative solutions are introduced in detail below.

[0378] For example, the resource scheduling solution based on time series prediction uses time series models such as ARIMA and LSTM to model and predict the historical memory usage data of database instances. According to the predicted memory usage, resource allocation is adjusted in advance, and memory resources are dynamically allocated or released to meet the needs of the instance. Based on the containerization and elastic scaling solution, the database instance is encapsulated in a container to isolate resources and improve the flexibility of deployment. According to the load situation, containers are dynamically created or destroyed, the number of instances and resource allocation are adjusted, and elastic scaling of resources is achieved.

[0379] For another example, based on the static resource isolation strategy, fixed memory resources are reserved for each database instance, and the resource usage of the instance is strictly limited, and overselling is not allowed. Through operating system-level restrictions, resource isolation between instances is ensured to avoid resource competition. The multi-dimensional resource joint scheduling solution considers the usage and demand of multiple resources such as CPU, memory, and disk I / O when scheduling resources. Use multi-objective optimization algorithms to find the scheduling balance point in each resource dimension to maximize overall resource utilization. Use the prediction method of advanced machine learning models and use deep learning models (such as deep neural networks and reinforcement learning) to model and predict the memory usage of instances. The model can continue to learn and self-optimize as data increases and the environment changes.

[0380] In summary, the above methods still have the following shortcomings: Most solutions fail to identify and specially handle transient instances that cause OOM risks, making it difficult to reduce OOM risks while improving resource utilization. Many solutions use fixed or single scheduling strategies that cannot be dynamically adjusted according to instance characteristics, and resource utilization and service quality cannot be balanced. Some advanced prediction or scheduling methods are complex to implement, have high requirements for computing resources and data, and are difficult to deploy in actual production environments. Many solutions do not consider countermeasures when OOM risks occur, and the system's robustness and stability are insufficient.

[0381] However, the embodiment of the present application adopts different scheduling strategies for different types of instances through accurate instance classification, which not only improves resource utilization but also reduces OOM risk. Optimize memory retention, reduce resource waste, and meet the requirements of service reliability. When prediction or scheduling errors occur, timely measures can be taken to prevent OOM errors and ensure service stability. The overall design is reasonable, the algorithm complexity is moderate, and it is easy to deploy and apply in an actual production environment. In terms of improving resource utilization, reducing OOM risks, and ensuring SLO achievement, it has obvious technical advantages and can better meet the needs of cloud database systems.

[0382] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 2 The memory resource scheduling method shown is a memory resource scheduling device.

[0383] Fig.11 is a schematic diagram of a scheduling device for memory resources according to an embodiment of the present application, such as Fig.11 As shown, the memory resource scheduling device 1100 may include: a first acquisition unit 1102 , a first identification unit 1104 , a first determination unit 1106 and a first execution unit 1108 .

[0384] The first acquiring unit 1102 is configured to acquire at least one instance in the database.

[0385] The first identification unit 1104 is configured to identify the instance type of the instance.

[0386] The first determining unit 1106 is configured to determine a resource scheduling strategy corresponding to the instance type.

[0387] The first execution unit 1108 is used to perform resource scheduling operations on the instance according to the resource scheduling policy.

[0388] Here, the first acquisition unit 1102, the first identification unit 1104, the first determination unit 1106 and the first execution unit 1108 correspond to steps S202 to S208 in the embodiment, and the four units are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1404) and processed by one or more processors (e.g., processors 1402a, 1402b..., 1402n), and the above units can also be run as part of the device in the computer terminal A provided in the following embodiment.

[0389] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 3 The memory resource scheduling method shown is a memory resource scheduling device.

[0390] Fig.12 is a schematic diagram of another memory resource scheduling device according to an embodiment of the present application, such as Fig.12 As shown, the memory resource scheduling device 1200 may include: a second acquisition unit 1202 , a second identification unit 1204 , a second determination unit 1206 and a second execution unit 1208 .

[0391] The second acquisition unit 1202 is configured to acquire at least one instance in the cloud database system.

[0392] The second identification unit 1204 is configured to identify the instance type of the instance.

[0393] The second determining unit 1206 is configured to determine a resource scheduling strategy corresponding to the instance type.

[0394] The second execution unit 1208 is used to execute resource scheduling operations on the instance according to the resource scheduling policy.

[0395] It should be noted that the second acquisition unit 1202, the second identification unit 1204, the second determination unit 1206 and the second execution unit 1208 correspond to steps S302 to S308 in the above embodiment, and the four units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1404) and processed by one or more processors (e.g., processors 1402a, 1402b..., 1402n), and the above units can also be run as part of the device in the computer terminal A provided in the following embodiment.

[0396] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 4 The memory resource scheduling method shown is a memory resource scheduling device.

[0397] Fig.13 is a schematic diagram of another memory resource scheduling device according to an embodiment of the present application, such as Fig.13 As shown, the memory resource scheduling device 1300 may include: a first calling unit 1302 , a third identifying unit 1304 , a third determining unit 1306 , a third executing unit 1308 and a second calling unit 1310 .

[0398] The first calling unit 1302 is configured to obtain an instance type of at least one instance in the database by calling a first interface.

[0399] The third identification unit 1304 is configured to identify the instance type of the instance.

[0400] The third determining unit 1306 is configured to determine a resource scheduling strategy corresponding to the instance type.

[0401] The third execution unit 1308 is used to perform a resource scheduling operation on the instance according to the resource scheduling policy to obtain a resource scheduling result.

[0402] The second calling unit 1310 is used to output the resource scheduling result by calling the second interface.

[0403] Here, the first calling unit 1302, the third identifying unit 1304, the third determining unit 1306, the third executing unit 1308 and the second calling unit 1310 correspond to steps S402 to S410 in the above embodiment, and the five units are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above units can be hardware components or software components stored in a memory (e.g., memory 1404) and processed by one or more processors (e.g., processors 1402a, 1402b..., 1402n), and the above units can also be run as part of the device in the computer terminal A provided in the following embodiment.

[0404] In the scheduling device of the memory resource, if the scheduling of the memory resource is required, the instance can be obtained from the database deployed on the service platform. The change service of the memory resource required for the process of the acquired instance running on the service platform can be analyzed to identify the instance type of the instance. Through the change range of the required memory resources, the difference between the instances can be accurately grasped, and the resource scheduling strategy can be avoided from being too unified, so as to achieve the purpose of setting the scheduling strategy for instances of different instance types in a targeted manner. The resource scheduling strategy corresponding to the instance type can be determined, and the corresponding resource scheduling operation can be performed according to the resource scheduling strategy to schedule the memory resources to the instance from the service platform. In this embodiment, the instance can be classified according to the change range of the memory resources required by different instances, and adaptive scheduling can be performed for different types of instances in a targeted manner. The scheduling method of memory resources in the database combined with instance classification and adaptive scheduling realizes the dynamic optimization of resource scheduling, thereby overcoming the limitations of the related technologies based on historical data prediction, single binning algorithm, unified resource scheduling and memory overselling strategy, realizing the precision, dynamic and differentiation of resource scheduling, significantly improving the scheduling efficiency and resource utilization of memory resources, while reducing the OOM risk and ensuring the stability and performance of the service. This achieves the technical effect of improving the scheduling efficiency of memory resources and solves the technical problem of scheduling efficiency of memory resources.

[0405] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0406] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0407] In this embodiment, the computer terminal can execute the program code of the method in the method for scheduling memory resources.

[0408] Optionally, Fig.14 is a structural block diagram of a computer terminal according to an embodiment of the present application. Fig.14 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1402 , a memory 1404 and a transmission device 1406 .

[0409] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the scheduling method and device of memory resources in the embodiment of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the computer terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0410] Using an embodiment of the present application, a method for scheduling memory resources is provided. In an embodiment of the present application, if memory resources need to be scheduled, an instance can be obtained from a database deployed on a service platform. The change service of the memory resources required for the process of running the acquired instance on the service platform can be analyzed to identify the instance type of the instance. Through the change range of the required memory resources, the differences between the instances can be accurately grasped, and the resource scheduling strategy can be avoided from being too unified, so as to achieve the purpose of setting scheduling strategies for instances of different instance types in a targeted manner. The resource scheduling strategy corresponding to the instance type can be determined, and the corresponding resource scheduling operation can be performed according to the resource scheduling strategy to schedule memory resources to the instance from the service platform. In this embodiment, the instance classification can be performed according to the change range of the memory resources required by different instances, and adaptive scheduling can be performed in a targeted manner for different types of instances. Through the scheduling method of memory resources in the database that combines instance classification and adaptive scheduling, dynamic optimization of resource scheduling is achieved, thus overcoming the limitations of related technologies based on historical data prediction, single binning algorithm, unified resource scheduling and memory overselling strategy, etc., and achieving precise, dynamic and differentiated resource scheduling, significantly improving the scheduling efficiency and resource utilization of memory resources, while reducing OOM risks and ensuring the stability and performance of services. In addition, the technical effect of improving the scheduling efficiency of memory resources is achieved, and the technical problem of scheduling efficiency of memory resources is solved.

[0411] It can be understood by those skilled in the art that Fig.14The structure shown is for illustration only, and the computer terminal A may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, referred to as MID), a PAD, or other terminal devices. Fig.14 It does not limit the structure of the above-mentioned computer terminal A. For example, the computer terminal A may also include Fig.14 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Fig.14 Different configurations shown.

[0412] A person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0413] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the first embodiment.

[0414] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0415] Optionally, in this embodiment, a computer-readable storage medium is configured to store program codes for executing the above steps of the embodiment of the present application.

[0416] According to another aspect of an embodiment of the present application, a computing device is also provided. The computing device may include a memory and a processor: the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions. When the above-mentioned computer-executable instructions are executed by the processor, any one of the above-mentioned methods is implemented.

[0417] An embodiment of the present application may provide an electronic device, which may include a memory and a processor.

[0418] Fig.15It is a block diagram of an electronic device according to a method for scheduling memory resources in an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0419] like Fig.15 As shown, the device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1502 or a computer program loaded from a storage unit 1508 into a random access memory (RAM) 1503. In the RAM 1503, various programs and data required for the operation of the device 1500 can also be stored. The computing unit 1501, the ROM 1502, and the RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.

[0420] A number of components in the device 1500 are connected to the I / O interface 1505, including: an input unit 1506, such as a keyboard, a mouse, etc.; an output unit 1504, such as various types of displays, speakers, etc.; a storage unit 1508, such as a disk, an optical disk, etc.; and a communication unit 1509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1509 allows the device 1500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0421] The computing unit 1501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (Graphics Processing Unit, referred to as GPU), various dedicated artificial intelligence computing chips, various computing units running machine learning model algorithms, a digital signal processor (Digital Signal Processor, referred to as DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, such as the scheduling method of memory resources. For example, in some embodiments, the scheduling method of memory resources may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1500 via the ROM 1502 and / or the communication unit 1509. When the computer program is loaded into the RAM 1503 and executed by the computing unit 1501, one or more steps of the scheduling method of memory resources described above may be performed. Alternatively, in other embodiments, the computing unit 1501 may be configured to execute the scheduling method of memory resources in any other appropriate manner (for example, by means of firmware).

[0422] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method for scheduling memory resources of the embodiment of the present application is implemented.

[0423] According to an embodiment of the present application, a method for scheduling memory resources is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0424] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Fig.16 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for scheduling memory resources according to an embodiment of the present application, such as Fig.16As shown, the computer terminal 160 (or mobile device) may include one or more (shown in the figure as 1602a, 1602b, ..., 1602n) processors 1602 (the processor 1602 may include but is not limited to a processing device such as a microprocessor (Microcontroller Unit, referred to as MCU) or a programmable logic device (Field Programmable Gate Array, referred to as FPGA)), a memory 1604 for storing data, and a transmission device 1606 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art will understand that Fig.16 The structure shown is only for illustration and does not limit the structure of the above electronic device. Fig.16 More or fewer components as shown, or with Fig.16 Different configurations shown.

[0425] Fig.16 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the above-mentioned computer terminal 160 (or mobile device), but also as an exemplary block diagram of the above-mentioned server. In an optional embodiment, Fig.17 The block diagram shows the use of the above Fig.16 The computer terminal 160 (or mobile device) is shown as an embodiment of a computing node in the computing environment 1701 .

[0426] Fig.17 is a structural block diagram of a computing environment of a method for scheduling memory resources according to an embodiment of the present application, such as Fig.17 As shown, computing environment 1701 includes multiple computing nodes (such as servers) running on a distributed network (shown as 1710-1, 1710-2, ...) in the figure. The computing nodes all contain local processing and memory resources, and end users 1702 can remotely run applications or store data in computing environment 1701. Applications can be provided as multiple services 1720-1, 1720-2, 1720-3 and 1720-4 in computing environment 1701, representing services "F", "G", "I" and "H" respectively.

[0427] The end user 1702 can provide and access services through a web browser or other software application on the client, and in some embodiments, the end user 1702's provision and / or request can be provided to the entry gateway 1730. The entry gateway 1730 can include a corresponding agent to handle the provision and / or request for the service (one or more services provided in the computing environment 1701).

[0428] Services are provided or deployed based on various virtualization technologies supported by the computing environment 1701. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine, and execute programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system so that multiple workloads can run on a single operating system instance.

[0429] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Fig.17 As shown, service 1720-2 can be equipped with one or more Pods 1740-1, 1740-2, ..., 1740-N (collectively referred to as Pods). Pods can include a proxy 1745 and one or more containers 1742-1, 1742-2, ..., 1742-M (collectively referred to as containers). One or more containers in a Pod process requests related to one or more corresponding functions of the service, and the proxy 1745 generally controls network functions related to the service, such as routing, load balancing, etc. Other services can also be equipped with Pods similar to Pods.

[0430] During operation, executing a user request from the end user 1702 may require calling one or more services in the computing environment 1701, and executing one or more functions of a service may require calling one or more functions of another service. Fig.17 As shown, service "F" 1720-1 receives a user request from end user 1702 from ingress gateway 1730, service "F" 1720-1 may call service "G" 1720-2, and service "G" 1720-2 may request service "I" 1720-3 to perform one or more functions.

[0431] The computing environment described above can be a cloud computing environment, where the allocation of resources is managed by the cloud service provider, allowing the development of functions without considering the implementation, adjustment or expansion of servers. The computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be divided into a set of functions that can be automatically and independently scaled, rather than expanding a single hardware device to handle potential loads.

[0432] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0433] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0434] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory for short), an optical fiber, a portable compact disk read-only memory (CD-ROM for short), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0435] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD), a monitor for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a path ball), through which the user may provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0436] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0437] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0438] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0439] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0440] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0441] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0442] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0443] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory, random access memory, mobile hard disk, disk or optical disk, etc., which can store program code.

[0444] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A resource scheduling method, characterized in that: include: Acquire at least one instance in a database, wherein the database is deployed on a service platform, and the instance is used to provide a service function of the service platform; Identifying an instance type of the instance, wherein the instance type is divided according to a change in memory resources required by the instance during operation on the service platform; Determine a resource scheduling policy corresponding to the instance type, wherein different instance types correspond to different resource scheduling policies, and the resource scheduling policy is used to represent a rule for scheduling memory resources from the service platform to the instance; According to the resource scheduling policy, a resource scheduling operation is performed on the instance.

2. The method according to claim 1, characterized in that Identify the instance type of the instance, including: The instance is classified by using an instance classification model to obtain the instance type, wherein the instance classification model is obtained by training a Markov chain model using instance type labels of instance samples.

3. The method according to claim 2, characterized in that The instance type label includes a transient instance type label and a stable instance type label, wherein the transient instance type label is used to indicate that a change value of the memory utilization of the instance sample exceeds a change threshold, and the stable instance type label is used to indicate that a change value of the memory utilization of the instance sample does not exceed the change threshold. The instance is classified by using an instance classification model to obtain the instance type, including: Extracting runtime features and non-runtime features from the instance, wherein the runtime features at least include memory utilization of the instance, and the non-runtime features at least include memory allocation of the instance; The runtime features and the non-runtime features are classified and processed using the instance classification model to obtain the instance type, wherein the instance type is a transient instance type or a stable instance type, the transient instance type is used to indicate that the change value of the memory utilization of the instance exceeds the change threshold, and the stable instance type is used to indicate that the change value of the memory utilization of the instance does not exceed the change threshold.

4. The method according to claim 3, characterized in that Using the instance classification model, classifying the runtime features and the non-runtime features to obtain the instance type includes: Converting the runtime features and non-runtime features into target features that match the instance classification model; The target feature is classified using the instance classification model to obtain the instance type.

5. The method according to claim 3, characterized in that: The method further comprises: Training an initial Markov chain model using the transient instance type label and the stable instance type label; Using the trained initial Markov chain model, predicting the instance type of the unlabeled instance sample to obtain a prediction result; Generate a pseudo instance type label based on the prediction result and a state transition matrix of the instance sample, wherein the state transition matrix is ​​used to represent the probability of the instance sample performing state transition between adjacent time periods; The trained initial Markov chain model is trained using the pseudo instance type label to obtain the Markov chain model.

6. The method according to claim 3, characterized in that Determining a resource scheduling strategy corresponding to the instance type includes: In response to the instance type being the stable instance type, the resource scheduling policy corresponding to the stable instance type is determined according to the resource demand information of the instance, wherein the resource demand information is used to represent the actual memory resources required by the instance during the operation of the instance on the service platform, and the resource scheduling policy is used to represent the rules for scheduling memory resources from the service platform to the instance according to the memory overselling policy.

7. The method according to claim 6, characterized in that The method further comprises: In response to the instance type being the stable instance type, determining a memory usage pattern of the instance; Determining an instance size determination policy corresponding to the memory usage pattern, wherein the instance size determination policy is used to represent a rule for determining an instance size of the instance; Determine the instance size according to the instance size determination strategy; The resource requirement information that satisfies the instance size is determined.

8. The method according to claim 3, characterized in that Determining a resource scheduling strategy corresponding to the instance type includes: In response to the instance type being the transient instance type, the resource scheduling policy corresponding to the transient instance type is determined according to the resource requirement information of the instance, wherein the resource requirement information is used to represent the actual memory resources required by the instance during its operation on the service platform, and the resource scheduling policy is used to represent the rule of scheduling memory resources from the service platform to the instance that are greater than or equal to the actual memory resources and less than or equal to the physical memory of the service platform.

9. The method according to claim 1, characterized in that: The method further comprises: Determining current characteristic information of the instance, wherein the current characteristic information is used to represent current characteristics of the instance running on the service platform; The resource scheduling policy is adjusted using the current characteristic information and the instance type, wherein the adjusted resource scheduling policy matches the current characteristic information.

10. The method according to claim 1, characterized in that The method further comprises: Determine the total resource requirement information of the instance sets under different accounts respectively, wherein the total resource requirement information is used to represent the sum of actual memory resources required by different instances in the instance set during the operation of the service platform; Determine the maximum total resource demand information among the different total resource demand information corresponding to the different accounts; The maximum total resource requirement information is retained.

11. The method according to any one of claims 1 to 10, characterized in that The method further comprises: Determining memory utilization of the instance in the service platform; In response to the memory utilization being greater than a memory utilization threshold, memory resources scheduled to the instance are adjusted, or a migration operation is performed on the instance.

12. The method according to claim 11, characterized in that Determine the probability of achieving the service level objective (SLO) of the service function; The target-reaching probability is input into a logistic regression model for analysis to obtain the memory utilization threshold, wherein the logistic regression model is used to represent the nonlinear relationship between different memory utilizations of the instance and different target-reaching probabilities corresponding to the service function.

13. A resource scheduling method, characterized in that: Applied to resource management systems, including: Acquire at least one instance in a cloud database system, wherein the cloud database system is deployed on a cloud service platform, and the instance is used to provide a service function of the cloud service platform; Identifying an instance type of the instance, wherein the instance type is divided according to a change in memory resources required by the instance during operation on the cloud service platform; Determine a resource scheduling policy corresponding to the instance type, wherein different instance types correspond to different resource scheduling policies, and the resource scheduling policy is used to represent a rule for scheduling memory resources from the cloud service platform to the instance; According to the resource scheduling policy, a resource scheduling operation is performed on the instance.

14. The method according to claim 13, characterized in that Identify the instance type of the instance, including: The instance classification model in the resource management system is called to classify the instance to obtain the instance type, wherein the instance classification model is obtained by training a Markov chain model using instance type labels of instance samples.

15. The method according to claim 13, characterized in that Determining a resource scheduling strategy corresponding to the instance type includes: In response to the instance type being a stable instance type, determining the resource scheduling policy corresponding to the stable instance type according to the resource demand information of the instance, wherein the stable instance type is used to indicate that a change value of the memory utilization of the instance does not exceed a change threshold, the resource demand information is used to indicate actual memory resources required by the instance during operation on the cloud service platform, and the resource scheduling policy is used to indicate a rule for scheduling the memory resources from the cloud service platform to the instance according to a memory overselling policy; In response to the instance type being a transient instance type, the resource scheduling policy corresponding to the transient instance type is determined according to the resource demand information of the instance, wherein the transient instance type is used to indicate that the change value of the memory utilization of the instance exceeds the change threshold, the resource demand information is used to indicate the actual memory resources required by the instance during its operation on the cloud service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources from the cloud service platform to the instance that are greater than or equal to the actual memory resources and less than or equal to the physical memory of the cloud service platform.

16. A resource scheduling method, characterized in that: include: Acquire the instance type of at least one instance in the database by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter includes the instance type, the database is deployed on a service platform, and the instance is used to provide a service function of the service platform; Identifying an instance type of the instance, wherein the instance type is divided according to a change in memory resources required by the instance during operation on the service platform; Determine a resource scheduling policy corresponding to the instance type, wherein different instance types correspond to different resource scheduling policies, and the resource scheduling policy is used to represent a rule for scheduling memory resources from the service platform to the instance; According to the resource scheduling strategy, a resource scheduling operation is performed on the instance to obtain a resource scheduling result; The resource scheduling result is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the resource scheduling result.

17. The method according to claim 16, characterized in that Identify the instance type of the instance, including: The instance is classified by using an instance classification model to obtain the instance type, wherein the instance classification model is obtained by training a Markov chain model using instance type labels of instance samples.

18. The method according to claim 16, characterized in that Determining a resource scheduling strategy corresponding to the instance type includes: In response to the instance type being a stable instance type, determining the resource scheduling policy corresponding to the stable instance type according to the resource demand information of the instance, wherein the stable instance type is used to indicate that a change value of the memory utilization of the instance does not exceed a change threshold, the resource demand information is used to indicate actual memory resources required by the instance during operation on the service platform, and the resource scheduling policy is used to indicate a rule for scheduling the memory resources from the service platform to the instance according to a memory overselling policy; In response to the instance type being a transient instance type, the resource scheduling policy corresponding to the transient instance type is determined according to the resource demand information of the instance, wherein the transient instance type is used to indicate that the change value of the memory utilization of the instance exceeds the change threshold, the resource demand information is used to indicate the actual memory resources required by the instance during its operation on the service platform, and the resource scheduling policy is used to indicate a rule for scheduling memory resources from the service platform to the instance that are greater than or equal to the actual memory resources and less than or equal to the physical memory of the service platform.

19. A resource scheduling system, characterized in that: include: A classifier is used to obtain at least one instance in a database, wherein the database is deployed on a service platform and the instance is used to provide a service function of the service platform; identify an instance type of the instance, wherein the instance type is divided according to a change range of memory resources required by the instance during operation on the service platform; A scheduler is used to determine a resource scheduling policy corresponding to the instance type, wherein different instance types correspond to different resource scheduling policies, and the resource scheduling policy is used to represent the rules for scheduling memory resources from the service platform to the instance; and perform resource scheduling operations on the instance according to the resource scheduling policy.

20. A computing device, characterized in that include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 18 when running.

21. An electronic device, characterized in that: include: A memory storing an executable program; A processor, connected to the memory via a bus, and configured to run the program, wherein the program executes the method described in any one of claims 1 to 18 when running.

22. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 18.

23. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Resource scheduling method and device, electronic equipment and computer readable storage medium

    CN114675936A

  • Resource scheduling method and device and electronic equipment

    CN117112213A

  • Instance scheduling method and device, equipment and storage medium

    CN117369959A

  • Micro-service resource scheduling method and device, electronic equipment and storage medium

    CN118796372A

  • Resource allocation method and network based on dynamic programming, and storage medium and processor

    WO2024114483A2

Cited By

  • Task processing method and device for multi-party security computing

    CN120567798A

  • Cluster dynamic resource overselling method based on multi-tenant resource entropy

    CN120743433A

  • A multi-tenant resource entropy-based cluster dynamic resource overselling method

    CN120743433B

  • High-concurrency service flexible scheduling method and system for new retail platform

    CN121579175A