Multi-tenant cloud native database horizontal extension prediction method, system and device and medium
By using the Transformer network in a multi-tenant cloud native database for resource load prediction and optimizing the K8s HPA strategy, the problems of elastic scaling, waste of resources and low prediction accuracy in the existing technology are solved, and more efficient resource management and system stability are achieved.
Patent Information
- Application Number
- CN202510229796.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the elastic scaling method has problems such as delay in response, waste of resources and low prediction accuracy, and it is difficult to effectively deal with complex and changing loads.
The Transformer network is used to build a resource load prediction model, predict through multi-channel information mixing, dynamic dimensional adjustment and multi-dimensional feature alignment, and optimize the K8s HPA strategy based on the prediction results to achieve automated expansion and scaling.
It improves the response speed of elastic scaling, reduces resource waste, and enhances the prediction accuracy of future load changes, thereby achieving more accurate resource supply adjustments.
Smart Images

Figure CN120104451A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing technology, and in particular to a method, system, device and medium for predicting horizontal expansion of a multi-tenant cloud native database. Background Art
[0002] With the explosive growth of the number of users and data, traditional relational databases are unable to meet the increasing concurrency and complex and changeable business traffic. For this reason, distributed databases are introduced. In traditional technologies, in order to cope with the surge in business traffic, database administrators (DBAs) and other professionals are required to perform performance checks on the running database system by observing and monitoring or receiving system alarms. After determining that expansion is needed, the distributed database is expanded manually, through scripts, platform operations, etc., and after the business peak, it is necessary to shrink the capacity, resulting in high operating costs and untimely expansion and shrinking.
[0003] In addition, with the rapid development of cloud computing technology, container technology has become a widely used lightweight virtualization technology. The elastic scaling technology in the container cloud environment is particularly important. It can dynamically adjust the computing resources of the container according to the actual resource demand, thereby ensuring the stability and efficiency of the system. The existing elastic scaling methods are mainly implemented by monitoring the system load and the thresholds set by the user, such as container management platforms such as Docker and Kubernetes. In addition, there are also related studies that propose a model-based prediction method to achieve elastic scaling of containers. The elastic scaling method based on model prediction is divided into two main steps: first, predicting the workload, and second, performing corresponding expansion or reduction operations based on the prediction results. This model-based prediction method allows resource configuration to be planned in advance to cope with future workload changes, thereby effectively preventing violations of the service level agreement (SLA). Accurately predicting future workloads and performing appropriate elastic scaling operations accordingly are the key challenges to achieving active elastic scaling. At present, there are some shortcomings in the elastic scaling methods in the container cloud environment: (1) Response delay: Traditional elastic scaling methods are all responsive scaling methods based on threshold settings. The scaling operation is triggered only when the load reaches the specified threshold. However, this threshold-based responsive scaling method has the problem of elastic lag, which will cause response delay and affect user experience; (2) Resource waste: Traditional threshold-based elastic scaling methods need to set thresholds based on experience. In order to ensure the stability of the business, a higher threshold is usually set. However, if the threshold is set too high, it will cause resource waste and low overall resource utilization; (3) Low prediction accuracy: In cloud computing environments, workload fluctuations are usually very complex and accompanied by high uncertainty. Therefore, accurate prediction of future loads poses a great challenge.
[0004] From the above description, we can see that one of the important features of multi-tenant databases in a cloud-native environment is scalability. However, most elastic scaling technologies find it difficult to make effective scaling decisions for complex and changing loads. If load changes can be predicted in advance, resource supply can be accurately adjusted.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0007] The disclosed embodiments provide a method for predicting horizontal expansion of a multi-tenant cloud native database to solve the problems of lag and resource waste in elastic scaling in the prior art.
[0008] In some embodiments, the method includes: a multi-tenant cloud native database horizontal expansion prediction method, including:
[0009] Step A: Multi-source data collection and preprocessing: collecting workload data of database instances and preprocessing the collected raw data;
[0010] Step B: Resource load prediction: A resource load prediction model is constructed using a Transformer network. The preprocessed workload data is input into the Transformer network. After multi-channel information mixing, dynamic size adjustment, and multi-dimensional feature alignment, the Transformer network predicts the resource load and outputs the prediction result.
[0011] Step C: Horizontally elastically scale out scheduling. Optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
[0012] As a further improvement, in step A, the preprocessing includes cleaning, alignment and normalization, the raw data is cleaned to remove outliers and fill missing values, the data frequency is unified through timestamp alignment, and the features of each dimension are converted into comparable scales using normalization or standardization methods.
[0013] As a further improvement, in step B, the process of multi-channel information mixing is: first, the data of each channel is passed to an independent encoder module respectively, the dependencies within each channel are captured through the self-attention mechanism, and information interaction is performed in the middle layer, and the information of each channel is fused together using a multi-head cross-attention mechanism.
[0014] As a further improvement, in step B, dynamic size adjustment is to adjust the length of the input sequence according to the characteristics of the data, the data characteristics include time granularity and data scale, and the dynamic size adjustment uses a sliding window of variable length to adapt to the long-term and short-term dependency requirements of different time series, or through stratified sampling technology, different window sizes are used on data at different levels to capture multi-level information.
[0015] As a further improvement, in step B, multi-dimensional feature alignment is performed through feature engineering and regularization methods, interpolation or time alignment methods are used to ensure the consistency of each feature on the time axis, and the weights of different features are adjusted through L2 regularization to ensure the balance between features.
[0016] As a further improvement, in step B, the Transformer network is trained using general time series data and proprietary time series data specific to the database instance load, and the Transformer network performs small sample / zero sample learning.
[0017] As a further improvement, step C includes: if the prediction result is that the system load will increase by more than the set value within the future time T, the capacity expansion is triggered in advance before the actual load arrives; resource scheduling and deployment, Kubernetes dynamically adjusts the number of Pod copies and updates the load balancing configuration; the prediction service implements distributed deployment to prevent single point failures, and combines load balancing and fault tolerance mechanisms, in the Kubernetes environment, with the help of Service Mesh to achieve traffic management and monitoring between services.
[0018] The present invention also discloses a multi-tenant cloud native database horizontal expansion prediction system, comprising:
[0019] The data collection and preprocessing module collects workload data of database instances and preprocesses the collected raw data;
[0020] The resource load prediction module is constructed using a Transformer network and is configured to perform multi-channel information mixing, dynamic size adjustment, and multi-dimensional feature alignment on the data, and then predict the resource load and output the prediction result;
[0021] The horizontal elastic expansion scheduling module is configured to optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
[0022] The present invention also discloses a multi-tenant cloud native database horizontal expansion prediction device, including a processor and a memory storing program instructions, wherein the processor is configured to execute the multi-tenant cloud native database horizontal expansion prediction method as described above when running the program instructions.
[0023] The present invention also discloses a storage medium storing program instructions, which, when running, execute the multi-tenant cloud native database horizontal expansion prediction method as described above.
[0024] The multi-tenant cloud native database horizontal expansion prediction method, system, device and medium provided by the embodiments of the present disclosure can achieve the following technical effects: The time series prediction technology based on the Transformer architecture is the core of the present invention. It combines the public general time series data and the proprietary time series data specific to the resource load, and improves the time series data pattern learning and representation capabilities of the large model through model training and small sample / zero sample learning technology. Training with advanced machine learning models aims to improve the learning and representation capabilities of large models for time series data patterns. This technology is particularly suitable for elastic scaling decisions in multi-tenant database environments because it can predict load changes in advance, thereby adjusting resource supply more accurately. Through the elastic scaling strategy, the number of database instances is adjusted based on the demand forecast results to ensure that resource supply is within a reasonable range. Compared with the most widely used scaling strategy in kubernetes, the elastic scaling method in this article avoids the lag and resource waste of elastic scaling.
[0025] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:
[0027] Figure 1 is a flow chart of the method described in Example 1;
[0028] Figure 2 is a functional block diagram of the system described in Example 2;
[0029] Figure 3 It is a principle block diagram of the device described in Example 3. DETAILED DESCRIPTION
[0030] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0031] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.
[0032] Unless otherwise stated, the term "plurality" means two or more.
[0033] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.
[0034] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.
[0035] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0036] Example 1
[0037] This embodiment discloses a method for horizontal expansion of a multi-tenant cloud native database, which includes three stages: in the first stage, multi-channel data (such as CPU, memory, disk I / O, etc.) is collected and cleaned, aligned and normalized to ensure the quality of time series data; in the second stage, a time series prediction model is built based on the Transformer architecture, and multi-channel information mixing, self-attention mechanism and dynamic window adjustment strategy are used to capture complex time series patterns, and multi-dimensional feature alignment is used to improve the information interactivity and prediction accuracy of the model; in the third stage, the Kubernetes HPA (horizontal Pod automatic expansion) strategy is optimized based on the prediction results, and the dynamic expansion and contraction of service instances are realized in combination with custom expansion indicators, effectively improving resource utilization and system responsiveness.
[0038] Specific as Figure 1 As shown, the method comprises the following steps:
[0039] Step A: Multi-source data collection and preprocessing: collect workload data of database instances and preprocess the collected raw data. In this step, the collected workload data includes CPU utilization, memory usage, disk I / O, and network traffic. The preprocessing of raw data includes cleaning, alignment, and normalization. The raw data is cleaned to remove outliers and fill in missing values. The data frequency is unified through timestamp alignment, and the normalization or standardization method is used to convert the features of each dimension into a comparable scale to ensure the quality and consistency of the time series data and provide high-quality input for subsequent predictive modeling.
[0040] Step B: Resource load prediction: A resource load prediction model is constructed using a Transformer network. The preprocessed workload data is input into the Transformer network. After multi-channel information mixing, dynamic size adjustment, and multi-dimensional feature alignment, the Transformer network predicts the resource load and outputs the prediction result.
[0041] In this embodiment, since the time series data usually comes from multiple channels (such as CPU utilization, memory usage, disk I / O, etc.), each channel may have a different time series pattern. Based on the Transformer model, these channels are passed into independent encoder modules respectively, and the dependencies within each channel are captured through the self-attention mechanism, and information is exchanged in the middle layer. Specifically, a multi-head cross-attention mechanism is used to fuse the information of each channel together to enhance the model's ability to understand data of different dimensions.
[0042] Different data channels may have different time granularities and data scales, and traditional fixed-size windows may not be able to handle these differences efficiently. Dynamic resizing algorithms can adjust the length of the input sequence according to the characteristics of the data. For example, use sliding windows of variable length to adapt to the long-term and short-term dependency requirements of different time series, or use stratified sampling techniques to use different window sizes on data at different levels to capture multi-level information. Use dynamic window adjustment strategies (such as sliding windows or stratified sampling) to optimize the model's adaptability to different time scales. Sliding windows help capture short-term dependencies, while stratified sampling enables more sophisticated modeling of different time periods in the data (such as peak and off-peak periods). Through flexible window adjustment, the model can adapt to different feature patterns in different time periods, thereby improving overall prediction accuracy.
[0043] In time series data, multiple feature dimensions may have different timestamps and scales. How to align these features is the key. Through feature engineering and regularization methods, we can ensure the alignment of features of different dimensions and avoid information loss. For example:
[0044] Use interpolation or time alignment methods to ensure the consistency of each feature on the time axis;
[0045] The weights of different features are adjusted through L2 regularization to ensure the balance between features so that the model will not be biased towards a certain feature during training;
[0046] The alignment module ensures that the relationship between different features can be effectively captured in the same time dimension through joint training.
[0047] In this embodiment, the Transformer network is trained using general time series data and proprietary time series data specific to the database instance load, and the Transformer network performs small sample / zero sample learning to improve the time series data pattern learning and representation capabilities of the large model.
[0048] Step C: Horizontally elastically scale out scheduling. Optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
[0049] Kubernetes' HPA (Horizontal Pod Autoscaler) has been widely used in cloud computing environments, and its scalability can be optimized based on custom indicators (such as load prediction results). Kubernetes HPA (Horizontal Pod Autoscaler) triggers model inference service instance Pods to perform horizontal scaling operations based on model load prediction results. Kubernetes HPA is a mechanism that dynamically adjusts the number of Pod replicas based on resource usage (such as CPU, memory) or custom indicators. By combining deep learning-based load prediction, this method can enhance the foresight of traditional HPA, so that the elastic scaling strategy not only depends on the current load status, but also makes intelligent decisions based on future trends.
[0050] For example, if the prediction service indicates that the system load will increase significantly in the next 30 minutes, HPA can trigger capacity expansion in advance before the actual load arrives, thereby ensuring that the system can run stably during peak hours. Resource scheduling and deployment, Kubernetes dynamically adjusts the number of Pod replicas and updates the load balancing configuration. The prediction service needs to be deployed in a distributed manner to prevent single point failures, and combined with load balancing and fault tolerance mechanisms, in the Kubernetes environment, with the help of Service Mesh (such as Istio), traffic management and monitoring between services can be achieved, while improving the security and observability of the system.
[0051] Service Mesh is a software layer specifically designed to handle communication between services. It is usually composed of containerized microservices. As applications expand and the number of microservices increases, it becomes more difficult to monitor service performance. Service Mesh provides new features such as monitoring, logging, tracing, and traffic control, independent of the code of each service.
[0052] The main functions of a service mesh include:
[0053] Service Discovery: Service mesh provides automated service discovery, reducing the operational burden of managing service endpoints. It uses a service registry to dynamically discover and track all services within the mesh, enabling services to communicate with each other seamlessly.
[0054] Load balancing: Service mesh uses various algorithms (such as polling, least connections, or weighted load balancing) to intelligently distribute requests to multiple service instances, improve resource utilization, and ensure high availability and scalability.
[0055] Traffic management: Service mesh provides advanced traffic management capabilities such as traffic splitting, request mirroring, and canary deployments, allowing fine-grained control over request routing and traffic behavior.
[0056] Security: The service mesh provides secure communication capabilities such as mutual TLS encryption, authentication, and authorization to ensure data confidentiality and integrity.
[0057] Monitoring: Service mesh provides comprehensive monitoring and observability capabilities to help developers understand the health, performance, and behavior of services, supporting troubleshooting and performance optimization.
[0058] Example 2
[0059] This embodiment discloses a multi-tenant cloud native database horizontal expansion prediction system, including:
[0060] The data collection and preprocessing module collects the workload data of the database instance and preprocesses the collected raw data.
[0061] The workload data collected by the data collection and preprocessing module include CPU utilization, memory usage, disk I / O, and network traffic. The preprocessing of raw data includes cleaning, alignment, and normalization. The raw data is cleaned to remove outliers and fill in missing values, the data frequency is unified through timestamp alignment, and the normalization or standardization method is used to convert the features of each dimension into a comparable scale to ensure the quality and consistency of the time series data and provide high-quality input for subsequent predictive modeling.
[0062] The resource load prediction module adopts the Transformer network to form the resource load prediction module, which is configured to perform multi-channel information mixing, dynamic size adjustment and multi-dimensional feature alignment on the data, and then predict the resource load and output the prediction result.
[0063] Since time series data usually comes from multiple channels (such as CPU utilization, memory usage, disk I / O, etc.), each channel may have different time series patterns. Based on the Transformer model, these channels are passed to independent encoder modules respectively, and the dependencies within each channel are captured through the self-attention mechanism, and information is exchanged in the middle layer. Specifically, the multi-head cross-attention mechanism is used to fuse the information of each channel together to enhance the model's ability to understand data of different dimensions.
[0064] Different data channels may have different time granularities and data scales, and traditional fixed-size windows may not be able to handle these differences efficiently. Dynamic resizing algorithms can adjust the length of the input sequence according to the characteristics of the data. For example, use sliding windows of variable length to adapt to the long-term and short-term dependency requirements of different time series, or use stratified sampling techniques to use different window sizes on data at different levels to capture multi-level information. Use dynamic window adjustment strategies (such as sliding windows or stratified sampling) to optimize the model's adaptability to different time scales. Sliding windows help capture short-term dependencies, while stratified sampling enables more sophisticated modeling of different time periods in the data (such as peak and off-peak periods). Through flexible window adjustment, the model can adapt to different feature patterns in different time periods, thereby improving overall prediction accuracy.
[0065] In time series data, multiple feature dimensions may have different timestamps and scales. How to align these features is the key. Through feature engineering and regularization methods, we can ensure the alignment of features of different dimensions and avoid information loss. For example:
[0066] Use interpolation or time alignment methods to ensure the consistency of each feature on the time axis;
[0067] The weights of different features are adjusted through L2 regularization to ensure the balance between features so that the model will not be biased towards a certain feature during training;
[0068] The alignment module ensures that the relationship between different features can be effectively captured in the same time dimension through joint training.
[0069] In this embodiment, the Transformer network is trained using general time series data and proprietary time series data specific to the database instance load, and the Transformer network performs small sample / zero sample learning to improve the time series data pattern learning and representation capabilities of the large model.
[0070] The horizontal elastic expansion scheduling module is configured to optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
[0071] Kubernetes' HPA (Horizontal Pod Autoscaler) has been widely used in cloud computing environments, and its scalability can be optimized based on custom indicators (such as load prediction results). Kubernetes HPA (Horizontal Pod Autoscaler) triggers model inference service instance Pods to perform horizontal scaling operations based on model load prediction results. Kubernetes HPA is a mechanism that dynamically adjusts the number of Pod replicas based on resource usage (such as CPU, memory) or custom indicators. By combining deep learning-based load prediction, the forward-looking nature of traditional HPA can be enhanced, so that elastic scaling strategies not only rely on the current load status, but also make intelligent decisions based on future trends.
[0072] For example, if the prediction service indicates that the system load will increase significantly in the next 30 minutes, HPA can trigger capacity expansion in advance before the actual load arrives, thereby ensuring that the system can run stably during peak hours. Resource scheduling and deployment, Kubernetes dynamically adjusts the number of Pod replicas and updates the load balancing configuration. The prediction service needs to be deployed in a distributed manner to prevent single point failures, and combined with load balancing and fault tolerance mechanisms, in the Kubernetes environment, with the help of Service Mesh (such as Istio), traffic management and monitoring between services can be achieved, while improving the security and observability of the system.
[0073] Example 3
[0074] This embodiment discloses a multi-tenant cloud native database horizontal expansion prediction device, including a processor (processor) 304 and a memory (memory) 301. Optionally, the device may also include a communication interface (Communication Interface) 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logic instructions in the memory 301 to execute X of the above embodiment.
[0075] In addition, the logic instructions in the memory 301 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0076] The memory 301 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes the functional application and data processing by running the program instructions / modules stored in the memory 301, that is, the multi-tenant cloud native database horizontal expansion prediction method in the above embodiment is implemented.
[0077] The memory 301 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.
[0078] Example 4
[0079] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned multi-tenant cloud native database horizontal expansion prediction method.
[0080] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0081] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes, or a transient storage medium.
[0082] The above description and accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent possible changes only. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the scope of protection. As used in the description in the text, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.
[0083] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians may clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0084] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
Claims
1. A method for predicting horizontal expansion of a multi-tenant cloud native database, characterized in that: include: Step A: Multi-source data collection and preprocessing: collecting workload data of database instances and preprocessing the collected raw data; Step B: Resource load prediction: A resource load prediction model is constructed using a Transformer network. The preprocessed workload data is input into the Transformer network. After multi-channel information mixing, dynamic size adjustment, and multi-dimensional feature alignment, the Transformer network predicts the resource load and outputs the prediction result. Step C: Horizontally elastically scale out scheduling. Optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
2. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: In step A, the preprocessing includes cleaning, alignment and normalization. The raw data is cleaned to remove outliers and fill missing values, the data frequency is unified through timestamp alignment, and the features of each dimension are converted into comparable scales using normalization or standardization methods.
3. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: In step B, the process of multi-channel information mixing is as follows: first, the data of each channel is respectively passed into an independent encoder module, the dependencies within each channel are captured through the self-attention mechanism, and information interaction is performed in the middle layer, and the information of each channel is fused together using a multi-head cross-attention mechanism.
4. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: In step B, dynamic size adjustment is to adjust the length of the input sequence according to the characteristics of the data, which include time granularity and data scale. Dynamic size adjustment uses a sliding window of variable length to adapt to the long-term and short-term dependency requirements of different time series, or uses stratified sampling technology to use different window sizes on data at different levels to capture multi-level information.
5. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: In step B, multi-dimensional feature alignment is performed through feature engineering and regularization methods, interpolation or time alignment methods are used to ensure the consistency of each feature on the time axis, and the weights of different features are adjusted through L2 regularization to ensure the balance between features.
6. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: In step B, the Transformer network is trained using general time series data and proprietary time series data specific to the database instance load, and the Transformer network performs small sample / zero sample learning.
7. The method for predicting horizontal expansion of a multi-tenant cloud native database according to claim 1, characterized in that: The step C includes: if the prediction result is that the system load will increase by more than the set value within the future time T, the capacity expansion is triggered in advance before the actual load arrives; resource scheduling and deployment, Kubernetes dynamically adjusts the number of Pod copies and updates the load balancing configuration; the prediction service implements distributed deployment to prevent single point failures, and combines load balancing and fault tolerance mechanisms, and in the Kubernetes environment, uses Service Mesh to implement traffic management and monitoring between services.
8. A multi-tenant cloud native database horizontal expansion prediction system, characterized in that: include: The data collection and preprocessing module collects workload data of database instances and preprocesses the collected raw data; The resource load prediction module is constructed using a Transformer network and is configured to perform multi-channel information mixing, dynamic size adjustment, and multi-dimensional feature alignment on the data, and then predict the resource load and output the prediction result; The horizontal elastic expansion scheduling module is configured to optimize the K8s HPA strategy based on the prediction results to achieve automatic expansion and reduction of model inference service instances.
9. A multi-tenant cloud native database horizontal expansion prediction device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the multi-tenant cloud native database horizontal expansion prediction method as described in any one of claims 1 to 7 when running the program instructions.
10. A storage medium storing program instructions, characterized in that: When the program instructions are run, they execute the multi-tenant cloud native database horizontal expansion prediction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Dynamic resource allocation and performance optimization method for multi-tenant system
CN120631594A
A multi-tenant system dynamic resource allocation and performance optimization method
CN120631594B
High-concurrency data processing method based on cloud computing
CN120804117A