Method for dynamically expanding and shrinking database calculation engine nodes during execution

Through the dynamic scaling method, the resource configuration of the database computing engine node is automatically adjusted, which solves the problem of manual operation or restarting services in the existing technology, and improves the flexibility and response speed of the database system.

CN120066781APending Publication Date: 2025-05-30NANJING LONGYUAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510135317.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When responding to dynamic scaling needs, existing database systems need to manually operate or restart services, which affects the flexibility and response speed of the database system.

Method used

Provide a method to query the interface of the cluster scaling suspending policy, generate and collect core indicators in the cluster, determine whether the cluster needs to suspend or expand, and receive expansion and suspend requests for legality checksum execution.

Benefits of technology

It realizes automation and seamless connection of resource adjustments, reduces manual intervention, improves the flexibility and response speed of the database system, optimizes resource utilization and enhances system scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066781A_ABST
    Figure CN120066781A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a method for dynamic expansion and contraction during execution of database calculation engine nodes, which comprises the following steps of: providing an interface for querying a cluster expansion and contraction suspension strategy; generating, collecting and calculating a core index in the cluster, and judging whether the cluster needs to be suspended and scaled or not according to a current cluster scaling suspension strategy provided by the metadata; receiving a scaling request and a hanging request of the computing cluster, and carrying out legality verification; by means of the mode, automation and seamless connection of resource adjustment are achieved, manual intervention is reduced, and the flexibility and the response speed of a database system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for dynamically scaling in and out during the execution of a database computing engine node. Background Art

[0002] With the advent of the big data era, database systems are facing the pressure of data growth. Currently, most traditional database systems adopt a static resource configuration scheme, that is, computing resources such as CPU, memory, storage, etc. are pre-allocated according to the estimated business load at the initial stage of system deployment. However, this static configuration method exposes many deficiencies in actual applications:

[0003] On the one hand, during high-load periods, the statically allocated resources may quickly reach the bottleneck, resulting in a decrease in processing speed and an extension of query response time, seriously affecting user experience and business continuity;

[0004] On the other hand, during low-load periods, a large amount of resources are in an idle state and cannot be fully utilized, causing resource waste and cost increase.

[0005] Especially in the environment of cloud computing and distributed systems, the performance, scalability, and reliability of database systems have become problems that need to be solved urgently. When existing databases respond to dynamic scaling in and out requirements, they often require manual operations or restart services, seriously affecting the flexibility and response speed of database systems.

[0006] Therefore, it is very necessary to propose a method that can support dynamic scaling in and out of database computing engine nodes during execution. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for dynamically scaling in and out during the execution of a database computing engine node, aiming to solve the technical problem that existing databases often require manual operations or restart services when responding to dynamic scaling in and out requirements, seriously affecting the flexibility and response speed of database systems.

[0008] To achieve the above purpose, a method for dynamically scaling in and out during the execution of a database computing engine node adopted by the present invention includes the following steps:

[0009] Provide an interface for querying the scaling suspension policy of the cluster;

[0010] Generate and collect core metrics within the computing cluster, and determine whether the cluster needs to be suspended and scaled according to the current scaling suspension policy of the cluster provided by the metadata;

[0011] Receive the scaling in / out and suspension requests of the computing cluster, and perform legality verification;

[0012] Execute the suspension and scaling actions.

[0013] Among them, in the step of generating and collecting the core metrics in the computing cluster and determining whether the cluster needs to be suspended, scaled in or out according to the current scaling and suspension policy of the cluster provided by the metadata:

[0014] In determining whether the cluster needs to be suspended:

[0015] When the duration of the cluster being in a state without running jobs and queued jobs, it is determined that the cluster needs to be suspended.

[0016] Among them, in the step of when the duration of the cluster being in a state without running jobs and queued jobs, it is determined that the cluster needs to be suspended:

[0017] Expose the metrics to be suspended through the exporter method. After capturing the metrics, compare with the suspension policy created by the user and perform subsequent actions.

[0018] Among them, in the step of generating and collecting the core metrics in the computing cluster and determining whether the cluster needs to be suspended, scaled in or out according to the current scaling and suspension policy of the cluster provided by the metadata:

[0019] In determining whether the cluster needs to be scaled in or out:

[0020] Set the threshold of the utilization rate for scaling in, calculate the average utilization rate of the cluster, and compare the average utilization rate of the cluster with the threshold of the utilization rate for scaling in.

[0021] Among them, in the step of generating and collecting the core metrics in the computing cluster and determining whether the cluster needs to be suspended, scaled in or out according to the current scaling and suspension policy of the cluster provided by the metadata:

[0022] In determining whether the cluster needs to be scaled in or out:

[0023] When the average utilization rate of the cluster is greater than the threshold of the utilization rate for scaling in, it is determined that the cluster needs to be scaled in.

[0024] Among them, in the step of receiving the requests for scaling in, out and suspension of the computing cluster and performing the legality check, the legality check includes:

[0025] Request parameter check, permission check, cluster status check, resource check.

[0026] Among them, in the step of request parameter check:

[0027] Check whether the parameters in the received requests for scaling in, out and suspension are complete and whether the formats are correct.

[0028] Among them, in the step of permission check:

[0029] Verify whether the user and the system initiating the request have the permission to perform this operation.

[0030] Among them, in the step of cluster status verification:

[0031] Check whether the current status of the cluster allows scaling and suspension operations.

[0032] Among them, in the step of resource verification:

[0033] In the scaling operation, verify whether the cluster has sufficient resources to support the configuration after scaling.

[0034] A method for dynamic scaling of a database computing engine node during execution, by providing an interface for querying the cluster scaling and suspension policy; generating and collecting core metrics in the computing cluster, and determining whether the cluster needs to be suspended and scaled according to the current cluster scaling and suspension policy provided by the metadata; receiving the scaling and suspension requests of the computing cluster, performing legality verification; executing the suspension and scaling actions; realizing the automation and seamless connection of resource adjustment, reducing manual intervention, and improving the flexibility and response speed of the database system. Brief Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is a flowchart of the steps of the method for dynamic scaling of a database computing engine node during execution according to the present invention.

[0037] Figure 2 It is a schematic diagram of Curve 1 of the present invention.

[0038] Figure 3 It is a schematic diagram of Curve 2 of the present invention.

[0039] Figure 4 It is a schematic diagram of Curve 3 and Curve 4 of the present invention.

[0040] Figure 5 It is a schematic diagram of the principle of the system for dynamic scaling of a database computing engine node during execution according to the present invention.

[0041] Figure 6 It is a schematic diagram of the principle of the electronic device according to the present invention.

[0042] 501 - Metadata service module, 502 - Computing cluster module, 503 - Infrastructure service module, 504 - Task scheduler. Detailed Embodiments

[0043] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.

[0044] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0045] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0046] Please refer to Figures 1 to 4 , the present invention provides a method for dynamically scaling a database computing engine node during execution, including the following steps:

[0047] S100: Provide an interface for querying the cluster scaling suspension policy;

[0048] In this embodiment, the suspension action is to immediately suspend the entire cluster and release resources or return them to the resource pool. The scaling-up action is to create or obtain nodes of the corresponding specification from the resource pool and add the nodes to the specified computing cluster. The addition method is to ensure the discovery.uri address in the node's config.properties and point it to the address of the coordinator in the cluster to be joined. The scaling-down action is to strip the specified nodes out of the cluster and release or return them to the resource pool. The stripping method is to ensure that the discovery.uri address in the node's config.properties points to an empty address. First, an interface for querying the cluster scaling suspension policy needs to be obtained to lay a foundation for subsequent judgments.

[0049] S200: Generate and collect core metrics within the computing cluster, and determine whether the cluster needs to be suspended and scaled according to the current cluster scaling suspension policy provided by the metadata;

[0050] In this embodiment, the definition of the core metrics is:

[0051] cluster-usage-avg: The average utilization rate of the computing engine within the last minute.

[0052] Definition of cluster average utilization rate (U_avg): Sum(U1, U2,... Un) / n, where U1, U2,.. Un represent the instantaneous cluster utilization rates for the completed samples within the time span, and n represents the number of sampling times within that time span.

[0053] Definition of instantaneous cluster utilization rate (U): Sum(L1, L2,..., Ln) / Sum(number of vCores of worker nodes).

[0054] Definition of instantaneous load (L) of a worker node: The CPU load value of the worker node.

[0055] cluster-queue-job-count: The current number of jobs queued in the computing cluster.

[0056] The reason for using the average utilization rate within the last minute is as follows: By using the sliding window algorithm, the spikes caused by the fluctuations in the instantaneous utilization rate curve are flattened, which effectively reduces the scaling jitter caused by the spikes.

[0057] Snowflake incorporates the queued job count metric into the scale-out assertion logic. Consider the following three scenarios:

[0058] 1. In the scenario of non-skewed load, when the computing bottleneck is in the CPU. Whether there is queuing is positively correlated with the cluster load, and the same effect can be achieved by asserting both metrics.

[0059] 2. In the scenario of non-skewed load, when the computing bottleneck is not in the CPU. Whether there is queuing is not correlated with the cluster load. By asserting the number of queued jobs, the job processing speed can be improved well. The cluster average utilization rate needs to consider other load factors, such as: IO load, etc.

[0060] 3. In the scenario of skewed load, since the cluster has sufficient idle resources, all jobs can be scheduled and will not be blocked. However, the operators scheduled to the nodes with high load in the jobs may be suspended. The two assertion logics have the same effect.

[0061] From the above inference, the average cluster utilization rate is used to measure whether the computing engine is optimized enough to achieve a linear increase in job processing speed and the average cluster utilization rate. The core purpose of whether to expand the cluster is to improve the parallel processing ability of the cluster and process more jobs in the same amount of time. However, due to various performance optimization issues, the computing engine often fails to achieve this ideal linear effect. Therefore, using the average cluster utilization rate to determine whether to expand may not be the best choice. So we determine the expansion rule based on the number of queued jobs; and determine the scaling-down rule based on the average cluster utilization rate. The calculation logic of the scaling-down water level line (Lin) remains unchanged. The core purpose is to ensure that the average cluster utilization rate after scaling down does not exceed 0.9. If it exceeds, it is very likely to increase the number of queued jobs.

[0062] Among them, when judging whether the cluster needs to be suspended:

[0063] If the duration during which the cluster is in a state without running jobs and queued jobs, it is judged that the cluster needs to be suspended;

[0064] Expose the metrics that need to be suspended through the exporter method. After capturing the metrics, compare with the suspension policy created by the user and perform subsequent actions.

[0065] Among them, when judging whether the cluster needs to be scaled up or down:

[0066] Set the scaling-down utilization rate threshold, calculate the average cluster utilization rate, and compare the average cluster utilization rate with the scaling-down utilization rate threshold; when the average cluster utilization rate is greater than the scaling-down utilization rate threshold, it is judged that the cluster needs to be scaled down.

[0067] During scaling down: Define the scaling-down threshold Lin and the cluster safe utilization rate threshold Lout. The computing engine is a cluster with an MPP architecture. The more computing nodes, the faster the computing speed can be improved, and the load can be almost evenly distributed among the working nodes. So under the condition that the external read and write load remains unchanged, the following general conclusions can be obtained:

[0068] After expansion, the average cluster utilization rate index will partially decrease;

[0069] After scaling down, the average cluster load will partially increase;

[0070] The derivation of the relationship between Lin and Lout is as follows:

[0071] Known:

[0072] 1. The cluster read and write load drops from X1 to X2, triggering the scaling-down logic, reducing from N working nodes to N - 1;

[0073] 2. After scaling down, the X2 read and write load remains stable;

[0074] 3. The computing power of cluster working nodes is linearly superimposed;

[0075] 4. Each time the cluster scales in or out, only one node is added or removed;

[0076] 5. The utilization threshold Lin for cluster scaling in and the utilization value Lout for cluster scaling out;

[0077] 6. The number of working nodes is: [1, 32];

[0078] After the cluster scales in or out, the cluster water level changes when the cluster computing power changes. Deduce the mathematical relationship between Lin and Lout to ensure that the cluster will not scale in and out repeatedly during scaling.

[0079] Assume that when the cluster utilization rate just reaches Lin, the cluster triggers scaling in;

[0080] The number of cluster computing nodes decreases from N to N - 1.

[0081] The cluster computing power decreases by (N - 1) / N times. After scaling in, the cluster utilization rate becomes Lin*(N - 1) / N.

[0082] Assume that the computing power of the cluster before scaling in is 1, and the cluster utilization rate after scaling in is a;

[0083] After the cluster scales in, the read-write load to be processed remains unchanged;

[0084] Then the computing power after scaling in is (N - 1) / N;

[0085] 1*Lin = a*(N - 1) / N;

[0086] a = (Lin*N) / (N - 1);

[0087] The cluster utilization rate after scaling in shall not exceed Lout;

[0088] Lout > (Lin*N) / (N - 1);

[0089] The value range of N is [2, 32];

[0090] Conclusion: When N = 2, Lout > 2Lin; when N = 32, Lout > 32 / 31Lin (approximately Lout > 1.032Lin).

[0091] Using the pandas plotting tool to obtain Figure 2 Curve 1, as Figure 2 shown. To make the cluster more stable during scaling in, it is necessary to make the cluster utilization rate a = Lout + Ldert after scaling in. Let Ldert = 0.1, then we get:

[0092] Lout = (Lin * N) / (N - 1) + 0.1;

[0093] Obtained using the pandas plotting tool Figure 3 Curve 2, as Figure 3 shown. When N = 2, Lout = 2.1Lin; when N = 32, Lout = 1.132Lin.

[0094] By backward deduction, assuming Lout = 0.9, that is, the cluster has a 90% utilization rate for capacity expansion. Then substitute into the formula Lin = (N - 1)(Lout - 0.1) / n to calculate the ideal value of Lin, as shown by Figure 4 curves 3 and 4 shown. Substitute into the above calculation formula to obtain the threshold of the scale-up / scale-down water level line for [2, 32].

[0095] S300: Receive the scale-up / scale-down and suspension requests of the computing cluster and perform legality verification;

[0096] In this embodiment, the legality verification includes:

[0097] Request parameter verification, permission verification, cluster status verification, resource verification.

[0098] Request parameter verification:

[0099] Verification content: First, check whether the parameters in the received scale-up / scale-down and suspension requests are complete and whether the format is correct. For example, whether the request contains the necessary cluster identifier, operation type (scale-up / scale-down or suspension), target node number, or suspension duration, etc.

[0100] Verification method: Verify through predefined parameter formats and rules using means such as regular expressions and type checking.

[0101] Judgment result: If the parameters are incomplete or the format is incorrect, directly return an error message and reject subsequent operations.

[0102] Permission verification:

[0103] Verification content: Verify whether the user and system initiating the request have the permission to perform this operation. For example, check whether the user belongs to the user group with the permission to scale up / scale down or suspend the cluster.

[0104] Verification method: Verify user permissions by querying the user permission database or access control list (ACL).

[0105] Judgment result: If the user has no permission, return an error message indicating insufficient permission and reject subsequent operations.

[0106] Cluster status verification:

[0107] Verification content: Check whether the current status of the cluster allows scaling and suspension operations. For example, whether the cluster is in a healthy state, whether there are ongoing maintenance tasks, etc.

[0108] Verification method: Obtain the current status of the cluster by querying the cluster management system or status monitoring tool.

[0109] Judgment result: If the cluster status does not allow the operation, return a status error message and reject the subsequent operations.

[0110] Resource verification:

[0111] Verification content: In the scaling operation, verify whether the cluster has sufficient resources, such as CPU, memory, and storage space, to support the scaled configuration. In the suspension operation, verify whether the cluster can still meet the minimum resource requirements after suspension.

[0112] Verification method: Evaluate the current and future resource requirements through the resource management system or resource monitoring tool.

[0113] Judgment result: If the resources are insufficient, return an error message indicating insufficient resources and reject the subsequent operations.

[0114] Corresponding operations for judgment results

[0115] Legal request: If all verifications pass, confirm that the request is legal and continue with the subsequent scaling or suspension operations.

[0116] Illegal request: If any verification fails, return the corresponding error message according to the specific reason for the verification failure and terminate the subsequent operations. At the same time, logs can be recorded for subsequent analysis and troubleshooting.

[0117] S400: Execute suspension and scaling actions.

[0118] In this embodiment, after receiving the scaling and suspension requests of the computing cluster, it is necessary to perform legality verification and execute the suspension and scaling actions after the verification is completed.

[0119] Among them, add a ScaleManager module in trino-main to manage the automatic suspension and scaling logic of the cluster. This module obtains the node addresses and health status from the InteralNodeManager, and adds the statistical logic of the core indicator cluster-idle-seconds to the QueryManagerStats in the DispatchManager.

[0120] Three periodic timers need to be added inside this module, and the function of each timer is as follows:

[0121] 1. The cluster-usage-avg metric timing collector. Periodically obtain the NodeStatus of each worker node respectively, and obtain the processCPULoad metric therein. And maintain the metrics in a sliding window, and the sliding window maintains data within 1 minute.

[0122] 2. The automatic suspension judgment timer. Periodically obtain the core metric cluster-idle-seconds from the DispatcherManage, and align with the latest cluster suspension policy to assert whether the cluster needs to be suspended. If so, request the infrastructure service to suspend the cluster.

[0123] 3. The automatic scaling judgment timer. It is internally implemented by a finite state machine (FSM). Periodically obtain the core metrics cluster-queue-job-count and cluster-usage-avg and input them to the state machine. The state machine decides whether to scale.

[0124] When scaling down, first mark a node offline in the DispatcherManager to ensure that the DispatcherManager will not schedule new tasks to the marked-offline node in the future. Then wait for a period of time for the existing tasks in the node to complete, configure a fixed time, such as 5 minutes, and forcefully shut down the marked-offline node after timeout. Two points need to be noted here: 1. The marked-offline nodes will not be counted into the cluster-usage-load metric. 2. If the cluster needs to scale up during the waiting process of the marked node going offline. Then the marked node can be directly restored to a normal working node.

[0125] The automatic scaling judgment timer forcibly shuts down the node by requesting a specified interface in the infrastructure service.

[0126] In the present invention, first provide an interface for querying the cluster scaling and suspension policy. Then generate and collect the core metrics within the computing cluster, and judge whether the cluster needs to be suspended and scaled according to the current cluster scaling and suspension policy provided by the metadata. Then receive the scaling and suspension requests of the computing cluster and perform legality verification.

[0127] Finally, perform the suspension and scaling actions. The following advantages are achieved through the above method:

[0128] First, improve the response ability and performance of the system: The dynamic scaling ability allows the database system to automatically adjust resource allocation according to the real-time load situation. During high-load periods, the system can quickly increase computing resources, such as CPU, memory, etc., thus avoiding resource bottlenecks, maintaining the stability of processing speed and query response time, and significantly improving user experience and business continuity.

[0129] II. Optimize resource utilization and enhance system scalability: During low-load periods, the system can automatically reduce resource allocation to avoid unnecessary idling of resources. This intelligent resource management strategy can significantly reduce resource waste and maximize cost-effectiveness. Especially in a cloud computing environment, it can effectively control operating costs. The dynamic scaling mechanism enables the database system to easily handle future business growth requirements without complex manual configuration or system upgrades. As the business volume increases, the system can automatically expand to ensure that performance is not affected while maintaining a high degree of flexibility and scalability.

[0130] III. Improve system reliability and stability: Compared with traditional methods of manual scaling or restarting services, automated dynamic scaling reduces human intervention and errors, simplifies the operation and maintenance management process of the database system. Operation and maintenance personnel do not need to frequently monitor system load and manually adjust resources, reducing the risk of system interruption due to improper operations. In addition, dynamically adjusting resources can more effectively handle sudden traffic or failures, improving the fault tolerance and stability of the system.

[0131] Corresponding to the foregoing embodiments of the method for dynamic scaling during the execution of a database computing engine node, the present application also provides an embodiment of a system for dynamic scaling during the execution of a database computing engine node.

[0132] Figure 5 is a system block diagram of a system for dynamic scaling during the execution of a database computing engine node shown according to an exemplary embodiment. Referring to Figure 5 , the system may include: a metadata service module 501, a computing cluster module 502, an infrastructure service module 503, and a task scheduler 504; where:

[0133] The metadata service module 501 is configured to provide an interface for querying the cluster scaling suspension policy.

[0134] The computing cluster module 502 is configured to generate and collect core metrics within the computing cluster, and determine whether the cluster needs to be suspended and scaled according to the current cluster scaling suspension policy provided by the metadata.

[0135] The infrastructure service module 503 is configured to receive scaling and suspension requests of the computing cluster and perform legality verification.

[0136] The task scheduler 504 is configured to execute suspension and scaling actions.

[0137] In this embodiment, the metadata service module 501 provides an interface for querying the cluster scaling suspension policy; the computing cluster module 502 generates and collects core metrics within the computing cluster, and determines whether the cluster needs to be suspended or scaled according to the current cluster scaling suspension policy provided by the metadata; the infrastructure service module 503 receives the scaling and suspension requests of the computing cluster and performs legality verification; the task scheduler 504 executes the suspension and scaling actions; realizing the automation and seamless connection of resource adjustment, reducing manual intervention, and improving the flexibility and response speed of the database system.

[0138] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0139] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0140] Correspondingly, the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for dynamic scaling during the execution of the database computing engine node as described above. As Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where the system for dynamic scaling during the execution of the database computing engine node provided by the embodiment of the present invention is located. In addition to Figure 6 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated herein.

[0141] Correspondingly, the present application also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the method for dynamic scaling during the execution of a database computing engine node as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0142] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0143] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for dynamically scaling a database computing engine node during execution, characterized in that: The steps include: Provides an interface for querying cluster expansion and contraction suspension policies; Generate and collect core indicators in the computing cluster, and determine whether the cluster needs to be suspended and expanded based on the current cluster expansion and suspension strategy provided by the metadata; Receive scaling and suspension requests from computing clusters and perform legitimacy verification; Execute suspension and scaling actions.

2. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 1, characterized in that: In the step of generating and collecting the core indicators in the computing cluster and judging whether the cluster needs to be suspended and scaled according to the current cluster scaling suspension policy provided by the metadata: When determining whether the cluster needs to be suspended: When the cluster is in a state of no running jobs and queued jobs for a certain period of time, it is determined that the cluster needs to be suspended.

3. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 2, characterized in that: When the cluster is in a state of no running jobs and queued jobs for a certain period of time, the step of determining whether the cluster needs to be suspended is: Expose the indicators that need to be suspended through exporter. After capturing the indicators, compare them with the suspension strategy created by the user to perform subsequent actions.

4. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 1, characterized in that: In the step of generating and collecting the core indicators in the computing cluster and judging whether the cluster needs to be suspended and scaled according to the current cluster scaling suspension policy provided by the metadata: To determine whether the cluster needs to be expanded or reduced: Set the utilization rate threshold for capacity reduction, calculate the average utilization rate of the cluster, and compare the average utilization rate of the cluster with the utilization rate threshold for capacity reduction.

5. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 4, characterized in that: In the step of generating and collecting the core indicators in the computing cluster and judging whether the cluster needs to be suspended and scaled according to the current cluster scaling suspension policy provided by the metadata: To determine whether the cluster needs to be expanded or reduced: When the average cluster usage is greater than the shrinking utilization threshold, it is determined that the cluster needs to be shrunk.

6. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 1, characterized in that: In the step of receiving the scaling or suspension request of the computing cluster and performing the legality check, the legality check includes: Request parameter verification, permission verification, cluster status verification, and resource verification.

7. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 6, characterized in that: In the step of request parameter verification: Check whether the parameters in the received scaling and suspension requests are complete and the format is correct.

8. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 6, characterized in that: In the permission verification step: Verify that the user and system initiating the request have permission to perform the operation.

9. The method for dynamic scaling of database computing engine nodes during execution as claimed in claim 6, characterized in that: In the cluster status verification step: Check whether the current cluster status allows scaling and suspension operations.

10. The method for dynamic scaling of database computing engine nodes during execution according to claim 6, characterized in that: In the resource verification step: During a scale operation, verify that the cluster has sufficient resources to support the scaled configuration.

Citation Information

Patent Citations

  • Capacity expansion and shrinkage control method and device based on Kubernetes cluster and electronic equipment

    CN112506444A

  • Small data volume job processing method and device based on big data engine

    CN115309560A

  • Pod cluster capacity expansion and contraction method, system and device and storage medium

    CN115509745A

  • Cluster expansion and contraction method and device, electronic equipment and computer readable medium

    CN118714012A

  • Capacity expansion method and device based on key authorization, equipment and storage medium

    CN119276616A