Elastic capacity expansion and contraction method and related device

By periodically obtaining the internal monitoring indicator value of Doris cluster and calculating the composite indicator value, the traditional scaling method has solved the problem of lag and misjudgment response, realizing the timely and accurate scaling of Doris cluster, improving load response capabilities and resource utilization efficiency.

CN119988046AInactive Publication Date: 2025-05-13BEIJING SOHU NEW MEDIA INFORMATION TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510481192.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional Doris cluster scaling method has problems of response lag and misjudgment, which leads to inadequate timely and accurate scaling decisions.

Method used

By periodically obtaining multiple internal monitoring indicator values ​​of the Doris cluster, calculating the composite indicator values, and scaling processing is performed when preset conditions are met to ensure the real-time and accuracy of decisions.

Benefits of technology

It realizes timely and accurate scaling operations, and improves the load response capability and resource utilization efficiency of the Doris cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988046A_ABST
    Figure CN119988046A_ABST
Patent Text Reader

Abstract

The invention discloses an elastic capacity expansion and contraction method and a related device, and relates to the technical field of data processing, and the method comprises the steps: periodically obtaining a plurality of internal monitoring index values of a Dores cluster, and for each period, obtaining a composite index value corresponding to the period according to the plurality of internal monitoring index values obtained by the period. And if the composite index values corresponding to the periods in the target duration meet the capacity expansion and contraction conditions, carrying out capacity expansion and contraction processing corresponding to the capacity expansion and contraction conditions on the Dores cluster. According to the method, the real-time load condition in the Dores cluster can be reflected more accurately and timely through the internal monitoring index, and the real-time performance of capacity expansion and contraction decision response is ensured; and the plurality of internal monitoring index values are processed into the composite index value of which the change degree is positively correlated with the number of the internal monitoring index values deviating from the corresponding baseline value, so that the accuracy of capacity expansion and shrinkage decision response is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an elastic expansion and contraction method and related devices. Background Art

[0002] In actual applications, Doris clusters often need to be expanded or reduced to adapt to changes in business volume. The traditional Doris cluster expansion and reduction method is to collect general resource indicators such as CPU (Central Processing Unit) utilization, memory utilization, disk IO (Input / Output) to determine the load status, and use weighted summation to trigger expansion and reduction operations.

[0003] However, general resource indicators will only change in value after the business load changes, which has a certain lag, resulting in a delayed response to scaling decisions. In addition, the weighted summation method may trigger scaling operations when a single indicator changes, which can easily lead to misjudgment of scaling responses. Summary of the invention

[0004] In view of the above problems, the present application provides an elastic expansion and contraction method and related devices to achieve the purpose of timely and accurate expansion and contraction operations. The specific scheme is as follows:

[0005] The first aspect of the present application provides an elastic expansion and contraction method, comprising:

[0006] Periodically obtain multiple internal monitoring indicator values ​​of the Doris cluster, wherein the multiple internal monitoring indicator values ​​refer to the indicator values ​​of the multiple internal monitoring indicators, and the internal monitoring indicator refers to an indicator for observing the actual operating status of the Doris cluster within the process of the Doris cluster;

[0007] For each cycle, a composite indicator value corresponding to the cycle is obtained according to the multiple internal monitoring indicator values ​​obtained in the cycle, wherein the degree of change of the composite indicator value is positively correlated with the number of deviations of the internal monitoring indicator value from the corresponding baseline value;

[0008] When the composite index value corresponding to the target period meets the preset expansion and contraction condition, detecting whether the composite index values ​​corresponding to each period within the target duration starting from the target period all meet the expansion and contraction condition;

[0009] If so, the Doris cluster is expanded or reduced in accordance with the expansion or reduction conditions.

[0010] In a possible implementation, the multiple internal monitoring indicators include one or more of the following indicators: a query delay indicator, a transaction delay indicator, a compaction_score indicator, and a cache usage indicator, the compaction_score indicator includes a first-stage file compression queue accumulation indicator and a second-stage file compression queue accumulation indicator, and the cache usage indicator includes a segment cache usage indicator and an index cache usage indicator;

[0011] The periodic acquisition of multiple internal monitoring indicator values ​​of the Doris cluster includes:

[0012] Periodically obtaining the respective index values ​​of the query delay index, the transaction delay index, the first-stage file compression queue accumulation index, the second-stage file compression queue accumulation index, the segment cache utilization index, and the index cache utilization index;

[0013] For each cycle, performing a first preprocessing on the index values ​​of the first-stage file compression queue accumulation index and the second-stage file compression queue accumulation index to obtain the index value of the compaction_score index;

[0014] For each cycle, a second preprocessing is performed on the respective index values ​​of the segment cache usage index and the index cache usage index to obtain the index value of the cache usage index.

[0015] In a possible implementation, obtaining the composite indicator value corresponding to the period according to the multiple internal monitoring indicator values ​​acquired in the period includes:

[0016] Obtaining the baseline values ​​and weights of the multiple internal monitoring indicators, wherein the baseline value refers to the internal monitoring indicator value of the Doris cluster under an ideal operating state, and the weight represents the sensitivity of the internal monitoring indicator to the scaling decision;

[0017] Performing indicator smoothing processing on the multiple internal monitoring indicator values ​​obtained in the period respectively to obtain multiple indicator smoothing values ​​corresponding to the period;

[0018] Normalizing the multiple indicator smoothing values ​​corresponding to the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain multiple normalized indicator values ​​corresponding to the period;

[0019] Perform logarithmic transformation on multiple normalized index values ​​corresponding to the period to obtain multiple logarithmic values ​​corresponding to the period;

[0020] Performing weighted summation according to the respective weights of the multiple internal monitoring indicators and the multiple logarithmic values ​​corresponding to the period to obtain a weighted summation value corresponding to the period;

[0021] Perform exponential operation on the weighted sum value corresponding to the period to obtain the composite index value corresponding to the period.

[0022] In a possible implementation, the baseline values ​​of the multiple internal monitoring indicators are updated after each scaling of the scaling decision system;

[0023] Among them, when the scaling decision system is not scaled up or down, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators at the time of system startup; after each scaling of the scaling decision system, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators after the system is stable.

[0024] In a possible implementation, the weights of the multiple internal monitoring indicators are updated once each time a weight adjustment cycle is reached;

[0025] Among them, the weights of each of the multiple internal monitoring indicators in the latter weight adjustment cycle are obtained by adjusting and normalizing the weights of each of the multiple internal monitoring indicators in the previous weight adjustment cycle using the gradient descent method, and the weights of each of the multiple internal monitoring indicators in the first weight adjustment cycle are all preset initial values.

[0026] In a possible implementation, the step of performing scaling processing on the Doris cluster corresponding to the scaling condition includes:

[0027] If the expansion and contraction condition is an expansion condition, the number of Doris nodes in the Doris cluster is increased by calling the Kubernetes interface through the deployment scheduler;

[0028] If the expansion and contraction condition is a contraction condition, a candidate node is selected in the Doris cluster by the deployment scheduler, and after all the data of the candidate node is moved to other Doris nodes in the Doris cluster, the candidate node is deleted.

[0029] A second aspect of the present application provides an elastic expansion and contraction device, comprising:

[0030] An indicator acquisition module, used to periodically acquire multiple internal monitoring indicator values ​​of the Doris cluster, wherein the multiple internal monitoring indicator values ​​refer to the respective indicator values ​​of multiple internal monitoring indicators, and the internal monitoring indicator refers to an indicator for observing the actual operating status of the Doris cluster within the process of the Doris cluster;

[0031] An indicator compound module, for obtaining, for each period, a composite indicator value corresponding to the period according to the multiple internal monitoring indicator values ​​obtained in the period, wherein the degree of change of the composite indicator value is positively correlated with the number of deviations of the internal monitoring indicator value from the corresponding baseline value;

[0032] The expansion and contraction judgment module is used to detect whether the composite index values ​​corresponding to each period within the target duration starting from the target period all meet the expansion and contraction conditions when the composite index value corresponding to the target period meets the preset expansion and contraction conditions;

[0033] The expansion and contraction operation module is used to perform expansion and contraction processing corresponding to the expansion and contraction conditions on the Doris cluster when the expansion and contraction judgment module detects that the composite indicator values ​​corresponding to each period within the target duration starting from the target period meet the expansion and contraction conditions.

[0034] A third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the elastic expansion and contraction method of the first aspect or any implementation of the first aspect.

[0035] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0036] The memory is used to store computer programs;

[0037] The processor is used to execute the computer program so that the electronic device can implement the elastic expansion and contraction method of the above-mentioned first aspect or any implementation method of the first aspect.

[0038] In a fifth aspect, the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the elastic expansion and contraction method of the above-mentioned first aspect or any implementation method of the first aspect.

[0039] By means of the above technical solution, the elastic scaling method provided by the present application periodically obtains multiple internal monitoring indicator values ​​of the Doris cluster, and for each cycle, obtains the composite indicator value corresponding to the cycle according to the multiple internal monitoring indicator values ​​obtained in the cycle. When the composite indicator value corresponding to the target cycle meets the preset scaling conditions, it is detected whether the composite indicator values ​​corresponding to each cycle within the target duration starting from the target cycle meet the scaling conditions. If so, the Doris cluster is scaled up or down corresponding to the scaling conditions. The internal monitoring indicators of the present application are indicators that observe the actual operating status of the Doris cluster from the process of the Doris cluster. Therefore, the internal monitoring indicators can better reflect the real-time load situation inside the Doris cluster, and elastic scaling is performed based on the indicator values ​​of the internal monitoring indicators, ensuring the real-time response of the scaling decision.

[0040] Furthermore, given that in actual scaling operations, most of the internal monitoring indicator values ​​deviate from the corresponding baseline values ​​at the same time, the present application processes multiple internal monitoring indicator values ​​into composite indicator values ​​in which the degree of change is positively correlated with the number of internal monitoring indicator values ​​that deviate from the corresponding baseline values, so that when a single internal monitoring indicator value deviates from the corresponding baseline value, the degree of change of the composite indicator value is small, and when multiple internal monitoring indicator values ​​deviate from the corresponding baseline value at the same time, the degree of change of the composite indicator value is large. As a result, scaling operations can be more accurately identified based on the composite indicator values, thereby improving the accuracy of scaling decision responses. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0042] Figure 1 A schematic diagram of a system architecture provided for this application;

[0043] Figure 2 A schematic diagram of an optional hardware structure of the terminal 100 provided in this application;

[0044] Figure 3 A schematic diagram of the structure of a server 200 provided in this application;

[0045] Figure 4 A schematic diagram of a flow chart of an elastic expansion and contraction method provided in this application;

[0046] Figure 5 A schematic diagram of the structure of an elastic expansion and contraction device provided in this application;

[0047] Figure 6A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0049] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0050] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any deformation of "including" and "having" are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0051] Currently, with the rapid development of cloud native technology and big data analysis, distributed database systems are widely used in massive data processing, real-time query and high-concurrency scenarios. Doris cluster is a massively parallel processing (MPP) database cluster for real-time analysis. With its efficient data storage, query optimization and distributed computing capabilities, it has important application value in the fields of Internet, finance, e-commerce, etc.

[0052] Usually, during business peaks, such as online advertising, e-commerce promotions and other scenarios, user visits and query requests increase sharply. The scaling decision system based on the Doris cluster needs to be expanded to increase computing and storage resources to ensure that the query response time and data writing speed meet the SLA (Service-Level Agreement) requirements; during off-peak periods, some Doris nodes are idle, and timely scaling can reduce resource waste and operation and maintenance costs.

[0053] As introduced in the background technology, the traditional Doris cluster scaling method uses general resource indicators to determine the load status, and then uses a weighted sum method to trigger the scaling operation. There are problems of delayed response and misjudgment of scaling decisions.

[0054] In order to solve the problems of delayed response and misjudgment of scaling decisions in the prior art, the present application provides an elastic scaling method, which can be applied to the scenario of elastic scaling of Doris clusters.

[0055] Optionally, the elastic expansion and contraction method of the present application can be applied to Figure 1 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 A server is included as an example for explanation).

[0056] The terminal 100 can be used alone to execute the elastic expansion and contraction method provided in the embodiment of the present application. In addition, the terminal 100 and the server 200 can also be used together to execute the elastic expansion and contraction method provided in the embodiment of the present application.

[0057] Next describe Figure 1 The product form of the mid-terminal 100;

[0058] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.

[0059] Figure 2 An optional hardware structure diagram of the terminal 100 is shown.

[0060] refer to Figure 2 As shown, the terminal 100 may include a radio frequency unit 110, a first memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160, a speaker 161, a microphone 162, an earphone jack 163 (optional), a first processor 170, an external interface 180, a power supply 190 and other components. Those skilled in the art will appreciate that Figure 2 These are merely examples of terminals or multi-function devices and do not constitute limitations on the terminals or multi-function devices, which may include more or fewer components than those shown in the figures, or combinations of certain components, or different components.

[0061] The input unit 130 can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 and / or other input devices 132. The touch screen 131 can collect the user's touch operations on or near it (such as the user's operation on or near the touch screen using any suitable object such as fingers, joints, stylus, etc.), and drive the corresponding connection device according to a pre-set program. The touch screen can detect the user's touch action on the touch screen, convert the touch action into a touch signal and send it to the first processor 170, and can receive and execute the command sent by the first processor 170; the touch signal at least includes the touch point coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, the touch screen can be implemented using multiple types such as resistive, capacitive, infrared and surface acoustic wave. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.

[0062] Among them, other input devices 132 can receive input data and so on.

[0063] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, interactive interfaces, file display and / or playback of any multimedia file. In the embodiment of the present application, the display unit 140 can be used to display various interactive interfaces, processing results, etc. in the elastic expansion and contraction method.

[0064] The first memory 120 can be used to store instructions and data. The first memory 120 can mainly include an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files, texts, etc.; the instruction storage area can store software units such as operating systems, applications, instructions required for at least one function, or subsets and extensions of software units. It can also include a non-volatile random access memory; provide the first processor 170 with hardware, software and data resources including management computing and processing equipment, and support control software and applications. It is also used for the storage of multimedia files, and the storage of running programs and applications.

[0065] The first processor 170 is the control center of the terminal 100. It uses various interfaces and lines to connect various parts of the entire terminal 100. By running or executing instructions stored in the first memory 120 and calling data stored in the first memory 120, it executes various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the first processor 170 may include one or more processing units; preferably, the first processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the first processor 170. In some embodiments, the first processor 170 and the first memory 120 may be implemented on a single chip, and in some embodiments, the first processor 170 and the first memory 120 may also be implemented on separate chips. The first processor 170 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing and processing equipment, read and process data in the software, especially read and process data and programs in the first memory 120, so that each functional module therein performs corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0066] Among them, the first memory 120 can be used to store software codes related to the elastic scaling method, the first processor 170 can execute the steps of the elastic scaling method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.

[0067] The radio frequency unit 110 (optional) can be used for receiving and sending information or receiving and sending signals during a call, for example, after receiving the downlink information of the base station, it is sent to the first processor 170 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LowNoiseAmplifier, LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (Global System of Mobile Communication, GSM), General Packet Radio Service (General PacketRadio Service, GPRS), Code Division Multiple Access (Code Division Multiple Access, CDMA), Wideband Code Division Multiple Access (Wideband Code Division Multiple Access, WCDMA), Long Term Evolution (Long Term Evolution, LTE), email, Short Messaging Service (SMS), etc.

[0068] In this embodiment of the present application, the RF unit 110 can send data to the server 200 and receive processing results sent by the server 200.

[0069] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.

[0070] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the first processor 170 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.

[0071] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to communicate with other devices, or to connect a charger to charge the terminal 100 .

[0072] Although not shown, the terminal 100 may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied in the following embodiments. Figure 2 In the terminal 100 shown.

[0073] Next describe Figure 1 The product form of the server 200;

[0074] Figure 3 A structural diagram of a server 200 is provided, such as Figure 3 As shown, the server 200 includes a first bus 201, a second processor 202, a communication interface 203, and a second memory 204. The second processor 202, the second memory 204, and the communication interface 203 communicate with each other via the first bus 201.

[0075] The first bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The first bus 201 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0076] The second processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0077] The second memory 204 may include a volatile memory, such as a random access memory (RAM). The second memory 204 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0078] The second memory 204 may be used to store software codes related to the elastic scaling method, and the second processor 202 may execute the steps of the elastic scaling method of the chip, and may also schedule other units to implement corresponding functions.

[0079] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the first processor 170 in the above-mentioned terminal 100 and the second processor 202 in the server 200 can be hardware circuits (such as application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), general-purpose processors, DSPs, microprocessors or microcontrollers, etc.), or a combination of these hardware circuits. For example, the first processor 170 and the second processor 202 can be hardware systems with the function of executing instructions, such as CPU, DSP, etc., or hardware systems without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.

[0080] Next, the elastic expansion and contraction method provided in the embodiment of the present application is introduced. The method is applied to a computer device as an example. The computer device can be specifically Figure 1 The terminal 100 in the system, or the system consisting of the terminal 100 and the server 200, wherein the expansion and contraction decision system is installed on the terminal 100. Figure 4 The elastic expansion and contraction method may specifically include the following steps:

[0081] Step S401: periodically obtain multiple internal monitoring indicator values ​​of the Doris cluster.

[0082] The above-mentioned multiple internal monitoring indicator values ​​refer to the respective indicator values ​​of multiple internal monitoring indicators. The internal monitoring indicator refers to an indicator for observing the actual operating status of the Doris cluster within the process of the Doris cluster. The multiple includes 2 or more.

[0083] In this embodiment, an indicator collection cycle can be preset, and the internal monitoring indicator value is obtained once every indicator collection cycle. For example, if the indicator collection cycle is 30 seconds, the internal monitoring indicator value can be obtained every 30 seconds.

[0084] It can be understood that the Doris cluster includes at least one Doris node, and the process of "periodically obtaining multiple internal monitoring indicator values ​​of the Doris cluster" in this embodiment includes: periodically obtaining multiple internal monitoring indicator values ​​of each Doris node.

[0085] In a specific implementation, this embodiment can configure monitoring indicator exposure rules on each Doris node to configure which internal monitoring indicators can be collected or which internal monitoring indicators are not collected. Further, the internal monitoring indicators in the monitoring indicator exposure rules are exposed through the IP:PORT / metrics interface, and the interface address of the IP:PORT / metrics interface is registered in the configuration file of Prometheus. Thus, this embodiment can periodically query multiple internal monitoring indicator values ​​of the Doris cluster through the PQL (Prometheus Query Language) provided by Prometheus.

[0086] Here, the PQL provided by Prometheus is a powerful query language specially designed for retrieving, aggregating and manipulating time series data stored in the Prometheus server; Prometheus is a combination of open source monitoring, alarm, and time series databases. It stores monitoring indicators as time series data to achieve real-time analysis of system operation status, execution time, number of calls, and other functions.

[0087] Step S402: For each cycle, a composite indicator value corresponding to the cycle is obtained according to multiple internal monitoring indicator values ​​obtained in the cycle.

[0088] In this embodiment, each of the multiple internal monitoring indicator values ​​corresponds to a baseline value, and the baseline value refers to the internal monitoring indicator value of the Doris cluster under an ideal operating state.

[0089] Considering that in actual scaling operations, most of the internal monitoring indicator values ​​often deviate from the corresponding baseline values ​​at the same time, there are very few scenarios where scaling is required when a single internal monitoring indicator value deviates from the corresponding baseline value. Based on this, in order to capture the nonlinear linkage effect between multiple internal monitoring indicators in actual scaling operations, this embodiment can process multiple internal monitoring indicator values ​​into composite indicator values.

[0090] The degree of change of the above composite indicator value is positively correlated with the number of internal monitoring indicator values ​​deviating from the corresponding baseline values, that is, the more the number of internal monitoring indicator values ​​deviating from the corresponding baseline values ​​is, the greater the degree of change of the composite indicator value is.

[0091] Optionally, the direction of change of the composite index value may be consistent with or opposite to the direction in which the majority of internal monitoring indicators deviate from the corresponding baseline values, and this application does not make any specific limitation.

[0092] Step S403: When the composite indicator value corresponding to the target period meets the preset expansion and contraction condition, it is detected whether the composite indicator values ​​corresponding to each period within the target duration starting from the target period all meet the expansion and contraction condition.

[0093] In a specific implementation, this embodiment can determine whether each composite indicator value satisfies the preset expansion and contraction conditions each time a composite indicator value is obtained. If so, the period corresponding to the composite indicator value is determined as the target period, and then the composite indicator values ​​corresponding to each period within the target duration starting from the target period (including the composite indicator value corresponding to the target period) are continuously detected to see whether they all satisfy the same expansion and contraction conditions.

[0094] The above-mentioned expansion and contraction conditions include expansion conditions and / or contraction conditions. The expansion conditions can be, for example, , the shrinkage condition can be, for example, ,here, Represents the composite index value, Indicates the expansion threshold. Indicates the shrinking threshold, >1, <1.

[0095] Of course, both the expansion conditions and the reduction conditions can be other conditions, which are not specifically limited in this application, and the number of expansion conditions and reduction conditions is not limited in this application.

[0096] For example, the composite indicator value corresponding to the target period is the value obtained based on multiple internal monitoring indicator values ​​collected at time t0. The composite indicator value corresponding to time t0 meets the expansion condition. , the target duration is 3 minutes, and the indicator collection cycle is 30 seconds. Then, we can sequentially detect whether the composite indicator values ​​corresponding to time t1 (t0+30 seconds), time t2 (t0+60 seconds), time t3 (t0+90 seconds), time t4 (t0+120 seconds), time t5 (t0+150 seconds), and time t6 (t0+180 seconds) all meet the requirements. .

[0097] In this embodiment, the target duration is set to perform continuous multi-cycle detection, thereby avoiding misjudgment of the scaling decision response caused by instantaneous fluctuations, and improving the accuracy of the scaling decision response to a certain extent.

[0098] Step S404: If yes, then the Doris cluster is scaled up or down according to the scaling conditions.

[0099] Specifically, if the composite index values ​​corresponding to each period within the target duration meet the same expansion condition, the Doris cluster is expanded corresponding to the same expansion condition; if the composite index values ​​corresponding to each period within the target duration meet the same reduction condition, the Doris cluster is reduced corresponding to the same reduction condition. That is, when If the duration exceeds the target, the expansion operation is triggered. If the duration exceeds the target, the scaling-in operation is triggered.

[0100] In summary, the elastic scaling method provided by the present application periodically obtains multiple internal monitoring indicator values ​​of the Doris cluster, and for each cycle, obtains the composite indicator value corresponding to the cycle based on the multiple internal monitoring indicator values ​​obtained in the cycle. When the composite indicator value corresponding to the target cycle meets the preset scaling conditions, it is detected whether the composite indicator values ​​corresponding to each cycle within the target duration starting from the target cycle meet the scaling conditions. If so, the Doris cluster is scaled up or down corresponding to the scaling conditions. The internal monitoring indicators of the present application are indicators that observe the actual operating status of the Doris cluster from the process of the Doris cluster. Therefore, the internal monitoring indicators can better reflect the real-time load situation inside the Doris cluster, and elastic scaling is performed based on the indicator values ​​of the internal monitoring indicators, ensuring the real-time response of the scaling decision.

[0101] Furthermore, given that in actual scaling operations, most of the internal monitoring indicator values ​​deviate from the corresponding baseline values ​​at the same time, the present application processes multiple internal monitoring indicator values ​​into composite indicator values ​​in which the degree of change is positively correlated with the number of internal monitoring indicator values ​​that deviate from the corresponding baseline values, so that when a single internal monitoring indicator value deviates from the corresponding baseline value, the degree of change of the composite indicator value is small, and when multiple internal monitoring indicator values ​​deviate from the corresponding baseline value at the same time, the degree of change of the composite indicator value is large. As a result, scaling operations can be more accurately identified based on the composite indicator values, thereby improving the accuracy of scaling decision responses.

[0102] In some embodiments of the present application, the process of “periodically obtaining multiple internal monitoring indicator values ​​of the Doris cluster” in the above step S401 is introduced.

[0103] In order to more accurately reflect the load of the Doris cluster, this embodiment can select internal monitoring indicators unique to the Doris cluster. Based on this, the multiple internal monitoring indicators can include one or more of the following indicators:

[0104] (1) Query delay index: Taking Doris 2.1 as an example, the query delay index specifically refers to the query delay index of Doris at the TP98 percentile, in milliseconds (ms), which means that within a 5-minute window, the response speed of the first 98% of the query statements is less than this value.

[0105] The code implementation of the query latency indicator is: doris_fe_query_latency_ms{quantile="0.98"}, denoted as QLM.

[0106] (2) Transaction delay index: In Doris version 2.1, the transaction delay index specifically refers to the import transaction execution delay of Doris at the TP98 percentile, in milliseconds, indicating that within a 5-minute window, the time from the start of task loading to the end of the task for the first 98% of Stream Load transactions is less than this value.

[0107] Stream Load refers to a streaming loading task, which is a data loading method in Doris. For example, data is continuously written from the Kafka message queue to Doris. This process will start multiple Stream Load tasks in Doris, loading a part of the data into the Doris table each time.

[0108] The code implementation of the transaction latency indicator is: doris_fe_txn_publish_latency_ms{quantile="0.98"}, denoted as TPLM.

[0109] (3) compaction_score indicator: Doris continuously merges small files into orderly large files through the compaction mechanism. The compaction_score indicator can reflect the load pressure of the current merging work.

[0110] Compaction operation is a mechanism in Doris that continuously combines small files into large files. In the process of loading data, many small data block files will be generated. These data files will be gradually merged into large block files before they are officially written to the disk. Base compaction and Cumulative compaction are two different stages of the compaction operation. Cumulative compaction is the first compaction, and Base compaction is the second compaction of files that have been compacted for the first time. During the compaction process, if there are too many small files to be processed, they will accumulate in the queue to form pressure and slow down the entire system. Based on this, the compaction_score indicator can be expressed by the two indicators TBMCS and TCMCS.

[0111] TBMCS refers to the second-stage file compression queue accumulation index, which indicates the pressure score value of the current Doris Base Compaction operation, and the code is implemented as doris_be_tablet_base_max_compaction_score; TCMCS refers to the first-stage file compression queue accumulation index, which indicates the pressure score value of the current Doris Cumulative Compaction operation, and the code is implemented as doris_be_tablet_cumulative_max_compaction_score.

[0112] (4) Cache usage index: Doris's cache usage index reflects the overall pressure of the Doris cluster. When Doris queries data, it reads data from the storage medium. If it directly accesses the storage every time, the performance will be very poor. Therefore, Doris adds a cache mechanism. If a piece of data is repeatedly accessed, it can be directly retrieved from the cache later. If the cache usage rate is relatively high, it will increase the pressure on cluster garbage collection and increase the risk of memory overflow.

[0113] Optionally, the cache usage index includes the segment cache usage index and the index cache usage index. The index cache usage index refers to the memory used by the current index cache / all memory allocated to the index cache. If the index cache usage is high, it means that the access is scattered and the pressure on multi-tenants is high. The segment cache usage index refers to the memory used by the data segment cache / the memory allocated to the segment cache. If the segment cache usage is high, it means that the access is scattered and the pressure on multi-tenant access is high.

[0114] The code implementation of the index cache usage ratio indicator is: doris_be_cache_usage_ratio{ name="IndexPageCache"}, denoted as CURI.

[0115] The code implementation of the segment cache usage ratio indicator is: doris_be_cache_usage_ratio{ name="SegmentCache"}, denoted as CURS.

[0116] Based on this, the above-mentioned process of "periodically obtaining multiple internal monitoring indicator values ​​of the Doris cluster" can include: periodically obtaining the respective indicator values ​​of the query delay indicator QLM, the transaction delay indicator TPLM, the first-stage file compression queue accumulation indicator TCMCS, the second-stage file compression queue accumulation indicator TBMCS, the segment cache utilization indicator CURS and the index cache utilization indicator CURI; for each cycle, performing a first preprocessing on the respective indicator values ​​of the first-stage file compression queue accumulation indicator TCMCS and the second-stage file compression queue accumulation indicator TBMCS to obtain the indicator value of the compaction_score indicator; for each cycle, performing a second preprocessing on the respective indicator values ​​of the segment cache utilization indicator CURS and the index cache utilization indicator CURI to obtain the indicator value of the Cache utilization indicator.

[0117] Specifically, for each cycle, the index value of the query latency index QLM can be obtained, recorded as QL (QueryLatency); the index value of the transaction latency index TPLM can be obtained, recorded as TL (Transaction Latency); the index value of the first-stage file compression queue accumulation index TCMCS and the index value of the second-stage file compression queue accumulation index TBMCS are obtained, and the first preprocessing is performed to obtain the index value of the compaction_score index, recorded as CL (Compation Load); the index value of the segment cache usage index CURS and the index value of the index cache usage index CURI are obtained, and the second preprocessing is performed to obtain the index value of the cache usage index, recorded as CU (CacheUsage).

[0118] Optionally, the first preprocessing formula may be: ,in, Indicates the accumulation index of the second-stage file compression queue The weight coefficient of Indicates the first-stage file compression queue accumulation index The weight coefficient of , optional, It can take any value in the range of 0.1~0.2.

[0119] Optionally, the second preprocessing formula may be: ,in, Indicates the index cache usage indicator The weight coefficient of Indicates the segment cache usage indicator The weight coefficient of , optional, It can take any value in the range of 0.2~0.3.

[0120] It should be noted that the above-mentioned multiple internal monitoring indicators, first preprocessing, and second preprocessing are only examples and are not intended to limit the present application.

[0121] The embodiment of the present application adopts internal monitoring indicators that are unique to the Doris cluster and related to the Doris operating mechanism. The operating status of Doris can be observed from within Doris, which can more accurately reflect the load situation of the Doris cluster and improve the accuracy of the expansion and contraction decision response.

[0122] In some embodiments of the present application, the process of step S402 of "obtaining a composite indicator value corresponding to the period according to multiple internal monitoring indicator values ​​obtained in the period" is described.

[0123] In order to achieve the effect that "the degree of change of the composite index value is positively correlated with the number of deviations of the internal monitoring index value from the corresponding baseline value", this embodiment can be implemented in a variety of ways.

[0124] In a possible implementation, multiple internal monitoring indicator values ​​obtained in the period can be multiplied to obtain a composite indicator value corresponding to the period. Through the multiplication process, the degree of deviation of multiple internal monitoring indicator values ​​from the corresponding baseline value can be magnified, achieving the effect of "the degree of change of the composite indicator value is positively correlated with the number of internal monitoring indicator values ​​deviating from the corresponding baseline value".

[0125] However, the product operation has the problem of large amount of calculation and time consumption. In order to avoid direct product operation, the present application provides another implementation method, which is as follows.

[0126] As described above, the units of the multiple internal monitoring indicator values ​​may be different. In order to make the indicators in different units comparable, this embodiment needs to normalize the multiple internal monitoring indicator values ​​first.

[0127] Based on this, this embodiment can obtain the baseline values ​​of multiple internal monitoring indicators, and normalize the multiple internal monitoring indicator values ​​obtained in the period according to the baseline values ​​of the multiple internal monitoring indicators to obtain multiple normalized indicator values ​​corresponding to the period. Here, the baseline value refers to the internal monitoring indicator value of the Doris cluster under an ideal operating state.

[0128] Optionally, the baseline values ​​of the multiple internal monitoring indicators may be fixed, for example, always being the indicator values ​​of the multiple internal monitoring indicators at the startup time of the scaling decision system.

[0129] Preferably, the baseline values ​​of the multiple internal monitoring indicators are updated after each expansion or contraction of the expansion or contraction decision system, wherein, when the expansion or contraction decision system is not expanded or contracted, the baseline values ​​of the multiple internal monitoring indicators are the indicator values ​​of the multiple internal monitoring indicators at the system startup time; after each expansion or contraction of the expansion or contraction decision system, the baseline values ​​of the multiple internal monitoring indicators are the indicator values ​​of the multiple internal monitoring indicators after the system is stable.

[0130] More specifically, when the scaling decision system is started for the first time, since the scaling decision system has not yet performed scaling operations, in order to be able to perform elastic scaling within the time period from the startup of the scaling decision system to the first scaling, this embodiment can determine the indicator values ​​of multiple internal monitoring indicators at the startup time of the scaling decision system as the baseline values ​​of the multiple internal monitoring indicators.

[0131] Each time automatic expansion or contraction occurs subsequently, the expansion or contraction decision system will gradually stabilize within a period of time after the expansion or contraction. Then, this embodiment can determine the indicator values ​​of multiple internal monitoring indicators after the system is stable as the baseline values ​​of multiple internal monitoring indicators.

[0132] Since expansion or reduction will cause data imbalance between Doris nodes, Doris will automatically rebalance the data. At this time, the system automatically migrates and replicates data and will be in a high-load state for a short time. When the data copies of each backend are evenly distributed, we can consider that the system has reached a stable state.

[0133] Based on this, the optional process of judging the stability of the system may include: using the SQL instruction SHOWREPLICA DISTRIBUTION on Doris to obtain the percentage value of each Doris node in the Doris cluster (i.e., the Percent value of the backend), and then determining whether the scaling decision system is in a stable state based on the percentage value of each Doris node.

[0134] For example, use the SQL command SHOW REPLICA DISTRIBUTION on Doris to obtain the following data:

[0135] “BackendID | Percent

[0136] -------------------------------

[0137] Backend1 | 7.2%

[0138] Backend2 | 9.3%

[0139] Backend3 | 15.5%

[0140] …”.

[0141] The above data represents the data distribution of each backend (one backend represents one Doris node).

[0142] Assume that there are 5 backends, backend1 to backend5, then the average value of the distribution balance of the scaling decision system is: Avg = 100% / 5 = 20%. Assume that the preset requirement is that the maximum deviation between each backend does not exceed 5%, that is, all backends fall between [Avg-2.5%, Avg + 2.5%], that is, the percentage value of each backend is between 17.5% and 22.5%, then the scaling decision system is considered stable.

[0143] It should be noted that the above process of determining whether the scaling decision system is in a stable state is only an example and is not intended to limit the present application.

[0144] Considering that the indicator values ​​of Doris' internal monitoring indicators may have short-term mutations and drastic fluctuations, in addition, the entire link also includes various external systems such as K8S clusters and Prometheus. If the external systems fluctuate, the indicator values ​​of the internal monitoring indicators will also be affected. In order to avoid the impact of mutations and drastic fluctuations on the accuracy of the scaling decision response, optionally, the above-mentioned process of "normalizing the multiple internal monitoring indicator values ​​obtained in the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain the multiple normalized indicator values ​​corresponding to the period" may include: performing indicator smoothing on the multiple internal monitoring indicator values ​​obtained in the period to obtain the multiple indicator smoothing values ​​corresponding to the period, and normalizing the multiple indicator smoothing values ​​corresponding to the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain the multiple normalized indicator values ​​corresponding to the period.

[0145] Optionally, the process of "performing indicator smoothing processing on multiple internal monitoring indicator values ​​obtained for the period respectively to obtain multiple indicator smoothed values ​​corresponding to the period" may include: obtaining multiple indicator smoothed values ​​of the period before the period, and obtaining multiple indicator smoothed values ​​corresponding to the period based on the multiple indicator smoothed values ​​of the period before the period and the multiple internal monitoring indicator values ​​obtained in the period.

[0146] Optionally, the process of "obtaining multiple indicator smoothing values ​​corresponding to the period according to multiple indicator smoothing values ​​of the period before the period and multiple internal monitoring indicator values ​​obtained in the period" can be implemented using the following formula (1).

[0147] Formula (1);

[0148] in, Indicates Internal monitoring indicators obtained in a cycle (i.e., this cycle) The indicator value of Indicates Internal monitoring indicators corresponding to the cycle (i.e. this cycle) The indicator smoothing value of ; Indicates Internal monitoring indicators corresponding to the cycle (that is, the cycle before this cycle) The indicator smoothing value of ; ; Indicates the smoothing coefficient, and its value range is 0.9~0.95.

[0149] Optionally, the process of "normalizing the multiple indicator smoothing values ​​corresponding to the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain the multiple normalized indicator values ​​corresponding to the period" can refer to the following formula (2).

[0150] Formula (2);

[0151] in, Indicates internal monitoring indicators The calculation process of the indicator smoothing value can refer to the previous formula (1) The calculation process of Indicates internal monitoring indicators Baseline value of Indicates internal monitoring indicators The normalized index value of .

[0152] Above Ideally, the value is 1. Any deviation from 1 indicates that the load of the scaling decision system is abnormal.

[0153] It should also be noted that the above formula (2) is applicable to any period. However, the baseline values ​​of internal monitoring indicators obtained in different periods may be the same or different. The specific acquisition process is described in the previous article and will not be repeated here.

[0154] Furthermore, this embodiment can obtain a composite index value corresponding to the period according to a plurality of normalized index values ​​corresponding to the period.

[0155] In an optional embodiment, the process of "obtaining a composite index value corresponding to the period according to multiple normalized index values ​​corresponding to the period" may include: performing logarithmic transformation on the multiple normalized index values ​​corresponding to the period to obtain multiple logarithmic values ​​corresponding to the period, obtaining the weights of the multiple internal monitoring indicators, performing weighted summation on the weights of the multiple internal monitoring indicators and the multiple logarithmic values ​​corresponding to the period to obtain a weighted summation value corresponding to the period, performing exponential operation on the weighted summation value corresponding to the period to obtain a composite index value corresponding to the period. The process may refer to the following formula (3).

[0156] Formula (3);

[0157] Indicates the composite index value; Indicates internal monitoring indicators The weight of (it should be noted that the weight here has been normalized); Represents exponential operation with e as base; It means taking logarithm; .

[0158] Weights of the above internal monitoring indicators Characterizing internal control indicators Sensitivity to scaling decisions.

[0159] Optionally, the weights of each of the above-mentioned multiple internal monitoring indicators are updated once each weight adjustment cycle is reached; wherein, the weights of each of the multiple internal monitoring indicators in the latter weight adjustment cycle are obtained by adjusting and normalizing the weights of each of the multiple internal monitoring indicators in the previous weight adjustment cycle using the gradient descent method, and the weights of each of the multiple internal monitoring indicators in the first weight adjustment cycle are all preset initial values.

[0160] Optionally, the weight adjustment period is greater than the indicator collection period mentioned above, for example, the indicator collection period is 30 seconds and the weight adjustment period is 3 minutes.

[0161] by For example, the above preset initial value can be 0.25, that is, in the first weight adjustment cycle, the weights of the four indicators QL, TL, CL and CU are each 0.25, indicating that the four indicators are equally important.

[0162] In the subsequent process, dynamic adjustments can be made in each weight adjustment cycle according to changes in the current business scenario, such as using the gradient descent method to perform the following dynamic adjustments.

[0163] First, this embodiment can calculate the loss value , to measure the deviation between the current scaling decision and the target performance, the loss value The calculation formula is as follows:

[0164] Formula (4);

[0165] in, represents the target performance, that is, the composite index value of the Doris cluster under the ideal operating state. In this embodiment, The value can be 1.

[0166] Next, the loss value obtained by combining formula (4) , the following formula (5) is used to adjust the weights of multiple internal monitoring indicators in the previous weight adjustment cycle.

[0167] Formula (5);

[0168] in, Represents the weight of the internal monitoring indicator in the T+1th weight adjustment cycle (i.e. the next weight adjustment cycle) (it should be noted that the weight here has not been normalized); represents the weight of the internal monitoring indicator in the Tth weight adjustment cycle (i.e., the previous weight adjustment cycle) (it should be noted that the weight here has been normalized); represents the adjustment coefficient; , ; .

[0169] It is worth noting that in the above formulas (4) and (5) It refers to the composite index value obtained based on multiple internal monitoring index values ​​obtained in the corresponding weight adjustment period. For example, when the weights of multiple internal monitoring indexes in the first weight adjustment period are adjusted and normalized in the second weight adjustment period, It refers to the composite index value obtained from multiple internal monitoring index values ​​obtained in the second weight adjustment cycle. When the weights of multiple internal monitoring indexes in the second weight adjustment cycle are adjusted and normalized in the third weight adjustment cycle, It refers to the composite indicator value obtained from multiple internal monitoring indicator values ​​obtained within the third weight adjustment cycle, and so on.

[0170] It should also be noted that With the previous They have the same meaning and both represent the normalized weights. However, since the calculation process of formula (5) does not consider the normalization problem, No normalization was performed.

[0171] The formula for weight normalization can refer to the following formula (6).

[0172] Formula (6);

[0173] in, represents the weight calculated by formula (5) ; , .

[0174] It should be noted that the above , This is only an example and is not intended to limit the present application.

[0175] The dynamic weight adjustment mechanism in this embodiment can ensure that during the operation of the Doris cluster-based scaling decision system, the weight distribution of each internal monitoring indicator can be automatically optimized according to historical feedback.

[0176] In this embodiment, a composite index value is obtained by weighted geometric mean calculation of multiple memory monitoring index values. The composite index value is proportional to the multiple memory monitoring index values. The increase of any memory monitoring index value will cause the composite index value to increase. However, the increase of a single memory monitoring index value has little effect on the amplitude of the composite index value. However, if multiple memory monitoring index values ​​increase at the same time, since formula (3) is the product form of the internal monitoring index values, the composite index value will increase rapidly. This can better reflect the overall operating status of the scaling decision system, avoid being affected by a single memory monitoring index value, and amplify the response under the synergistic effect of multiple indicators, thereby achieving the goal of quickly and smoothly adjusting resources and significantly improving the system response speed and the stability of the Doris cluster.

[0177] In some other embodiments of the present application, the process of the above step S404 "performing expansion and contraction processing on the Doris cluster corresponding to the expansion and contraction conditions" is introduced.

[0178] Optionally, the process of "scaling the Doris cluster corresponding to the scaling conditions" may include: if the scaling condition is an expansion condition, the number of Doris nodes in the Doris cluster is increased by calling the Kubernetes interface through the deployment scheduler; if the scaling condition is a shrinking condition, the candidate node is selected in the Doris cluster through the deployment scheduler, and the candidate node is deleted after all the data of the candidate node is moved to other Doris nodes in the Doris cluster.

[0179] More specifically, in this embodiment, a custom resource (Custom Resource Definition, CRD) can be predefined, and expansion and contraction related fields can be defined in the CRD, such as the expected number of BEs, operation type (such as expansion type, contraction type), execution time window, expansion threshold, contraction threshold, etc.

[0180] An optional CRD code implementation is as follows:

[0181] "apiVersion: doris.example.com / v1

[0182] kind: DorisClusterScaling

[0183] metadata:

[0184] name: doris-scaling-task

[0185] spec:

[0186] action: "scale-out" # or scale-in

[0187] desiredReplica: 5

[0188] minInterval: 600

[0189] threshold:

[0190] scaleOut: 1.05

[0191] scaleIn: 0.95

[0192] duration: 180".

[0193] Among them, kind represents the custom Doris scaling task; spec.action represents the scaling operation, scale-out represents the expansion operation, and scale-in represents the reduction operation; spec.desiredReplica represents the target number of replicas, and 5 represents the expansion to 5 Doris nodes; spec.minInterval represents the minimum time interval, which represents the time interval between two operations; spec.threshold.scaleOut represents the expansion threshold, for example, when the composite index value is continuously greater than 1.05, the expansion is triggered; spec.threshold.scaleIn represents the reduction threshold, for example, when the composite index value is continuously less than 0.95, the reduction is triggered; spec.duration represents the weight adjustment period.

[0194] When the expansion conditions are met in step S403 of the previous text, the number of Doris nodes in the Doris cluster can be increased by calling the Kubernetes interface through the deployment scheduler (i.e., K8S Operator). For example, one Doris node is added each time the expansion is triggered.

[0195] More specifically, the number of Doris BE Pods can be increased by calling the K8S API (Application Programming Interface) through the deployment scheduler Operator. Here, Doris BE (Backend) is the backend node in the Apache Doris (formerly known as Palo) architecture, which is mainly responsible for data storage, query execution and result return. Doris BE works with FE (Frontend) nodes to provide users with high-performance real-time analysis services. In the Kubernetes environment, Doris BE is usually deployed in the form of Pod. Pod is the smallest deployable unit in Kubernetes and contains an environment for running one or more containers. Doris BE Pod usually contains one or more Doris BE containers, which run on the same physical or virtual node and share network namespaces and storage volumes.

[0196] When the previous step S403 meets the scaling-down condition, the candidate node can be selected in the Doris cluster by deploying the scheduler, and the candidate node can be deleted after all the data of the candidate node is moved to other Doris nodes in the Doris cluster.

[0197] More specifically, you can first select at least one candidate node from the Doris cluster, for example, randomly select a candidate node, mark the selected candidate node as "to be offline" in the Doris cluster, and then wait for the tablet to automatically migrate the data on the candidate node marked "to be offline" to other Doris nodes, and poll the BE status at the same time until the tablet returns to 0, and then call the kubernetes API through the Operator to delete the Pod corresponding to the candidate node.

[0198] This embodiment combines the composite indicator value mentioned above with the Kubernetes Operator mechanism, realizes automated operation through custom resource CRD, and comprehensively solves the shortcomings of the existing technology in load detection, data balancing and system stability.

[0199] The elastic expansion and contraction device provided in the embodiment of the present application is described below. The elastic expansion and contraction device described below and the elastic expansion and contraction method described above can be referred to each other.

[0200] See also Figure 5 , Figure 5 A schematic diagram of the structure of an elastic expansion and contraction device provided in an embodiment of the present application.

[0201] like Figure 5 As shown, the device may include:

[0202] The indicator acquisition module 501 is used to periodically obtain multiple internal monitoring indicator values ​​of the Doris cluster, where the multiple internal monitoring indicator values ​​refer to the indicator values ​​of the multiple internal monitoring indicators, and the internal monitoring indicator refers to the indicator for observing the actual running status of the Doris cluster inside the process of the Doris cluster;

[0203] The index compound module 502 is used to obtain, for each period, a composite index value corresponding to the period according to multiple internal monitoring index values ​​obtained in the period, wherein the degree of change of the composite index value is positively correlated with the number of deviations of the internal monitoring index value from the corresponding baseline value;

[0204] The expansion / contraction judgment module 503 is used to detect whether the composite index values ​​corresponding to each period within the target duration starting from the target period all meet the expansion / contraction conditions when the composite index value corresponding to the target period meets the preset expansion / contraction conditions;

[0205] The scaling operation module 504 is used to perform scaling processing corresponding to the scaling conditions on the Doris cluster when the scaling judgment module detects that the composite index values ​​corresponding to each period within the target duration starting from the target period meet the scaling conditions.

[0206] In a possible implementation, the above-mentioned multiple internal monitoring indicators include one or more of the following indicators: a query delay indicator, a transaction delay indicator, a compaction_score indicator and a cache utilization indicator, the compaction_score indicator includes a first-stage file compression queue accumulation indicator and a second-stage file compression queue accumulation indicator, and the cache utilization indicator includes a segment cache utilization indicator and an index cache utilization indicator.

[0207] Based on this, when the above indicator acquisition module periodically obtains multiple internal monitoring indicator values ​​of the Doris cluster, it can be specifically used for:

[0208] Periodically obtain the respective indicator values ​​of the query delay indicator, the transaction delay indicator, the first-stage file compression queue accumulation indicator, the second-stage file compression queue accumulation indicator, the segment cache utilization indicator, and the index cache utilization indicator;

[0209] For each cycle, the first preprocessing is performed on the index values ​​of the first-stage file compression queue accumulation index and the second-stage file compression queue accumulation index to obtain the index value of the compaction_score index;

[0210] For each cycle, a second preprocessing is performed on the respective index values ​​of the segment cache usage index and the index cache usage index to obtain the index value of the cache usage index.

[0211] In a possible implementation, when the above-mentioned indicator composite module obtains the composite indicator value corresponding to the period according to the multiple internal monitoring indicator values ​​obtained in the period, it can be specifically used to:

[0212] Get the baseline values ​​and weights of multiple internal monitoring indicators. The baseline value refers to the internal monitoring indicator value of the Doris cluster under ideal operating conditions, and the weight represents the sensitivity of the internal monitoring indicator to the scaling decision.

[0213] Performing indicator smoothing processing on multiple internal monitoring indicator values ​​obtained in the period respectively to obtain multiple indicator smoothing values ​​corresponding to the period;

[0214] Normalizing the multiple indicator smoothing values ​​corresponding to the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain multiple normalized indicator values ​​corresponding to the period;

[0215] Perform logarithmic transformation on multiple normalized index values ​​corresponding to the period to obtain multiple logarithmic values ​​corresponding to the period;

[0216] Performing weighted summation according to respective weights of the plurality of internal monitoring indicators and the plurality of logarithmic values ​​corresponding to the period, to obtain a weighted summation value corresponding to the period;

[0217] Perform exponential operation on the weighted sum value corresponding to the period to obtain the composite index value corresponding to the period.

[0218] In one possible implementation, the baseline value of each of the above-mentioned multiple internal monitoring indicators is updated after each expansion or contraction of the expansion or contraction decision system; wherein, when the expansion or contraction decision system is not expanded or contracted, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators at the time of system startup; and after each expansion or contraction of the expansion or contraction decision system, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators after the system is stable.

[0219] In one possible implementation, the weights of each of the above-mentioned multiple internal monitoring indicators are updated once each weight adjustment cycle is reached; wherein, the weights of each of the multiple internal monitoring indicators in the latter weight adjustment cycle are obtained by adjusting and normalizing the weights of each of the multiple internal monitoring indicators in the previous weight adjustment cycle using the gradient descent method, and the weights of each of the multiple internal monitoring indicators in the first weight adjustment cycle are all preset initial values.

[0220] In a possible implementation, when the above-mentioned expansion and contraction operation module performs expansion and contraction processing on the Doris cluster corresponding to the expansion and contraction conditions, it can be specifically used to:

[0221] If the scaling condition is the expansion condition, the number of Doris nodes in the Doris cluster is increased by calling the Kubernetes interface through the deployment scheduler;

[0222] If the expansion condition is a shrinking condition, the candidate node is selected in the Doris cluster by deploying the scheduler, and the candidate node is deleted after all the data of the candidate node is moved to other Doris nodes in the Doris cluster.

[0223] The present application also provides an electronic device in an embodiment. Figure 6 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0224] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a second bus 604. An input / output (I / O) interface 605 is also connected to the second bus 604.

[0225] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0226] A computer program product is also provided in an embodiment of the present application, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any elastic expansion and contraction method provided in the embodiment of the present application.

[0227] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any elastic scaling method provided in the embodiment of the present application.

[0228] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0229] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0230] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0231] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. A method for elastic expansion and contraction, characterized in that: include: Periodically obtain multiple internal monitoring indicator values ​​of the Doris cluster, wherein the multiple internal monitoring indicator values ​​refer to the indicator values ​​of the multiple internal monitoring indicators, and the internal monitoring indicator refers to an indicator for observing the actual operating status of the Doris cluster within the process of the Doris cluster; For each cycle, a composite indicator value corresponding to the cycle is obtained according to the multiple internal monitoring indicator values ​​obtained in the cycle, wherein the degree of change of the composite indicator value is positively correlated with the number of deviations of the internal monitoring indicator value from the corresponding baseline value; When the composite index value corresponding to the target period meets the preset expansion and contraction condition, detecting whether the composite index values ​​corresponding to each period within the target duration starting from the target period all meet the expansion and contraction condition; If so, the Doris cluster is expanded or reduced in accordance with the expansion or reduction conditions.

2. The elastic expansion and contraction method according to claim 1, characterized in that: The multiple internal monitoring indicators include one or more of the following indicators: a query delay indicator, a transaction delay indicator, a compaction_score indicator, and a cache utilization indicator, wherein the compaction_score indicator includes a first-stage file compression queue accumulation indicator and a second-stage file compression queue accumulation indicator, and the cache utilization indicator includes a segment cache utilization indicator and an index cache utilization indicator; The periodic acquisition of multiple internal monitoring indicator values ​​of the Doris cluster includes: Periodically obtaining the respective index values ​​of the query delay index, the transaction delay index, the first-stage file compression queue accumulation index, the second-stage file compression queue accumulation index, the segment cache utilization index, and the index cache utilization index; For each cycle, performing a first preprocessing on the index values ​​of the first-stage file compression queue accumulation index and the second-stage file compression queue accumulation index to obtain the index value of the compaction_score index; For each cycle, a second preprocessing is performed on the respective index values ​​of the segment cache usage index and the index cache usage index to obtain the index value of the cache usage index.

3. The elastic expansion and contraction method according to claim 1 or 2, characterized in that: The step of obtaining a composite indicator value corresponding to the period according to the multiple internal monitoring indicator values ​​obtained in the period includes: Obtaining the baseline values ​​and weights of the multiple internal monitoring indicators, wherein the baseline value refers to the internal monitoring indicator value of the Doris cluster under an ideal operating state, and the weight represents the sensitivity of the internal monitoring indicator to the scaling decision; Performing indicator smoothing processing on the multiple internal monitoring indicator values ​​obtained in the period respectively to obtain multiple indicator smoothing values ​​corresponding to the period; Normalizing the multiple indicator smoothing values ​​corresponding to the period according to the respective baseline values ​​of the multiple internal monitoring indicators to obtain multiple normalized indicator values ​​corresponding to the period; Perform logarithmic transformation on multiple normalized index values ​​corresponding to the period to obtain multiple logarithmic values ​​corresponding to the period; Performing weighted summation according to the respective weights of the multiple internal monitoring indicators and the multiple logarithmic values ​​corresponding to the period to obtain a weighted summation value corresponding to the period; Perform exponential operation on the weighted sum value corresponding to the period to obtain the composite index value corresponding to the period.

4. The elastic expansion and contraction method according to claim 3, characterized in that: The baseline values ​​of the multiple internal monitoring indicators are updated after each expansion or contraction of the expansion or contraction decision system; Among them, when the scaling decision system is not scaled up or down, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators at the time of system startup; after each scaling of the scaling decision system, the baseline value of each of the multiple internal monitoring indicators is the indicator value of each of the multiple internal monitoring indicators after the system is stable.

5. The elastic expansion and contraction method according to claim 3, characterized in that: The weights of the multiple internal monitoring indicators are updated every time a weight adjustment cycle is reached; Among them, the weights of each of the multiple internal monitoring indicators in the latter weight adjustment cycle are obtained by adjusting and normalizing the weights of each of the multiple internal monitoring indicators in the previous weight adjustment cycle using the gradient descent method, and the weights of each of the multiple internal monitoring indicators in the first weight adjustment cycle are all preset initial values.

6. The elastic expansion and contraction method according to claim 1, characterized in that: The step of performing expansion and contraction processing on the Doris cluster corresponding to the expansion and contraction conditions includes: If the expansion and contraction condition is an expansion condition, the number of Doris nodes in the Doris cluster is increased by calling the Kubernetes interface through the deployment scheduler; If the expansion and contraction condition is a contraction condition, a candidate node is selected in the Doris cluster by the deployment scheduler, and after all the data of the candidate node is moved to other Doris nodes in the Doris cluster, the candidate node is deleted.

7. An elastic expansion and contraction device, characterized in that: include: An indicator acquisition module, used to periodically acquire multiple internal monitoring indicator values ​​of the Doris cluster, wherein the multiple internal monitoring indicator values ​​refer to the respective indicator values ​​of multiple internal monitoring indicators, and the internal monitoring indicator refers to an indicator for observing the actual operating status of the Doris cluster within the process of the Doris cluster; An indicator compound module, for obtaining, for each period, a composite indicator value corresponding to the period according to the multiple internal monitoring indicator values ​​obtained in the period, wherein the degree of change of the composite indicator value is positively correlated with the number of deviations of the internal monitoring indicator value from the corresponding baseline value; The expansion and contraction judgment module is used to detect whether the composite index values ​​corresponding to each period within the target duration starting from the target period all meet the expansion and contraction conditions when the composite index value corresponding to the target period meets the preset expansion and contraction conditions; The expansion and contraction operation module is used to perform expansion and contraction processing corresponding to the expansion and contraction conditions on the Doris cluster when the expansion and contraction judgment module detects that the composite indicator values ​​corresponding to each period within the target duration starting from the target period meet the expansion and contraction conditions.

8. A computer program product, characterized in that It includes computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the elastic expansion and contraction method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the elastic expansion and contraction method as described in any one of claims 1 to 6.

10. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the elastic scaling method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Dynamic capacity expansion and contraction method and device

    CN117632897A

  • Power consumer energy consumption stability analysis method based on autoregression algorithm

    CN117633710A

  • Method, system, medium and equipment for constructing composite index for expanding and shrinking container

    CN117667387A