Method, computer device, and computer program for accurately predicting and utilizing future system usage
A time series prediction model for cluster systems using input indices with few sudden changes accurately predicts future system performance, enhancing resource management and stability by correlating input and system performance indices.
Patent Information
- Application Number
- JP2024220443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-17
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2044-12-17
AI Technical Summary
Existing cluster monitoring technologies struggle with accurate prediction of system performance indices due to sudden changes, leading to inefficiencies in resource management and maintenance.
A method using a time series prediction model trained on input indices with few sudden change elements to predict future system performance by correlating input indices with system performance indices, reducing outliers and improving accuracy.
Enables stable system management and reduced operational resources by predicting future system performance through input index changes, facilitating automated resource operations like equipment addition or removal.
Smart Images

Figure 2025104293000001_ABST
Abstract
Description
Technical Field
[0001] The following description relates to a technique for predicting system usage.
Background Art
[0002] A cluster is a system that connects independent computer systems such as personal computers by network equipment for special purposes, and is widely used as a solution for providing high performance computing and high availability computing.
[0003] This has the advantage that upgrades and expansions can be achieved simply by replacing or adding some nodes or networks without changing the entire system, providing flexible scalability. Due to such advantages, clusters are constructed and used in enterprises, research institutes, university laboratories, etc. that require complex scientific computing problems according to the required computing requirements.
[0004] For such clusters, monitoring technologies are required to continuously observe and maintain the normal operation and quality of the clusters.
[0005] For example, Patent Document 1 (publication date: June 16, 2009) discloses a large-scale cluster monitoring technology that can actively take measures against configuration information when a node failure occurs in a cluster system environment and can continuously improve availability.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] As an input index having a correlation with the system performance index, a time series prediction model is constructed using an input index with few sudden change elements.
[0008] Rather than directly predicting the system performance index, the input index at a future time point is predicted using a prediction model, and the system performance index based on the correlation is calculated.
[0009] The future system performance index is utilized in conjunction with a system that automates resource operation such as the addition or removal of equipment (machine).
Means for Solving the Problems
[0010] A method for predicting the future system usage amount of a computer device including at least one processor, the method comprising: constructing, by at least one processor, a prediction model for an input index using time series data of the input index stored during a first fixed period; obtaining, by at least one processor, a correlation between the input index and the system performance index using time series data of the input index and the system performance index stored during a second fixed period; predicting, by at least one processor, the input index at at least one future time point using the prediction model; and calculating, by at least one processor, the system performance index at the future time point using the predicted input index and the correlation.
[0011] According to one aspect, the input index is characterized in that, among the indexes to which the system resources respond, an index with a change within a third fixed period less than a threshold is selected.
[0012] According to another aspect, user traffic corresponding to log data imported from the outside is used as the input index, and CPU usage or disk usage is used as the system performance index.
[0013] According to another aspect, the step of obtaining the correlation relationship is characterized by obtaining the correlation relationship between the input index and the system performance index using a linear regression model.
[0014] According to another aspect, the step of obtaining the correlation relationship may include the step of determining a linear regression coefficient value and a linear regression intercept value by linear regression on the time series data of the second fixed period.
[0015] According to another aspect, the linear regression intercept value is characterized in that a value close to 0 is determined.
[0016] According to another aspect, the step of constructing the prediction model may include the step of updating the prediction model using the data stored after the previous prediction time point for the input index.
[0017] According to another aspect, the step of obtaining the correlation relationship may include the step of updating the correlation relationship between the input index and the system performance index at a certain period.
[0018] According to another aspect, the step of obtaining the correlation relationship is characterized in that, for some types according to the type of the system performance index, the correlation relationship between the input index and the system performance index is obtained using the value obtained by integrating the time series data of the input index over time.
[0019] According to another aspect, the future system usage prediction method may further include the step of recommending a component configuration at a future time point using the calculated system performance index by at least one processor.
[0020] According to another aspect, the step of calculating the future system performance index calculates the future system performance index of each component constituting the cluster, and the step of recommending the future component configuration recommends the number of equipment units of the corresponding component based on the future system performance index of each component.
[0021] According to another aspect, the future system usage prediction method may further include a step of calculating, by at least one processor, the system performance index of each component constituting a new cluster by utilizing the correlation between the received input index and the received input index when the received input index is provided for constructing a new cluster, and a step of calculating and recommending the required number of equipment units of each component based on the system performance index of each component by at least one processor.
[0022] According to still another aspect, the future system usage prediction method may further include a step of providing, by at least one processor, a time point at which the calculated system performance index exceeds a threshold range in the future.
[0023] A computer program for causing a computer device to execute the future system usage prediction method is provided.
[0024] A computer device is provided, which includes at least one processor configured to execute instructions readable by the computer device, and the at least one processor constructs a prediction model of the input index by using time series data of the input index stored during a first fixed period, obtains the correlation between the input index and the system performance index by using time series data of the input index and the system performance index stored during a second fixed period, predicts the input index at a future time point by using the prediction model for at least one future time point, and calculates the system performance index at the future time point by utilizing the correlation with the predicted input index.
Advantages of the Invention
[0025] According to an embodiment of the invention, as an input index having a correlation with a system performance index, by constructing a prediction model for predicting a change in the input index by causing a time series prediction model to learn an input index with few sudden change factors, outliers in the learning data can be reduced and the accuracy of prediction can be improved.
[0026] According to an embodiment of the invention, instead of directly predicting the system performance index, the change in the input index is predicted using a prediction model, and the system performance index based on the correlation is calculated using the predicted value of the input index, so that the future system performance can be easily predicted only by the input index.
[0027] According to an embodiment of the invention, by linking and utilizing the future system performance index to a system that automates resource operations such as equipment installation and exclusion, the system can be stably maintained and managed, and the operation resources required for calculating the number of equipment when configuring a new cluster can be reduced.
Brief Description of the Drawings
[0028]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0029] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0030] Embodiments of the present invention relate to a technique for predicting system usage.
[0031] Embodiments including the matters specifically disclosed in this specification use input indicators with few sudden change elements as input indicators having a correlation with system performance indicators to train a time series prediction model to predict changes in the input indicators, thereby calculating future system performance indicators based on the predicted changes in the input indicators and providing an accurate prediction of future system usage.
[0032] The future system usage prediction device according to an embodiment of the present invention may be realized by at least one computer device, and the future system usage prediction method according to an embodiment of the present invention may be executed by at least one computer device included in the future system usage prediction device. At this time, in the computer device, a computer program according to an embodiment of the present invention may be installed and executed, and the computer device may execute the future system usage prediction method according to an embodiment of the present invention according to the control of the executed computer program. The above-described computer program may be recorded on a computer-readable recording medium in order to be combined with the computer device to cause the computer to execute the future system usage prediction method.
[0033] FIG. 1 is a diagram showing an example of a network environment in an embodiment of the present invention. The network environment of FIG. 1 shows an example including a plurality of electronic devices 110, 120, 130, 140, a plurality of servers 150, 160, and a network 170. Such FIG. 1 is merely an example for explaining the invention, and the number of electronic devices and the number of servers should not be limited as in FIG. 1. Further, the network environment of FIG. 1 is merely an example for explaining an environment applicable to the present embodiment, and the environment applicable to the present embodiment should not be limited to the network environment of FIG. 1.
[0034] The plurality of electronic devices 110, 120, 130, 140 may be fixed terminals or mobile terminals realized by a computer device. Examples of the plurality of electronic devices 110, 120, 130, 140 include smartphones, mobile phones, navigation devices, PCs (personal computers), notebook PCs, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablets, and the like. As an example, in FIG. 1, a smartphone is shown as an example of the electronic device 110. However, in the embodiments of the present invention, the electronic device 110 may mean one of various physical computer systems that substantially utilize a wireless or wired communication method and can communicate with other electronic devices 120, 130, 140 and / or servers 150, 160 via the network 170.
[0035] The communication method is not limited, and it may include not only a communication method that utilizes a communication network that the network 170 can include (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network), but also short-range wireless communication between devices. For example, the network 170 may include any one or more of networks such as a PAN (personal area network), a LAN (local area network), a CAN (campus area network), a MAN (metropolitan area network), a WAN (wide area network), a BBN (broadband network), and the Internet. Further, the network 170 may include any one or more of network topologies including a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree, or a hierarchical network, but is not limited thereto.
[0036] Each of the servers 150 and 160 may be implemented by one or more computer devices that communicate with a plurality of electronic devices 110, 120, 130, and 140 via a network 170 to provide instructions, code, files, content, services, and the like. For example, the server 150 may be a system that provides services (such as, for example, a log integration management service) to a plurality of electronic devices 110, 120, 130, and 140 connected via the network 170.
[0037] FIG. 2 is a block diagram showing an example of a computer device according to an embodiment of the present invention. Each of the plurality of electronic devices 110, 120, 130, and 140 and each of the servers 150 and 160 described above may be implemented by the computer device 200 shown in FIG. 2.
[0038] As shown in FIG. 2, such a computer device 200 may include a memory 210, a processor 220, a communication interface 230, and an input / output interface 240. The memory 210 is a computer-readable recording medium and may include a RAM (random access memory), a ROM (read only memory), and a persistent mass storage device such as a disk drive. Here, a persistent mass storage device such as a ROM or a disk drive may be included in the computer device 200 as a separate persistent recording device distinct from the memory 210. Also, an operating system and at least one program code may be recorded in the memory 210. Such software components may be loaded into the memory 210 from a computer-readable recording medium different from the memory 210. Such another computer-readable recording medium may include computer-readable recording media such as a floppy (registered trademark) drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, and the like. In other embodiments, the software components may be loaded into the memory 210 through the communication interface 230 that is not a computer-readable recording medium. For example, the software components may be loaded into the memory 210 of the computer device 200 based on a computer program installed by a file received via the network 170.
[0039] The processor 220 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processor 220 by the memory 210 or the communication interface 230. For example, the processor 220 may be configured to execute instructions received according to program code recorded in a recording device such as the memory 210.
[0040] The communication interface 230 may provide a function for the computer device 200 to communicate with other devices (for example, the recording device described above) via the network 170. As an example, requests, instructions, data, files, etc. generated by the processor 220 of the computer device 200 according to program codes recorded in a recording device such as the memory 210 may be transmitted to other devices via the network 170 under the control of the communication interface 230. Conversely, signals, instructions, data, files, etc. from other devices may be received by the computer device 200 through the communication interface 230 of the computer device 200 via the network 170. Signals, instructions, data, etc. received through the communication interface 230 may be transmitted to the processor 220 or the memory 210, and files, etc. may be recorded in a recording medium (the permanent recording device described above) that the computer device 200 may further include.
[0041] The input / output interface 240 may be means for interfacing with the input / output device 250. For example, the input device may include devices such as a microphone, a keyboard, or a mouse, and the output device may include devices such as a display or a speaker. As another example, the input / output interface 240 may be means for interfacing with a device in which functions for input and output are integrated into one, such as a touch screen. The input / output device 250 may be configured as one device together with the computer device 200.
[0042] Also, in other embodiments, the computer device 200 may include fewer or more components than the components in FIG. 2. However, it is not necessary to clearly show most of the conventional components in the figure. For example, the computer device 200 may be implemented to include at least a part of the input / output device 250 described above, or may further include other components such as a transceiver or a database.
[0043] Hereinafter, specific embodiments of the technology for accurate prediction and utilization of future system usage amounts will be described.
[0044] This embodiment is directed to a cluster system environment and includes a technique for predicting system usage and a technique for utilizing the prediction result of system usage in system operation.
[0045] In this specification, a cluster means a system that receives and processes a large amount of data, and FIG. 3 shows an example of a cluster system.
[0046] Referring to FIG. 3, the integrated log management system 300 is a system that receives logs 301 of various services and supports storage to the real-time indexing & search system 310 and the data storage platform 320. That is, the integrated log management system 300 is a platform that collects logs 301 generated by various services in one place and provides services so that they can be used for problem tracking and data analysis.
[0047] The integrated log management system 300 is composed of an open-source framework that processes a large amount of data in a distributed environment, and stores it by dividing or replicating it on multiple disks of a plurality of clustered computers.
[0048] Such a cluster of the integrated log management system 300 is composed of a plurality of components, and each component is composed of at least one or more devices (for example, servers, etc.).
[0049] In the following embodiments, the integrated log management system 300 composed of a plurality of components will be described as a typical example of a cluster, but it should not be limited thereto, and any system environment that consumes resources for input data given from the outside can be widely applied.
[0050] The computer device 200 according to this embodiment may include a future system usage prediction device implemented by a computer. The processor 220 of the computer device 200 may be implemented by components for executing the following future system usage prediction method. According to the embodiment, the components of the processor 220 may be selectively included in or excluded from the processor 220. Also, according to the embodiment, the components of the processor 220 may be separated or combined for the expression of the functions of the processor 220.
[0051] Such a processor 220 and the components of the processor 220 may control the computer device 200 to execute the steps included in the following future system usage prediction method. For example, the processor 220 and the components of the processor 220 may be implemented to execute instructions by the code of the operating system included in the memory 210 and the code of at least one program.
[0052] Here, the components of the processor 220 may be representations of different functions that are executed by the processor 220 according to the instructions provided by the program code recorded in the computer device 200.
[0053] The processor 220 may read the necessary instructions from the memory 210 in which the instructions related to the control of the computer device 200 are loaded. In this case, the read instructions may include instructions for controlling the processor 220 to execute the steps described below.
[0054] The steps included in the future system usage prediction method described below may be executed in an order different from that shown in the figure, and some of the steps may be omitted or additional processes may be further included.
[0055] FIG. 4 is a flowchart showing an example of a method that can be executed by a computer device in an embodiment of the present invention.
[0056] Referring to FIG. 4, at stage 410, the processor 220 may construct a prediction model for the input metric using the time-series data of the pre-selected input metric. The processor 220 may use the time-series data of the input metric stored for a relatively long period as learning data, and generate a time-series prediction model for the input metric by having this learned by a time-series prediction library (such as Prophet, etc.).
[0057] In the present embodiment, as an index to which the system resources respond, an external index with few sudden changes may be selected as the input index. An index for analyzing the input index and the output index of the system for a long period and using an index whose short-term change is less than a threshold as learning data for the prediction model may be selected. System performance indexes such as CPU and disk usage have characteristics in which large changes occur in a short period due to parameter tuning, rollback, or failure situations, etc. Using such indexes as learning data leads to large prediction errors. On the other hand, user traffic, which is log data imported from the outside and is one of the input indexes, does not have large short-term changes. Therefore, using this as learning data for the prediction model can reduce outliers in the learning data and improve the accuracy of prediction. As an example, the processor 220 may call the traffic time-series data stored for a long period and construct a prediction model for predicting traffic changes by having the called traffic data learned by a time-series prediction library.
[0058] In stage 420, the processor 220 may determine the correlation between the input metric and the system performance metric. The system performance metric is an output metric having characteristics that respond to the input metric, and may mean, for example, an output metric indicating system performance such as CPU usage (number of utilized cores per second), disk usage (capacity of the disk in use), memory usage (capacity of the memory in use), etc. System performance metrics such as CPU usage, disk usage, and memory usage are related to user traffic, which is one of the input metrics. For example, in the case of a search service, the correlation between search traffic and the CPU / disk usage of the search server may be utilized, and in the case of a blog service, the correlation between usage traffic such as blog search, posting, and neighbor notification and the CPU / disk usage of each blog server may be utilized. In the case of a shopping service, the correlation between usage traffic such as shopping search, purchase, and shipping inquiry and the CPU / disk usage of each shopping server may be utilized. The processor 220 calls the time-series data of the input metric and the system performance metric for a short period (e.g., three days) respectively, and uses a linear regression model to derive the correlation between the input metric and the system performance metric. In this embodiment, it has been described that the correlation between the input metric and the system performance metric is derived by linear regression, but it should not be limited thereto, and any technique that can define the correlation between the input metric and the system performance metric from the time-series data is applicable.
[0059] In step 430, after predicting the input metrics at a future time point (e.g., 7 days later, 30 days later, etc.) using a prediction model for at least one future time point, the processor 220 may calculate the system performance metrics at the future time point by utilizing the predicted input metrics and the correlation relationship obtained in step 420. In this embodiment, instead of directly predicting system performance metrics with many sudden change factors, the future change of input metrics with few sudden change factors is predicted using the prediction model constructed in step 410. At this time, the processor 220 may apply the correlation relationship between the input metrics and the system performance metrics to the input metrics at the future time point predicted using the prediction model to calculate the system performance metrics at the future time point. That is, the processor 220 predicts in advance the change of the input metrics at each future time by time series prediction, and predicts the future change of the system performance metrics by utilizing the predicted input metrics and the correlation relationship obtained in advance.
[0060] In step 440, the processor 220 may recommend the component configuration at the corresponding time point based on the system performance metrics at the future time point. As an example, the processor 220 predicts the system performance metrics at the future time point of each component constituting the cluster in a situation where the trend for the input metrics is maintained, and when the future predicted value of the system performance metrics exceeds the threshold range, recommends adding additional equipment for the corresponding component, and when the future predicted value of the system performance metrics is less than the threshold range, may recommend removing some of the equipment for the corresponding component. In order to configure an efficient cluster, the number of equipment to be added or removed is calculated and recommended based on the predicted future change value of the system performance metrics. As another example, when a predicted value of the input metrics generated by the new linkage due to adding a cluster is given during new linkage due to adding a cluster, the processor 220 may recommend the number of equipment for each component of the corresponding cluster. When constructing a new cluster, the number of equipment required for each component is calculated and recommended.
[0061] FIG. 5 to FIG. 6 are diagrams showing examples of input metrics and system performance metrics in an embodiment of the present invention.
[0062] Most services have the need to provide stable services by deploying more equipment when the resource usage is too high according to the resource usage of each piece of equipment, and the need to reduce operating costs by removing some of the existing equipment when the resource usage is too low.
[0063] To achieve stable system operation and reduction of operation resources, a precise prediction of future system usage is required. For this purpose, it is important to build a model with high prediction accuracy and to consider the correlation between input indicators and system performance indicators.
[0064] Referring to FIG. 5, in this embodiment, user traffic 501 indicating the amount of log capture by service users (e.g., BPS (byte / s), QPS (query / s), etc.) may be utilized as an input indicator, and the CPU usage 502 of components that respond to such user traffic 501 may be utilized as a system performance indicator. According to the time-series data of user traffic 501 and CPU usage 502 stored for a long time, it can be seen that the relationship between user traffic 501 and CPU usage 502 is extremely high.
[0065] Referring to FIG. 6, not only the CPU usage 502, but also the disk usage 602 corresponds to the system performance indicator that responds to the user traffic 501. Based on the time-series data of user traffic 501 and disk usage 602 stored for a long time, it can be grasped that the value 601 obtained by integrating the BPS corresponding to the user traffic 501 over time is similar to the disk usage 602. Thereby, it can be seen that the relationship between user traffic 501 and disk usage 602 is extremely high.
[0066] FIG. 7 is a flowchart showing the prediction process of CPU usage in an embodiment of the present invention.
[0067] Referring to FIG. 7, the processor 220 may call the time-series data of user traffic stored for a relatively long period as the learning data of the prediction model (step 701).
[0068] The processor 220 may generate a time-series prediction model for user traffic by causing the traffic data called in step 701 to be learned by a time-series prediction library (e.g., Prophet, etc.) (step 702).
[0069] The processor 220 may call the time-series data (X) of user traffic stored for a relatively short period compared to the model learning data (step 703).
[0070] The processor 220 may call the time-series data (Y) of CPU usage stored for a relatively short period compared to the model learning data (step 704).
[0071] The processor 220 may use the time-series data of user traffic and CPU usage called in steps 703 and 704 respectively, and use a linear regression model to obtain the correlation between user traffic and CPU usage (Y = aX + b) (step 705). Here, a represents the linear regression coefficient value, and b represents the linear regression intercept value. a and b may be determined by linear regression on the time-series data for a short period. For example, the time-series data stored for three days may be used to determine a and b, and this period can be adjusted as needed.
[0072] In the correlation between user traffic and CPU usage, b may be determined to be a value close to 0. Although it may make sense to set b to 0, since the CPU usage in the idle state is not always 0, b is not set to 0.
[0073] After predicting the user traffic (X’) at a future time point (t) using the time series prediction model constructed in step 702, the processor 220 may calculate the CPU usage (Y’) at the future time point (t) by utilizing the future prediction value (X’) for the user traffic and the correlation obtained in step 705 (step 706).
[0074] When additional equipment is added to or removed from a component, or server tuning is performed, etc., the actual CPU usage may change significantly in a short period. When using the CPU usage as learning data, such changes may increase the prediction error, but it is extremely rare for user traffic to change in a short period. In this embodiment, a time series prediction model is constructed using user traffic with few sudden change factors. Furthermore, by obtaining the correlation between user traffic and the CPU usage of the equipment, a relationship such as this much CPU will be used at this time of such traffic can be derived. By using the prediction model for user traffic, it is possible to predict how much (Δ) the user traffic (X) will change in the future, and from the correlation between user traffic and CPU usage, it can be predicted that if the user traffic (X) changes by Δ, this much (aΔ) CPU will be used. FIG. 8 is a diagram showing an example of predicting CPU usage in one embodiment of the present invention. In FIG. 8, 801 represents the CPU usage (Y) based on the correlation (aX + b) with user traffic, 802 represents the value (Y’) obtained by multiplying the value (X’) predicted by the time series prediction model that learned the user traffic (X) by a and adding b, 803 represents the prediction error (upper / lower) of the time series prediction model, and 804 represents the actually measured CPU usage.
[0075] The learning process (step 702) and the correlation analysis process (step 705) of the prediction model may be executed at regular intervals. At this time, when user traffic (X) is added after the previous prediction time point, in order to reflect the added data in the prediction model, model learning may be periodically performed to update the prediction model. Also, even if the user traffic (X) is the same, if the CPU usage (Y) increases or decreases due to parameter tuning or rollback, etc., in order to reflect such changes in the prediction process (step 706), it is necessary to periodically recalculate the correlation between the user traffic (X) and the CPU usage (Y). For example, in the relationship Y = aX + b between the user traffic (X) and the CPU usage (Y), even if the traffic is the same, due to parameter tuning, the correlation may change from a = 0.1, b = 1 before tuning to a = 0.05, b = 1 after tuning. By periodically updating the correlation between the user traffic (X) and the CPU usage (Y), it is possible to respond to changes in the relationship and more accurately predict the CPU usage at a future time point.
[0076] The change at a future time point of the disk usage, which is another system performance indicator, can also be predicted by the same method as the prediction process of the CPU usage described in FIG. 7. However, the correlation between the user traffic and the disk usage is obtained by calling the time-series data of the user traffic for a short period and using the value integrated over time.
[0077] The correlation between the user traffic (X) and the disk usage (Y) may be defined as disk usage (Y) = a × [integrated value of the user traffic (X) stored during the retention hour × number of copies] + b. In the correlation between the user traffic and the disk usage, a may be determined to be a value close to 1. Depending on the embodiment, it may also make sense to set a to 1 and b to 0.
[0078] By determining the correlation between user traffic and disk usage, it is possible to derive the relationship that this much disk will be used for such traffic. By using a prediction model for user traffic, in the future, it is possible to predict how much (Δ) the user traffic (X) will change, and from the correlation between user traffic and disk usage, it is possible to predict that if the user traffic (X) changes by Δ, this much (aΔ) disk will be used.
[0079] FIG. 9 is a diagram showing an example of a configuration file for predicting resource usage in one embodiment of the present invention. FIG. 9 shows a configuration file for predicting the available resources of a system called Logiss, which is a log integration management platform published by NAVER, as an example of a cluster. The future system usage prediction technology of the present invention as described above can be abstracted like the configuration file of FIG. 9 so that it can be used in other services and additional models can be added.
[0080] For example, referring to FIG. 9, for predicting the available resources of the Logiss system, it may include a query file 901 for calculating user traffic, a query file 902 for calculating CPU usage, a model file 903 for analyzing the correlation between user traffic and CPU usage, a file 904 for setting the query execution target for which the available resources are to be predicted, an integer indicating the section 905 for executing the linear regression of the two query files 901 and 902 specified by the model file 903, an integer indicating the period 906 for comparing the actual CPU usage with the value (Y = aX + b) obtained by linear regression, an integer indicating the minimum / maximum range 907 of the learning data used for traffic prediction, and the like. The configuration file for predicting disk usage is also similar or identical to the configuration file of FIG. 9, and high reusability can be provided by the content of the abstracted configuration file.
[0081] In short, the correlation between user traffic (X) and CPU / disk usage (Y) can be obtained from time-series data stored in a short period, and based on this, the change in CPU / disk usage due to future traffic changes can be calculated. At this time, instead of directly predicting the usage of the CPU and disk, only the change in traffic is predicted, and then the correlation (coefficient and intercept) is used to calculate the change in CPU / disk usage due to the predicted traffic change, thereby reducing the outliers in the learning data.
[0082] If an indicator with few sudden change factors is selected as the input indicator and the correlation with the system performance indicator is obtained, this can be applied to various services. At this time, in order to apply the prediction of resource usage to various services, the settings for queries, models, TSDB targets, etc. may be abstracted.
[0083] FIGS. 10 to 12 are diagrams showing an example of recommending a component configuration based on the prediction of resource usage in an embodiment of the present invention.
[0084] The processor 220 may calculate future system performance indicators for each component constituting the cluster by using the prediction model for the input indicator and the correlation of the input indicator, and then recommend each component configuration based on the future system performance indicators.
[0085] To facilitate the prediction of the resource usage of each component, each component may be quantified as a unit system. For example, a capacity of 4 CPUs, 500G (disk), and 32G (memory) may be defined as one unit system, and the total required resources may be expressed by the number of quantified unit systems.
[0086] For the recommendation of component configuration, it may include content such as recommending the number of equipment to be added or removed when it is predicted that the future CPU / disk usage will exceed or be less than a predetermined threshold range, and according to the correlation between user traffic and CPU / disk usage, a certain traffic (for example, maximum 800 MiB / s + ingestion volume of 10 TiB per day) can withstand a specific level of system resources (for example, with a maximum usage of 50% for CPU and 70% for disk). The recommended number of equipment required to achieve this is also included.
[0087] The processor 220 may link each component constituting the cluster to a system that automates operations such as equipment addition and removal, and provide recommendation information regarding component configuration through the administrator screen.
[0088] Figure 10 shows an availability status administrator screen 1000 that recommends component configuration according to the current availability status of the integrated log management system 300. Referring to Figure 10, when the processor 220 inputs the specification information 1010 of a specific cluster and the target resource usage 1020 into the availability status administrator screen 1000, for each component, it recommends the predicted value 1030 for the resource usage at the future time point 1001 and the required number of equipment units 1040 for each component based on the predicted value 1030 for the target resource usage 1020. Therefore, by inputting equipment specifications such as the number of CPU cores and disk capacity, it is possible to calculate and recommend how many units of equipment should be added to align the resource usage within the target range desired by the operator in preparation for future traffic changes. For example, the equipment operation automation system can appropriately adjust and operate the component configuration by receiving, through an API, the number of equipment units to be added or removed recommended when targeting the system usage rate within the system predicted by this system to be 60% or less and 5% or more within 7 days or 30 days.
[0089] FIG. 11 shows a capacity prediction administrator screen 1100 that recommends a component configuration according to the capacity prediction based on the traffic addition of the integrated log management system 300. Referring to FIG. 11, when the processor 220 inputs the specification information 1110 of a specific cluster, the target resource usage 1120, and the amount of traffic 1150 to be added to the capacity prediction administrator screen 1100, for each component, it recommends the predicted value 1130 at the future time point 1101 for the resource usage due to traffic addition and the required number of equipment units 1140 of each component based on the predicted value 1130 for the target resource usage 1120. Even in an environment where new traffic is added, it is possible to calculate and recommend how many units of equipment should be invested to match the resource usage within the target range desired by the operator in preparation for future traffic changes.
[0090] In the above-described embodiments, the system performance indicators of each component are predicted for the future time points 1001 and 1101, and the component configuration at the corresponding time points is recommended. However, the present invention is not limited to this. Depending on the embodiment, the system performance indicators for future time points are predicted at regular intervals, and a UI that allows the administrator to confirm on the administrator screen the future time points (i.e., the time points expected to exceed or be less than the threshold range) when the predicted values of the system performance indicators exceed the threshold range (e.g., the target resource usage, etc.) is provided. That is, in the process of utilizing the system capacity prediction result, the administrator can be guided to the future time points when the system capacity exceeds the pre-determined threshold range, and support can be provided so that the number of equipment units of each component can be adjusted in advance during the time period when the traffic is low.
[0091] FIG. 12 shows a new cluster construction administrator screen 1200 that recommends a component configuration for constructing a new cluster of the integrated log management system 300. Referring to FIG. 12, when the processor 220 inputs the specification information 1210 of the cluster to be added, the target resource usage 1220, and the scale of the traffic expected to be received (for example, the traffic value received at peak time, the amount of traffic imported in a day, etc.) 1230 into the new cluster construction administrator screen 1200, it may recommend the required number of equipment units 1240 of each component capable of processing the scale of the traffic expected to be received within the target resource usage 1220. Based on the correlation relationship of CPU usage (number of cores used per second) = aХ(traffic value at peak time) + b, a and b are calculated for each component. When the traffic value at peak time received by the corresponding cluster is input, the maximum CPU usage is calculated, and the number of cores used per second is calculated. Therefore, it is possible to calculate and recommend how many units of equipment should be installed to keep the maximum CPU usage rate below the target range. Similarly, based on the correlation relationship of disk usage = aХ(integrated value of traffic × number of replicas) + b, the integrated value of traffic is calculated from the amount of traffic imported by the cluster in a day, and the total disk usage is calculated from this. As a result, it becomes possible to calculate how many units of equipment should be installed to keep the disk usage below the target range.
[0092] Thus, according to the embodiments of the present invention, as an input index having a correlation with the system performance index, an input index with few sudden change elements is used to train a time series prediction model to predict the change of the input index, thereby constructing a prediction model. This can reduce the outliers in the training data and improve the accuracy of the prediction. Also, according to the embodiments of the present invention, instead of directly predicting the system performance index, the change of the input index is predicted using the prediction model, and the system performance index based on the correlation is calculated using the predicted value of the input index, so that the future system performance can be easily predicted only by the input index. Furthermore, according to the embodiments of the present invention, the future system performance index is utilized in conjunction with a system that automates resource operations such as equipment installation and removal, so that the system can be stably maintained and managed, and the operation resources for calculating the number of equipment when configuring a new cluster can be reduced.
[0093] The above-described apparatus may be implemented by hardware components, software components, and / or a combination of hardware components and software components. For example, the apparatus and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an ALU (arithmetic logic unit), a digital signal processor, a microcomputer, an FPGA (field programmable gate array), a PLU (programmable logic unit), a microprocessor, or various devices capable of executing instructions and responding. The processing device may execute an operating system (OS) and one or more software applications running on the OS. Further, the processing device may access data, record, manipulate, process, and generate data in response to the execution of the software. For the sake of convenience of understanding, it may be described as if one processing device is used, but those skilled in the art will understand that the processing device may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing device may include a plurality of processors or one processor and one controller. Also, other processing configurations, such as a parallel processor, are possible.
[0094] The software may include a computer program, code, instructions, or a combination of one or more of these, and may configure the processing device to operate as desired or command the processing device independently or collectively. The software and / or data may be embodied in any type of machine, component, physical device, computer recording medium, or device for interpreting by the processing device or providing instructions or data to the processing device. The software may be distributed over a computer system connected by a network and recorded or executed in a distributed state. The software and data may be recorded on one or more computer-readable recording media.
[0095] The method according to the embodiment may be implemented in the form of program instructions executable by various computer means and recorded on a computer-readable medium. At this time, the medium may be one that continuously records a program executable by a computer, or may be one that temporarily records it for execution or download. Further, the medium may be various recording means or storage means in a form in which a single or a plurality of hardware are combined, and may be not only a medium directly connected to a certain computer system, but also one that is distributed and exists on a network. Examples of the medium include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to record program instructions such as ROMs, RAMs, and flash memories. Further, as examples of other media, it also includes an app store that distributes an application, a recording medium and a storage medium managed by a site, a server, etc. that supply and distribute other various software.
[0096] As described above, the embodiments have been described based on limited embodiments and drawings. However, those skilled in the art will be able to make various modifications and variations from the above description. For example, even if the described technology is executed in an order different from the described method, and / or the components such as the described system, structure, device, circuit, etc. are combined or combined in a form different from the described method, or are opposed or replaced by other components or equivalents, appropriate results can be achieved.
[0097] Therefore, even if they are different embodiments, as long as they are equivalent to the scope of the claims, they belong to the scope of the appended claims.
Description of Reference Numerals
[0098] 300: Integrated Log Management System 310: Real-Time Indexing & Search System 320: Data storage platform
Claims
1. A method for predicting future system usage of a computer device including at least one processor, comprising: constructing a prediction model for the input metric using time series data of the input metric stored during a first fixed period by the at least one processor; obtaining a correlation relationship between the input metric and the system performance metric using time series data of the input metric and the system performance metric stored during a second fixed period by the at least one processor; predicting the input metric at at least one future time point using the prediction model by the at least one processor; and calculating the system performance metric at the future time point using the predicted input metric and the correlation relationship by the at least one processor A method for predicting future system usage, comprising the steps described above.
2. Among the metrics to which the system resources respond, the input metric is selected such that the change over a third fixed period is less than a threshold. The method for predicting future system usage according to claim 1, characterized in that.
3. User traffic corresponding to log data imported from the outside is used as the input metric, CPU usage or disk usage is used as the system performance metric The method for predicting future system usage according to claim 1, characterized in that.
4. The step of obtaining the correlation relationship is obtaining the correlation relationship between the input metric and the system performance metric using a linear regression model The method for predicting future system usage according to claim 1, characterized in that.
5. The step of obtaining the correlation relationship is determining a linear regression coefficient value and a linear regression intercept value by linear regression on the time series data of the second fixed period The method for predicting future system usage according to claim 1, comprising the steps described above.
6. The linear regression intercept value is determined to be a value close to 0 The method for predicting future system usage according to claim 5, characterized in that.
7. The step of constructing the prediction model is updating the prediction model using data stored after the previous prediction time point for the input metric The method for predicting future system usage according to claim 1, comprising the steps described above.
8. The step of obtaining the correlation relationship is Updating the correlation between the input index and the system performance index at a certain period The future system usage prediction method according to claim 1, comprising the above step
9. The step of obtaining the correlation includes In some types according to the type of the system performance index, using the value obtained by integrating the time series data of the input index over time to obtain the correlation between the input index and the system performance index The future system usage prediction method according to claim 1, characterized by the above
10. The future system usage prediction method includes The step of recommending the component configuration at the future time point by using the calculated system performance index by the at least one processor The future system usage prediction method according to claim 1, further comprising the above step
11. The step of calculating the system performance index at the future time point includes Calculating the system performance index at the future time point of each component constituting the cluster The step of recommending the component configuration at the future time point includes Calculating and recommending the number of installed units of the corresponding component based on the system performance index at the future time point for each component The future system usage prediction method according to claim 10, characterized by the above
12. The future system usage prediction method includes When a received scheduled input index is given for constructing a new cluster by the at least one processor, calculating the system performance index of each component constituting the new cluster by using the received scheduled input index and the correlation, and Calculating and recommending the required number of installed units of each component based on the system performance index of each component by the at least one processor The future system usage prediction method according to claim 1, further comprising the above steps
13. The future system usage prediction method includes The step of providing, by the at least one processor, a time point at which the calculated system performance index exceeds the threshold range among the future time points The future system usage prediction method according to claim 1, further comprising the above step
14. A computer program for causing a computer device to execute the future system usage prediction method according to any one of claims 1 to 13
15. At least one processor realized to execute instructions readable by a computer device Including by the at least one processor, construct a prediction model for the input metric using the time-series data of the input metrics stored during a first fixed period, obtain the correlation between the input metric and the system performance metric using the time-series data of the input metric and the system performance metric stored during a second fixed period, predict the input metric at at least one future time point using the prediction model, calculate the system performance metric at the future time point using the predicted input metric and the correlation A computer device characterized by the above.
16. Among the metrics to which the system resources respond, a metric whose change during a third fixed period is less than a threshold is selected as the input metric. The computer device according to claim 15, characterized by the above.
17. by the at least one processor, update the prediction model using the data stored after the previous prediction time point for the input metric, update the correlation between the input metric and the system performance metric at regular intervals The computer device according to claim 15, characterized by the above.
18. by the at least one processor, In some types according to the type of the system performance metric, obtain the correlation between the input metric and the system performance metric using the value obtained by integrating the time-series data of the input metric over time. The computer device according to claim 15, characterized by the above.
19. by the at least one processor, calculate the system performance metric at the future time point for each component constituting the cluster, For each component, calculate and recommend the number of installed units of the corresponding component based on the system performance metric at the future time point. The computer device according to claim 15, characterized by the above.
20. by the at least one processor, Provide a time point at which the calculated system performance metric exceeds the threshold range among the future time points. The computer device according to claim 15, characterized by the above.
Citation Information
Patent Citations
Server performance prediction method and apparatus
CN109117352A
Techniques for modifying a cluster computing environment
JP2023548405A
Method and system for predicting the channel usage
WO2014102318A1
Analysis device, recording medium, and analysis method
WO2015098302A1
Large scale cluster monitoring system, and automatic building and restoration method thereof
KR1020090061522A