Database tuning method using deep genetic algorithm
The database tuning method employs a deep genetic algorithm to optimize performance knobs in LSM-Tree databases, addressing the complexity of knob tuning and improving performance and storage efficiency.
Patent Information
- Application Number
- PCT/KR2023/020236
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-12
AI Technical Summary
Tuning multiple performance knobs in key-value databases based on Log Structured Merge Tree (LSM-Tree) is challenging due to numerous potential configuration combinations and trade-offs, which affects system performance and storage device lifespan.
A database tuning method using a deep genetic algorithm that generates a combined workload by matching performance control parameters, internal metrics, and external metrics, and applies a genetic algorithm to optimize these parameters through deep learning.
This method improves database performance by deriving optimized parameters, reducing tuning time, and enhancing system efficiency while minimizing write amplification and storage wear.
Smart Images

Figure KR2023020236_12062025_PF_FP_ABST
Abstract
Description
Database Tuning Method Using Deep Genetic Algorithms
[0001] The technical field to which the present invention belongs relates to a method for tuning a database based on a log-structured merge tree using a deep genetic algorithm.
[0002] The material described in this paragraph merely provides background information for the present embodiment and does not constitute prior art.
[0003] Key-value databases are useful for handling unstructured data, such as sensor data and social network data. Key-value databases primarily use log-structured merge trees.
[0004] The Log Structured Merge Tree (LSM-Tree) is designed for workloads that perform sequential write operations. The LSM-Tree structure consists of a single in-memory data structure and an append-based data structure for storage across multiple blocks (e.g., disks).
[0005] LSM-Tree efficiently handles frequent insertions and updates in key-value databases. It is a write-friendly structure that stores data in log format first, deferring merging operations such as data sorting and update processing on the log. However, later merging operations cause write amplification, impacting system performance and storage device lifespan.
[0006] LSM-Tree writes data sequentially, not in an arbitrary order. When retrieving data, it's impossible to know where a given piece of data is located within the tree, so it must be searched sequentially from the top level. Even if the data is not on disk, all files at all levels must be read.
[0007] Database systems typically have numerous tools that database administrators must configure to achieve high performance. RocksDB uses a log-structured merge tree to achieve fast data write performance. RocksDB includes numerous performance tuning parameters (knobs) related to writes and space amplification, which are important performance metrics. Tuning these database knobs can yield significant performance improvements. However, tuning multiple knobs simultaneously is a daunting task, with numerous potential configuration combinations and tradeoffs.
[0008] (Patent Document 1) U.S. Patent Publication No. US 2017-0344619 (November 30, 2017)
[0009] (Patent Document 2) Korean Patent Publication No. KR 10-2016-0121819 (October 21, 2016)
[0010] (Patent Document 3) United States Patent Publication No. US 2018-0121121 (May 3, 2018)
[0011] Research related to the present invention was conducted under a national research and development project. The national research and development project identification number is 1711103018, the project number is 2017-0-00477-004, the responsible ministry is the Ministry of Science and ICT, the project management agency is the National IT Industry Promotion Agency (NIPA), and the research project name is the Information and Communications Broadcasting Research and Development Project.
[0012] The main purpose of the embodiments of the present invention is to define a new representation of a workload to find the most similar workload in a database, apply a genetic algorithm to the combined workload, and derive optimized parameters through deep learning.
[0013] Additional, unspecified, purposes of the present invention can be additionally considered within the scope that can be easily inferred from the detailed description and effects thereof below.
[0014] According to one aspect of the present embodiment, a database tuning method is provided, comprising: a step of creating a data repository for a plurality of basic workloads in which a performance control parameter (knob) of a workload, an internal metric representing basic representation information including a basic configuration of the workload, and an external metric representing the performance of the workload are triple-matched; a step of creating statistical representation information using the internal metric for the plurality of basic workloads and a sample-selected target workload; a step of generating a combined workload by synthesizing the plurality of basic workloads by calculating a similarity distance based on the statistical representation information and adjusting a scale based on the similarity distance; and a step of outputting an optimal configuration of the workload from the combined workload through an optimization model.
[0015] The above database tuning method may include a step of inputting a candidate configuration of the workload through an external metric prediction model learned using the combined workload and evaluating the candidate configuration based on a fitness score indicating the performance of the database.
[0016] The composition of the above workload may include an index, a key size, a value size, a number of entries, a read / write ratio, update information, or a combination thereof.
[0017] The internal metrics may include the basic configuration, a database cache miss indicating that the requested data does not exist, or a combination thereof.
[0018] The external metrics may include total execution time, data processing speed, data write amplification factor, space amplification factor, or a combination thereof.
[0019] The above statistical representation information may include mean, variance, first quartile, second quartile, and third quartile.
[0020] To calculate the above similarity distance, the Mahalanobis distance, which indicates how many times the standard deviation the distance from the mean is, can be applied.
[0021] The step of generating the combined workload comprises obtaining a similarity distance matrix by calculating the similarity distance between the statistical expression information for the target workload and the statistical expression information for the plurality of basic workloads, and adjusting the scale based on the maximum similarity distance according to the similarity distance matrix to synthesize the plurality of basic workloads, and the combined workload may be triple-matched with the performance adjustment parameter, the internal metric, and the external metric.
[0022] The above fitness score can be calculated by calculating the proportion of the default and predicted values of the total execution time, the default and predicted values of the data processing speed, the default and predicted values of the data write amplification factor, and the default and predicted values of the space amplification factor, applying weights, and then adding them together.
[0023] The step of outputting the optimal configuration of the workload may include applying a genetic algorithm to the optimization model, setting a gene group such that genes of the genetic algorithm correspond to performance control parameters of the workload, and chromosomes of the genetic algorithm correspond to basic configurations of the workload, determining a rank of the gene group according to the fitness score calculated using the external metric prediction model, classifying the gene group by considering the rank of the gene group, and evolving the classified gene group through multiple generations through a genetic operation that applies crossover and mutation to the performance control parameters, thereby finding the optimal configuration.
[0024] The above external metric prediction model can receive the performance adjustment parameters and the basic configuration as input and produce the fitness score through a structure in which multiple linear layers are connected.
[0025] According to another aspect of the present embodiment, a database tuning device including a processor is provided, wherein the processor creates a data repository for a plurality of basic workloads in which a performance control parameter (knob) of a workload, an internal metric representing basic expression information including a basic configuration regarding the workload, and an external metric representing the performance of the workload are triple-matched, and statistical expression information is created using the internal metric for the plurality of basic workloads and a sampled target workload, and calculates a similarity distance based on the statistical expression information and adjusts a scale based on the similarity distance to create a combined workload that synthesizes the plurality of basic workloads, and outputs an optimal configuration of the workload through an optimization model from the combined workload.
[0026] The processor can input a candidate configuration of the workload using an external metric prediction model learned using the combined workload and evaluate the candidate configuration based on a fitness score indicating the performance of the database.
[0027] The above processor applies a genetic algorithm to the above optimization model, sets a gene group such that genes of the genetic algorithm correspond to performance control parameters of the workload, and chromosomes of the genetic algorithm correspond to basic configurations of the workload, determines a rank of the gene group according to the fitness score calculated using the external metric prediction model, classifies the gene group by considering the rank of the gene group, and evolves the classified gene group through multiple generations through a genetic operation that applies crossover and mutation to the performance control parameters, thereby finding the optimal configuration.
[0028] As described above, according to the embodiments of the present invention, there is an effect of improving database performance by defining a new representation for workload to find the most similar workload in a database, applying a genetic algorithm to the combined workload, and deriving optimized parameters through deep learning.
[0029] Even if the effect is not explicitly mentioned herein, the effect and its provisional effect described in the following specification expected by the technical features of the present invention are treated as described in the specification of the present invention.
[0030] Figure 1 is a diagram illustrating a database based on a log structure merge tree performing a data compaction operation.
[0031] FIG. 2 and FIG. 3 are block diagrams illustrating a database tuning method according to one embodiment of the present invention.
[0032] FIG. 4 and FIG. 5 are diagrams illustrating the configuration of a workload processed by a database tuning method according to one embodiment of the present invention.
[0033] FIG. 6 is a diagram illustrating basic representation information of a workload processed by a database tuning method according to one embodiment of the present invention.
[0034] FIG. 7 is a diagram illustrating statistical representation information of a workload processed by a database tuning method according to one embodiment of the present invention.
[0035] FIG. 8 is a diagram illustrating a combined workload processed by a database tuning method according to one embodiment of the present invention.
[0036] FIG. 9 and FIG. 10 are diagrams illustrating an external metric prediction model processed by a database tuning method according to one embodiment of the present invention.
[0037] FIG. 11 and FIG. 12 are diagrams illustrating an optimization model processed by a database tuning method according to one embodiment of the present invention.
[0038] FIG. 13 is a drawing illustrating a database tuning device according to another embodiment of the present invention.
[0039] Figures 14 to 19 illustrate the results of simulation experiments performed according to embodiments.
[0040] Hereinafter, in describing the present invention, if it is judged that related known functions are obvious to those skilled in the art and may unnecessarily obscure the gist of the present invention, a detailed description thereof will be omitted, and some embodiments of the present invention will be described in detail through exemplary drawings.
[0041] Figure 1 is a diagram illustrating a database based on a log structure merge tree performing a data compaction operation.
[0042] Representative databases that use LSM-Tree include RocksDB.
[0043] When an insertion operation is performed, LSM-Tree first stores data in a memory area. Once the data reaches a certain memory capacity, the memory contents are flushed to disk. The flushed data is merge-sorted with the existing data stored on disk and written to disk. When each level of the disk area exceeds a threshold, a merge-sort is performed to create a lower level.
[0044] LSM-Tree-based databases store data in key-value format. When a data insertion request comes into an LSM-Tree-based database, a log is first written to a log file before the data is written to memory. After the log is written, the data is stored in a memtable in the memory area. When write requests continue and the memtable reaches a certain capacity, the memtable is converted to an immutable memtable (read-only memtable). When the immutable memtable becomes full, it is flushed to the block (disk) area.
[0045] When a flush operation is performed, the memtable files are sorted by key order and converted into SST (Storted String Table) files. SST files contain multiple blocks. Examples of blocks include a data block that stores data, an index block that indexes the location of the data block, and a footer block that processes the location of the index block.
[0046] SST files are updated through disk compaction. Once created, an SST file may not disappear. Lower-level SST files may contain older data than higher-level SST files.
[0047] If a system failure or power outage occurs during a transaction, any data remaining in the buffer and not yet reflected on disk will be lost. When the database recovers after a system reboot, it uses a log that records the update operations performed by the transaction. One such logging method is the Write-Ahead-Logging (WAL) rule. WAL records relevant logs in a log file before data changes made by a transaction are written to disk.
[0048] An LSM-Tree-based database performs two commands: a flush command that moves data from memory to disk, and a compaction command that adjusts the levels on disk.
[0049] When a flush command is executed, the immutable memtable is converted to a single SST file. When a large amount of data is input at once, the flush rate is intentionally throttled to balance the flush rate and maintain the SST file's level capacity limit. This intentional delay is called a "write stall."
[0050] Each level of an LSM-Tree has distinct characteristics. Compaction costs and disk writes are concentrated at specific levels. The hierarchical storage structure causes data to accumulate progressively from higher levels to lower levels. The lifetime of an SST file represents the number of times compaction is performed during its existence at that level. If an SST file is not deleted at a specific level during compaction, its lifetime is considered long.
[0051] Examining the results of each level of compaction command execution shows that upper level SST files are created and deleted during the short time it takes to perform data compaction.
[0052] This embodiment defines a new representation for workloads to quickly and accurately find the most similar workload in a log-structured merge tree-based database, applies a genetic algorithm to the combined workloads, and derives an optimized function through deep learning to improve database performance.
[0053] FIG. 2 and FIG. 3 are block diagrams illustrating a database tuning method according to one embodiment of the present invention.
[0054] The database tuning method can be performed by a database tuning device.
[0055] In step S10, a step may be performed to create a data repository for multiple basic workloads in which a performance control parameter (knob) of the workload, an internal metric representing basic representation information including a basic configuration regarding the workload, and an external metric representing the performance of the workload are triple-matched.
[0056] In step S20, a step of generating statistical representation information using internal metrics for multiple base workloads and sample-selected target workloads can be performed.
[0057] Step S30 may be performed to generate a combined workload by synthesizing multiple base workloads by calculating a similarity distance based on statistical representation information and adjusting the scale based on the similarity distance. The combined workload may be considered the target workload.
[0058] In step S40, a step of outputting an optimal configuration of the workload through an optimization model from the combined workload can be performed.
[0059] In step S50, a step may be performed to input a candidate configuration of a workload using an external metric prediction model learned using a combined workload and evaluate the candidate configuration based on a fitness score indicating the performance of the database.
[0060] Figure 4 illustrates the configuration of a basic workload, and Figure 5 illustrates the configuration of a target workload.
[0061] The basic configuration can include workload index, key size, value size, number of entries, read / write ratio, update information, or a combination of these.
[0062] The base workload is a workload whose configuration is randomly generated, and the target workload is a workload selected based on speed.
[0063] Figures 6 and 7 illustrate internal metrics of the workload, with Figure 6 illustrating basic representation information and Figure 7 illustrating statistical representation information.
[0064] Internal metrics may include default configuration, database cache misses indicating that requested data does not exist (e.g., rocksdb.block.cache.miss), or a combination of these.
[0065] Statistical representation information can include the mean, variance, first quartile, second quartile, and third quartile. Applying statistical representation information can simplify the workload representation and reduce the dimensionality of the search space.
[0066] FIG. 8 is a diagram illustrating a combined workload processed by a database tuning method according to one embodiment of the present invention.
[0067] A combined workload can have triple matching of performance tuning parameters, internal metrics, and external metrics.
[0068] Generating a combined workload involves calculating the similarity distances between the statistical representation information for the target workload and the statistical representation information for multiple base workloads, obtaining a similarity distance matrix, and then synthesizing the multiple base workloads by scaling based on the maximum similarity distance according to the similarity distance matrix. Calculating the similarity distance can be achieved using the Mahalanobis distance, which represents the standard deviation times the distance from the mean.
[0069] FIG. 9 and FIG. 10 are diagrams illustrating an external metric prediction model processed by a database tuning method according to one embodiment of the present invention.
[0070] External metrics may include total execution time (TIME), data throughput rate (RATE), write amplification factor (WAF), space amplification (SA), or a combination thereof.
[0071] TIME represents the time interval from the start to the end of database tuning (e.g. RocksDB tuning).
[0072] RATE represents the number of operations RocksDB processes per second, such as compactions, compressions, reads, and writes.
[0073] WAF is a performance metric that represents the ratio of the amount of data written to storage (physical data size) to the amount of data written to the database (logical data size). WAF represents the amount of additional data written while using RocksDB.
[0074]
[0075] SA is measured by the size of the data actually recorded in the LSM-Tree. Since the LSM-Tree contains both valid and invalid SST files, the size of the data recorded in the LSM-Tree can be used as an effective indicator of the size of the actual physical space utilized by RocksDB.
[0076] The external metric prediction model can receive performance adjustment parameters and basic configuration as input and produce the fitness score through a structure in which multiple linear layers are connected.
[0077] The fitness score produced by the external metric prediction model can be calculated by adding together the default and predicted values of the total execution time, the default and predicted values of the data processing speed, the default and predicted values of the data write amplification factor, and the default and predicted values of the spatial amplification factor after applying weights to them.
[0078]
[0079] FIG. 11 and FIG. 12 are diagrams illustrating an optimization model processed by a database tuning method according to one embodiment of the present invention.
[0080] Genetic algorithms can be applied to optimization models to output the optimal configuration of workloads.
[0081] The genes of the genetic algorithm correspond to the performance control parameters of the workload, and the chromosomes of the genetic algorithm can set up a group of genes to correspond to the basic configuration of the workload.
[0082] The ranking of gene groups can be determined based on the fitness score calculated using an external metric prediction model, and the gene groups can be classified by considering the ranking of the gene groups.
[0083] The optimal configuration can be found by evolving a group of genes classified through genetic operations that apply crossover and mutation to performance control parameters over several generations.
[0084] According to another embodiment of the present invention, the external metric prediction model of the present invention can calculate a fitness score using total execution time, data processing speed, data write amplification factor, or space amplification factor. The fitness score may be a value calculated in advance by adding the relative ratio values of TIME, RATE, WAF, and SA values according to weights, but may also be a value adaptively determined according to the usage environment of the database. For example, the processor can adaptively determine the weights in calculating the fitness score according to the update speed of the LSM-Tree-based database and / or the structure of the hierarchical layer. In addition, when the file size stored in the database is large, an intermediate layer may be added, and the processor can calculate the fitness score by applying different combinations of weights when there are changes in the file size, update speed, and hierarchical connection structure according to a user request, or when there is a delay in the update speed due to computer performance reasons.
[0085] FIG. 13 is a drawing illustrating a database tuning device according to another embodiment of the present invention.
[0086] The database tuning device (110) includes at least one processor (120), a computer-readable storage medium (130), and a communication bus (170).
[0087] The processor (120) may control the operation of the database tuning device (110). For example, the processor (120) may execute one or more programs stored in a computer-readable storage medium (130). The one or more programs may include one or more computer-executable instructions, and the computer-executable instructions, when executed by the processor (120), may be configured to cause the database tuning device (110) to perform operations according to the exemplary embodiment.
[0088] A computer-readable storage medium (130) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (140) stored in the computer-readable storage medium (130) includes a set of instructions executable by the processor (120). In one embodiment, the computer-readable storage medium (130) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that can be accessed by the database tuning apparatus (110) and capable of storing desired information, or a suitable combination thereof.
[0089] A communication bus (170) interconnects various other components of the database tuning device (110), including the processor (120) and the computer-readable storage medium (140).
[0090] The database tuning device (110) may also include one or more input / output interfaces (150) and one or more communication interfaces (160) that provide interfaces for one or more input / output devices. The input / output interfaces (150) and the communication interfaces (160) are connected to a communication bus (170). The input / output devices (not shown) may be connected to other components of the database tuning device (110) via the input / output interfaces (150).
[0091] The processor of the database tuning device can create a data repository for multiple basic workloads that are triple-matched with performance adjustment parameters (knobs) of the workload, internal metrics representing basic representation information including basic configuration information about the workload, and external metrics representing the performance of the workload.
[0092] The processor of the database tuning device can generate statistical representation information using internal metrics for multiple base workloads and sampled target workloads.
[0093] The processor of the database tuning device can generate a combined workload that synthesizes multiple primary workloads by calculating a similarity distance based on statistical representation information and adjusting a scale based on the similarity distance.
[0094] A database tuning device is provided, characterized in that the processor of the database tuning device outputs an optimal configuration of the workload through an optimization model from the combined workload.
[0095] The processor of the database tuning device can input candidate configurations of the workload using an external metric prediction model learned using the combined workload and evaluate the candidate configurations based on a fitness score representing the performance of the database.
[0096] The processor of the database tuning device applies a genetic algorithm to the optimization model, and the genes of the genetic algorithm correspond to the performance control parameters of the workload, and the chromosomes of the genetic algorithm set up a gene group corresponding to the basic configuration of the workload, and determines the rank of the gene group according to the fitness score calculated using an external metric prediction model, classifies the gene group by considering the rank of the gene group, and evolves the classified gene group through multiple generations through a genetic operation that applies crossover and mutation to the performance control parameters to find the optimal configuration.
[0097] Figures 14 to 19 illustrate the results of simulation experiments performed according to embodiments. Figure 14 shows overall performance, Figure 15 shows performance depending on whether workloads are combined and whether internal metrics are reduced, Figure 16 shows performance depending on the number of knobs, Figure 17 shows performance depending on the weight of the fitness score, Figure 18 shows performance depending on cosine similarity and Mahalanobis distance, and Figure 19 shows performance depending on Bayesian optimization.
[0098] The evaluation score can be assigned by calculating the total execution time (TIME), data processing speed (RATE), data write amplification factor (WAF), and space amplification factor (SA) using the default value (D) and the actual value (A).
[0099]
[0100] According to this embodiment, a new statistical representation for workloads is defined to find the most similar workload in the database, the Mahalanobis distance is applied to determine the similarity distance, a genetic algorithm is applied to the combined workload, and an optimized function is derived through deep learning, thereby reducing tuning time and improving database performance.
[0101] Components within a database can be separated and connected, and multiple components can be interconnected and implemented as at least one module. Components are connected to communication paths that connect software or hardware modules within the device and operate organically with each other. These components communicate using one or more communication buses or signal lines.
[0102] The database may be implemented within logic circuits using hardware, firmware, software, or a combination thereof, or may be implemented using a general-purpose or special-purpose computer. The device may be implemented using hardwired devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. Furthermore, the device may be implemented as a system-on-chip (SoC) including one or more processors and controllers.
[0103] A database can be installed on a computing device equipped with hardware elements, either in the form of software, hardware, or a combination thereof. A computing device can encompass a variety of devices, including, in whole or in part, communication devices such as communication modems for communicating with various devices or wired or wireless communication networks, memory for storing data for executing programs, and a microprocessor for executing programs, performing calculations, and executing commands.
[0104] Although FIG. 2 describes each process as being executed sequentially, this is merely an example, and those skilled in the art may modify and apply various modifications and variations, such as changing the order described in FIG. 5, executing one or more processes in parallel, or adding other processes, without departing from the essential characteristics of the embodiment of the present invention.
[0105] The operations according to the present embodiments may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium refers to any medium that participates in providing commands to a processor for execution. The computer-readable medium may include program commands, data files, data structures, or a combination thereof. For example, the computer program may be a magnetic medium, an optical recording medium, a memory, etc. The computer program may be distributed on a network-connected computer system, and the computer-readable code may be stored and executed in a distributed manner. Functional programs, codes, and code segments for implementing the present embodiments may be easily inferred by programmers in the technical field to which the present embodiments belong.
[0106] These examples are intended to illustrate the technical concepts of this embodiment, and the scope of the technical concepts of this embodiment is not limited by these examples. The scope of protection of this embodiment should be interpreted by the claims below, and all technical concepts within the scope equivalent thereto should be interpreted as being included within the scope of the rights of this embodiment.
Claims
1. In the database tuning method, A step of creating a data repository for a plurality of basic workloads that are matched with a performance control parameter (knob) of the workload, an internal metric representing basic expression information including a basic configuration regarding the workload, and an external metric representing the performance of the workload; A step of generating statistical expression information using the internal metrics for the above-mentioned plurality of basic workloads and sample-selected target workloads; A step of generating a combined workload by calculating a similar distance based on the statistical expression information and adjusting a scale based on the similar distance to synthesize the plurality of basic workloads; and A database tuning method comprising the step of outputting an optimal configuration of the workload through an optimization model from the combined workload.
2. In the first paragraph, the step of outputting the optimal configuration of the workload is: A step of applying a genetic algorithm to the above optimization model, A database tuning method, characterized by including a step of receiving a candidate configuration of the workload through an external metric prediction model learned using the combined workload and evaluating the candidate configuration based on a fitness score representing the performance of the database.
3. In paragraph 2, A database tuning method, characterized in that the basic configuration of the above workload includes a workload index, a key size, a value size, a number of entries, a read / write ratio, update information, or a combination thereof.
4. In paragraph 2, A method of tuning a database, wherein the internal metrics include a database cache miss indicating that the base configuration, requested data does not exist, or a combination thereof.
5. In paragraph 2, A database tuning method, wherein the external metrics include total execution time, data processing speed, data write amplification factor, space amplification factor, or a combination thereof.
6. In paragraph 2, A database tuning method, characterized in that the above statistical expression information includes mean, variance, first quartile, second quartile, and third quartile.
7. In paragraph 2, A database tuning method characterized in that calculating the above similarity distance applies the Mahalanobis distance, which indicates how many times the standard deviation the distance from the mean is.
8. In paragraph 2, The steps for generating the above combined workload are: By calculating the similarity distance between the statistical expression information for the target workload and the statistical expression information for the plurality of basic workloads, a similarity distance matrix is obtained, and the plurality of basic workloads are synthesized by adjusting the scale based on the maximum similarity distance according to the similarity distance matrix. A database tuning method, wherein the combined workload is triple-matched with the performance adjustment parameter, the internal metric, and the external metric.
9. In paragraph 5, A database tuning method, characterized in that the fitness score is calculated by applying weights to the default and predicted values of the total execution time, the default and predicted values of the data processing speed, the default and predicted values of the data write amplification factor, and the default and predicted values of the space amplification factor and then adding them together.
10. In paragraph 2, The step of outputting the optimal configuration of the above workload is: Applying a genetic algorithm to the above optimization model, The genes of the above genetic algorithm correspond to the performance control parameters of the workload, and the chromosomes of the above genetic algorithm set a group of genes to correspond to the basic configuration of the workload. The ranking of the gene group is determined based on the fitness score calculated using the external metric prediction model. Classify the above gene groups by considering the ranking of the above gene groups, A database tuning method characterized by finding the optimal configuration by evolving the classified gene group through multiple generations through genetic operations that apply crossover and mutation to the performance adjustment parameters.
11. In paragraph 10, The above external metric prediction model is, A database tuning method characterized by receiving the above performance adjustment parameters and the basic configuration and calculating the fitness score through a structure in which a plurality of linear layers are connected.
12. In a database tuning device including a processor, The above processor, Create a data repository for multiple basic workloads that are matched with performance control parameters (knobs) of the workload, internal metrics representing basic representation information including basic configurations of the workload, and external metrics representing the performance of the workload. Generate statistical representation information using the internal metrics for the above multiple basic workloads and sample-selected target workloads, Based on the above statistical expression information, a similar distance is calculated, and a scale based on the similar distance is adjusted to generate a combined workload that synthesizes the above multiple basic workloads. A database tuning device characterized in that it outputs an optimal configuration of the workload through an optimization model from the above combined workload.
13. In paragraph 12, The above processor, Applying a genetic algorithm to the above optimization model, A database tuning device characterized in that it receives a candidate configuration of the workload through an external metric prediction model learned using the above combined workload and evaluates the candidate configuration based on a fitness score representing the performance of the database.
14. In paragraph 13, The above processor, Applying a genetic algorithm to the above optimization model, The genes of the above genetic algorithm correspond to the performance control parameters of the workload, and the chromosomes of the above genetic algorithm set a group of genes to correspond to the basic configuration of the workload. The ranking of the gene group is determined based on the fitness score calculated using the external metric prediction model. Classify the above gene groups by considering the ranking of the above gene groups, A database tuning device characterized in that it finds the optimal configuration by evolving the classified gene group through several generations through genetic operations that apply crossover and mutation to the performance adjustment parameters.
Citation Information
Patent Citations
Database tuning device, database tuning method, and program
JP4810113B2
Database management method, database management device, and database management program
JP5304950B2
System and method for managing database
KR101628097B1
Image processing apparatus, learning method of feature extractor, updating method of identifier, and image processing method
KR1020240085178A
Database performance tuning framework
US20180322154A1