Key-Value Storage Layer Merging Optimization Method and Device Based on Rime Application Characteristics

By introducing buffers and hot and cold record tables into the key-value storage system of Rime input method, the search and write process is optimized, and the appropriate SST file is selected for merging when merging between layers, the problems of low write efficiency and write amplification are solved, and the system performance is improved.

CN119960703BActive Publication Date: 2025-06-13HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510443647.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the Rime input method, the point query operation before writing key-value pairs is inefficient, and the write amplification problem caused by inter-layer merging is serious, which affects performance.

Method used

By introducing key-value pair buffers and SST file hot and cold record tables in the key-value storage system, the data search and writing process is optimized, the number of disk queries is reduced, and the appropriate SST file is selected for merging during the inter-layer merging process, reducing write amplification.

Benefits of technology

Improve write efficiency, reduce write amplification during inter-layer merging, reduce system delay, and improve overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960703B_ABST
    Figure CN119960703B_ABST
Patent Text Reader

Abstract

A method and device for optimizing the merging between key-value storage layers based on Rime application characteristics, which relate to the field of computer storage, including: for a lookup operation, sequentially search for data in the user database with the pinyin input by the user as the key prefix until the target data is found; if the target data is read, update the key-value pair buffer and record the read times of the SST file; when the user selects a Chinese character / word, sequentially query the key-value pair buffer and the user database through the complete key composed of pinyin and the Chinese character / word until the target data is found; for a write operation, the data is first stored in a writable memory table, and when it reaches the threshold, it is converted into a read-only memory table and written to the disk in the form of an SST file. After writing the SST file, check the layer size, and when the limit is exceeded, trigger an inter-layer merging operation. During the merging process, split the SST file according to the key prefix, and when a new SST file is written, check the SST file cold-hot record table, and preferentially merge the hot-range data to ensure data orderliness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer storage, and particularly to a method and device for optimizing the merging between key-value storage layers based on Rime application characteristics. Background Art

[0002] With the advent of the big data era and the further development of the digital economy, the integration of the Internet and the real economy has promoted the explosive growth of data. Data sources are diverse and in various forms, including structured, semi-structured, and unstructured data. Especially with the rapid development of technologies such as the Internet of Things and cloud computing, the growth of unstructured data has been significant, making it increasingly difficult for traditional relational databases (RDBMS) to meet the growing storage and processing requirements. To address this challenge, non-relational databases (NoSQL) have emerged. With advantages such as a flexible storage model, high availability, and easy horizontal scalability, NoSQL has become one of the mainstream technologies for big data storage and fast access. In NoSQL systems, the log-structured merge (LSM) tree, as an efficient storage engine, is adopted by popular key-value storage systems such as LevelDB, RocksDB, and Cassandra, and these key-value storage systems are widely used in various workloads.

[0003] Rime input method uses LevelDB as the storage engine for its user data. When a user inputs pinyin, Rime first uses the input pinyin as the key prefix to search for matching key-value pairs in the user database of LevelDB. The system will continuously read up to 10 key-value pairs with the same prefix, parse out the Chinese characters in these keys, as well as information such as the corresponding selection times and timestamps. By calculating the selection times and timestamps, Rime sorts the Chinese characters and presents the sorted results to the user. If there are no matching key-value pairs in the user database, Rime will query the system database and return the corresponding Chinese characters. When the user selects a certain Chinese character, Rime takes the pinyin and the selected Chinese character as a new key, first searches in the user database to see if there is already data corresponding to this key. If found, Rime updates information such as the selection times and timestamps based on the existing data; otherwise, it initializes this information. Finally, Rime writes the new key-value pair into LevelDB.

[0004] In LevelDB, the written key-value pairs are first stored in the writable memory table. To avoid write stalls, when the writable memory table reaches its capacity limit, a new writable memory table is created to receive subsequent data, and the old writable memory table is converted into a read-only memory table. After being sorted by key values, it is persistently stored on disk in the form of an SST file. LevelDB uses a hierarchical structure to manage SST files. Whenever a new SST file is written to a level, the system checks whether the capacity of the current level exceeds the limit. If it does, LevelDB will adopt an inter-level merge strategy, read an SST file in this level and the SST files in the next level that overlap with its range into memory, perform a merge operation, and then write the new SST file to the next level. LevelDB can effectively handle the writing and querying of user data, but in practical applications, there are some problems. For example, for a point query operation before writing a key-value pair, the queried data may have been loaded into memory during the previous lookup based on the same key prefix, so there is no need to perform a disk lookup again. And in the input method application, there is a read skew, where a small amount of data is frequently read and written, and it needs to be effectively organized. At the same time, frequent write operations will trigger a large number of inter-level merges, resulting in a serious write amplification problem. Assume that the capacity of each level is 10 times that of the previous level. When an SST file in level i needs to be merged into level i + 1, in the worst case, 10 files in level i + 1 need to be read for merge sorting, and then these data are written back to level i + 1, resulting in a write amplification factor of 10 times. In this case, when moving an SST file from level L0 to level L6, the write amplification may be as high as 50 times, thus affecting the performance of the Rime input method. Summary of the Invention

[0005] The main objective of the present invention is to overcome the above-mentioned defects in the prior art, and propose a method and device for optimizing inter-level merging of key-value storage based on Rime application characteristics, which can reduce the point query operation before writing data, thereby improving the writing efficiency. At the same time, it can reduce the write amplification generated during the inter-level merging process, and consider the strategy of preferentially merging hot data to ensure that frequently accessed data can be processed preferentially, thereby reducing system latency and improving overall performance.

[0006] The present invention adopts the following technical solutions:

[0007] On the one hand, a method for optimizing inter-level merging of key-value storage based on Rime application characteristics, which is used for a key-value storage system, includes:

[0008] Data search step: Use the pinyin input by the user using the Rime input method as the key prefix to search in the memory table. If the target data is not found in the memory table, search each level in the disk in turn. After finding and reading the target data, temporarily store the read target data in the key-value pair buffer, and update the "read count" field in the SST file hot and cold record table to mark that the SST file has been read. According to the Chinese character / word selected by the user using the Rime input method, form a complete key with the pinyin and the Chinese character / word, and then search in turn whether there is a key-value pair with the same key in the key-value pair buffer and the user dictionary. If found, update the information according to the existing data. Otherwise, initialize the relevant information to obtain a new key-value pair and write the new key-value pair into the user dictionary.

[0009] Data writing step: For the key-value pairs written using the Rime input method, first save them in the writable memory table. When the size of the writable memory table reaches the first threshold, convert the writable memory table into a read-only memory table and write its content to the disk in the form of an SST file. When writing to the disk, check whether the size of any level exceeds the second threshold limit. If so, trigger an inter-level merge operation. For the level that exceeds the second threshold, select all SST files in that level for merging and write the merged SST file to the next level. During the merge process, if the size of the written SST file exceeds the preset file size, do not close the file immediately, but read the next key-value pair and determine whether its key prefix is the same as the key prefix of the current key-value pair. If the same, continue writing the data to the current SST file. If the size of the written SST file does not exceed the preset file size, determine whether the remaining data volume is less than the SST file size. If less, write all the remaining key-value pairs to the current SST file. If not less, close the current SST file and generate a new SST file for subsequent data writing. When the merged new SST file is written to the next level, determine whether there is a hot SST file in the next level. If there is, and its key range overlaps with the range of the newly written SST file, merge the hot SST file with the new SST file, calculate the read count of the merged file, and record the read count in the SST file hot and cold record table. For the old SST files participating in the merge, mark them for deletion in the corresponding SST file hot and cold record table.

[0010] Preferably, the SST files in each level are unordered. When searching, construct a table iterator for all SST files in each level until the target data is found.

[0011] Preferably, the updated information includes the selection count and the timestamp.

[0012] Preferably, the hot SST file is an SST file whose read count is greater than the average read count of the same level within a certain time range.

[0013] Preferably, after merging the hot SST file and the new SST file, it further includes: writing the merged file to the same level as the hot SST file.

[0014] Preferably, the number of times of reading the merged file is equal to the average value of the number of times of reading the hot SST files participating in the merge.

[0015] Preferably, the SST file cold and hot record table consists of four fields: level number, file number, number of reads, and whether deleted; among them, the level number refers to the numbers 1-6 of the L1-L6 levels; the file number refers to the serial number of the SST file; the number of reads refers to the frequency of reading the SST file, which is used to evaluate the heat of the file; whether deleted refers to whether the SST file has been deleted. When the size of the SST file cold and hot record table reaches the third threshold, the records with the "whether deleted" field being true in the table will be cleared.

[0016] On the other hand, a key-value storage inter-layer merge optimization device based on Rime application characteristics, the device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the device executes the key-value storage inter-layer merge optimization method.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] (1) The present invention can reduce the number of disk queries and improve the writing efficiency: when querying data according to the key prefix, all the data queried from the user database is temporarily stored in the key-value pair buffer. In the point query operation before writing the key-value pair, query the key-value pair buffer first, which can effectively reduce the number of disk queries and speed up the writing efficiency;

[0019] (2) The present invention can reduce write amplification and improve writing performance: during the inter-layer merge process, when the Li of the i-th layer reaches the capacity limit, select to merge all the SST files of the i-th layer and write them to the L(i + 1) of the i + 1-th layer, avoiding duplicate writing of data in the i + 1-th layer. When the i-th layer does not reach the capacity limit, only merge the data whose range overlaps with the hot SST file, which can effectively reduce write amplification and improve write performance;

[0020] (3) The present invention can improve data reading efficiency and reduce system latency: divide the SST file range according to the key prefix, reduce the probability of range overlap between SST files, reduce the number of SST files queried when reading data, and at the same time, preferentially merge hot SST files, maintain the order of hot data, improve the query term hit rate, and reduce system latency. Brief Description of the Drawings

[0021] Figure 1System structure of the key-value storage layer merging optimization method based on Rime application characteristics according to an embodiment of the present invention

[0022] Figure 2 Data reading flowchart of the key-value storage layer merging optimization method based on Rime application characteristics according to an embodiment of the present invention;

[0023] Figure 3 Data writing flowchart of the key-value storage layer merging optimization method based on Rime application characteristics according to an embodiment of the present invention. Detailed implementation manners

[0024] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0025] The key-value storage layer merging optimization method based on Rime application characteristics proposed by the present invention includes the following steps.

[0026] For the lookup operation, the key-value storage system uses the pinyin input by the user through Rime as the key prefix and first looks it up in the in-memory table. If the target data is not found in the in-memory table, it then sequentially looks up each level on disk. Since in this optimization method, the SST files in each level are unordered, when looking up, table iterators need to be constructed for all SST files in each level until the target data is found. After finding and reading the target data, the read data is temporarily stored in the key-value pair buffer, and the "read count" field in the SST file cold-hot record table is updated to mark that the SST file has been read. When the user selects a Chinese character / word, the system first forms a complete key based on the pinyin and the Chinese character / word, and then sequentially checks whether there is a key-value pair with the same key in the key-value pair buffer and the user dictionary. If found, the selected count, timestamp, and other information are updated based on the existing data; otherwise, the relevant information is initialized to obtain a new key-value pair. Finally, the system writes the new key-value pair into the user dictionary. For the key-value pairs written through Rime, they are first saved in the writable in-memory table. When the size of the writable in-memory table reaches the first threshold, the writable in-memory table is converted into a read-only in-memory table, and its content is written to disk in the form of an SST file. When writing to disk, it is necessary to check whether the size of any level exceeds its limit. If so, an inter-level merge operation is triggered. For the level that exceeds the second threshold, all SST files in that level are selected for merging, and the merged SST file is written to the next level. During the merging process, when the size of the written SST file exceeds the preset file size, the file is not immediately closed. Instead, the next key-value pair is read, and it is determined whether its key prefix is the same as that of the current key-value pair. If the same, the data is continued to be written to the current SST file; otherwise, it is determined whether the remaining data volume is less than the SST file size. If less, all the remaining key-value pairs are written to the current SST file; otherwise, the current file is closed, and a new SST file is generated for subsequent data writing. When the newly merged SST file is written to the next level, it is also necessary to check whether there is a hot SST file in the next level. A hot SST file refers to an SST file whose access count within a certain time range is greater than the average access count of the same level. If it exists and its key range overlaps with the range of the newly written SST file, the hot SST file is merged with the new SST file, and the merged file is written to the same level as the hot SST file to ensure the orderliness of hot data. In addition, for the SST file generated after merging, the key-value storage system calculates the read count based on its original key range, that is, the read count of the merged SST file is equal to the average of the read counts of the SST files whose ranges overlap with it during the merge. The read count is recorded in the SST file cold-hot record table, and for the old SST files participating in the merge, they are marked for deletion in the corresponding SST file cold-hot record table.

[0027] Such as Figure 1The figure shows the system structure diagram of the key-value storage inter-layer merging optimization method based on the Rime application characteristics.

[0028] Figure 1 The key-value storage inter-layer merging optimization method based on the Rime application characteristics includes a series of memory structures, such as writable memory tables, read-only memory tables, key-value pair buffers, and SST file hot and cold record tables, as well as the LSM tree structure in the disk. As data is continuously written, it will trigger the inter-layer merging operation based on key prefix splitting. When writing new files to the next level during the merging process, it may also trigger the inter-layer merging operation based on hot data priority to maintain the order of hot data. Among them, the size of the SST files in layer L0 is the same as the size of the memory table, and the file sizes of layers L1 to L6 are greater than or equal to the second threshold. The purpose is to make the number of key prefixes included in the key range of the SST files after splitting according to the key prefix as small as possible, thereby reducing the probability of range overlap between SST files and reducing the query cost.

[0029] Specifically, the key-value pair buffer is located in the memory and is used to temporarily store all the key-value pairs currently queried in the user database (the user database can be, for example, leveldb, and the memory table and disk LSM tree are components in leveldb) with the user input pinyin as the key prefix. The SST file hot and cold record table consists of four fields: layer number, file number, read count, and whether deleted. Among them, the layer number refers to the numbers 1 - 6 of layers L1 - L6, the file number refers to the serial number of the SST file, the read count refers to the frequency of the SST file being read, which is used to evaluate the heat of the file. When file merging occurs, the read count of the new SST file generated after merging is determined by the average read count of the hot SST files participating in the merging. Whether deleted is used to indicate whether the SST file has been deleted. When the size of the SST file hot and cold record table reaches the third threshold, the key-value storage system clears the records with the "whether deleted" field being true in the table.

[0030] Such as Figure 2 As shown, it is the data reading flow chart of the key-value storage inter-layer merging optimization method based on the Rime application characteristics, including the following steps.

[0031] 101. The Rime input method sends a read request to the key-value storage system and proceeds to 102.

[0032] 102. Determine whether the key to be searched is a complete key containing pinyin and Chinese characters. If so, proceed to 103; otherwise, proceed to 106.

[0033] 103. Search for the target data in the key-value pair buffer. If found, proceed to 104; otherwise, proceed to 105.

[0034] 104. The key-value storage system completes the read request and returns the target data to the input method, or returns that the target data is not found.

[0035] 105. Search for the target data level by level in the in-memory table and the disk LSM tree. After the search ends, go to 104.

[0036] 106. Search for the target data in the in-memory table. If found, go to 107; otherwise, go to 108.

[0037] 107. The key-value storage system completes the read request, returns the target data to the input method, and temporarily stores the data in the key-value pair buffer, or returns that the target data is not found.

[0038] 108. Search for the target data layer by layer in the disk LSM tree. After the search ends, go to 107.

[0039] As Figure 3 shown is the data writing flowchart of the key-value storage layer merging optimization method based on the Rime application characteristics, including the following steps.

[0040] 201. The Rime input method sends a write request to the key-value storage system and goes to 202.

[0041] 202. Write the key-value pair into the writable in-memory table, and determine whether the size of the writable in-memory table reaches the limit. If so, go to 203; otherwise, go to 211.

[0042] 203. Create a new writable in-memory table for data writing, convert the old writable in-memory table into a read-only in-memory table, organize it into an SST file, and write it into the disk L0 layer. Determine whether the number of files in the L0 layer reaches the limit. If so, go to 204; otherwise, go to 211.

[0043] 204. Merge all SST files in this layer to obtain a new SST file, and write the new SST file to the next level and go to 205.

[0044] 205. Determine whether the merge is completed. If so, go to 206; otherwise, go to 212.

[0045] 206. Calculate the popularity of the current new SST file, add the information of the new SST file with non-zero popularity to the SST file cold and hot record table, and mark the old SST files participating in the merge as yes in the whether to delete field and go to 207.

[0046] 207. Determine whether the capacity of the Li (i is 1-6) layer reaches the limit. If so, go to 204; otherwise, go to 208.

[0047] 208. Determine whether there is a hot SST file in the Li layer that overlaps with the range of the new SST file. If so, go to 209; otherwise, go to 211.

[0048] At 209, the hot SST files with overlapping ranges are merged with the newly written SST file into a new file, and the merged new SST file is written back to the Li layer, then to 210.

[0049] At 210, calculate the read count of the merged new SST file, and add the information of this file (layer number, file number, read count, and whether it is deleted) to the SST file cold and hot record table. At the same time, mark the old files participating in the merge as "yes" in the whether deleted field, then to 211.

[0050] At 211, the key-value storage system completes the input method write request and returns a write success message.

[0051] At 212, write data to the current SST file, and determine whether the size of the current SST file has reached the limit. If so, go to 213; otherwise, go to 205.

[0052] At 213, read the next key-value pair, and determine whether its key prefix is the same as the prefix of the current key. If so, go to 205; otherwise, go to 214.

[0053] At 214, determine whether the remaining data volume is less than the size of an SST file. If so, go to 205; otherwise, go to 215.

[0054] At 215, close the current SST file and create a new SST file, then to 205.

[0055] In addition, this embodiment also discloses a key-value storage inter-layer merge optimization device based on Rime application characteristics. The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the device executes the key-value storage inter-layer merge optimization method based on Rime application characteristics.

[0056] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection of the present invention.

Claims

1. A key-value storage inter-layer merging optimization method based on Rime application characteristics, used in a key-value storage system, characterized in that: include: In the data search step, the pinyin input by the user using the Rime input method is used as the key prefix and searched in the memory table. If the target data is not found in the memory table, each level in the disk is searched in turn; After the target data is found and read, the read target data is temporarily stored in the key-value pair buffer, and the "Number of Reads" field in the hot and cold record table of the SST file is updated to mark the SST file as read; according to the Chinese characters / words selected by the user using the Rime input method, the pinyin and Chinese characters / words are combined into a complete key, and the key-value pairs with the same key are searched in the key-value pair buffer and the user dictionary in turn. If found, the information is updated according to the existing data, otherwise the relevant information is initialized, and a new key-value pair is obtained and written into the user dictionary; Data writing step: For the key-value pairs written by Rime input method, they are first saved in a writable memory table. When the size of the writable memory table reaches a first threshold, the writable memory table is converted to a read-only memory table, and its content is written to the disk in the form of an SST file. When writing to disk, check whether the size of a layer exceeds the second threshold limit. If so, trigger the inter-layer merge operation. For the layer that exceeds the second threshold, select all SST files in the layer for merging, and write the merged SST files to the next layer. During the merging process, if the size of the written SST file exceeds the preset file size, the file is not closed immediately. Instead, the next key-value pair is read and its key prefix is ​​determined to be the same as the key prefix of the current key-value pair. If the same, the data continues to be written to the current SST file. If the size of the written SST file does not exceed the preset file size, it is determined whether the remaining data volume is less than the SST file size. If it is, the remaining key-value pairs are written into the current SST file. If not, the current SST file is closed and a new SST file is generated for subsequent data writing. When the merged new SST file is written to the next level, it is determined whether there is a hot SST file in the next level. If so, and its key range overlaps with the range of the newly written SST file, the hot SST file is merged with the new SST file, the number of reads of the merged file is calculated, and the number of reads is recorded in the SST file hot and cold record table. For the old SST files involved in the merge, they are marked for deletion in the corresponding SST file hot and cold record table.

2. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: The SST files in each level are unordered. When searching, a table iterator is built for all SST files in each level until the target data is found.

3. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: The update information includes the number of selections and a timestamp.

4. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: The hot SST file is an SST file that is read more times than the average number of times read at the same layer within a certain time range.

5. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: After merging the hot SST file with the new SST file, the method further includes: writing the merged file to the same level as the hot SST file.

6. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: The number of reads of the calculated merged file is equal to the average number of reads of the hot SST files involved in the merge.

7. The key-value storage inter-layer merging optimization method based on Rime application characteristics according to claim 1 is characterized in that: The SST file hot and cold record table consists of four fields: level number, file number, number of reads and whether to delete; among them, the level number refers to the numbers 1-6 of the L1-L6 layers; the file number refers to the serial number of the SST file; the number of reads refers to the frequency with which the SST file is read, which is used to evaluate the popularity of the file; whether to delete refers to whether the SST file has been deleted. When the size of the SST file hot and cold record table reaches the third threshold, the records in the table where the whether to delete field is true will be cleared.

8. A key-value storage inter-layer merging optimization device based on Rime application characteristics, characterized in that: The device comprises a memory and a processor; the memory stores codes, and the processor is configured to execute the codes. When the codes are executed, the device executes the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Time series data storage method and system based on LSM tree key value separation

    CN114780530A

  • Key value storage and read-write method based on memory table index and iterator reduction mechanism

    CN118092812A