Data processing method and device, data query method and device, equipment, medium and product

Through the hybrid-separated storage architecture and data split retention strategy, the problems of write latency and query efficiency of the log structure merge tree storage system are solved, and efficient write and query performance is achieved.

CN120386794AActive Publication Date: 2025-07-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510873037.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing data storage systems based on log structure merging trees are difficult to meet the needs of low write latency and improve query efficiency at the same time. A fully hybrid storage architecture leads to low read amplification and query efficiency, while a fully separate storage architecture increases write latency and processor overhead.

Method used

Adopt a hybrid-separated storage architecture, the hybrid layer adopts a hybrid storage method, and the separation layer adopts a volume-based storage method, combining the splitting strategy of the target volume data and the retention strategy of other data, simplifying the write path and improving query accuracy.

Benefits of technology

Reduces write latency, improves write throughput and query efficiency, reduces storage fragmentation, and saves storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386794A_ABST
    Figure CN120386794A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, a data query method and device, equipment, a medium and a product, and can be applied to the technical field of computers. The data processing method comprises the steps that in response to a write-in instruction or a disk refreshing instruction, to-be-stored data indicated by the write-in instruction is written into an active memory table stored in a memory, or to-be-stored data indicated by the disk refreshing instruction is written into a first file stored in a first-level external memory, and the active memory table stores multiple pieces of data according to the sequence of multiple keys, the data comprises a key with a volume identifier, and the first file corresponding to the active memory table comprises at least one piece of data with a plurality of volume identifiers; and in response to a merging operation between the (n-1) th-level external memory and the nth-level external memory, obtaining target volume data with a target volume identifier in the plurality of volume identifiers, and writing the target volume data with the target volume identifier into an nth file which is stored in the nth-level external memory and corresponds to the target volume identifier, n being an integer greater than 1.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technologies, and particularly to a data processing method, a data query method, a device, a device, a medium, and a product. Background Art

[0002] With the development of computer storage technologies, data storage systems based on Log Structured MergeTree (LSM-Tree) are becoming increasingly popular and can be applied to key-value storage databases, distributed file systems, or time series databases, etc., especially suitable for large-scale data storage systems that need to optimize both write throughput and query latency.

[0003] To meet the respective requirements of write operations and query operations, it is necessary to reduce write latency and improve query efficiency. Summary of the Invention

[0004] In view of the above problems, the embodiments of the present application provide a data processing method, a data query method, a device, a device, a medium, and a product for reducing write latency and improving query efficiency.

[0005] According to a first aspect of the embodiments of the present application, a data processing method is provided. The method may include: in response to a write instruction or a disk flushing instruction, writing the data to be stored indicated by the write instruction into the active memory table stored in the memory, or writing the data to be stored indicated by the disk flushing instruction into the first file stored in the first-level external memory, where the active memory table stores multiple data sorted by multiple keys, the data includes keys with volume identifiers, and the first file corresponding to the active memory table includes at least one data for each of the multiple volume identifiers; in response to a merge operation between the (n-1)-th level external memory and the n-th level external memory, obtaining target volume data with a target volume identifier among the multiple volume identifiers, and writing the target volume data with the target volume identifier into the n-th file corresponding to the target volume identifier stored in the n-th level external memory, where n is an integer greater than 1.

[0006] According to a second aspect of the embodiments of the present application, a data query method is provided. The method may include: obtaining a query instruction, where the query instruction includes a query key with a volume identifier to be queried; in response to not querying a query value matching the query key from the active memory table and the inactive memory table stored in the memory, and at least one first file stored in the first-level external memory according to the query key, querying a query value matching the query key from the n-th external memory file corresponding to the volume identifier to be queried stored in the n-th level external memory, where n is an integer greater than 1.

[0007] According to a third aspect of the embodiments of the present application, a data processing device is provided. The device may include: a first writing module configured to, in response to a writing instruction or a disk flushing instruction, write the data to be stored indicated by the writing instruction into an active memory table stored in the memory, or write the data to be stored indicated by the disk flushing instruction into a first file stored in the first-level external memory, wherein the active memory table stores multiple data sorted by multiple keys, the data includes keys with volume identifiers, and the first file corresponding to the active memory table includes at least one data for each of the multiple volume identifiers; an obtaining module configured to, in response to a merge operation between the (n-1)-th level external memory and the n-th level external memory, obtain target volume data with a target volume identifier among the multiple volume identifiers; and a second writing module configured to write the target volume data with the target volume identifier into an n-th file corresponding to the target volume identifier and stored in the n-th level external memory, where n is an integer greater than 1.

[0008] According to a fourth aspect of the embodiments of the present application, a data query device is provided. The device may include: an obtaining module configured to obtain a query instruction, wherein the query instruction includes a query key with a volume identifier to be queried; a query module configured to, in response to failing to query a query value matching the query key from an active memory table and an inactive memory table stored in the memory, and at least one first file stored in the first-level external memory according to the query key, query a query value matching the query key from an n-th external memory file corresponding to the volume identifier to be queried and stored in the n-th level external memory, where n is an integer greater than 1.

[0009] According to a fifth aspect of the present application, an electronic device is provided, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0010] According to a sixth aspect of the present application, a computer-readable storage medium is further provided, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0011] According to a seventh aspect of the present application, a computer program product is further provided, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented. Description of the Drawings

[0012] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other purposes, features and advantages of the present application will become clearer. In the drawings:

[0013] Figure 1Shows a schematic diagram of a storage device for storing log-structured merge data;

[0014] Figure 2 Shows a schematic diagram of the principle of a data processing method based on a hybrid-separated storage architecture according to an embodiment of the present application;

[0015] Figure 3 Shows a schematic diagram of the principle of a splitting strategy for target volume data and a retention strategy for other data according to an embodiment of the present application;

[0016] Figure 4 Shows a system architecture diagram of a data processing system applicable to a data processing method or a data query method according to an embodiment of the present application;

[0017] Figure 5 Shows a flowchart of a data processing method according to an embodiment of the present application;

[0018] Figure 6 Shows an application schematic diagram of a splitting strategy for target volume data and a retention strategy for other data according to an embodiment of the present application;

[0019] Figure 7 Shows a flowchart of a data query method according to an embodiment of the present application;

[0020] Figure 8 Shows a block diagram of a data processing device according to an embodiment of the present application;

[0021] Figure 9 Shows a block diagram of a data query device according to an embodiment of the present application;

[0022] Figure 10 Shows a block diagram of an electronic device suitable for implementing a data processing method or a data query method according to an embodiment of the present application. Detailed implementation manners

[0023] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0024] The terms used herein are for describing specific embodiments only and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] In the case of using expressions such as "at least one of A, B, or C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, having A, B, and C, etc.).

[0027] In the embodiments of the present application, "indicating" may include direct indication, indirect indication, explicit indication, or implicit indication. When describing that a certain indication information is used to indicate A, it can be understood that the indication information carries A, directly indicates A, or indirectly indicates A.

[0028] In the embodiments of the present application, the various numerical numbers involved are only for the convenience of description and are not used to limit the protection scope of the present application. The magnitudes of the serial numbers involved in the embodiments of the present application do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic. For example, terms such as "first", "second", "third", "fourth", and other various term numbers in the specification, claims, and drawings of the embodiments of the present application (if any) can be used to distinguish similar objects and do not have to be used to describe a specific order or sequence. Among them, such terms can be interchanged under appropriate circumstances.

[0029] If there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments of the embodiments of the present application are consistent and can be cross-referenced. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.

[0030] For the convenience of clearly describing the technical solutions of the embodiments of the present application, some terms involved in the embodiments of the present application are described below.

[0031] 1. Write amplification, read amplification, spatial amplification, write-intensive applications, read-intensive applications

[0032] Write amplification can refer to the ratio between the amount of data actually written to a storage device and the amount of data truly written to the storage device in the case of writing data.

[0033] Read amplification can refer to the ratio between the amount of data actually read from a storage device and the amount of data truly read from the storage device in the case of querying data.

[0034] Space amplification can refer to the ratio between the actually occupied storage space and the logical size of the valid data.

[0035] Write-intensive applications can refer to applications where write operations dominate the input / output operations. For example, write-intensive applications can include logging applications. Write-intensive applications require reducing write amplification and increasing write throughput.

[0036] Read-intensive applications can refer to applications where read operations dominate the input / output operations.

[0037] 2. Off-site update, in-place update

[0038] Off-site update can refer to the situation where the update to data does not act on the data itself, but writes the data to be updated to a new storage location. The data is retained and marked as invalid. Off-site update is suitable for write-intensive scenarios.

[0039] In-place update can refer to the situation where the update to data directly acts on the data itself, that is, the data can be updated according to the data to be updated. The data is modified. The data to be updated and the storage location of the data are the same.

[0040] 3. Volume (or storage volume), volume identifier, key with volume identifier, volume data

[0041] A volume can refer to a logical partition unit for data storage. A volume identifier can indicate a volume. A key with a volume identifier can refer to a key that can include a volume identifier. The volume identifier is a component of the key. Volume data can include at least one data with a volume identifier. Different volume data have different volume identifiers. For example, a key can be represented as "<volume identifier, key identifier>".

[0042] 4. Log-structured merge tree, memory, external storage, active memory table, inactive memory table, nth-level external storage, nth file

[0043] The log-structured merge tree converts random writes into sequential writes through a memory structure, has high write performance, and uses logs to achieve data persistence. It is friendly to write-intensive applications and can efficiently process large-scale data. The storage device used to store the log-structured merge tree can include memory and external storage. For ease of understanding, the following is combined with Figure 1 for illustration.

[0044] Figure 1Shows a schematic diagram of a storage device for storing log-structured merge data.

[0045] As Figure 1 shown, the storage device may include memory and external storage. The memory can be used to store an active memory table (i.e., MemTable) and at least one inactive memory table (i.e., Immutable MemTable).

[0046] The external storage may include multiple levels of external storage. For example, the external storage may include a first-level external storage (i.e., L0), a second-level external storage (i.e., L1),..., an nth-level external storage (i.e., Ln-1),.... n can be an integer greater than or equal to 1. The nth-level external storage can be used to store at least one nth file. The number of files stored in different levels of external storage may be different. The number of nth files stored in the nth-level external storage may be greater than the number of (n - 1)th files stored in the (n - 1)th-level external storage. The file format of the nth file may be a sorted string table (SSTable).

[0047] 5. Hybrid storage, volume grouping, hybrid layer, separation layer

[0048] Hybrid storage may refer to storing data sorted by a global key without distinguishing volume identifiers. In the embodiments of the present application, the storage methods of the active memory table, the inactive memory table, and the first file may be hybrid storage. The keys in the active memory table, the inactive memory table, and the first file may be referred to as global keys.

[0049] Volume-based storage may refer to storing data sorted by a local key while distinguishing volume identifiers. In the embodiments of the present application, the storage method of the nth file may be volume-based storage. The keys in the nth file may be referred to as local keys. n can be an integer greater than 1.

[0050] The hybrid layer may refer to the layer that stores data based on the hybrid storage method. In the embodiments of the present application, the hybrid layer may include memory and the first-level external storage.

[0051] The separation layer may refer to the layer that stores data based on the volume-based storage method. In the embodiments of the present application, the separation layer may include the second-level external storage and the external storage levels above the second level.

[0052] 6. Skip list (i.e., Skip List)

[0053] Skip list is a probabilistic data structure. A skip list maintains an ordered structure similar to a multi-level linked list. When data to be stored is written, new nodes will be added to the skip list. The height of a skip list node is at least 1, and the value can be determined according to probability. The higher the level, the lower the probability of generating a node. Therefore, the number of nodes at higher levels is small, and there are more other nodes between adjacent higher-level nodes. When performing a query operation, the skip list starts querying from the top layer. The upper linked list structure in the skip list can be used as an index for key queries. By skipping unmatched nodes through higher-level nodes, the number of nodes to be queried is reduced, accelerating the query process.

[0054] With the development of computer storage technology, data storage systems based on log-structured merge trees are becoming more and more popular. They can be applied to key-value storage databases, distributed file systems, or time-series databases, etc., and are especially suitable for large-scale data storage systems that need to optimize both write throughput and query latency simultaneously.

[0055] A data storage system based on a log-structured merge tree may include the following storage architectures.

[0056] As an implementation, a fully hybrid storage architecture. For example, from the active memory table to the nth file, there is no distinction between volumes or columns. Since query operations need to traverse the files stored in each level of external memory respectively, the read amplification problem is relatively serious, the query efficiency is low, and the input / output overhead is large.

[0057] As another implementation, a fully separated storage architecture. For example, from the active memory table to the nth file, volumes or columns are distinguished. Since when performing a write operation, it is necessary to first determine the data classification of the data to be stored, that is, determine which volume or column to write the data to be stored, the write latency and processor overhead are increased.

[0058] In the process of implementing the inventive concept of this application, it is found that it is difficult to simultaneously meet low write latency and improve query efficiency. For this reason, the embodiments of this application propose a hybrid-separated storage architecture, that is, a hybrid storage method can be adopted in the hybrid layer and a per-volume storage method can be adopted in the separated layer. Hybrid storage may refer to not distinguishing volume identifiers, and data is stored according to a global key. The hybrid layer may include memory and the first-level external memory. Per-volume storage may refer to distinguishing volume identifiers, and data is stored according to a local key. The separated layer may include the second-level external memory and the external memory levels above the second-level external memory. Thus, query efficiency can be improved on the basis of improving write performance. For the sake of easy understanding, the following will be described in conjunction with Figure 2 for illustration.

[0059] Figure 2 shows a schematic diagram of the principle of a data processing method based on a hybrid-separated storage architecture according to an embodiment of this application.

[0060] As Figure 2As shown, in the hybrid-separated storage architecture, the hybrid layer may include memory and a first-level external storage. The hybrid layer may store data in a hybrid storage manner. The memory may be used to store an active memory table and an inactive memory table. The first-level external storage may be used to store at least one first file.

[0061] The active memory table may include multiple data. The data may include keys with volume identifiers, that is, the volume identifier is a component of the key. Thus, the active memory table may include at least one data for each of multiple volume identifiers. The active memory table may store multiple data sorted by multiple keys. The keys in the active memory table may be referred to as global keys. For example, Figure 2 the active memory table may include 8 data. Each of the 8 data may have "volume identifier V1", "volume identifier V2", "volume identifier V1", "volume identifier V1", "volume identifier V1", "volume identifier V3", "volume identifier V1", and "volume identifier V1".

[0062] The inactive memory table may be obtained by performing a state transition operation on the active memory table. The inactive memory table has the same stored data and storage method as the active memory table, that is, the inactive memory table may store multiple data sorted by multiple keys with volume identifiers. Similarly, the keys in the inactive memory table may be referred to as global keys.

[0063] The data stored in the first file may be obtained by performing a disk flushing operation on the inactive memory table. The first file has the same stored data and storage method as the inactive memory table, that is, the first file may store multiple data sorted by multiple keys with volume identifiers.

[0064] In the hybrid-separated storage architecture, the separated layer may include a second-level external storage, …, an n-level external storage, …. n may be an integer greater than 1. The separated layer may store data in a volume-based storage manner. The n-level external storage may be used to store at least one nth file. At least one nth file may have its own nth file identifier. At least one nth file identifier may each have a target volume identifier among multiple volume identifiers, that is, the nth file identifier corresponding to the target volume identifier may have the target volume identifier. Thus, the nth file corresponding to the target volume identifier may store target volume data with the target volume identifier. The keys in the nth file may be referred to as local keys.

[0065] For example, Figure 2The target volume identifier can be "volume identifier V1", "volume identifier V2", or "volume identifier V3". Thus, the nth file corresponding to the target volume identifier "volume identifier V1" can store the target volume data with "volume identifier V1". The nth file corresponding to the target volume identifier "volume identifier V2" can store the target volume data with "volume identifier V2". The nth file corresponding to the target volume identifier "volume identifier V3" can store the target volume data with "volume identifier V3".

[0066] Since the hybrid layer stores data using a hybrid storage method (i.e., a fully flattened storage method), data belonging to different volume data is physically stored continuously and shares the same set of index structures. Therefore, the write operation does not require prior data classification, and the data to be stored can be directly appended to the current active memory table, that is, the data to be stored can be written to the current active memory table using the off-site update method. Thus, the write path is simplified, the write latency is reduced, and the write throughput is increased. Since the separation layer stores data using a per-volume storage method, when the query operation accesses the secondary external memory and the external memory above the secondary external memory, it is only necessary to query the nth file corresponding to the volume identifier to be queried, without querying all the nth files, achieving accurate positioning of the nth file corresponding to the volume identifier to be queried that needs to be accessed. Thus, the input / output overhead is reduced and the query efficiency is improved.

[0067] In addition, it is also found that due to the fully hybrid storage architecture, the hot data with high-frequency access is scattered and stored in different nth files, thus reducing the cache hit rate. Since the fully separated architecture may cause the active memory tables of different columns or volumes to simultaneously meet the conditions for being converted into inactive memory tables to perform the disk flushing operation, write amplification is caused. Since the secondary external memory and the external memory above the secondary external memory in the embodiments of the present application use a per-volume storage method, the hot data with high-frequency access can be centrally stored, thus increasing the cache hit rate. Since the active memory table and the inactive memory table in the embodiments of the present application are hybridly stored, write amplification is alleviated.

[0068] In addition, it is also found that there is a problem of storage fragmentation caused by premature splitting of highly hybrid small data blocks. The small data blocks can refer to volume data with a stored data volume less than or equal to the first predetermined data volume.

[0069] To this end, the embodiments of the present application propose a merging operation between the first-level external memory and the second-level external memory, and a splitting strategy for the target volume data can be adopted, that is, writing the target volume data determined from multiple volume data stored in the first-level external memory to the second-level external memory. The target volume data may refer to volume data that meets a predetermined condition. The predetermined condition may refer to a condition adapted to reduce storage fragmentation. For example, the predetermined condition may include at least one of a first predetermined data volume or a predetermined data volume ratio. The data volume of the target volume data may be greater than or equal to the first predetermined data volume or the data volume ratio of the target volume data may be greater than or equal to at least one of the predetermined data volume ratios. Thereby, the probability of occurrence of the storage fragmentation problem caused by prematurely splitting highly mixed small data blocks is reduced. In addition, space amplification can also be reduced, saving storage space.

[0070] In addition, the embodiments of the present application also propose a retention strategy for other data that does not meet the predetermined conditions, that is, other data in the multiple volume data can be retained in the first-level external memory. Since the first-level external memory stores other data, the data volume stored in the first-level external memory is reduced.

[0071] To facilitate understanding of the splitting strategy for the target volume data and the retention strategy for other data, the following will be combined with Figure 3 for illustration.

[0072] Figure 3 FIG. shows a schematic diagram of the principle of the splitting strategy for the target volume data and the retention strategy for other data according to the embodiments of the present application.

[0073] As Figure 3 shown, at least one data with multiple volume identifiers can be obtained from at least one first file stored in the first-level external memory to obtain multiple volume data. Then, the target volume data that meets the predetermined conditions is determined from the multiple volume data, and the target volume data is written to the second file with the target volume identifier stored in the second-level external memory, while other data in the multiple volume data can be retained in the first-level external memory. For example, a merging operation can be performed on other data in the multiple volume data to obtain first merged data. The first merged data can be merged into the target first file stored in the first-level external memory.

[0074] The above describes the inventive concept of the embodiments of the present application. The following will specifically describe the data processing method and data query method provided by the embodiments of the present application with reference to the accompanying drawings.

[0075] First, the data processing system applicable to the data processing method will be described in combination with Figure 4 Next, taking the system architecture shown in Figure 4 as an example, in combination with Figure 5 and Figure 6, a specific description of the data processing method provided by the embodiments of the present application is given. Among them, Figure 5 A description of the data processing method is given. Figure 6 A description of the splitting strategy for target volume data and the retention strategy for other data is given. Then, in combination with Figure 7 A specific description of the data query method provided by the embodiments of the present application is given.

[0076] It should be noted that the embodiments of the present application can be implemented independently or in combination with each other. For the same or similar concepts or processes, they will not be elaborated in some embodiments.

[0077] Figure 4 The system architecture diagram of a data processing system applicable to the data processing method or the data query method according to the embodiments of the present application is shown.

[0078] As Figure 4 shown, the data processing system may include a computing device and a storage device. In some embodiments, the computing device may be a server, a server cluster or a distributed system composed of multiple servers. The computing device may also be a cloud service cluster that provides basic cloud computing services such as cloud storage, cloud services, cloud databases, cloud computing, cloud functions, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, big data or artificial intelligence platforms. The embodiments of the present application do not limit this.

[0079] The computing device may include a memory and a processor. The memory may be used to store an active memory table and an inactive memory table. The memory may include at least one of the following: DRAM (Dynamic Random Access Memory), NVM (Non-Volatile Memory), off-heap memory (i.e., Off-Heap) or persistent memory (Persistent Memory, PMEM), etc.

[0080] The storage device may include an external memory. The external memory may include at least one of the following: solid state drive (SSD), hard disk drive (HDD), NVMe (Non-Volatile Memory Express), or persistent memory. Optionally, the external memory may also be a disk array. The external memory may include multiple levels of external memory, that is, the first-level external memory, the second-level external memory,..., the nth-level external memory,.... n may be an integer greater than 1. The nth-level external memory may be used to store the nth file. n may be an integer greater than 1.

[0081] For the description of the hybrid layer and the separation layer, please refer to the corresponding part of the above description, which will not be elaborated here.

[0082] The computing device can be used to execute the data processing method or the data query method described in the embodiments of the present application. The data processing method may include a write operation of writing data to be stored into the active memory table, a status switching operation for the active memory table, a disk flushing operation for the inactive memory table, and a merging operation between the n-1th level external memory and the nth level external memory. Please refer to the following specific descriptions of each operation, which will not be elaborated here.

[0083] It should be noted that the data processing method or the data query method described in the embodiments of the present application may also be executed by other devices different from the computing device, and the embodiments of the present application do not limit this. The execution subjects of the data processing method and the data query method may be the same or different. Based on the technical solutions provided in the embodiments of the present application, more or fewer components or functions may be configured for the data processing system according to business requirements.

[0084] Next, based on Figure 4 the described system architecture, combined with Figure 5 、 Figure 6 and Figure 7 the data processing method and the data query method of the embodiments of the present application will be specifically described.

[0085] Figure 5 FIG. shows a flowchart of the data processing method according to the embodiments of the present application.

[0086] As Figure 5 shown, the method includes operations S510~S520. Operation S510 involves a write operation and a disk flushing operation. Operation S520 involves a merging operation.

[0087] In operation S510, in response to a write instruction, write the data to be stored indicated by the write instruction into the active memory table stored in the memory, or in response to a disk flushing instruction, write the data to be stored indicated by the disk flushing instruction into the first file stored in the first level external memory.

[0088] In operation S520, in response to the merging operation between the n-1th level external memory and the nth level external memory, obtain the target volume data with the target volume identifier among the multiple volume identifiers, and write the target volume data with the target volume identifier into the nth file corresponding to the target volume identifier stored in the nth level external memory.

[0089] The data to be stored can be data that needs to be written into the active memory table or the first file. When the data to be stored is data that needs to be written into the active memory table, the write instruction can include the data to be stored. Optionally, the write instruction can include a first access path, whereby the first access path can be accessed to obtain the data to be stored. When the data to be stored is data that needs to be written into the first file, the disk flushing instruction can include data in the inactive memory table. Optionally, the disk flushing instruction can include a second access path, whereby the inactive memory table can be accessed according to the second access path. Thus, the data in the inactive memory table can be obtained. The data in the inactive memory table can be the data to be stored.

[0090] The target volume data can refer to volume data that meets a predetermined condition. The predetermined condition can include at least one predetermined evaluation metric. Optionally, the predetermined condition can refer to the condition for a pre-trained classification model to determine whether the volume data is target volume data.

[0091] The storage methods of the active memory table and the first file can be hybrid storage. The storage method of the nth file can be volume-based storage. n can be an integer greater than 1. Hybrid storage can mean that volume identification data is not distinguished and data is stored sorted by a global key. Volume-based storage can mean that volume identification is distinguished and data is stored sorted by a local key. The following will separately describe hybrid storage and volume-based storage.

[0092] The active memory table can include multiple pieces of data. The data can include a key with a volume identification, that is, the volume identification is a component of the key. Thus, the active memory table can include at least one piece of data for each of multiple volume identifications. Multiple pieces of data can be sorted according to the sorting of multiple keys to obtain the multiple pieces of data stored in the active memory table. The multiple pieces of data are stored in the active memory table sorted by multiple keys. The keys in the active memory table can be called global keys.

[0093] The inactive memory table can be obtained by performing a state switching operation on the active memory table. For example, if the data volume of the active memory table is equal to a third predetermined data volume, the status flag of the active memory table can be configured as a read-only flag to obtain the inactive memory table. Thus, the inactive memory table can be an active memory table with a status flag of read-only and a data volume equal to the third predetermined data volume. The stored data and storage method of the inactive memory table are the same as those of the active memory table, that is, the inactive memory table can store multiple pieces of data sorted by multiple keys with volume identifications. Similarly, the keys in the inactive memory table can be called global keys.

[0094] The data stored in the first file can be obtained by performing a disk flushing operation on the inactive memory table. The number of the first files can be at least one. If the number of the first files is one, the first file has the same storage data and storage method as the inactive memory table, that is, the first file can store multiple data sorted by multiple keys with volume identifiers. If the number of the first files is multiple, the first files can store some data in the inactive memory table and have the same storage method as the inactive memory table. Similarly, the keys in the first files can be called global keys.

[0095] The nth-level external storage can store at least one nth file. The nth file can be obtained by performing a merge operation between the (n - 1)th level and the nth level. At least one nth file can have its own nth file identifier. At least one nth file identifier can each have a target volume identifier among multiple volume identifiers, that is, the nth file identifier corresponding to the target volume identifier can have the target volume identifier. Thus, the nth file corresponding to the target volume identifier can store the target volume data with the target volume identifier. n can be an integer greater than 1. The target volume data can include multiple data with the target volume identifier. The data can include keys with the target volume identifier. The keys in the nth file can be called local keys.

[0096] Based on the above, the storage path of the data stored in the first file can be: active memory table -> inactive memory table -> at least one first file. Since the active memory table can include at least one data with each of the M volume identifiers, therefore, if there is one first file, the first file corresponding to the active memory table can include at least one data with each of the M volume identifiers. If there are multiple first files, the first files corresponding to the active memory table can store some data in the active memory table, and the some data can be data with each of the M volume identifiers. Optionally, the some data can be data with each of some of the M volume identifiers. M can be an integer greater than 1.

[0097] According to an embodiment of the present application, since the active memory table can store multiple data sorted by multiple keys, and the data includes keys with volume identifiers, a data storage method of the active memory table using a hybrid storage method is implemented. Based on this, when writing the data to be stored indicated by the write instruction into the active memory table, since there is no need to classify the data in advance, the write latency is reduced and the write throughput is increased, thereby improving the write performance. In addition, since the first file corresponding to the active memory table includes at least one data for each of the multiple volume identifiers, a data storage method of the first file in the first-level external memory using a hybrid storage method is implemented. Since the nth file corresponding to the target volume identifier in the nth-level external memory is used to store the target volume data with the target volume identifier, the volume-based storage of the nth file is implemented. On this basis, since the nth file uses a volume-based data storage method, the query value matching the query key can be queried from the nth external memory file corresponding to the query volume identifier stored in the nth-level external memory according to the query key, without querying the nth file with other volume identifiers. Therefore, the query accuracy is improved, the query efficiency is increased, and the input / output overhead is reduced. n can be an integer greater than 1.

[0098] The following specifically describes the write operation, state transition operation, disk flushing operation, and merge operation between the first-level external memory and the second-level external memory involved in the hybrid layer, as well as the merge operation between the (n - 1)th-level external memory and the nth-level external memory involved in the separation layer, where n ∈ {3, 4,...}.

[0099] I. Hybrid layer

[0100] 1.1. Write operation involved in the hybrid layer

[0101] To simplify the write path, the data to be stored can be directly appended to the active memory table, without first determining the data to be stored in the volume data according to the volume identifier to be stored, then determining the local write position from the data to be stored in the volume according to the local key to be stored, and then writing the data to be stored to the local write position. Thus, the data to be stored is written to the data to be stored in the volume of the active memory table, thereby reducing the storage overhead generated by data classification (i.e., determining the data to be stored in the volume), and improving the storage efficiency and write throughput.

[0102] The direct appending of the data to be stored to the active memory table can be achieved in the following way.

[0103] As an implementation, the write instruction may include data to be stored. The data to be stored may include a value to be stored and a key to be stored with a volume identifier to be stored. For example, the key to be stored may be represented as "<volume identifier to be stored, key identifier to be stored>". The computing device may receive the write instruction. According to the key to be stored with the volume identifier to be stored included in the write instruction, determine the global write position from the active memory table. The key to be stored may be referred to as the global key to be stored. Thus, the data to be stored can be written to the global write position. For example, the data to be stored can be written to the global write position using an atomic operation, so as to implement writing the data to be stored to the active memory table. In addition, the computing device may also write the data to be stored to a write-ahead log (WAL) stored in external memory for maintaining data persistence.

[0104] As an implementation, the computing device may include a write thread. Thus, the write operation can be performed using the write thread, that is, the write thread can receive the write instruction, and according to the key to be stored included in the write instruction, determine the global write position from the active memory table

[0105] The data structure of the active memory table can be configured according to actual business requirements and is not limited here. For example, the data structure of the active memory table can be a skip list (i.e., Skip List), a balanced binary search tree, or a B+ tree (i.e., B+ Tree), etc. The balanced binary search tree may include at least one of the following: a red-black tree or an AVL tree (i.e., Adelson-Velsky and Landis Tree). Here, taking the data structure of the active memory table as a skip list as an example, further description will be made for writing the data to be stored to the active memory table.

[0106] The skip list may include multiple layers of linked lists. The current layer linked list may be a subset of the next layer linked list. The lowest layer linked list may include multiple data stored in the active memory table. The higher layer linked lists can determine whether the data is located in the higher layer linked lists through a randomization method. The skip list enables query operations to skip some data, thus accelerating the query. A node may include data and a pointer to the same key in the next layer linked list. The data may include a value and a key with a volume identifier.

[0107] As an implementation, start traversing from the highest-level linked list of the skip list. For the current-level linked list among multiple level-linked lists, traverse multiple nodes of the current-level linked list. For the current node among multiple nodes, when the key of the current node is equal to the key to be stored, update the value of the current node according to the value to be stored. When the key of the current node is greater than the key to be stored or when the current node reaches the end of the current-level linked list, move to the next-level linked list, and the next-level linked list can be used as the new current-level linked list. Repeat the above traversal operation until reaching the lowest-level linked list. Determine the predecessor node of the position to be written in the lowest-level linked list. Generate a new node, and the new node can include the data to be stored. Randomly determine the level to which the new node belongs and insert the new node into each level.

[0108] 1.2. State transition operation

[0109] The computing device can obtain the data volume stored in the active memory table. When the data volume stored in the active memory table is equal to the third predetermined data volume, configure the status flag of the active memory table as a read-only flag. Thus, the active memory table is transformed into an inactive memory table. As an implementation, the computing device can include a writing thread. Thus, the writing thread can be used to perform the state transition operation, that is, the writing thread configures the status flag of the active memory table as a read-only flag in response to the data volume stored in the active memory table being equal to the third predetermined data volume, resulting in an inactive memory table. In addition, the writing thread can also be used to store the inactive memory table in a dynamic array.

[0110] 1.3. Disk flushing operation

[0111] The disk flushing operation is a serialization operation from the data structure of the inactive memory table to the data structure of the external storage. The storage medium changes from the inactive memory table in the memory to the nth file in the external storage. The stored data and storage method of the inactive memory table and the first file are the same, that is, the inactive memory table stores multiple data sorted by multiple keys. The first file can also store multiple data sorted by multiple keys. The data structures of the same data in the inactive memory table and the first file can be different but the content is the same. The disk flushing operation can achieve the persistence of memory data, retain the mixed data, and change the storage medium.

[0112] As an implementation, the disk flushing instruction can be generated in the following way: the disk flushing instruction can be generated in response to the number of inactive memory tables being equal to the predetermined number. Optionally, the disk flushing instruction can be generated in response to the data volume stored in the first-level external storage being equal to the fourth predetermined data volume. Optionally, the disk flushing instruction can be generated in response to the data volume stored in the active memory table being greater than the third predetermined data volume. Optionally, the disk flushing instruction can be generated in response to a trigger operation of the operation body. The operation body can include a user or a stylus. The predetermined number, the third predetermined data volume, and the fourth predetermined data volume can be configured according to actual business requirements and are not limited here.

[0113] Accordingly, the computing device can access the inactive memory table in response to the disk flushing instruction, according to the second access path indicated by the disk flushing instruction. Obtain the data to be stored from the inactive memory table, and write the data to be stored into the first file according to a predetermined data format. For example, the computing device can sequentially scan the inactive memory table to obtain the data in the inactive memory table. The data in the inactive memory table is the data to be stored. Serialize the obtained data into a binary format to obtain the data stored in the first file. As an implementation, the data structure of the inactive memory table can be a skip list. Accordingly, the computing device can sequentially scan the inactive memory table according to the skip list order to obtain the data in the inactive memory table.

[0114] As an implementation, the computing device can include a disk flushing thread. Accordingly, use the disk flushing thread to perform the disk flushing operation, that is, use the disk flushing operation to access the inactive memory table in response to the disk flushing instruction, according to the second access path indicated by the disk flushing instruction. Obtain the data to be stored from the inactive memory table, and write the data to be stored into the first file according to a predetermined data format. In addition, use the disk flushing thread to remove the inactive memory table from the dynamic array in response to completing the disk flushing operation for the inactive memory table, realizing the management of the inactive memory table using dynamic data.

[0115] 1.4. Merging operation between the first-level external storage and the second-level external storage

[0116] The computing device can respond to the merging operation between the first-level external storage and the second-level external storage, obtain at least one data with multiple volume identifiers from at least one first file to obtain multiple volume data. At least one first file can be stored in the first-level external storage. Then determine the target volume data that meets the predetermined conditions from the multiple volume data, and write the target volume data into the second file with the target volume identifier stored in the second-level external storage, while the other data in the multiple volume data can be retained in the first-level external storage. The target volume data can be referred to as the dominant volume data. The splitting strategy for the target volume data and the retention strategy for the other data will be described below with reference to the accompanying drawings.

[0117] 1.4.1. Splitting strategy for the target volume data (or dominant volume data)

[0118] Regarding how to obtain multiple volume data from at least one first file, it can be implemented in the following ways.

[0119] As an implementation, the data may include keys with volume identifiers. The computing device may obtain at least one data with each of a plurality of volume identifiers from at least one first file, and obtain volume data with each of the plurality of volume identifiers. For any volume identifier among the plurality of volume identifiers included in the at least one first file, at least one data with the volume identifier may be obtained from the at least one first file, whereby volume data with the volume identifier is obtained. Optionally, at least one data with the volume identifier may be obtained from the at least one first file, and a deduplication operation may be performed on the at least one data with the volume identifier to obtain a volume identifier with the volume data. In the above manner, volume data with each of the plurality of volume identifiers can be obtained.

[0120] Regarding how to determine the target volume data from multiple volume data, it can be implemented in the following manner.

[0121] 1.4.1.1. Based on the manner of satisfying at least one predetermined evaluation condition

[0122] The predetermined condition may include at least one predetermined evaluation index. The at least one evaluation index may include at least one of the following: the first predetermined data volume, the predetermined data volume ratio, or the predetermined access frequency. The first predetermined data volume, the predetermined data volume ratio, and the predetermined access frequency can be configured according to actual business requirements and are not limited herein. For example, in order to match the system load, at least one of the first predetermined data volume, the predetermined data volume ratio, or the predetermined access frequency may be adjusted according to at least one of the load information or the inter - volume data distribution entropy. The load information may include at least one of the following: the data volume stored in the first - level external memory, the first resource utilization rate, or the first input - output delay. The inter - volume data distribution entropy may indicate the degree of data distribution uniformity among multiple volume data.

[0123] As an implementation, in response to at least one of the data volume stored in the first - level external memory being greater than or equal to the second predetermined data volume, the inter - volume data distribution entropy being greater than or equal to the predetermined distribution entropy, the first resource utilization rate being greater than or equal to the predetermined utilization rate, or the first input - output delay being greater than or equal to the predetermined delay, at least one of the first predetermined data volume, the predetermined data volume ratio, or the predetermined access frequency is reduced.

[0124] As an implementation, in response to at least one of the data volume stored in the first - level external memory being less than the second predetermined data volume, the inter - volume data distribution entropy being less than the predetermined distribution entropy, the first resource utilization rate being less than the predetermined utilization rate, or the first input - output delay being less than the predetermined delay, at least one of the first predetermined data volume, the predetermined data volume ratio, or the predetermined access frequency is increased.

[0125] A computing device may evaluate multiple volume data according to at least one predetermined evaluation metric to obtain the evaluation results of each of the multiple volume data. The evaluation result of the volume data may indicate whether the volume data meets at least one predetermined evaluation metric.

[0126] Since the predetermined evaluation metric for evaluating whether the volume data is the target volume data can be adjusted according to at least one of the load information or the inter - volume data distribution entropy, the adaptive adjustment of the predetermined evaluation instruction is realized. Thus, the determination accuracy of the target volume data is improved.

[0127] In response to the evaluation results of each of the multiple volume data including an evaluation result indicating that the volume data meets at least one predetermined evaluation metric, the volume data can be used as the target volume data. For any one of the multiple volume data, in response to the evaluation result of the volume data indicating that the volume data meets at least one predetermined evaluation metric, the volume data can be used as the target volume data. In response to the evaluation result of the volume data indicating that the volume data does not meet at least one predetermined evaluation metric, that is, in response to the evaluation result of the volume data indicating that the volume data does not meet any of the predetermined evaluation metrics, the volume data is not used as the target volume data. Thus, the multiple volume data may not include the target volume data. The multiple volume data may also include one target volume data or multiple target volume data.

[0128] When at least one of the predetermined evaluation metrics includes at least one of the following: the first predetermined data volume, the predetermined data volume ratio, or the predetermined access frequency, the evaluation result of the target volume data may indicate at least one of the following: the data volume of the target volume data is greater than or equal to the first predetermined data volume, the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio, or the access frequency of the target volume data within a predetermined time period is greater than or equal to the predetermined access frequency. The data volume ratio of the target volume data may be the ratio between the data volume of the target volume data and the total data volume of the multiple volume data.

[0129] Thus, the evaluation result of the target volume data can indicate that: the data volume of the target volume data is greater than or equal to the first predetermined data volume. Optionally, the evaluation result of the target volume data can indicate that: the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio. Optionally, the evaluation result of the target volume data can indicate that: the access frequency of the target volume data within a predetermined time period is greater than or equal to the predetermined access frequency. Optionally, the evaluation result of the target volume data can indicate that: the data volume of the target volume data is greater than or equal to the first predetermined data volume and the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio. Optionally, the evaluation result of the target volume data can indicate that: the data volume of the target volume data is greater than or equal to the first predetermined data volume and the access frequency of the target volume data within the predetermined time period is greater than or equal to the predetermined access frequency. Optionally, the evaluation result of the target volume data can indicate that: the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio and the access frequency of the target volume data within the predetermined time period is greater than or equal to the predetermined access frequency. Optionally, the evaluation result of the target volume data can indicate that: the data volume of the target volume data is greater than or equal to the first predetermined data volume, the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio, and the access frequency of the target volume data within the predetermined time period is greater than or equal to the predetermined access frequency.

[0130] The first predetermined data volume is used to reduce the probability of small data blocks being split, and the predetermined data volume ratio is used to reduce the probability of storage fragmentation. Thus, the probability of the storage fragmentation problem caused by premature splitting of highly mixed small data blocks is reduced.

[0131] For the case including multiple target volume data, the target volume data can be written to the second file stored in the second-level external memory in the following manner.

[0132] As an implementation method, the computing device can write the respective target volume data to multiple second files stored in the second-level external memory. The multiple second files each have different target volume identifiers. The target volume data can have a target volume identifier.

[0133] As another implementation method, the computing device can determine one target volume data from multiple target volume data. For example, one target volume data can be randomly determined from multiple target volume data. Optionally, one target volume data can be determined from multiple target volume data according to the priority of at least one predetermined evaluation index and the respective evaluation results of the multiple target volume data. The target volume data has a target volume identifier. The target volume data can be written to the second file with the target volume identifier stored in the second-level external memory.

[0134] 1.4.1.2. Based on the pre-trained classification model method

[0135] The target volume data can be determined by a computing device from multiple volume data using a pre-trained classification model. The model structure of the pre-trained classification model can be configured according to actual business requirements, which is not limited herein. Since the pre-trained classification model can achieve dynamic adaptation and self-learning and can upgrade static methods to intelligent and adaptive solutions, the accuracy of determining the target volume data is improved.

[0136] The pre-trained classification model can be obtained in the following manner.

[0137] Multiple sample volume feature data of multiple sample volume data can be obtained according to multiple sample volume data and sample system data. The sample volume feature data can include at least one of the following: the data volume included in the sample volume data, the data volume ratio of the sample volume data, the third resource utilization rate, or the third input / output latency. The data volume ratio of the sample volume data can be the ratio between the data volume of the sample volume data and the total data volume of the multiple sample volume data. The pre-trained classification model is obtained by training an initial classification model using the sample volume feature data of the multiple sample volume data.

[0138] For example, for any one of the multiple sample volume data, the sample volume feature data of the sample volume data can be input into the initial classification model to obtain the sample classification result of the sample volume data. Thus, the sample classification results of the multiple sample volume data can be obtained. Based on a predetermined loss function, a loss function value is obtained according to the sample classification results of the multiple sample volume data and the sample classification labels of the multiple sample volume data. The model parameters of the initial classification model are adjusted according to the loss function value until a predetermined end condition is satisfied, and the pre-trained classification model is obtained. The predetermined end condition can include at least one of the predetermined number of training rounds reaching the maximum number of training rounds or the loss function value being less than or equal to a predetermined threshold. The predetermined number of training rounds and the predetermined threshold can be configured according to actual business requirements, which is not limited herein.

[0139] Since the sample volume feature data includes at least one of the following: the data volume included in the sample volume data, the data volume ratio of the sample volume data, the third resource utilization rate, or the third input / output latency, the comprehensiveness of the sample volume feature data is improved, and all are feature data related to evaluating whether the volume data is target volume data. Therefore, using the sample volume feature data to train the initial classification model improves the accuracy of the classification result of the pre-trained classification model.

[0140] Accordingly, the computing device can obtain the volume feature data of each of the multiple volume data based on the multiple volume data and the system data. For any one of the multiple volume data, the volume feature data of the volume data is input into the pre-trained classification model to obtain the classification result of the volume data. Accordingly, the classification results of each of the multiple volume data can be obtained. The classification result can indicate whether the volume data is the target volume data, that is, the classification result can indicate that the volume data is the target volume data or the volume data is not the target volume data.

[0141] The computing device can, in response to the classification result of each of the multiple volume data including the classification result indicating that the volume data is the target volume data, regard the volume data as the target volume data. For example, for any one of the multiple volume data, in response to the classification result of the volume data indicating that the volume data is the target volume data, the volume data can be regarded as the target volume data. In response to the classification result of the volume data indicating that the volume data is not the target volume data, the volume data is not regarded as the target volume data. Accordingly, the multiple volume data may not include the target volume data. The multiple volume data may also include one target volume data or multiple target volume data.

[0142] For the case including multiple target volume data, the following method can be used to write the target volume data to the second file stored in the secondary external memory.

[0143] As an implementation, the computing device can write the respective target volume data to multiple second files stored in the secondary external memory. Each of the multiple second files has a different target volume identifier. The target volume data can have a target volume identifier.

[0144] As another implementation, the computing device can determine one target volume data from the multiple target volume data. For example, one target volume data can be randomly determined from the multiple target volume data. The target volume data can be written to the second file with the target volume identifier stored in the secondary external memory.

[0145] 1.4.2. Retention Policy for Other Data

[0146] The computing device can, in response to the merge operation on other data in at least one first file except the target volume data, obtain the first merged data, and then write the first merged data to the target first file stored in the primary external memory. The target first file can be the first file in at least one first file. Optionally, the target first file can be a newly created file outside at least one first file. The first merged data can include at least two data with different volume identifiers.

[0147] Since other data in at least one first file except for the target volume data is volume data that does not meet at least one predetermined evaluation metric or volume data determined not to be target volume data using a pre-trained classification model, a merging operation is performed on the other data to obtain first merged data, and the first merged data is written to the target first file stored in the first-level external storage. Thus, the target first file is retained in the first-level external storage, and the amount of data stored in the first-level external storage is reduced through the merging operation.

[0148] In addition, the computing device can also configure a predetermined identifier associated with the target first file to reduce the probability of being repeatedly processed. The predetermined identifier can be used as metadata. The predetermined identifier can indicate at least one of the round in which the target first file has participated in the merging operation or the merging source, and does not participate in the merging operation for a predetermined number of times after this round. Not participating in the merging operation for a predetermined number of times after this round can mean not participating in multiple rounds of merging operations after this round. The predetermined number of times can be configured according to actual business requirements and is not limited herein. For example, the predetermined number of times can be 5. The round in which the target first file has participated in the merging operation can be the second round, the merging source can be the first file SSTableA and the first file SSTableB, and the predetermined number of times can be 3. Thus, the target first file can not participate in the merging operations of the third round, the fourth round, and the fifth round.

[0149] To better understand the splitting strategy for target volume data and the retention strategy for other data in the embodiments of the present application, the following will be described in conjunction with Figure 6 for illustration.

[0150] Figure 6 FIG. shows an application schematic diagram of the splitting strategy for target volume data and the retention strategy for other data according to the embodiments of the present application.

[0151] As Figure 6 shown, the first file can include 8 pieces of data. The data can include keys with volume identifiers. The volume identifiers respectively possessed by the 8 pieces of data are "volume identifier V1", "volume identifier V2", "volume identifier V1", "volume identifier V1", "volume identifier V1", "volume identifier V3", "volume identifier V2", and "volume identifier V1". The predetermined evaluation metric can include a predetermined data volume ratio of 0.6.

[0152] For "volume identifier V1", "volume identifier V2", and "volume identifier V3", volume data with "volume identifier V1", volume data with "volume identifier V2", and volume data with "volume identifier V3" are obtained. Since the data volume ratio of the volume data with "volume identifier V1" is 0.625. The data volume ratio of the volume data with "volume identifier V2" is 0.25. The data volume ratio of the volume data with "volume identifier V3" is 0.125.

[0153] Since the data volume ratio of the volume data with the "volume identifier V1" is greater than the predetermined data volume ratio, the volume data with the "volume identifier V1" is the target volume data. The "volume identifier V1" is the target volume identifier. The target volume data can be written to the second file corresponding to the target volume identifier stored in the secondary external storage.

[0154] In response to the merge operation for other data in the first file except the target volume data, merged data is obtained and written to the target first file stored in the primary external storage.

[0155] In addition, to further improve the query efficiency, when there are similar first files stored in the primary external storage, if a predetermined identifier associated with the target first file indicates not to participate in the merge operation, a merge operation can also be performed between the target first file and the similar first files to obtain second merged data. The second merged data can be written to other first files stored in the primary external storage. The similar first files can have a similarity greater than or equal to the predetermined similarity with the target first file. The predetermined similarity can be configured according to actual business requirements and is not limited here. The other first files can be the target first file. Optionally, the other first files can be newly created files other than the target first file. The second merged data can include at least two data with different volume identifiers respectively.

[0156] II. Separation layer

[0157] In the separation layer, each volume data can be stored independently, and the volume data can be stored using a B+ tree. The nth file corresponding to the target volume identifier stores multiple data sorted by multiple keys with the target volume identifier. The data stored in the nth file corresponding to the target volume identifier all have the same target volume identifier. The key ranges between different nth files do not overlap. In addition, the computing device can also configure at least one of the storage medium or compression policy adapted to the nth file. For example, the compression policy can include at least one of the following: high compression ratio policy or low compression ratio policy. The file identifier of the nth file can have the target volume identifier. Since the compression policy and storage medium can be configured according to actual business requirements, different requirements are met and scalability is improved.

[0158] 2.1. Merge operation between the (n - 1)th level external storage and the nth level external storage involved in the separation layer

[0159] When n is an integer greater than 2, the computing device can, in response to the merge operation between the (n - 1)th level external storage and the nth level external storage, obtain the (n - 1)th target volume data with the target volume identifier among multiple volume identifiers from at least one (n - 1)th file and obtain the nth target volume data with the target volume identifier from at least one nth file.

[0160] The (n-1)th file can be stored in the (n-1)th level external storage. The nth file can be stored in the nth level external storage. There can be a key overlap range between the (n-1)th target volume data and the nth target volume data. For example, the key range of the (n-1)th target volume data can be the (n-1)th key range. The key range of the nth target volume data can be the nth key range. There is a key overlap range between the (n-1)th key range and the nth key range.

[0161] The computing device can obtain the target volume data with the target volume identifier based on the (n-1)th target volume data and the nth target volume data with the target volume identifier. For example, the computing device can perform a merging operation on the (n-1)th target volume data and the nth target volume data with the target volume identifier to obtain the target volume data with the target volume identifier.

[0162] Thus, the computing device can write the target volume data with the target volume identifier to the nth file corresponding to the target volume identifier stored in the nth level external storage.

[0163] Figure 7 The flowchart of the data query method according to the embodiment of the present application is shown.

[0164] As Figure 7 described, the method includes operations S710 to S720.

[0165] In operation S710, a query instruction is obtained.

[0166] In operation S720, in response to the query value matching the key to be queried not being queried from the active memory table and the non-active memory table stored in the memory, and at least one first file stored in the first level external storage, the query value matching the key to be queried is queried from the nth external storage file corresponding to the volume identifier to be queried stored in the nth level external storage.

[0167] n can be an integer greater than 1. In the data query method according to the embodiment of the present application, the data stored in the non-active memory table can be different from the data stored in the active memory table, that is, there are different parts between the data stored in the non-active memory table and the data stored in the active memory table. The active memory table can refer to the memory table that is currently in a modifiable state. The non-active memory table can refer to the memory table that is currently in a read-only state. The non-active memory table is not corresponding to the active memory table, that is, the non-active memory table is not obtained based on the active memory table.

[0168] The query instruction can include the key to be queried with the volume identifier to be queried. The computing device can determine whether the query value matching the key to be queried is queried in the order of the memory, the first level external storage to the last level external storage until the query value matching the key to be queried is queried or the query value matching the key to be queried is not queried from the last level external storage.

[0169] For example, a computing device may determine whether a query value matching the query key can be retrieved from an active memory table stored in memory based on the query key to be queried. In response to retrieving a query value matching the query key from the active memory table, a query result including the query value is obtained. In response to not retrieving a query value matching the query key from the active memory table, it may be determined based on the query value whether a query value matching the query key can be retrieved from an inactive memory table stored in memory. In response to retrieving a query value matching the query key from the inactive memory table, a query result including the query value is obtained. In response to not retrieving a query value matching the query key from the inactive memory table, it may be determined based on the query value whether a query value matching the query key can be retrieved from at least one first file stored in an external memory at the first level. In response to retrieving a query value matching the query key from at least one first file, a query result including the query value is obtained.

[0170] In response to not retrieving a query value matching the query key from at least one file, it is determined based on the query key whether a query value matching the query key can be retrieved from the nth external memory file corresponding to the query volume identifier stored in the external memory at the nth level. In response to retrieving a query value matching the query key from the nth external memory file corresponding to the query volume identifier stored in the external memory at the nth level, a query result including the query value is obtained. In response to not retrieving a query value matching the query key from the nth external memory file corresponding to the query volume identifier stored in the external memory at the nth level, n can be incremented by 1, and the process of determining based on the query key whether a query value matching the query key can be retrieved from the nth external memory file corresponding to the query volume identifier stored in the external memory at the nth level is repeated until a query value matching the query key is retrieved or a query value matching the query key is not retrieved from the external memory at the last level.

[0171] According to an embodiment of the present application, since the nth file stores data in a volume-based storage manner, therefore, based on the query key, a query value matching the query key can be retrieved from the nth external memory file corresponding to the query volume identifier stored in the external memory at the nth level, without having to query the nth file with other volume identifiers. Therefore, the query accuracy is improved, the query efficiency is increased, and the input / output overhead is reduced.

[0172] The above describes the data processing method and data query method provided by the embodiments of the present application.

[0173] Based on the above, the data processing method and data query method provided by the embodiments of the present application are described as a whole below.

[0174] In the embodiments of the present application, through the intelligent hierarchical design of the hybrid write layer combined with the separate storage layer, while achieving a large write throughput, the read performance is optimized, and the storage overhead is reduced. This can be reflected in the following aspects: Since data classification is not required in the hybrid layer, the write is faster, saving the classification overhead and increasing the write throughput. Since directed access to the data of the volume to be queried can be achieved in the separate layer, the input / output overhead is reduced. The space amplification is reduced through the dynamic splitting strategy, saving storage space. By adaptively adjusting the predetermined evaluation metrics to meet the data of the target volume according to the load information, the operation and maintenance are made more intelligent, which is suitable for large-scale production environments.

[0175] Regarding the splitting strategy for the data of the target volume, the first predetermined data volume and the predetermined quantity ratio are used to reduce the probability of occurrence of the storage fragmentation problem caused by premature splitting of highly mixed small data blocks. For example, the first predetermined data volume can be used to reduce the probability of splitting of small data blocks, and the predetermined data volume ratio can be used to reduce the probability of storage fragmentation. Regarding the retention strategy for other data, by retaining other data in the first-level external memory and participating in the subsequent round of merging operations when the conditions are met, the amount of data stored in the first-level external memory is reduced.

[0176] Based on the same inventive concept as the embodiments of the foregoing data processing method, the embodiments of the present application also provide a data processing device to implement the data processing method provided by the embodiments of the present application.

[0177] Figure 8 The block diagram of the data processing device according to the embodiments of the present application is shown.

[0178] As Figure 8 shown, the device 800 may include a first write module 810, an acquisition module 820, and a second write module 830.

[0179] The first write module 810 is configured to, in response to a write instruction or a disk flushing instruction, write the data to be stored indicated by the write instruction to the active memory table stored in the memory, or write the data to be stored indicated by the disk flushing instruction to the first file stored in the first-level external memory.

[0180] The acquisition module 820 is configured to, in response to the merge operation between the (n - 1)-th level external memory and the n-th level external memory, obtain the target volume data having the target volume identifier among the multiple volume identifiers.

[0181] The second write module 830 is configured to write the target volume data having the target volume identifier to the n-th file corresponding to the target volume identifier stored in the n-th level external memory.

[0182] According to an embodiment of the present application, the active memory table can store multiple data sorted by multiple keys. The data can include keys with volume identifiers. The first file corresponding to the active memory table can include at least one data for each of the multiple volume identifiers. n can be an integer greater than 1.

[0183] Based on the same inventive concept as the embodiment of the foregoing data query method, an embodiment of the present application further provides a data query device to implement the data query method provided by the embodiment of the present application.

[0184] Figure 9 A block diagram of a data query device according to an embodiment of the present application is shown.

[0185] As Figure 9 shown, the device 900 may include an acquisition module 910 and a query module 920.

[0186] The acquisition module 910 is configured to acquire a query instruction.

[0187] The query module 920 is configured to, in response to not querying a query value matching the key to be queried from the active memory table and the inactive memory table stored in the memory, and at least one first file stored in the first-level external memory, query a query value matching the key to be queried from the nth external memory file corresponding to the volume identifier to be queried and stored in the nth-level external memory.

[0188] According to an embodiment of the present application, the query instruction may include a key to be queried with a volume identifier to be queried. n can be an integer greater than 1.

[0189] According to an embodiment of the present application, any multiple modules among the first writing module 810, the obtaining module 820, and the second writing module 830 may be combined and implemented in one module, any multiple modules among the obtaining module 910 and the query module 920 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present application, at least one of the first writing module 810, the obtaining module 820, and the second writing module 830 may be at least partially implemented as a hardware circuit, and at least one of the obtaining module 910 and the query module 920 may be at least partially implemented as a hardware circuit. For example, a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way such as hardware or firmware that can integrate or package circuits can be used for implementation, or it can be implemented in any one of the three implementation manners of software, hardware, and firmware or in an appropriate combination of any several of them. Alternatively, at least one of the first writing module 810, the obtaining module 820, and the second writing module 830 may be at least partially implemented as a computer program module, and at least one of the obtaining module 910 and the query module 920 may be partially implemented as a computer program module. When the computer program module runs, it can execute the corresponding functions.

[0190] Figure 10 The block diagram of an electronic device suitable for implementing a data processing method or a data query method according to an embodiment of the present application is shown.

[0191] As Figure 10 shown, the electronic device 1000 according to an embodiment of the present application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a ROM 1002 (ROM is a read-only memory) or a program loaded from a storage section 1008 into a RAM 1003 (RAM is a random access memory). The processor 1001 may include, for example, a general microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), etc. The processor 1001 may also include on-board memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0192] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via the bus 1004. The processor 1001 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the programs may also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 may also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.

[0193] According to an embodiment of the present application, the electronic device 1000 may further include an I / O interface 1005 (I / O stands for input / output), and the I / O interface 1005 is also connected to the bus 1004. The electronic device 1000 may further include one or more of the following components connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, etc.; an output portion 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as a LAN card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage portion 1008 as needed.

[0194] The present application also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0195] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the above-described ROM 1002 and / or RAM 1003 and / or one or more memories other than ROM 1002 and RAM 1003.

[0196] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the data processing method or data query method provided by the embodiment of the present application.

[0197] When the computer program is executed by the processor 1001, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0198] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 1009, and / or be installed from the removable medium 1011. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0199] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or be installed from the removable medium 1011. When the computer program is executed by the processor 1001, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0200] In accordance with embodiments of the present application, program code for executing the computer programs provided by the embodiments of the present application can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0201] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0202] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.

[0203] The above describes the embodiments of the present application. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.

Claims

1. A data processing method, characterized in that, Including: In response to a write instruction or a disk flushing instruction, writing the data to be stored indicated by the write instruction into the active memory table stored in the memory, or writing the data to be stored indicated by the disk flushing instruction into the first file stored in the first-level external memory, wherein the active memory table stores multiple data sorted by multiple keys, the data includes keys with volume identifiers, and the first file corresponding to the active memory table includes at least one data for each of the multiple volume identifiers; and In response to a merge operation between the (n-1)-th level external memory and the n-th level external memory, obtaining target volume data with a target volume identifier among the multiple volume identifiers, and writing the target volume data with the target volume identifier into the n-th file corresponding to the target volume identifier stored in the n-th level external memory, where n is an integer greater than 1.

2. The method according to claim 1, wherein The obtaining target volume data with a target volume identifier among the multiple volume identifiers in response to a merge operation between the (n-1)-th level external memory and the n-th level external memory includes: In response to a merge operation between the first-level external memory and the second-level external memory, obtaining at least one data for each of the multiple volume identifiers from at least one of the first files, to obtain volume data for each of the multiple volume identifiers; and In response to the volume data including volume data that meets at least one predetermined evaluation metric, using the volume data that meets the at least one predetermined evaluation metric as the target volume data, or determining the target volume data from the multiple volume data using a pre-trained classification model according to the multiple volume data.

3. The method according to claim 2, wherein The using the volume data that meets the at least one predetermined evaluation metric as the target volume data in response to the volume data including volume data that meets at least one predetermined evaluation metric includes: Evaluating the multiple volume data according to the at least one predetermined evaluation metric to obtain evaluation results for each of the multiple volume data, wherein the evaluation results indicate whether the volume data meets the at least one predetermined evaluation metric; and In response to the evaluation results for each of the multiple volume data including an evaluation result indicating that the volume data meets the at least one predetermined evaluation metric, using the volume data as the target volume data.

4. The method according to claim 3, wherein The at least one predetermined evaluation metric includes at least one of the following: a first predetermined data volume, a predetermined data volume ratio, or a predetermined access frequency; The evaluation result of the target volume data indicates at least one of the following: the data volume of the target volume data is greater than or equal to the first predetermined data volume, the data volume ratio of the target volume data is greater than or equal to the predetermined data volume ratio, or the access frequency of the target volume data within a predetermined time period is greater than or equal to the predetermined access frequency, and the data volume ratio of the target volume data is the ratio between the data volume of the target volume data and the total data volume of the multiple volume data.

5. The method according to claim 4, characterized in that, The method further includes: Adjust at least one of the first predetermined data volume, the proportion of the predetermined data volume, or the predetermined access frequency according to at least one of the load information or the inter - volume data distribution entropy, where the load information includes at least one of the following: the data volume stored in the first - level external memory, the first resource utilization rate, or the first input - output delay, and the inter - volume data distribution entropy indicates the degree of uniformity of the distribution of the data among multiple volume data.

6. The method according to claim 5, wherein The adjusting at least one of the first predetermined data volume, the proportion of the predetermined data volume, or the predetermined access frequency according to at least one of the load information or the inter - volume data distribution entropy includes: In response to at least one of the data volume stored in the first - level external memory being greater than or equal to a second predetermined data volume, the inter - volume data distribution entropy being greater than or equal to a predetermined distribution entropy, the first resource utilization rate being greater than or equal to a predetermined utilization rate, or the first input - output delay being greater than or equal to a predetermined delay, reduce at least one of the first predetermined data volume, the proportion of the predetermined data volume, or the predetermined access frequency; and / or In response to at least one of the data volume stored in the first - level external memory being less than the second predetermined data volume, the inter - volume data distribution entropy being less than the predetermined distribution entropy, the first resource utilization rate being less than the predetermined utilization rate, or the first input - output delay being less than the predetermined delay, increase at least one of the first predetermined data volume, the proportion of the predetermined data volume, or the predetermined access frequency.

7. The method according to claim 2, wherein The determining the target volume data from multiple volume data according to the multiple volume data by using a pre - trained classification model includes: Obtain the volume feature data of each of the multiple volume data according to the multiple volume data and system data, where the volume feature data includes at least one of the following: the data volume included in the volume data, the proportion of the data volume of the volume data, the second resource utilization rate, or the second output - input delay, and the proportion of the data volume of the volume data is the ratio of the data volume of the volume data to the total data volume of the multiple volume data; Input the volume feature data of each of the multiple volume data into the pre - trained classification model respectively to obtain the classification result of each of the multiple volume data, where the classification result indicates whether the volume data is the target volume data; and In response to the classification results of each of the multiple volume data including a classification result indicating that the volume data is the target volume data, regard the volume data as the target volume data.

8. The method according to claim 7, wherein The pre - trained classification model is obtained by training an initial classification model with the sample volume feature data of each of the multiple sample volume data, and the sample volume feature data includes at least one of the following: the data volume included in the sample volume data, the proportion of the data volume of the sample volume data, the third resource utilization rate, or the third output - input delay, and the proportion of the data volume of the sample volume data is the ratio of the data volume of the sample volume data to the total data volume of the multiple sample volume data.

9. The method according to any one of claims 2 to 8, characterized in that, The method further includes: In response to a merge operation on other data in at least one of the first files except the target volume data, obtain merged data; and Write the merged data to a target first file stored in the first - level external memory, where the target first file is the first file among at least one of the first files or the target first file is a newly created file outside at least one of the first files.

10. The method according to claim 9, wherein The method further includes: Configure a predetermined identifier associated with the target first file, where the predetermined identifier indicates at least one of the round in which the target first file has participated in the merge operation or the merge source, and does not participate in the merge operation for a predetermined number of times after the round.

11. The method according to claim 10, wherein The method further includes: In response to a similar first file stored in the first - level external memory having a similarity greater than or equal to a predetermined similarity with the target first file, when the predetermined identifier associated with the target first file indicates non - participation in the merge operation, perform a merge operation between the target first file and the similar first file.

12. The method according to any one of claims 1 to 8, characterized in that, In response to a disk - flushing instruction, write the data to be stored indicated by the disk - flushing instruction to a first file stored in the first - level external memory, including: Use a disk - flushing thread to obtain, in response to the disk - flushing instruction, the data to be stored indicated by the disk - flushing instruction from an inactive memory table stored in the memory, where the storage data and storage method of the inactive memory table are the same as those of the active memory table; and Use the disk - flushing thread to write the data to be stored indicated by the disk - flushing instruction to a first file stored in the first - level external memory.

13. The method according to claim 12, wherein The method further includes: Use a writing thread to configure the status identifier of the active memory table as a read - only identifier when the amount of data stored in the active memory table is equal to a third predetermined data amount, obtaining the inactive memory table.

14. The method according to claim 12, wherein The method further includes: Use a writing thread to store the inactive memory table in a dynamic array; and Use the disk - flushing thread to remove the inactive memory table from the dynamic array in response to completion of the disk - flushing operation for the inactive memory table.

15. The method according to any one of claims 1 to 8, characterized in that, In response to a writing instruction, write the data to be stored indicated by the writing instruction to an active memory table stored in the memory, including: Obtain the writing instruction, where the writing instruction includes the data to be stored, and the data to be stored includes a value to be stored and a key to be stored with a volume identifier to be stored; Determine a global write position from the active memory table in the memory according to the key to be stored; and Write the data to be stored indicated by the writing instruction to the global write position.

16. The method according to claim 15, wherein The writing the data to be stored to the global write position includes: Use an atomic operation to write the data to be stored to the global write position.

17. The method according to any one of claims 1 to 8, characterized in that The obtaining the target volume data having a target volume identifier among multiple volume identifiers in response to a merge operation between the (n - 1) - level external memory and the n - level external memory includes: In response to a merge operation between the (n - 1)-th level external memory and the n-th level external memory, obtain the (n - 1)-th target volume data with the target volume identifier among multiple said volume identifiers from at least one (n - 1)-th file and obtain the n-th target volume data with the target volume identifier from at least one n-th file, where the (n - 1)-th file is stored in the (n - 1)-th level external memory, the n-th file is stored in the n-th level external memory, and n is an integer greater than 2; and Obtain the target volume data with the target volume identifier based on the (n - 1)-th target volume data and the n-th target volume data with the target volume identifier.

18. The method according to any one of claims 1 to 8, characterized in that, Further includes: Configure at least one of a storage medium or a compression policy adapted to the n-th file, where the file identifier of the n-th file has the target volume identifier.

19. A data query method, characterized in that, Includes: Obtain a query instruction, where the query instruction includes a query key with a volume identifier to be queried; And In response to not querying a query value matching the query key from an active memory table and an inactive memory table stored in memory, and at least one first file stored in the first-level external memory according to the query key, query a query value matching the query key from an n-th external memory file corresponding to the volume identifier to be queried and stored in the n-th level external memory, where n is an integer greater than 1.

20. A data processing device, characterized in that, Includes: A first writing module configured to, in response to a writing instruction or a disk flushing instruction, write the data to be stored indicated by the writing instruction to an active memory table stored in memory, or write the data to be stored indicated by the disk flushing instruction to a first file stored in the first-level external memory, where the active memory table stores multiple data sorted by multiple keys, the data includes keys with volume identifiers, and the first file corresponding to the active memory table includes at least one data for each of the multiple volume identifiers; and An obtaining module configured to, in response to a merge operation between the (n - 1)-th level external memory and the n-th level external memory, obtain target volume data with the target volume identifier among multiple said volume identifiers; and a second writing module configured to write the target volume data with the target volume identifier to an n-th file corresponding to the target volume identifier and stored in the n-th level external memory, where n is an integer greater than 1.

21. A data query device, characterized in that, Includes: An obtaining module configured to obtain a query instruction, where the query instruction includes a query key with a volume identifier to be queried; And A query module configured to, in response to not querying a query value matching the query key from an active memory table and an inactive memory table stored in memory, and at least one first file stored in the first-level external memory according to the query key, query a query value matching the query key from an n-th external memory file corresponding to the volume identifier to be queried and stored in the n-th level external memory, where n is an integer greater than 1.

22. An electronic device, characterized in that, Includes: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 18 or the steps of the method according to claim 19.

23. A computer-readable storage medium, characterized in that, A computer program or instructions are stored thereon, characterized in that when the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 18 or the steps of the method according to claim 19 are implemented.

24. A computer program product, characterized in that, It includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 18 or the steps of the method according to claim 19 are implemented.

Citation Information

Patent Citations

  • Database management method and database system

    CN108319602A

  • Data storage method and device based on LSM, storage medium and computer equipment

    CN111352908A

  • Log merge tree key value storage system and related method and related equipment

    CN114840134A

  • Key value storage method and device based on key value separation and hybrid storage medium

    CN119046289A

  • Method for partitioning memory, electronic device, and storage medium

    WO2024235184A1