Method and apparatus for compressing a directory of extendible hash with raw key support
Patent Information
- Application Number
- KR1020250158905
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-10-29
Smart Images

Figure 112025120582155-PAT00011_ABST
Abstract
Description
Technology Field
[0001] Below, a compression technique for a low-key supported scalable hash table is provided. Background Technology
[0003] Computer systems need to select a database with an appropriate data structure for processing data according to a specific workload. Databases can include hierarchical databases, relational databases, object-oriented databases, XML databases, and Key-Value Stores (KVS). Among these, a Key-Value Store is a type of non-relational database that stores data using simple key-value methods and can also be referred to as a NoSQL database. Examples of Key-Value Stores include Dynamo, Level DB, Rocks DB, Cassandra, and HB BASE. Computer systems utilizing this approach can store data in the Key-Value Store as a collection of key-value pairs where the key serves as a unique identifier. Due to their high partitionability in data processing, these Key-Value Stores can manage data more flexibly and efficiently than relational databases.
[0004] The background technology described above is possessed or acquired by the inventor in the process of deriving the content of the disclosure of the present application, and cannot necessarily be considered as prior art disclosed to the general public prior to the filing of this application. Prior art literature
[0006] Korean Registered Patent Publication No. 2085132 (published February 28, 2020) presents an efficient Kuku hash method using classification of hash functions within a bucket. Korean Registered Patent Publication No. 2189398 (published December 4, 2020) presents a method and system for efficient range search using locality-maintaining hashing.
[0007] A method for compressing a directory of an extensible hash capable of supporting low keys according to one embodiment may include the steps of: linearly searching through all items of the directory and generating a directory bit string that encodes bits according to whether the searched item was previously pointed to; performing directory encoding that changes the bit representation of the directory bit string into an integer representation; and generating a compressed directory that compresses the directory based on the directory bit string in which the directory encoding was performed so as to represent only information about the actual existing bucket.
[0008] The step of generating the directory bit string may include generating a directory bit string in which the bit is encoded as '1' if the searched item points to a bucket that has not been pointed to before, or the bit is encoded as '0' if the searched item points to a bucket that has been pointed to before.
[0009] The above compressed directory may have an index where the number of '1' representations of the counted directory bit string is an index.
[0010] The step of performing the above directory encoding is, when '0' appears 64 or more times consecutively in the bit representation of the directory bit string, converting the bit representation for a predetermined number of '0's into a 64-bit integer representation ( It may include a step of generating a bit string.
[0011] The above integer representation can be identified by using the preceding 3-bit sequence of the above integer representation as a mark.
[0012] The above 3-bit sequence may be '001'.
[0013] The above method may further include the step of expanding the compressed directory based on whether the local depth of the bucket where the overflow occurred is the same as the predetermined global depth when an overflow occurs in any one of at least one bucket connected to the compressed directory after the compressed directory is created.
[0014] The step of expanding the compressed directory may include, if the predetermined global depth and the local depth of the bucket are not the same, the step of creating an additional bucket that divides the information about the bucket where the overflow occurred, the step of creating an additional compressed directory in the compressed directory that points to the additional bucket, and the step of changing the information of the directory bit string at the location corresponding to the additional compressed directory; or, if the predetermined global depth and the local depth of the bucket are the same, the step of creating an additional bucket that divides the information about the bucket where the overflow occurred, the step of creating an additional compressed directory in the compressed directory that points to the additional bucket, the step of increasing the global depth to increase the number of the directory bit strings, and the step of changing the information of the increased directory bit string at the location corresponding to the additional compressed directory.
[0015] An apparatus for compressing a directory of a scalable hash capable of supporting low keys according to one embodiment may include a processor that linearly searches all items of the directory and generates a directory bit string that encodes bits according to whether the searched item has been previously pointed to, performs directory encoding that changes the bit representation of the directory bit string into an integer representation, and generates a compressed directory that compresses the directory so as to represent only information about the actual existing bucket based on the directory bit string in which the directory encoding was performed.
[0016] The processor may generate a directory bit string in which, when generating the directory bit string, if the searched item points to a bucket that has not previously pointed to, the bit is encoded as '1', or if the searched item points to a bucket that has previously pointed to, the bit is encoded as '0'.
[0017] The above compressed directory may have an index where the number of '1' representations of the counted directory bit string is an index.
[0018] In performing the directory encoding, the processor, when '0' appears 64 or more times consecutively in the bit representation of the directory bit string, converts the bit representation for a predetermined number of '0's into a 64-bit integer representation ( It can generate a bit string.
[0019] The above integer representation can be identified by using the preceding 3-bit sequence of the above integer representation as a mark.
[0020] The above 3-bit sequence may be '001'.
[0021] The processor may expand the compressed directory after the compressed directory is created, if an overflow occurs in any one of at least one bucket connected to the compressed directory, depending on whether the predetermined global depth and the local depth of the bucket where the overflow occurred are the same.
[0022] When expanding the compressed directory, the processor may, if the predetermined global depth and the local depth of the bucket are not the same, create an additional bucket that divides the information about the bucket where the overflow occurred, create an additional compressed directory that points to the additional bucket in the compressed directory, and change the information of the directory bit string at the location corresponding to the additional compressed directory, or if the predetermined global depth and the local depth of the bucket are the same, create an additional bucket that divides the information about the bucket where the overflow occurred, create an additional compressed directory that points to the additional bucket in the compressed directory, increase the global depth to increase the number of the directory bit strings, and change the information of the increased directory bit string at the location corresponding to the additional compressed directory. Brief explanation of the drawing
[0024] FIG. 1 illustrates a block diagram of a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 2 illustrates, exemplarily, the structure of a virtual component within a device that compresses a directory of scalable hashes capable of low-key support according to one embodiment. FIG. 3 exemplarily illustrates an index of a compressed directory in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 4 exemplarily illustrates a directory bit string structure and a naive bit string structure in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 5 schematically illustrates a naive bit string structure with added marks in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 6 schematically illustrates a decoding operation in a device that compresses a directory of scalable hashes capable of low-key support according to one embodiment. FIG. 7 exemplarily illustrates the operation of expanding a directory bit string in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 8 exemplarily illustrates the operation of doubling a directory bit string in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment. FIG. 9 illustrates a flowchart of a method for compressing a directory of scalable hashes capable of low-key support according to one embodiment. Specific details for implementing the invention
[0025] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, actual implementations are not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or substitutions included in the technical concept described by the embodiments.
[0026] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.
[0027] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or coupled with that other component, or that there may be other components in between.
[0028] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0029] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.
[0030] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.
[0031] For in-memory data management systems, such as in-memory databases and key-value stores, the efficiency of the index structure is critical, as it can significantly impact the system's final latency. Furthermore, even in disk-based data management systems, the index structure (e.g., the software layer) is a critical factor due to the continuous advancement of high-speed storage systems.
[0032] Index structures such as B+ trees and hash indexes are adopted in data management systems to support the three major operations: searching, inserting, and scanning. However, hash indexes have the problem that while they are very effective for point queries searching for individual keys, they are not suitable for range queries querying key ranges. In contrast, B+ trees have the problem that while they are effective for range queries, they are not suitable for point queries.
[0033] Learned indexes, another index structure, excel for all three major tasks but may require training in the form of bulk loading. Depending on the characteristics of complex and dynamic real-world training datasets, determining how much to train and when to retrain can become a difficult obstacle to overcome in bulk loading training. These obstacles can lead to performance degradation in learned indexes.
[0034] In addition, the Extendable Hash in a hash index may include a dynamic indexing structure that can expand and contract according to changes in workload. Accordingly, the Extendable Hash has the characteristic of preventing frequent re-hashing caused by collisions in the hash index while still providing fast O(1) search speeds. Due to these characteristics, the Extendable Hash is widely used in various fields, such as database and storage systems. The Extendable Hash operates based on a directory, and a directory label can be determined using some bits of the result of a hash function. Each entry in the directory contains a pointer pointing to a segment, and the size of the directory can always be maintained in units of powers of 2. A single segment can contain multiple buckets. Additionally, a single segment has a fixed number of entries, and when all entries are full, the segment is split to handle collisions. During this process, doubling may occur as needed, causing the directory to expand to twice its size. In this case, a structure emerges where multiple directory entries share a single segment, which can lead to pointer waste.
[0035] Meanwhile, because hash indexes index by transforming raw keys using a hash function, they can evenly redistribute the distribution of key values and mitigate key concentration in specific ranges. However, since this structure does not preserve the original order of keys, it has the problem of not being able to efficiently process range queries. As a solution to this problem, a method that uses raw keys directly without passing them through a hash function can be proposed. However, in this case, the likelihood of key distribution being concentrated in specific ranges increases, leading to data concentration in specific segments. Consequently, this can induce repetitive segment splitting and directory doubling. As this abnormally increases the directory size, it can eventually exceed the limits of system memory (DRAM), potentially leading to out-of-memory (OOM) errors and system failures. Therefore, hash indexes and scalability hashes contain structural limitations, such as the inefficient use of directory pointers, excessive memory usage due to uneven key distribution, and the lack of support for range queries.
[0036] Accordingly, an apparatus and method for compressing a directory of an extensible hash capable of supporting low keys according to one embodiment generates a directory bit string that encodes bits for the directory, performs directory encoding that changes the representation of the directory bit string, and generates a compressed directory that compresses the directory, thereby preventing excessive directory expansion and memory exhaustion problems and maintaining system stability, and supporting range queries even in an environment where the distribution of keys is not uniform.
[0037] FIG. 1 illustrates a block diagram of a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0038] The device (100) may include a processor (110) and a memory (120).
[0039] The processor (110) can linearly search through all items (e.g., entries) of the directory of the Extendible Hash and generate a directory bit string with bits encoded according to whether the searched item was previously pointed to. The (logical) directory of the Extendible Hash and the directory bit string may correspond to virtual components. In contrast, the directory bit string and compressed directory, which have undergone directory encoding as described below, may correspond to actual components. Virtual components refer to conceptual structures that exist only logically and are not physically stored in memory, while actual components refer to structures that are actually stored in memory. A detailed description of the directory bit string is provided later in FIG. 2.
[0040] The processor (110) can perform directory encoding that changes the bit representation of a directory bit string into an integer representation. This may mean that, for memory efficiency, the processor (110) performs additional encoding based on duplicate patterns in the directory bit string. A detailed description of directory encoding is provided later in FIG. 4.
[0041] The processor (110) can generate a compressed directory by compressing the directory based on a directory bit string in which directory encoding has been performed, so that only information about the actual existing buckets is represented. The compressed directory may store unique pointers to buckets that store key-value pairs in an array form. The pointers may be 8 bytes. A detailed description of the compressed directory is provided later in FIG. 7.
[0042] As such, the processor (110) operates based on a scalable hash structure and may operate in a single-threaded environment. This implies that for the processor (110), a single-threaded environment is an optimized environment compared to a multi-threaded environment, where potential consistency issues may occur. Additionally, the operation of the processor (110) may be performed in volatile memory where key-value pairs are stored, and / or in a storage device with a large capacity distinct from the volatile memory. The volatile memory and / or storage device may be implemented within the same package as the processor (110), but may also be implemented in a different package. Volatile memory may represent a memory device that stores data only when power is supplied and loses the stored data when the power supply is cut off. For example, the volatile memory may include at least one of Static Random Access Memory (SRAM) and Dynamic Random Access Memory (DRAM). Additionally, the storage (130) may include at least one of a solid state drive (SSD) and a hard disk drive (HDD).
[0043] According to various embodiments, the processor (110) may execute computer-readable code (e.g., software) stored in memory (120) and instructions triggered by the processor (110). The processor (110) may be a data processing device implemented in hardware having a circuit having a physical structure for executing desired operations. Desired operations may include, for example, code or instructions included in a program. The data processing device implemented in hardware may include, for example, a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an Application-Specific Integrated Circuit (ASIC), or a Field Programmable Gate Array (FPGA).
[0044] The memory (120) can store at least one instruction (e.g., a program) that can be executed by the processor (110).
[0045] FIG. 2 illustrates, exemplarily, the structure of a virtual component within a device that compresses a directory of scalable hashes capable of low-key support according to one embodiment.
[0046] The scalable hash table (210) may include a directory (211) and a bucket (212). The directory (211) may be represented as an array in which a prefix (e.g., Most Significant Bit (MSB)) or suffix (e.g., Least Significant Bit (LSB)) of an input key is used as a label (or index). At least one label of the directory (211) may be linked to one bucket (212). The number of most significant bits or least significant bits of the labels may be determined based on the Global Depth (GD). For example, the directory (211) has a Global Depth of 3, and the number of labels in the array is 2 3 There are 8, and each label can be represented by three bits.
[0047] The processor (110) can linearly traverse from the first item to the last item (e.g., MSB to LSB) of the extensible hash directory (211) (e.g., logical directory). The processor (110) can encode a bit as '1' (one) if the traversed item points to a bucket that has never been pointed to before, and encode a bit as '0' (zero) if the traversed item points to a bucket that has been pointed to before. The directory bit string (220) encoded as '1' or '0' can be 8 bits since there are 8 items. If the global depth is 4, the directory bit string (220) can be 16 bits since there are 16 items.
[0048] For example, the first label 000 pointing to bucket 1 (B1) of the directory (211) is the first entry and has never been pointed to before, so the first entry of the directory bit string (220) can be encoded as '1'. Also, the second label 001 pointing to bucket 1 (B1) points to the same bucket 1 (B1) as the first entry, so the second entry of the directory bit string (220) can be encoded as '0'. The third label 010 and the fourth label 011 each point to only one bucket (e.g., the second bucket (B2) and the third bucket (B3)), so the third and fourth entries of the directory bit string (220) can each be encoded as '1'. Since the fifth label, 100, pointing to bucket 4 (B4) is the first to point to bucket 4 (B4), the fifth entry of the directory bit string (220) can be encoded as '1'. Also, since the sixth label, 101, the seventh label, 110, and the eighth label, 111, pointing to bucket 4 (B4) are the same bucket 4 (B4) as the fifth label, 100, the sixth, seventh, and eighth entries of the directory bit string (220) can be encoded as '0'.
[0049] FIG. 3 exemplarily illustrates an index of a compressed directory in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0050] The processor (110) can find an index (331) of the compressed directory (330) based on the global depth (G) and a given key. First, the processor (110) finds an i (0) corresponding to the given key in the (logical) directory (310). <i≤2 GAn index (311) can be found. The processor (110) can count the number of occurrences of the bit representation '1' up to the i-th bit of the directory bit string (320). Accordingly, in the compressed directory (330), the number of '1' representations of the directory bit string (320) counted can be the compressed directory index (331).
[0051] For example, height is 100 In this case, the logical directory index (311) of index 100 in the (logical) directory (310) may be 5. The processor (110) may count the number of '1' representations in the directory bit string (320). Accordingly, the target compressed directory index (331) corresponding to the target logical directory index (311) of 5 may be 4. The bucket (e.g., bucket 4 (B4)) pointed to by the entry (e.g., 1XX) of the compressed directory (330) corresponding to the index (331) is the input key (e.g., 100 It may include key-value pairs corresponding to ).
[0052] FIG. 4 exemplarily illustrates a directory bit string structure and a naive bit string structure in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0053] Since the directory bit string (410) is a sparse representation having a long sequence of zero bits, the existing scalable hash structure can be represented with very little data capacity by compressing the same representation into a fixed number. The directory bit string (410) can be exemplified as having a total of 256 bits, assuming a global depth of 8. However, the numerical part exemplified is not limited to this and may be changed. When the processor (110) performs directory encoding, if '0' appears 64 or more times consecutively in the bit representation of the directory bit string (410), the bit representation for a predetermined number (e.g., 127) of '0's is converted into a 64-bit integer representation ( A bit string (420) can be generated. In contrast, the three '0's that follow can be represented in the bit string as they are without conversion to an integer representation because there are not 64 of them. For example, the bit representation for the 127 duplicate '0's occupying 127 bits in the directory bit string (410) can be represented as 1111111 in 64 bits (e.g., 127 in decimal). The number of bits in the directory bit string (410) is 256 bits, and the number of bits in the naive bit string (420) compressed into an integer representation can be 193 bits.
[0054] FIG. 5 schematically illustrates a naive bit string structure with added marks in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0055] The processor (110) may add a mark (511) within 64 bits so that an integer representation (512) is identified. Accordingly, the integer representation (512) can be identified by using the preceding 3-bit sequence of the integer representation (512) as the mark (511). Since the naive bit string doubles in size, the number of consecutive zeros cannot be even. Based on this, a number with an even number of short consecutive zeros (e.g., 001) that cannot appear in the bit representation following the first bit representation of the naive bit string, '1', can be used as the mark (511). Accordingly, the 3-bit sequence can be '001'. When decoding, the processor (110) can read the (64-3) bits following the 3-bit mark (511) of the 64-bit integer representation in the naive bit string as the integer representation (512).
[0056] FIG. 6 schematically illustrates a decoding operation in a device that compresses a directory of scalable hashes capable of low-key support according to one embodiment.
[0057] Decoding may refer to the process of finding the target compressed directory index (612) for the compressed directory using directory encoding to find the index of the bucket corresponding to a given key. In a naive bit string (610) containing marks and integer representations, the method of finding the compressed directory index (331) in FIG. 3 cannot be applied because the number of (logical) directories and the number of bits do not match 1:1. Accordingly, the naive bit string (610) may use the decoding method of FIG. 6. For example, if the index of the target (logical) directory is 131, the index (611) in the (logical) directory bit string may be counted by the number of integer representations (e.g., 1 or 127) rather than by the number of bits (e.g., 3 or 64). Accordingly, since the integer representation represents 127, the index for the integer representation may be 2-129. Therefore, the index (611) in the (logical) directory bit string of 131 may be the second '1' representation after the integer representation. The target compressed directory index (612) corresponding to the target logical directory index (611) of 131 may be 3.
[0058] FIG. 7 exemplarily illustrates the operation of expanding a directory bit string in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0059] The architecture implemented in the processor (110) may be a three-level structure including a directory bit string level where directory encoding is performed, a compressed directory level, and a bucket level. A directory bit string within the directory bit string level may represent the existence of each entry in the (logical) directory with 1 bit. A '0' in the directory bit string may mean that the entry does not exist in the compressed directory, and a '1' may mean that the entry exists in the compressed directory.
[0060] For example, the processor (110) can perform a decoding operation through a decoding algorithm implemented at the directory bit string level when a key is given. The processor (110) can check the existence of each entry by checking the location of the directory bit string corresponding to the given key through the decoding algorithm. Subsequently, if an entry for the given key exists, the processor (110) can find a pointer within the compressed directory and check the bucket pointed to by the pointer. This three-level structure improves the memory usage efficiency and stability of the system, and the decoding algorithm can support finding the desired bucket accurately without loss of compressed information.
[0061] A compressed directory can be compressed to correspond 1:1 with each bucket. Referring to FIG. 3, the processor (110) can compress two (logical) directories (310) for bucket 1 (B1) into a single compressed directory (330). Accordingly, if the labels of the (logical) directories (310) are 000 and 001, the label of the compressed directory (330) can be 00X. In FIG. 3, since the global depth is set to 3, the label can also be expressed as a three-digit number. Buckets 2 (B2) and 3 (B3) are already 1:1 corresponding, so no compression operation is necessary. The processor (110) can compress four (logical) directories (310) for bucket 4 (B4) into a single compressed directory (330). Accordingly, if the labels of the (logical) directory (310) are 100, 101, 110, and 111, the labels of the compressed directory (330) may be 1XX.
[0062] Additionally, after the compressed directory is created, the processor (110) may cause an overflow in any one of the buckets (710) connected to the compressed directory. At this time, the processor (110) may expand the compressed directory depending on whether the determined global depth is the same as the local depth of the bucket where the overflow occurred. The expansion of the compressed directory may include directory bit string expansion and directory bit string doubling. If the determined global depth is not the same as the local depth of the bucket, the processor (110) may increase only the local depth and perform a directory bit string expansion operation during the expansion operation of the compressed directory.
[0063] Overflow means that a bucket is full, so the processor (110) can create an additional bucket (712) (e.g., bucket 5 (B5)) that divides information (e.g., key-value pairs) about the bucket (710) where the overflow occurred. The processor (110) can create an additional compressed directory (721) (e.g., 11X) that points to the additional bucket (712) in the compressed directory. The processor (110) can change the information (731) of the directory bit string at the location corresponding to the additional compressed directory (721). The information (731) of the directory bit string can be changed from a '0' notation to a '1' notation.
[0064] FIG. 8 exemplarily illustrates the operation of doubling a directory bit string in a device for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment.
[0065] If the determined global depth and the local depth of the bucket are the same, the processor (110) can increase the global depth and also increase the local depth, and then perform a directory bit string doubling operation. Accordingly, the processor (110) can create an additional bucket (812) (e.g., Bucket 5 (B5)) that divides information about the bucket (810) where an overflow occurred. The processor (110) can create an additional compressed directory (821) (e.g., 0111) that points to the additional bucket (812) in the compressed directory. The processor (110) can increase the number of directory bit strings by increasing the global depth. In FIG. 8, the global depth is increased from 3 to 4, and the number of directory bit strings is increased from 8 to 16. The processor (110) can change the information (831) of the increased directory bit strings at the location corresponding to the additional compressed directory. The information (831) of the directory bit string can be changed from '0' notation to '1' notation.
[0066] Accordingly, the device (100) according to one embodiment can maximize memory efficiency by storing / managing only the pointers actually needed without duplication within the compressed directory. In addition, the device (100) can encode and decode the directory used in the existing extended hash into a concise bit string format. Through this, the device (100) can effectively mitigate the problem of directory space amplification that may occur when using the raw key as is.
[0067] FIG. 9 illustrates a flowchart of a method for compressing a directory of scalable hashes capable of low-key support according to one embodiment.
[0068] In step (910), the processor can linearly traverse all items in the directory and generate a directory bit string that encodes bits based on whether the traversed item was previously pointed to.
[0069] In step (920), the processor can perform directory encoding to change the bit representation of the directory bit string into an integer representation.
[0070] In step (930), the processor can create a compressed directory in which only information about the actual existing buckets is represented based on the directory bit string in which directory encoding has been performed.
[0071] In step (910), the processor may generate a directory bit string in which the bit is encoded as '1' if the searched item points to a bucket that has not been previously pointed to, or the bit is encoded as '0' if the searched item points to a bucket that has been previously pointed to. The compressed directory may be an index of the number of '1' representations in the counted directory bit string.
[0072] In step (920), the processor converts the bit representation of a predetermined number of '0's into a 64-bit integer representation (when '0' appears 64 or more times consecutively in the bit representation of the directory bit string) A bit string can be generated. For example, the number of bits in a directory bit string can be 256 bits, and the number of bits in a naive bit string can be 193 bits. However, this is not limited to this, and the number of bits in each bit string can have different bit values. An integer representation can be identified by the preceding 3-bit sequence of the integer representation as a mark. The 3-bit sequence can be '001'.
[0073] After step (930), the processor may expand the compressed directory if an overflow occurs in any one of at least one bucket associated with the compressed directory, depending on whether the determined global depth and the local depth of the bucket where the overflow occurred are the same.
[0074] When expanding a compressed directory, the processor may perform directory bit string expansion or directory bit string doubling. When expanding a directory bit string, if the determined global depth and the local depth of the bucket are not the same, the processor may create an additional bucket by dividing the information about the bucket where the overflow occurred. Subsequently, the processor may create an additional compressed directory in the compressed directory that points to the additional bucket, and may change the information of the directory bit string at the location corresponding to the additional compressed directory.
[0075] In directory bit string doubling, if the determined global depth and the local depth of the bucket are the same, the processor may create an additional bucket by dividing the information about the bucket where the overflow occurred, and add a directory pointing to the additional bucket to the compressed directory. Afterwards, the processor may increase the global depth to increase the number of directory bit strings and change the information of the increased directory bit strings at the location corresponding to the additional bucket.
[0076] Accordingly, a method for compressing a directory of scalable hashes capable of supporting low keys according to one embodiment can support range queries while maintaining the advantages of existing scalable hash structures to the fullest extent.
[0077] Furthermore, the compression rate of the method for compressing a directory of an extensible hash capable of supporting low keys according to one embodiment achieves a high level of compression rate. This can be confirmed by the result of a reduction in memory usage ranging from a minimum of 50.00% to a maximum of 99.80% in an environment tested by the inventor. However, this may vary depending on the characteristics of the workload. For example, the memory usage of a conventional extensible hash is based on the bytes of the pointer and the number of (logical) directories (e.g., 2 GIt can be based on the product of ). The memory usage of a method for compressing a directory of scalable hashes capable of low-key support according to one embodiment is the capacity required for directory encoding (e.g., 2 G-3 It can be based on the sum of the capacity less than or equal to ) and the capacity required by the compressed directory (e.g., the product of the pointer's bytes and buckets).
[0078] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0079] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.
[0080] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0081] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0082] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0083] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 A method for compressing a directory of scalable hashes capable of supporting low keys, performed by a processor, comprising: a step of linearly searching all items of the directory and generating a directory bit string by encoding bits according to whether the searched item has previously pointed to; a step of performing directory encoding to change the bit representation of the directory bit string into an integer representation; and a step of generating a compressed directory by compressing the directory based on the directory bit string after performing the directory encoding so as to represent only information about the actual existing buckets, wherein the step of generating the directory bit string includes generating a directory bit string by encoding a bit as '1' when the searched item points to a bucket that has not previously pointed to, or encoding a bit as '0' when the searched item points to a bucket that has previously pointed to, and wherein the compressed directory is an index of the number of '1' representations of the directory bit string counted. Claim 2 delete Claim 3 delete Claim 4 In claim 1, the step of performing the directory encoding comprises, when '0' appears 64 or more times consecutively in the bit representation of the directory bit string, converting the bit representation for a predetermined number of '0's into a 64-bit integer representation ( A method including the step of generating a bit string. Claim 5 In paragraph 4, the method wherein the integer representation is identified by the preceding 3-bit sequence of the integer representation as a mark. Claim 6 In paragraph 5, the method in which the above 3-bit sequence is '001'. Claim 7 A method according to claim 1, further comprising the step of expanding the compressed directory when, after the compressed directory is created, an overflow occurs in any one of at least one bucket connected to the compressed directory, depending on whether the local depth of the bucket where the overflow occurred is the same as a predetermined global depth. Claim 8 In claim 7, the step of expanding the compressed directory comprises: creating an additional bucket that divides information about the bucket where the overflow occurred when the previously determined global depth and the local depth of the bucket are not the same; creating an additional compressed directory in the compressed directory that points to the additional bucket; and changing information about the directory bit string at a location corresponding to the additional compressed directory; or, when the previously determined global depth and the local depth of the bucket are the same, creating an additional bucket that divides information about the bucket where the overflow occurred; creating an additional compressed directory in the compressed directory that points to the additional bucket; increasing the global depth to increase the number of directory bit strings; and changing information about the increased directory bit string at a location corresponding to the additional compressed directory. Claim 9 A device for compressing a directory of scalable hashes capable of supporting low keys, comprising a processor that linearly searches all items of the directory and generates a directory bit string by encoding bits according to whether the searched item has previously pointed to, performs directory encoding by changing the bit representation of the directory bit string into an integer representation, and generates a compressed directory by compressing the directory based on the directory bit string in which the directory encoding was performed so as to represent only information about the actual existing buckets, wherein, in generating the directory bit string, the processor generates a directory bit string in which, if the searched item points to a bucket that has not previously pointed to, the bit is encoded as '1', or if the searched item points to a bucket that has previously pointed to, the bit is encoded as '0', and the compressed directory is an index in which the number of '1' representations of the directory bit string counted. Claim 10 delete Claim 11 delete Claim 12 In claim 9, the processor, when performing the directory encoding, if '0' appears 64 or more times consecutively in the bit representation of the directory bit string, converts the bit representation for a predetermined number of '0's into a 64-bit integer representation ( ) Device that generates a bit string. Claim 13 In paragraph 12, the device, wherein the integer representation is identified by the preceding 3-bit sequence of the integer representation as a mark. Claim 14 In paragraph 13, the device, wherein the above 3-bit sequence is '001'. Claim 15 In claim 9, the processor expands the compressed directory based on whether the determined global depth and the local depth of the bucket where the overflow occurred are the same when an overflow occurs in any one of at least one bucket connected to the compressed directory after the compressed directory is created. Claim 16 In claim 15, the processor, when expanding the compressed directory, if the predetermined global depth and the local depth of the bucket are not the same, creates an additional bucket that divides the information about the bucket where the overflow occurred, creates an additional compressed directory in the compressed directory that points to the additional bucket, and changes the information of the directory bit string at the location corresponding to the additional compressed directory, or if the predetermined global depth and the local depth of the bucket are the same, creates an additional bucket that divides the information about the bucket where the overflow occurred, creates an additional compressed directory in the compressed directory that points to the additional bucket, increases the global depth to increase the number of the directory bit strings, and changes the information of the increased directory bit string at the location corresponding to the additional compressed directory.
Citation Information
Patent Citations
A method for controlling directory splits of theextendible hashing
KR1020010047384A
Parallel file system and method with extensible hashing
US5893086A
Compression of nodes in a trie structure
US6910043B2