Lucene Acceleration Optimization Method Based on Block Index

Through the block index storage solution, the optimized CK structure and bitmap storage are used to solve the problem of Lucene's over-order chaining under large data volumes, and the index storage efficiency and query speed are improved.

CN115563344BActive Publication Date: 2025-07-29NANJING FIBERHOME STARRYSKY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210974173.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-07-29
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

The existing Lucene technology has too long inverted chains corresponding to Term under large data volume, resulting in low index storage compression rate and poor query performance.

Method used

Using a block index-based storage scheme, the structure of the block index CK is composed of priority, subid and value, and the data is grouped and bitmap storage is used to reduce storage usage and random read times.

Benefits of technology

Significantly reduce data storage requirements, improve write performance, reduce the number of random reads of queries, and improve query performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563344B_ABST
    Figure CN115563344B_ABST
Patent Text Reader

Abstract

The present invention discloses a Lucene acceleration optimization method based on block indexing, and designs the structure of block index CK which consists of priority, subid and value; when writing the block index, the data to be indexed is grouped according to CK, and if the amount of data in the group exceeds the preset value, it is split and grouped, and the split group IDs form subid; after grouping the CK, the data groups are assembled into a document Document, and a Field is constructed for the same values of each index field, a bitmap is constructed according to the row number ID of each group, and CK and the bitmap are stored using the Payload structure to form the document Document, completing the establishment of the block index. The present invention solves the problems in index storage and query caused by the too long inverted list corresponding to Term under a large amount of data, and proposes a storage scheme of "block index" to reduce storage occupancy and the number of random reads, and improve the writing and query performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of index storage and retrieval. Background Art

[0002] Lucene is a high-performance and feature-rich search engine library implemented in Java. This technology is applicable to almost any application that requires structured search, full-text search, pagination, nearest neighbor search across high-dimensional vectors, spelling correction, or query suggestion.

[0003] Retrieval is generally divided into two processes: index creation (Indexing) and querying the index (Search).

[0004] Index creation: The process of extracting information from all structured and unstructured data in the real world and creating an index.

[0005] Querying the index: The process of searching the created index according to the user's query request and then returning the results.

[0006] 1. Writing Process

[0007] The most important index established by Lucene saves the mapping from strings to files, which is called an inverted index or reverse index. Assume there are N documents in the document collection. For convenience of representation, the document numbers are 1, 2, 3, …, N. The resulting inverted index structure is as Figure 1 shown: On the left, a series of strings are saved, which is called a dictionary. The strings that make up the dictionary are called Terms. Each Term points to a linked list of documents that contain this string, which is called an inverted list.

[0008] In Lucene, a series of files with special structures are used to save the index data, and different functions are distinguished according to the file suffix names. The main files are shown in Table 1:

[0009] Table 1 Index File List of Lucene

[0010]

[0011]

[0012] Now, if it is necessary to index fields F1 and F2 and store field F3 for querying to return the field value, the resulting Document information is as Figure 2 shown.

[0013] Assume the DocId automatically generated by Lucene is the row number in Table 2. Then the storage structures of fields F1 and F3 are as Figure 3 shown. The storage structure of field F2 is similar to that of field F1 and will not be listed here.

[0014] 2. Query Process

[0015] The query process of Lucene is generally divided into two stages: the Query stage and the Fetch stage. In the Query stage, DocIds that meet the conditions are found from the dictionary and the inverted index according to the query conditions. In this stage, the inverted index is mainly used, that is, index files with suffixes such as tip, tim, and doc; in the Fetch stage, the field value information required by the user is obtained from the index data according to the DocIds found in the Query stage. In this stage, the forward index is used, that is, index files with suffixes such as fdx and fdt. The query process is as Figure 4 shown, and the process explanation is as follows:

[0016] If there are valid segments in the index directory, a BulkScorer is generated for each segment, and the inverted list corresponding to the Term in the query condition is traversed according to this to obtain the DocIds that meet the conditions and collect them using a Collector

[0017] If the set of DocIds in the Collector is not empty, the doc() of the IndexSearcher class is called repeatedly to obtain detailed Document field value information

[0018] It can be seen from this that the prior art has the following deficiencies or disadvantages: When the data volume is very large and the frequency of a certain Term appearing in the document is very high, the inverted list corresponding to this Term becomes very long, reducing the compression ratio of index storage. When querying, the inverted list to be traversed becomes longer and the number of random reads of the fdt and fdx files in the Fetch stage is too high, resulting in poor query performance. Summary of the Invention

[0019] In view of the problems existing in the prior art, the present invention provides a Lucene acceleration optimization method based on block index, which solves the problems in the index storage and query solutions caused by the too long inverted list corresponding to the Term under the existing large data volume, and proposes a storage solution of "block index" to reduce storage occupancy and the number of random reads, and improve the write and query performance.

[0020] To solve the above technical problems, the present invention adopts the following technical solutions:

[0021] A Lucene acceleration optimization method based on block index, comprising the following steps:

[0022] Step S1, design of the data block index storage scheme

[0023] Set the structure of the block index CK, which consists of three parameters: priority, subid, and value of CK;

[0024] Step S2, Block Index Writing

[0025] First, when writing the block index, the data to be indexed is grouped by CK. If the amount of data in a group exceeds the preset value, it is split into groups, and the split group IDs form the sub-block identifier subid;

[0026] Then, the data with the same CK after grouping by CK is assembled into a document Document, and the same values of each index field are constructed into a Field; at the same time, a bitmap is constructed according to the row number ID of each group, and constructed by column for different values; CK and the bitmap are stored using the Payload structure of Lucene to form the document Document;

[0027] Finally, call the addDocuemnt() function of the index creator IndexWriter to complete the writing and establish;

[0028] Step S3, Block Index Query

[0029] First, execute the search() function of IndexSearch, rewrite the Query to generate Weight, and further extend the BulkScorer;

[0030] Then, traverse the inverted list corresponding to the conditional Term to obtain the DocID that meets the conditions;

[0031] Next, obtain the CK and the bitmap corresponding to the DocID that meets the conditions;

[0032] Furthermore, perform intersection and union on the bitmap values with the same value in the block index. If the bitmap after intersection and union is not empty, collect the corresponding CK and end.

[0033] Furthermore, the structure of the block index CK in Step S1 is encoded in the following storage format: priority + subid + value, where:

[0034] Priority priority, occupying 1 byte, the lowest two bits store the priority, and the value range is 0 to 2. When the highest bit is 1, it means there is a subid, and when the highest bit is 0, it means there is no subid;

[0035] Sub-block identifier subid, which is an optional item; stored using short, occupying 2 bytes;

[0036] CK value value, storing the specific value of CK, using the String type, occupying N bytes.

[0037] Furthermore, assume that the largest row number in the block index is highest and the number of rows is numWord. When When it is [condition], the bitmap in step S2 is stored according to the original value; otherwise, it is stored as a bitmap. The storage structure encoding for storing according to the original value and storing as a bitmap is as follows:

[0038] (1) Storing according to the original value

[0039] The storage format of the bitmap adopts the structure of BitsetMetaData + several RecordNums, where:

[0040] BitsetMetaData, the highest bit is 1, and the remaining bits store the number of bytes occupied by RecordNum;

[0041] RecordNum, which represents the row record of this Term in the block and is stored in the short type. Each RecordNum occupies 2 bytes;

[0042] (2) Storing the original value as a bitmap

[0043] The storage format of the bitmap adopts the structure of BitsetMetaData + Bitset, where:

[0044] BitsetMetaData, the highest bit is 10 and the remaining bits store the number of bytes occupied by Bitset;

[0045] Bitset, which represents the bitmap for storing row numbers and is stored as a Long array at the bottom layer; each Bits element is a Long type, and the last Bits only stores the byte whose low bit is not 0.

[0046] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0047] (1) Based on the block index storage mode, it can make full use of the similarity of data, significantly reduce data storage, and improve the writing performance. Through actual measurement and verification, the storage size occupies about 43% less storage than the original scheme, and the writing time-consuming is improved by about 53%.

[0048] (2) Due to the new index query mode, the number of random data reads is reduced, and the overall query performance is improved. Through experimental tests, the more obvious the query performance improvement is when the limit is larger. Description of the drawings

[0049] Figure 1 It is a schematic diagram of the native inverted index structure of the prior art Lucene described in the background technology;

[0050] Figure 2 It is a schematic diagram of the Document structure constructed by the original storage method in the background technology case;

[0051] Figure 3 Schematic diagram of the storage structure of the native F1 field and F3 field in the background technology case;

[0052] Figure 4 Schematic diagram of the process of the native query solution of Lucene in the background technology case;

[0053] Figure 5 Schematic diagram of the optimized inverted index table structure in the embodiment of the present invention;

[0054] Figure 6 Schematic diagram of the Document structure in the block index in the embodiment of the present invention;

[0055] Figure 7 Schematic diagram of the CK storage encoding structure in the embodiment of the present invention;

[0056] Figure 8 Schematic diagram of the inverted index chain structure of the F1 field in the block index in the embodiment of the present invention;

[0057] Figure 9 Schematic diagram of the block index query process in the embodiment of the present invention. Detailed implementation manners

[0058] The present invention will be further clarified below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention all fall within the scope defined by the appended claims of this application.

[0059] The core of the present invention is to construct an index structure in the form of "blocks" based on the idea of "block data", reduce the length of the inverted index chain corresponding to the Term, reduce the index storage scale, and improve the write and query performance. The detailed implementation manners are specifically introduced as follows:

[0060] 1) Scheme design

[0061] Under the new index storage scheme, the data is aggregated by CK, and CK is the identifier of the data block.

[0062] To avoid the problem of null value fields in the data, which makes it difficult to aggregate the data by a single field, a multi-level CK scheme is adopted to solve this problem, that is, the CK in the block index is multi-level. When the high-priority field has a value, it is aggregated according to the high priority. For the data with a null high-priority field, it is aggregated according to the CK key of the next priority.

[0063] To avoid too much data aggregated by a certain CK, subid is introduced to solve this problem.

[0064] Secondly, to describe which row in the block the specific value of each field exists in, a bitmap is introduced to solve this problem. Each bit in the bitmap represents whether the field value exists in the corresponding row.

[0065] In summary, the CK of the block index consists of three parameters: priority, subid (sub-block identifier), and value (the specific CK value). The data structure is defined as shown in Table 3.

[0066] Table 3 Data Structure of Block Index CK

[0067]

[0068] After aggregation, compared with the original storage structure, the inverted list can be optimally shortened by up to the maximum number of rows within a block (for example, when aggregating every 256 rows, it can be shortened by 256 times). Compared with Figure 1 the original inverted list, the optimized inverted list structure of the present invention is as Figure 5 shown.

[0069] 2) Writing of Block Index

[0070] When writing the block index, the data to be indexed is first grouped by CK, with each CK as a group. If the amount of data within a group exceeds the preset maximum number (for example, 256 rows), it is split, and the grouping ID forms the subid. Each group of data is assembled into a Document. The same values of each index field build a Field, and at the same time, a bitmap is built through the row numbers within the group. Different values are built as multi-value columns. The CK and the bitmap are stored using the Payload structure of Lucene. After constructing the Document, call the addDocuemnt() of the IndexWriter class to write.

[0071] For example, for the data in Table 2, if column F3 is selected as the CK, after aggregation according to the CK, it will be divided into two groups, that is, two Documents will be constructed; and since the information to be returned (column F3) is stored in the Payload structure as the CK, there is no StroedField in the block index, that is, there is no forward index information related to the field values. The assembled data is shown in Table 4.

[0072] Table 4 Assembled Data

[0073]

[0074] Since all values in column F3 of the original data are not empty, the constructed CK priorities are all the first priority, and the data within the group does not exceed the maximum row number limit, so the subids are all 0. The detailed information of the assembled Document is as Figure 6 shown.

[0075] Since the block index stores CK and the bitmap using the Payload structure, it is necessary to encode CK and the bitmap appropriately to save storage space as much as possible.

[0076] 1. CK Storage Encoding

[0077] In the block index, CK is encoded and stored in the Payload using a byte array. The encoding method is: priority(1 byte) + subid(2 bytes) + value(N bytes). Among them, subid is an optional item. When there is no subid, the highest bit of priority is 0. When there is a subid, the highest bit of priority is 1 and the subid is stored. The storage format is as Figure 7 shown.

[0078] priority: Occupies 1 byte. The lowest two bits store the priority, and the value range is 0 to 2. The highest bit being 1 indicates there is a subid, and 0 indicates there is no subid;

[0079] subid: This item exists when the highest bit of priority is 1. It is stored using a short and occupies 2 bytes;

[0080] value: Stores the specific value of CK, with the type being String and occupying N bytes.

[0081] 2. Bitmap Storage Encoding

[0082] The size of the bitmap depends on the number of rows of data in the block (for example, 256 rows). When directly using a certain bit to represent whether a certain row of data exists, there are the following situations that are relatively wasteful of storage space:

[0083] When there are only a dozen rows of data in the block, it needs to occupy the size of a long type, that is, 8 bytes;

[0084] When the data in the block is relatively scattered, it does not save space compared to storing the original value. For example, when a certain Term only exists in the 255th row in the block, it needs 4 long types for storage, that is, 32 bytes, but if the original value is stored, it only needs 2 bytes.

[0085] For the above reasons, there are the following two situations in the encoding design of the bitmap:

[0086] i. Store by original value

[0087] BitsetMetaData: The highest bit is 1, and the remaining bits store the number of bytes occupied by RecordNum, with the type being byte (see Figure 7 )

[0088] RecordNum: The in-block row records of this Term, stored using short, and each RecordNum occupies 2 bytes (see Figure 7 )

[0089] ii. Store the original value by bitmap storage

[0090] BitsetMetaData: The highest bit is 0, and the remaining bits store the number of bytes occupied by BitSet, with the type being byte

[0091] BitSet: The bitmap storing the row numbers, stored using a Long array at the bottom layer, each Bits element is a Long, and for the last LastBits, since the first few bits may be 0, only the low bytes that are not 0 are stored (for example, when LastBit needs to store 7, since the first 3 bytes of 7 are all 0, only the last byte needs to be stored).

[0092] Assume the maximum row number in the block is highest and the number of rows is numWord. Then the number of bytes required for the above two storage schemes is as follows:

[0093] Number of bytes required to store the original value: (numWord * 2)

[0094] Number of bytes required to store the bitmap: ((highest - 1) / 8) + 1

[0095] Therefore, when ((highest - 1) / 8) + 1 > (numWord * 2), store the original value; otherwise, store the bitmap.

[0096] Through the reassembly of bit data and the encoding of CK and the bitmap, the inverted index structure corresponding to the F1 field is as Figure 8 shown. It can be found that the inverted index corresponding to the F1 field has been shortened, and there is no need to store the F3 field separately. The storage of the F2 field is similar to that of the F1 field and will not be listed here.

[0097] 3) Block index query

[0098] The process of the query execution engine based on the new storage scheme is as Figure 9 shown.

[0099] Since CK and the bitmap are stored in the Payload in the block index, and the Payload is obtained during the traversal of the inverted list, there is no Fetch phase in the entire block index query similar to the original Lucene query, reducing the number of random reads. Moreover, because the block index is grouped and reassembled according to the data, the length of the inverted list is greatly shortened (as described above, if there is a block for every 256 pieces of data, the inverted list can be optimally shortened by 256 times), significantly reducing the time taken to traverse the inverted list.

[0100] Finally, the method of the present invention is tested, and the detailed test results are shown in Table 5: It is verified by actual measurement that the storage size occupies about 43% less storage than the original scheme, and the write time consumption is increased by about 53%. The detailed situation is shown in Table 5.

[0101] Table 5 Comparison of Write Storage Size and Time Consumption between Native Lucene Index and Block Index

[0102] Original Lucene Index Block Index Effect Storage Size 1.7G 979.8M Around 43% Writing Time Consumed 935424ms (about 15.6min) 437469ms (about 7.3min) 53%

[0103] Through experimental tests, it is found that the larger the limit, the more obvious the improvement in query performance (the original size of the test data is 57.6G, with a total of 26 fields and 10 fields indexed). This proves that the new index query mode of the present invention reduces the number of random data reads and improves the overall query performance. The detailed situation is shown in Table 6.

[0104] Table 6 Comparison of Query Performance between Native Lucene Index and Block Index

[0105] Original Lucene Index Block Index Effect limit1 3.5ms 4.4ms Basically the same limit100 9.8ms 5.5ms Around 44% limit500 20ms 7.4ms Around 63% limit1000 50ms 9ms Around 82% 。

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Lucene acceleration and optimization method based on block indexing, characterized in that It includes the following steps: Step S1, design of the data block index storage scheme Set the structure of the block index CK, which consists of three parameters: priority, subid, and value of CK; Step S2, writing the block index First, when writing the block index, the data to be indexed is grouped according to CK. If the amount of data in the group exceeds the preset value, it is split into groups, and the split group IDs form the subid; Then, the data with the same CK after grouping according to CK is assembled into a document Document, and the same values of each index field are used to construct a Field; at the same time, a bitmap is constructed according to the row number ID of each group, and constructed by column for different values; CK and the bitmap are stored using the Payload structure of Lucene to form the document Document; Finally, call the addDocuemnt() function of the index creator IndexWriter to complete the writing and establish; Step S3, querying the block index First, execute the search() function of IndexSearch, rewrite the Query to generate Weight, and further extend the BulkScorer; Then, traverse the inverted list corresponding to the conditional Term to obtain the DocID that meets the conditions; Next, obtain the CK and the bitmap corresponding to the DocID that meets the conditions; Furthermore, perform intersection and union on the bitmap values within the block index. If the resulting bitmap is not empty, collect the corresponding CK and end.

2. The Lucene acceleration optimization method based on block indexing according to claim 1, wherein: The structure of the block index CK in Step S1 is encoded using the following storage format: priority+subid+value, where: Priority priority, occupying 1 byte. The lowest two bits store the priority, and the value range is 0 to 2. When the highest bit is 1, it indicates the existence of subid, and when the highest bit is 0, it indicates the non-existence of subid; Sub-block identifier subid, which is an optional item; stored using short, occupying 2 bytes; Value of CK value, storing the specific value of CK, using the String type, occupying N bytes.

3. The Lucene acceleration optimization method based on block index according to claim 1, wherein: Assume that the maximum line number in the block index is highest and the number of lines is numWord. When occurs, the bitmap in step S2 is stored according to the original value; otherwise, it is stored according to the bitmap. The storage structure encoding for storing according to the original value and according to the bitmap is as follows: (1) Store according to the original value The storage format of the bitmap adopts the structure of BitsetMetaData + several RecordNums, where: BitsetMetaData, the highest bit is 1, and the remaining bits store the number of bytes occupied by RecordNum; RecordNum, indicating the row record of this Term existing in the block, stored in the short type, and each RecordNum occupies 2 bytes; (2) Store the original value according to the bitmap The storage format of the bitmap adopts the structure of BitsetMetaData + Bitset, where: BitsetMetaData, the highest bit is 10 and the remaining bits store the number of bytes occupied by Bitset; Bitset, indicating the bitmap storing the row number, stored at the bottom layer using a Long array; each Bits element is a Long type, and the last Bits only stores the byte with the low bit not equal to 0.

Citation Information

Patent Citations

  • Gene variation data distributed storage method and architecture

    CN108563923A

  • Detecting at least one predetermined pattern in stream of symbols

    US20160267142A1