Bitmap rapid operation method based on CPU hardware instruction

By integrating AVX512 and BMI2 instructions in CPU hardware instructions, the position of ‘1’ in the bitmap index is quickly positioned, and the problem of high time overhead when obtaining the ‘1’ position in the compressed bitmap in the prior art is solved, and efficient data acquisition and system throughput are achieved.

CN119988428AActive Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510063351.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art uses compression algorithms such as WAH to greatly reduce storage overhead while performing high performance, but when obtaining the position of ‘1’ in the compressed bitmap, it will lead to a lot of time overhead.

Method used

Using a fast bitmap operation method based on CPU hardware instructions, the positioning of all '1' in two types of words under WAH algorithm compression is quickly realized through the AVX512 and BMI2 instruction sets.

Benefits of technology

It significantly shortens the time for calling related data, improves the system's throughput, and greatly reduces the time overhead of data acquisition by making full use of CPU resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988428A_ABST
    Figure CN119988428A_ABST
Patent Text Reader

Abstract

The invention discloses a bitmap rapid operation method based on a CPU hardware instruction, and relates to the technical field of data processing. According to the quick bitmap operation method based on the CPU hardware instruction, analysis of stored data in a bitmap index subjected to WAH compression is achieved through the hardware instruction sets of AVX512 and BMI2, the positions and the total number of effective data in the current bitmap index are quickly obtained, 1 is quickly obtained and positioned, the time for calling related data is remarkably shortened, and the efficiency of operating the bitmap is improved. Related information of data can be quickly queried during database query, the throughput rate of a system can be remarkably improved, CPU resources are fully utilized, and compared with a traditional traversal method, larger-scale data can be processed at a time, the time expenditure for data acquisition is greatly reduced, and the data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of data processing, in particular to a bitmap fast operation method based on CPU hardware instructions. Background Art

[0002] A bitmap index is a special database indexing technology. Its index uses a bit array for storage and calculation operations. It uses a bitmap to indicate the existence of data in the database. Bitmap indexes are mainly created for columns with a large number of identical values, such as categories, operators, department IDs, warehouse IDs, etc. For each possible value, a bitmap index creates a bitmap. If there are n rows of data in a table, then there are n bits in the bitmap for a specific value, and each bit corresponds to a row in the table. If the value of the column in the row is the specific value, the corresponding bit is set to 1, otherwise it is 0.

[0003] Bitmap indexes have been widely used in scientific and commercial databases. They are particularly effective in performing certain types of queries, such as equality and selective range queries. The reason is that bitmap indexes can answer queries by performing bit-by-bit logical operations on bitmaps. Compression algorithms such as WAH can greatly reduce storage overhead while maintaining high performance. However, obtaining the position of the original '1' in the compressed bitmap in order to obtain relevant information will result in a lot of time overhead. For this reason, a fast bitmap operation method based on CPU hardware instructions is proposed. Through some instructions in AVX512 and BMI2, the location of all '1' in two types of words under WAH algorithm compression can be quickly realized, so that the event of calling related data is significantly shortened. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a method for fast bitmap operation based on CPU hardware instructions, which solves the problem that compression algorithms such as WAH can greatly reduce storage overhead while achieving high performance in bitmap indexing, but a large amount of time overhead is incurred when obtaining the position of '1' in the compressed bitmap to facilitate obtaining relevant information.

[0005] To achieve the above purpose, the present invention is implemented by the following technical scheme: a method for fast bitmap operation based on CPU hardware instructions, specifically comprising the following steps:

[0006] Step 1: Define the compressed bitmap index bitmap = {src1, src2, ..., src n}, define the 8-bit data to be read as word, define src i ={word1,word2,…,word 64}, src iFor the combination of the i-th 64 words in the bitmap index bitmap, define the number of all bits stored in the bitmap index as nbits, define the number of all bits stored in the bitmap index as nsets, define the word that stores a large number of consecutive '0' or '1' sequences in the bitmap index compression algorithm as fillword, and the word that stores mixed sequences as literal word;

[0007] Define mark as the value that marks fill word and literal word, which is located at the first position in the word. 1 represents fill word and 0 represents literal word. The binary format of literal word is 0xxxxxxx, and the last 7 bits are a mixed sequence of consecutive 01s. Define base as the data stored continuously in fill word. The base value is 0 or 1 and is located at the second position in fillword. Define a word with all 1s compressed in fill word as 1-fill word, whose binary format is 11xxxxxx, and a word with all 0s compressed as 0-fill word, whose binary format is 10xxxxxx. Define block as 7 consecutive 1s or 0s. Define the last 6 bits of data in fill word as nblock, where nblock is the number of blocks. Define the actual number of bits stored in fill word as nbit;

[0008] Step 2: Read the compressed bitmap index bitmap and call the read instruction in the AVX512 instruction set to read src from the bitmap in sequence i ;

[0009] Step 3: Use AVX512 instructions to distinguish between signed and unsigned integers, and identify src by calling mask comparison instructions i 1-fill word 64-bit mask fillmask i and the 64-bit mask litemask of the literal word i ;

[0010] Step 4: Define src i The number of '1's contained in the literal word stored in nsetInLite i , by calling the mask count instruction combined with litemask i Get nsetInLite in the marked literal word i ;

[0011] Step 5: Define src iThe 512 bits of nblcok that identifies the fill word is nblockInFill512 i ={fcnt i1 ,fcnt i2 ,…,fcnt i64}, calculate the number of blocks contained in each word nblockInFill512 i ;

[0012] Defining src i The 512 bits of nblock corresponding to all words are nblockInAll512 i ={acnt i1 ,acnt i2 ,…,acnt i64}, calculate src i The number of blocks corresponding to all words in nblockInAll512 i ;

[0013] Step 6: To avoid nblockInFill512 i The direct accumulation of 64 words in the call to accumulate will cause the data overflow of the int8_t storage space, and the nblockInFill512 i and nblockInAll512 i Split and get nblockInFill512A' i and nblockInFill512B' i ;

[0014] Step 7: Define nblockInFill512A i The sum of the values ​​of each word is totalblockInFillA i ,nblockInFill512B i The sum of the word values ​​is totalblockInFillB i , call _mm512_reduce_add_epi64 function to process nblockInFill512A' i Get totalblockInFillA i , call _mm512_reduce_add_epi64 function to process nblockInFill512B' i Get totalblockInFillB i , calculate src iThe total number of nblocks in the fill word;

[0015] Step 8: Define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the word values ​​is totalblockInFillB i , define nblockInFill512A i The data compressed by AVX instructions is nblockInAll64A i 、nblockInAll512B i The compressed data is defined as nblockInAll64B i , define nblockInAll64A i 、nblockInAll64B i The bit-extended data is nblockInAll512A' i and nblockInAll512B' i , define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the values ​​of each word is totalblockInAllB i After executing steps 6 and 7, get totalblockInAllA i and totalblockInAllB i , get src by adding the two i The total number of blocks, calculate the total number of bits nbits in the bitmap;

[0016] Step 9. Use the BMI2 instruction set to process the words in the bitmap in sequence, extract the last '1' bit of the current word, extract the position of the bit, and then remove the last '1' bit in the word. Loop through step 9 until the word is 0, and continue to call and process the next word.

[0017] The present invention is further configured as follows: the calculation method of nbit in step 1 includes:

[0018] nbit=nblock*7.

[0019] The present invention is further configured as follows: the read instruction in the AVX512 instruction set in step 2 is the _mm512_loadu_epi8 instruction, and each instruction operation processes src i The 8-bit word in .

[0020] The present invention is further configured as follows: the mask comparison instruction in step 3 is a _mm512_cmp_epu8_mask instruction, and its use method includes:

[0021] Call the _mm512_cmp_epu8_mask instruction to calculate word as uint8_t, compare each word with 192 (11000000), and mark the src i 1-fill word 64-bit mask fillmask i ;

[0022] By calling the _mm512_cmp_epi8_mask instruction, word is calculated as int8_t, and each word is compared with 0 (00000000), marking the src i Literal word 64-bit mask litemask i .

[0023] The present invention is further configured as follows: the mask counting instruction in step 4 is a _mm512_maskz_sub_epi8 instruction.

[0024] The present invention is further configured as follows: in step 5, the number of blocks nblockInFill512 contained in each word is calculated i The methods include:

[0025] Call _mm512_maskz_sub_epi8 instruction to fillmask i The indicated src i Modify the word in src i Delete the head corresponding to each word in to get fcnt j , you can get the number of blocks contained in each word nblockInFill512 i :

[0026] nblockInFill512 i =_mm512_maskz_sub_epi8(src i ,127).

[0027] The present invention is further configured as follows: in step 5, src is calculated iThe number of blocks corresponding to all words in nblockInAll512 i The methods include:

[0028] Call _mm512_maskz_set1_epi8 instruction to litemask i nblockInFill512 indicated i The literal word in is directly set to '1', and the src is obtained by calling the _mm512_and_epi32 instruction i The number of blocks corresponding to all words in nblockInAll512 i :

[0029] nblockInAll512 i =_mm512_maskz_set1_epi8(

[0030] nblockInFill512 i ,litemask i ).

[0031] The present invention is further configured as follows: in step 6, nblockInFill512A' is obtained. i and nblockInFill512B' i The methods include:

[0032] define nblockInFill512 i The first 32 words are nblockInFill512A i The last 32 words are nblockInFill512B i , define nblockInAll512 i The first 32 words are nblockInAll512A i The last 32 words are nblockInFill512B i ;

[0033] Call _mm512_mask_reduce_add_epi64 to nblockInFill512A i and nblockInFill512B i Accumulate as 4 64-bit words to get nblockInFill64A i and nblockInFill64B i :

[0034] nblockInFill64A i =

[0035] _mm512_mask_reduce_add_epi64(nblockInFill512 i ,15)

[0036] nblockInFill64B i =

[0037] _mm512_mask_reduce_add_epi64(nblockInFill512 i ,240)

[0038] Call _mm_cvtsi64_si128, _mm512_cvtepu8_epi32 function to nblockInFill64A i 、nblockInFill64B i Each 8-bit word is converted to 64 bits to obtain nblockInFill512A' i and nblockInFill512B' i .

[0039] The present invention is further configured as follows: in step eight, src is calculated i The ways to calculate the total number of nblocks of fill word in nsets include:

[0040]

[0041] The present invention is further configured as follows: the method of calculating the total number of bits nbits in the bitmap in step eight includes:

[0042]

[0043] The present invention provides a method for fast bitmap operation based on CPU hardware instructions, which has the following beneficial effects:

[0044] (1) The present invention implements the parsing of data stored in a WAH-compressed bitmap index through the hardware instruction set of AVX512 and BMI2, quickly obtains the position of valid data in the current bitmap index and its total number, quickly locates the '1' therein, significantly shortens the time for calling related data, and quickly queries the relevant information of the data when querying the database, which can significantly improve the system throughput.

[0045] (2) By making full use of CPU resources, the present invention can process larger-scale data at one time compared to traditional traversal methods, greatly reducing the time overhead of data acquisition and improving data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a word structure diagram of the compression algorithm in the embodiment of the present invention;

[0047] Figure 2 is an example diagram of WAH in an embodiment of the present invention;

[0048] Figure 3 Obtain example images for liteMask and nsetInLite in the embodiments of the present invention;

[0049] Figure 4 Obtain example diagrams for fillMask, nblockInFill512, and nblockInAll512 in the embodiments of the present invention;

[0050] Figure 5 Obtaining an example diagram for totalblock in an embodiment of the present invention;

[0051] Figure 6 This is an example diagram of BMI2 processing in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0053] See also Figure 1-6 The embodiment of the present invention provides the following technical solution: a method for fast bitmap operation based on CPU hardware instructions, specifically comprising the following steps:

[0054] Step 1: Define the compressed bitmap index bitmap = {src1, src2, ..., src n}, define the 8-bit data to be read as word, define src i ={word1,word2,…,word 64}, src i For the combination of the i-th 64 words in the bitmap index bitmap, define the number of all bits stored in the bitmap index as nbits, define the number of all bits stored in the bitmap index as nsets, define the word that stores a large number of consecutive '0' or '1' sequences in the bitmap index compression algorithm as fillword, and the word that stores mixed sequences as literal word;

[0055] Define mark as the value that marks the fill word and literal word, which is located at the first position in the word. 1 represents fill word and 0 represents literal word. The binary format of literal word is 0xxxxxxx, and the last 7 bits are a mixed sequence of consecutive 01s. Define base as the data stored continuously in the fill word. The base value is 0 or 1 and is located at the second position in the fill word. Define the word with all 1s compressed in the fill word as 1-fill word, and the binary format is 11xxxxxx. Define the word with all 0s compressed as 0-fill word, and the binary format is 10xxxxxx. Define block as 7 consecutive 1s or 0s. Define the last 6 bits of data in the fill word as nblock, where nblock is the number of blocks. Define the actual number of bits stored in the fill word as nbit. The calculation methods of nbit include:

[0056] nbit=nblock*7.

[0057] Step 2: Read the compressed bitmap index bitmap and call the _mm512_loadu_epi8 instruction in the AVX512 instruction set to read src from the bitmap in sequence i Each instruction operation processes src i The 8-bit word in .

[0058] Step 3: Use AVX512 instructions to distinguish whether there is a signed integer or not, and identify src by calling the _mm512_cmp_epu8_mask instruction i 1-fill word 64-bit mask fillmask i and the 64-bit mask litemask of the literal word i , including:

[0059] Call the _mm512_cmp_epu8_mask instruction to calculate word as uint8_t, compare each word with 192 (11000000), and mark the src i 1-fill word 64-bit mask fillmask i ;

[0060] By calling the _mm512_cmp_epi8_mask instruction, word is calculated as int8_t, and each word is compared with 0 (00000000), marking the src iLiteral word 64-bit mask litemask i .

[0061] Step 4: Define src i The number of '1's contained in the literal word stored in nsetInLite i , by calling the _mm512_maskz_sub_epi8 instruction combined with litemask i Get nsetInLite in the marked literal word i .

[0062] Step 5: Define src i The 512 bits of nblcok that identifies the fill word is nblockInFill512 i ={fcnt i1 ,fcnt i2 ,…,fcnt i64}, calculate the number of blocks contained in each word nblockInFill512 i , including:

[0063] Call _mm512_maskz_sub_epi8 instruction to fillmask i The indicated src i Modify the word in src i Delete the head corresponding to each word in to get fcnt j , you can get the number of blocks contained in each word nblockInFill512 i :

[0064] nblockInFill512 i =_mm512_maskz_sub_epi8(src i ,127).

[0065] Defining src i The 512 bits of nblock corresponding to all words are nblockInFill512 i ={scnt i1 ,scnt i2 ,…,scnt i64}, calculate src i The number of blocks corresponding to all words in nblockInAll512 i , including:

[0066] Call _mm512_maskz_set1_epi8 instruction to litemask i nblockInFill512 indicated i The literal word in is directly set to '1', and the src is obtained by calling the _mm512_and_epi32 instruction i The number of blocks corresponding to all words in nblockInAll512 i :

[0067] nblockInAll512 i =_mm512_maskz_set1_epi8(nblockInFill512 i ,litemask i ).

[0068] Step 6: To avoid nblockInFill512 i The direct accumulation of 64 words in the call to accumulate will cause the data overflow of the int8_t storage space, and the nblockInFill512 i and nblockInAll512 i Split and define nblockInFill512 i The first 32 words are nblockInFill512A i The last 32 words are nblockInFill512B i , define nblockInFill512 i The first 32 words are nblockInFill512A i The last 32 words are nblockInFill512B i ;

[0069] Call _mm512_mask_reduce_add_epi64 to nblockInFill512A i and nblockInFill512B i Accumulate as 4 64-bit words to get nblockInFill64A i and nblockInFill64B i :

[0070] nblockInFill64A i =

[0071] _mm512_mask_reduce_add_epi64(nblockInFill512 i ,15)

[0072] nblockInFill64B i =

[0073] _mm512_mask_reduce_add_epi64(nblockInFill512 i ,240) Call _mm_cvtsi64_si128, _mm512_cvtepu8_epi32 function to nblockInFill64A i 、nblockInFill64B i Each 8-bit word is converted into 64 bits to obtain nblockInFill512A' i and nblockInFill512B' i .

[0074] Step 7: Define nblockInFill512A i The sum of the values ​​of each word is totalblockInFillA i ,nblockInFill512B i The sum of the word values ​​is totalblockInFillB i , call _mm512_reduce_add_epi64 function to process nblockInFill512A' i Get totalblockInFillA i , call _mm512_reduce_add_epi64 function to process nblockInFill512B' i Get totalblockInFillB i , calculate src i The total number of nblocks of fill word in nsets is calculated by the following formula:

[0075]

[0076] Step 8: Define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the values ​​of each word is totalblockInAllB i, define nblockInFill512A i The data compressed by AVX instructions is nblockInAll64A i 、nblockInAll512B i The compressed data is defined as nblockInAll64B i , define nblockInAll64A i 、nblockInAll64B i The bit-extended data is nblockInAll512A' i and nblockInAll512B' i , define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the values ​​of each word is totalblockInAllB i After executing steps 6 and 7, get totalblockInAllA i and totalblockInAllB i , get src by adding the two i The total number of blocks is calculated, and the total number of bits nbits in the bitmap is calculated using the following methods:

[0077]

[0078] Step 9: Use the BMI2 instruction set to process the words in the bitmap in sequence, extract the last '1' bit of the current word, extract the position of the bit, and then remove the last '1' bit in the word. Loop through step 9 until the word is 0, and continue to call and process the next word, including:

[0079] Use the _blsi_u32 function in the BMI2 instruction set to extract tmp2 from tmp1. tmp2 only contains the last '1' in tmp1. Use anchor to record the starting position of the current word in the bitmap index, and use the _blsr_u32 instruction to clear the last '1' in tmp1.

[0080] tmp2 = _blsi_u32 (tmp1)

[0081] loc=anch or+1-pow(tmp2,2)

[0082] tmp1 = _blsr_u32 (tmp1).

[0083] In summary, AVX512 is used to process compressed bitmap data, 512-bit data is processed at a time, and the corresponding fillmask and litemask are generated according to the characteristics of the corresponding word. The number of bits and the number of 1s in the compressed bitmap are calculated through mask operation instructions. The risk of bit overflow during the calculation process is avoided through bit scaling instructions, which significantly improves the time cost of data statistics. BMI2 is used to process compressed bitmap data. For the literal word in the compression algorithm, the '1' can be quickly located and quickly obtained through the calculation formula, which improves the time cost.

[0084] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for fast bitmap operation based on CPU hardware instructions, characterized in that: The specific steps include: Step 1: Define the compressed bitmap index bitmap = {src1, src2, ..., src n }, define the 8-bit data to be read as word, define src i ={word1,word2,…,word 64 }, src i For the combination of the i-th 64 words in the bitmap index bitmap, define the number of all bits stored in the bitmap index as nbits, define the number of all bits stored in the bitmap index as nsets, define the word that stores a large number of consecutive '0' or '1' sequences in the bitmap index compression algorithm as fillword, and the word that stores mixed sequences as literal word; Define mark as the value that marks fill word and literal word, which is located at the first position in word. 1 represents fill word and 0 represents literal word. The binary format of literal word is 0xxxxxxx, and the last 7 bits are a mixed sequence of consecutive 01s. Define base as the data stored continuously in fill word. The base value is 0 or 1, which is located at the second position in fill word. Define a word with all 1s compressed in fill word as 1-fill word, whose binary format is 11xxxxxx, and a word with all 0s compressed as 0-fillword, whose binary format is 10xxxxxx. Define block as 7 consecutive 1s or 0s. Define the last 6 bits of data in fill word as nblock, where nblock is the number of blocks. Define the actual number of bits stored in fill word as nbit; Step 2: Read the compressed bitmap index bitmap and call the read instruction in the AVX512 instruction set to read src from the bitmap in sequence i ; Step 3: Use AVX512 instructions to distinguish between signed and unsigned integers, and identify src by calling mask comparison instructions i 1-fill word 64-bit mask fillmask i and the 64-bit mask litemask of the literal word i ; Step 4: Define src i The number of '1's contained in the literal word stored in nsetInLite i , by calling the mask count instruction combined with litemask i Get nsetInLite in the marked literal word i ; Step 5: Define src i The 512 bits of nblcok that identifies the fill word is nblockInFill512 i ={fcnt i1 ,fcnt i2 ,…,fcnt i64 }, calculate the number of blocks contained in each word nblockInFill512 i ; Defining src i The 512 bits of nblock corresponding to all words are nblockInAll512 i ={acnt i1 ,acnt i2 ,…,acnt i64 }, calculate src i The number of blocks corresponding to all words in nblockInAll512 i ; Step 6: nblockInFill512 i and nblockInAll512 i Split and get nblockInFill512A' i and nblockInFill512B' i ; Step 7: Define nblockInFill512A i The sum of the values ​​of each word is totalblockInFillA i ,nblockInFill512B i The sum of the values ​​of each word is totalblockInFillA i , call _mm512_reduce_add_epi64 function to process nblockInFill512A' i Get totalblockInFillA i , call _mm512_reduce_add_epi64 function to process nblockInFill512B' i Get totalblockInFillB i , calculate src i The total number of nblocks in the fill word; Step 8: Define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the word values ​​is totalblockInFillB i , define nblockInFill512A i The data compressed by AVX instructions is nblockInAll64A i 、nblockInAll512B i The compressed data is defined as nblockInAll64B i , define nblockInAll64A i 、nblockInAll64B i The bit-extended data is nblockInAll512A' i and nblockInAll512B' i , define nblockInAll512A i The sum of the values ​​of each word is totalblockInAllA i , nblockInAll512B i The sum of the values ​​of each word is totalblockInAllB i After executing steps 6 and 7, get totalblockInAllA i and totalblockInAllB i , get src by adding the two i The total number of blocks, calculate the total number of bits nbits in the bitmap; Step 9. Use the BMI2 instruction set to process the words in the bitmap in sequence, extract the last '1' bit of the current word, extract the position of the bit, and then remove the last '1' bit in the word. Loop through step 9 until the word is 0, and continue to call and process the next word.

2. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: The calculation method of nbit in step 1 includes: nbit=bblock*7.

3. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: The read instruction in the AVX512 instruction set in step 2 is the _mm512_loadu_epi8 instruction, and each instruction operation processes src i The 8-bit word in .

4. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: The mask comparison instruction in step 3 is the _mm512_cmp_epu8_mask instruction, and its usage method includes: Call the _mm512_cmp_epu8_mask instruction to calculate word as uint8_t, compare each word with 192 (11000000), and mark the src i 1-fillword 64-bit mask fillmask i ; By calling the _mm512_cmp_epi8_mask instruction, word is calculated as int8_t, and each word is compared with 0 (00000000), marking the src i Literal word 64-bit mask litemask i .

5. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: The mask counting instruction in step 4 is the _mm512_maskz_sub_epi8 instruction.

6. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: In step 5, the number of blocks contained in each word is calculated as nblockInFill512 i The methods include: Call _mm512_maskz_sub_epi8 instruction to fillmask i The indicated src i Modify the word in src i By deleting the header corresponding to each word in fcntj, we can get the number of blocks contained in each word nblockInFill512 i : nbockInFill512 i =_mm512_maskz_sub_epi8(src i ,127)。 7. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: In step 5, src is calculated i The number of blocks corresponding to all words in nblockInAll512 i The methods include: Call _mm512_maskz_set1_epi8 instruction to litemask i nblockInFill512 indicated i The literal word in is directly set to '1', and the src is obtained by calling the _mm512_and_epi32 instruction i The number of blocks corresponding to all words in nblockInAll512 i : nblockInAll512 i =_mm512_maskz_set1_epi8( nblockInFill512 i ,litemask i )。 8. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: In step 6, nblockInFill512A' is obtained. i and nblockInFill512B' i The methods include: define nblockInFill512 i The first 32 words are nblockInFill512A i The last 32 words are nblockInFill512B i , define nblockInAll512 i The first 32 words are nblockInAll512A i The last 32 words are nblockInFill512B i ; Call _mm512_mask_reduce_add_epi64 to nblockInFill512A i and nblockInFill512B i As 4 64-bit words, accumulate to get nblockInFill64A i and nblockInFill64B i : nblockInFill64A i = _mm512_mask_reduce_add_epi64(nblockInFill512 i ,15) nblockInFill64B i = _mm512_mask_reduce_add_epi64(nblockInFill512 i ,240) Call _mm_cvtsi64_si128, _mm512_cvtepu8_epi32 function to nblockInFill64A i 、nblockInFill64B i Each 8-bit word is converted to 64 bits to obtain nblockInFill512A' i and nblockInFill512B' i .

9. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: In step seven, src is calculated i The ways to calculate the total number of nblocks of fill word in nsets include:

10. The method for fast bitmap operation based on CPU hardware instructions according to claim 1, characterized in that: The method of calculating the total number of bits nbits in the bitmap in step eight includes:

Citation Information

Patent Citations

  • Hardware fingerprint information generation method and system based on national cryptographic algorithm

    CN111709044A

  • Bitmap index compression method oriented to genome variation data

    CN116230098A

  • Metadata query optimization method and terminal

    CN117520384A

  • Single table query processing method for activation function

    CN118277617A

  • Method for Performing Compressed Column Operations

    US20190332387A1