Data retrieval method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202310988218.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-07
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-08-07
AI Technical Summary
[0004]本发明的主要目的在于提供一种数据检索方法、装置、设备及存储介质,旨在解决现有技术数据检索效率以及准确率低的技术问题
[0062]本发明响应于索引创建指令,获取数据表的主键信息和数据表信息;基于所述主键信息和所述数据表信息对所述数据表进行分解,得到包括有多组主键指纹数据的数据表分组信息;通过预设划分规则将所述数据表分组信息更新至第一索引文件中,得到第一索引信息,其中,所述第一索引信息包括各个主键指纹所在的分段及所在分段的偏移量;基于所述第一索引信息将所述数据表的数据信息对应的文件信息和对应的主键指纹数据进行汇总,得到第二索引信息;基于所述第一索引信息和所述第二索引信息进行数据检索,通过对数据表进行分段,从而以主键指纹在分段中的位置实现两级索引创建,大大提高索引处理效率,通过主键指纹的索引机制,符合业务检索需求,提高数据检索效率和准确率。
Smart Images

Figure CN117056368B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of cloud computing and big data technology, and in particular to a data retrieval method, apparatus, device and storage medium. Background Technology
[0002] In scenarios involving real-time data transmission across systems, Kafka (a high-throughput distributed publish-subscribe messaging system) is typically used as the transmission channel. Message content is organized in JSON (JavaScript Object Notation) format. Businesses need to save historical messages for message reconciliation and business history backtracking. For indexing and retrieving massive amounts of data, object databases like MongoDB (a database based on distributed file storage) or text databases like Elasticsearch (a search and data analysis engine) are typically used for storage and management. Due to the large volume of data, cluster or distributed deployment is generally adopted, and relevant indexes are created based on search criteria.
[0003] Database systems typically use uncompressed storage, and the indexes created to improve retrieval speed also consume a significant amount of disk space. MongoDB meets retrieval needs through key-value indexes, but it struggles to support data volumes of hundreds of billions. While Elasticsearch can support hundreds of billions of data points, its keyword-based indexing principle means that search results cannot be sorted, leading to lower retrieval efficiency and difficulty in meeting project query requirements when there are many results. All of these solutions require the installation of a database system, demanding a high level of technical expertise, necessitating a professional DBA (Database Administrator) for daily maintenance. The high disk usage of massive amounts of data also results in large index data volumes, further increasing disk consumption. Summary of the Invention
[0004] The main objective of this invention is to provide a data retrieval method, apparatus, device, and storage medium, aiming to solve the technical problems of low data retrieval efficiency and accuracy in the prior art.
[0005] To achieve the above objectives, the present invention provides a data retrieval method, the method comprising the following steps:
[0006] In response to an index creation command, retrieve the primary key information and table information of the data table;
[0007] Based on the primary key information and the data table information, the data table is decomposed to obtain data table grouping information including multiple sets of primary key fingerprint data;
[0008] The data table grouping information is updated to the first index file by a preset partitioning rule to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment.
[0009] Based on the first index information, the file information corresponding to the data information of the data table and the corresponding primary key fingerprint data are summarized to obtain the second index information;
[0010] Data retrieval is performed based on the first index information and the second index information.
[0011] Optionally, the step of decomposing the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data includes:
[0012] Based on the data table information, obtain the file information corresponding to the data information of each data table;
[0013] The file information is split into multiple sub-file information;
[0014] A fingerprint is generated based on each primary key in the multiple sub-file information to obtain a primary key fingerprint, wherein each primary key in the primary key information corresponds to each of the sub-file information.
[0015] The primary key fingerprints belonging to the same file and the primary key fingerprints belonging to the same table are grouped and stored to obtain data table grouping information containing multiple groups of primary key fingerprint data.
[0016] Optionally, updating the data table grouping information to the first index file according to a preset partitioning rule to obtain the first index information includes:
[0017] Get the amount of data in the data table;
[0018] The segmentation rules for the first index file are set based on the amount of data.
[0019] The files in the data table are segmented according to the segmentation rules to obtain the first segmentation information;
[0020] Based on the grouping information in the data table, obtain the primary key fingerprint corresponding to each file in the first segment information;
[0021] Files belonging to the same primary key fingerprint are divided into the same segment to obtain the second segment information;
[0022] The grouping information of each table in the data table and the primary key fingerprint of each file are stored in the first index file using the second segment information to obtain the first index information.
[0023] Optionally, the step of summarizing the file information corresponding to the data information of the data table and the corresponding primary key fingerprint data based on the first index information to obtain the second index information includes:
[0024] Based on the first index information, obtain the file information and primary key fingerprint data corresponding to the data information of the data table;
[0025] The file number of each file is obtained based on the file information;
[0026] Obtain the segment number and offset of the primary key fingerprint data;
[0027] Obtain the byte length of the primary key fingerprint of the file in the data table, and use the byte length of the primary key fingerprint as the index length;
[0028] Calculate the fingerprint value of the primary key to obtain the primary key range;
[0029] The file number, segment number, offset, index length, and primary key range are summarized to obtain the second index information.
[0030] Optionally, before obtaining the primary key information and data table information of the data table, the method further includes:
[0031] Scan the original compressed data file to obtain data file information;
[0032] The file number, file name, start time, and end time of file storage are obtained based on the data file information.
[0033] A file index list is created using the file number, the file name, the start time, and the end time.
[0034] Accordingly, obtaining the primary key information and data table information of the data table includes:
[0035] The data file information is parsed to obtain the primary key information and data table information of the data table.
[0036] Optionally, the data retrieval based on the first index information and the second index information includes:
[0037] Retrieve search information and the list of primary keys to be searched;
[0038] The list of primary keys to be retrieved is preprocessed to obtain multiple fingerprints of primary keys to be retrieved;
[0039] The initial index segments are obtained by searching the range of multiple primary key fingerprints to be retrieved in the second index information.
[0040] The initial index segments are filtered using the file index list and the retrieval information to obtain filtered index segments;
[0041] The primary key fingerprint corresponding to the filtered index segment is obtained by retrieving the first index information using the offset and index length in the filtered index segment.
[0042] Compare the primary key fingerprint corresponding to the filtered index segment with the primary key fingerprint to be retrieved, and use the file corresponding to the primary key fingerprint that matches the primary key fingerprint as the initial data information.
[0043] Based on the initial data information, the target data information is obtained by searching the original compressed data file.
[0044] Optionally, the step of retrieving target data information from the original compressed data file based on the initial data information includes:
[0045] The data table for obtaining the initial data information;
[0046] Obtain the order of each primary key fingerprint in the data table to get the relative position of the primary key fingerprint;
[0047] Establish the correspondence between the primary key fingerprint and the relative position;
[0048] The initial data information is grouped based on a preset grouping structure and the corresponding relationship to obtain multiple groups of data information.
[0049] Based on the multiple sets of data information, the target data information is obtained by searching the original compressed data file.
[0050] Optionally, the step of retrieving target data information from the original compressed data file based on the multiple sets of data information includes:
[0051] The relative location and file information are obtained based on the multiple sets of data information described above;
[0052] Based on the relative position and the file information, a search is performed in the original compressed data file to obtain reference data information;
[0053] The primary key fingerprint of the reference data information is compared with the primary key fingerprint to be retrieved. The reference data information corresponding to the primary key fingerprint that is inconsistent with the primary key fingerprint to be retrieved is removed to obtain the target data information.
[0054] Furthermore, to achieve the above objectives, the present invention also proposes a data retrieval device, the data retrieval device comprising:
[0055] The retrieval module is used to retrieve the primary key information and data table information of the data table in response to the index creation command;
[0056] The decomposition module is used to decompose the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data;
[0057] The update module is used to update the data table grouping information to the first index file according to the preset partitioning rules to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment.
[0058] The aggregation module is used to aggregate the file information and the corresponding primary key fingerprint data corresponding to the data information of the data table based on the first index information to obtain the second index information;
[0059] The retrieval module is used to perform data retrieval based on the first index information and the second index information.
[0060] Furthermore, to achieve the above objectives, the present invention also proposes a data retrieval device, which includes: a memory, a processor, and a data retrieval program stored in the memory and executable on the processor, the data retrieval program being configured to implement the steps of the data retrieval method described above.
[0061] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a data retrieval program, which, when executed by a processor, implements the steps of the data retrieval method described above.
[0062] This invention, in response to an index creation command, obtains the primary key information and data table information of a data table; decomposes the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data; updates the data table grouping information to a first index file according to a preset partitioning rule to obtain first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment; summarizes the file information corresponding to the data information of the data table and the corresponding primary key fingerprint data based on the first index information to obtain second index information; performs data retrieval based on the first index information and the second index information. By segmenting the data table, a two-level index is created based on the position of the primary key fingerprint in the segment, greatly improving index processing efficiency. The primary key fingerprint indexing mechanism meets business retrieval needs and improves data retrieval efficiency and accuracy. Attached Figure Description
[0063] Figure 1 This is a schematic diagram of the structure of a data retrieval device in the hardware operating environment involved in the embodiments of the present invention;
[0064] Figure 2 This is a flowchart illustrating the first embodiment of the data retrieval method of the present invention;
[0065] Figure 3 This is a flowchart illustrating the second embodiment of the data retrieval method of the present invention;
[0066] Figure 4 This is a flowchart illustrating the third embodiment of the data retrieval method of the present invention;
[0067] Figure 5 This is a flowchart illustrating the fourth embodiment of the data retrieval method of the present invention;
[0068] Figure 6 This is a schematic diagram illustrating the data index creation principle of an embodiment of the data retrieval method of the present invention;
[0069] Figure 7 This is a flowchart illustrating the fifth embodiment of the data retrieval method of the present invention;
[0070] Figure 8 This is a schematic diagram illustrating the data retrieval principle of an embodiment of the data retrieval method of the present invention;
[0071] Figure 9 This is a structural block diagram of the first embodiment of the data retrieval device of the present invention.
[0072] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0073] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0074] Reference Figure 1 , Figure 1 This is a schematic diagram of the data retrieval device structure in the hardware operating environment involved in the embodiments of the present invention.
[0075] like Figure 1As shown, the data retrieval device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0076] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the data retrieval device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0077] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a data retrieval program.
[0078] exist Figure 1 In the data retrieval device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the data retrieval device of the present invention can be set in the data retrieval device, and the data retrieval device calls the data retrieval program stored in the memory 1005 through the processor 1001 and executes the data retrieval method provided in the embodiment of the present invention.
[0079] This invention provides a data retrieval method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the data retrieval method of the present invention.
[0080] In this embodiment, the data retrieval method includes the following steps:
[0081] Step S10: In response to the index creation command, obtain the primary key information and data table information of the data table.
[0082] It should be noted that the execution entity in this embodiment can be a data retrieval device, or other devices capable of performing the same or similar functions. This embodiment does not impose any limitations on this. The data retrieval device is equipped with an application for index creation and an application for data retrieval and querying, such as a file processing tool for index creation and data retrieval and querying. Thus, the file processing tool is used to create the index and perform data retrieval after index creation.
[0083] The index creation command is used to create an index for the currently stored data. By using the index creation command, an index can be created for the currently stored data, making it easier to directly query the data through the index and improving query efficiency.
[0084] A data table can be a table formed after data is stored in a database. Since the storage of massive amounts of data can consume a lot of disk space, the data can be compressed before index creation to reduce disk usage.
[0085] When aggregating incremental changes across the entire CRM (Customer Relationship Management) table, each piece of incremental data can be transmitted via Kafka in JSON format. One JSON object corresponds to one piece of data in the database table. After receiving the incremental data, business processing is performed, and then each message is saved to disk. Each table corresponds to a type of file, and each file is split into multiple fragments after reaching a certain size. During the archiving process, the fragments are compressed to facilitate subsequent index creation / updating based on the compressed files, thereby reducing disk usage and enabling fast data retrieval with low disk usage.
[0086] The data table contains multiple tables, each storing several files. Each file records a primary key. Therefore, the primary key information of the data table can be obtained by retrieving the data table. The data table information is the file information corresponding to the data information of each data table.
[0087] Step S20: Decompose the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data.
[0088] In practice, the data table can be decomposed using primary key information and data table information, thereby splitting the files in each data table and generating a primary key fingerprint for the primary key corresponding to each split file. This yields data table grouping information containing multiple sets of primary key fingerprint data, i.e., which primary key fingerprints each file in each target table has, and thus can be saved according to their natural order in the data files.
[0089] Step S30: Update the data table grouping information to the first index file according to the preset partitioning rules to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment.
[0090] It is understandable that the preset partitioning rules may include the set segment boundaries and primary key positions. By setting the segment boundaries and primary key positions, the data table grouping information of multiple sets of primary key fingerprint data is updated to the first index file. The first index file is used to store the data of the corresponding data table grouping information, thereby generating the first index information after storage. The first index information includes the segment where each primary key fingerprint is located and the offset of the segment. The offset is the specific position of the primary key fingerprint in the segment. The first index information is the primary key fingerprint level index.
[0091] Step S40: Based on the first index information, summarize the file information and the corresponding primary key fingerprint data corresponding to the data information of the data table to obtain the second index information.
[0092] It should be understood that after obtaining the first index information, all files in each table of the data table and their corresponding primary key fingerprint information can be summarized to obtain the second index information, which is a file-level index.
[0093] Step S50: Perform data retrieval based on the first index information and the second index information.
[0094] It should be noted that by adopting a two-level index structure of file-level index and primary key fingerprint index, the index processing efficiency and management capacity can be greatly improved, thereby easily supporting fast querying of massive amounts of data, improving data query efficiency, and conforming to the business query of the core business. By building an index mechanism on the basis of compressed data, less storage space is occupied, saving a lot of storage space and reducing storage resource consumption.
[0095] However, when data retrieval is required, the first index information and the second index information can be used to retrieve the target data and improve retrieval efficiency.
[0096] In this embodiment, in response to an index creation command, the primary key information and data table information of the data table are obtained. Based on the primary key information and the data table information, the data table is decomposed to obtain data table grouping information including multiple sets of primary key fingerprint data. The data table grouping information is updated to a first index file using a preset partitioning rule to obtain first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of that segment. Based on the first index information, the file information corresponding to the data information of the data table and the corresponding primary key fingerprint data are summarized to obtain second index information. Data retrieval is performed based on the first index information and the second index information. By segmenting the data table, a two-level index is created based on the position of the primary key fingerprint within the segment, greatly improving index processing efficiency. The primary key fingerprint indexing mechanism meets business retrieval needs, improving data retrieval efficiency and accuracy.
[0097] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the data retrieval method of the present invention.
[0098] Based on the first embodiment described above, step S20 of the data retrieval method in this embodiment specifically includes:
[0099] Step S201: Obtain the file information corresponding to the data information of each data table based on the data table information.
[0100] It should be noted that the data table information includes multiple data tables, and each data table stores multiple file data. Therefore, the file information corresponding to the data information of each data table can be obtained based on the data table information.
[0101] Step S202: The file information is segmented to obtain multiple sub-file information.
[0102] Specifically, by splitting the files in the file information, the files in each table can be divided into multiple sub-files, resulting in multiple sub-file information. For example, if the file included in Table 1 is file 1, then file 1 can be split into n sub-files.
[0103] Step S203: Generate a fingerprint based on each primary key in the multiple sub-file information to obtain a primary key fingerprint, wherein each primary key in the primary key information corresponds to each of the sub-file information.
[0104] It should be noted that a primary key fingerprint can be obtained by generating a fingerprint from the primary keys of each sub-file in the information of multiple sub-files. When a file is split into multiple sub-files, the primary key information includes multiple primary keys, and each primary key corresponds one-to-one with each sub-file. A primary key fingerprint can be obtained by generating a fingerprint for the primary key corresponding to each sub-file in file 1 using a 32-bit hash algorithm.
[0105] Step S204: Group and store the primary key fingerprints belonging to the same file and the primary key fingerprints belonging to the same table to obtain data table grouping information containing multiple groups of primary key fingerprint data.
[0106] In practical implementation, primary key fingerprints belonging to the same file and primary key fingerprints belonging to the same table can be grouped and stored to obtain data table grouping information containing multiple sets of primary key fingerprint data. For example, all primary key fingerprint information can be stored according to a two-level storage structure of table → file, thereby obtaining which primary key fingerprints each file of each table has, and obtaining data table grouping information containing multiple sets of primary key fingerprints.
[0107] For example, in CRM business, the data table that needs to be indexed stores customer-related information. The primary key information includes customer name, ID number, mobile phone number, service plan information, and whether the customer is a VIP. This primary key information can be obtained by analyzing the data table.
[0108] This embodiment obtains file information corresponding to the data information of each data table based on the data table information; it then segments the file information to obtain multiple sub-file information; it generates a fingerprint based on each primary key in the multiple sub-file information to obtain a primary key fingerprint, where each primary key in the primary key information corresponds to a specific sub-file information; it groups and stores the primary key fingerprints belonging to the same file and the primary key fingerprints belonging to the same table to obtain data table grouping information containing multiple sets of primary key fingerprint data. This allows for the grouping and organization of files in each table within the data table, generating a fingerprint for each primary key, and storing the primary key fingerprint in the corresponding file, thereby improving the efficiency of grouping.
[0109] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the data retrieval method of the present invention.
[0110] Based on the first embodiment described above, step S30 of the data retrieval method in this embodiment specifically includes:
[0111] Step S301: Obtain the amount of data in the data table.
[0112] It should be noted that when segmenting files in a data table, the amount of data needs to be considered. This can be done by analyzing the data in the table to determine the amount of data.
[0113] Step S302: Set the segmentation rules for the first index file based on the data volume.
[0114] In practice, the segmentation rule of the first index file can be based on a reasonable size as the segmentation limit. For example, if the data volume of the data table is 10GB, the segmentation rule limit of the first index file can be set to 2GB. Then the segmentation rule of the first index file is: segment every 2GB of data.
[0115] Step S303: Segment the files in the data table according to the segmentation rules to obtain the first segmentation information.
[0116] It is understandable that by setting segmentation rules, the data can be segmented according to the segmentation rules as the data volume continues to grow, thus obtaining the first segmentation rule after the segmentation rules.
[0117] Step S304: Obtain the primary key fingerprint corresponding to each file in the first segment information based on the grouping information of the data table.
[0118] In this embodiment, in addition to considering the amount of data during the segmentation process, the primary key factor can also be considered to make the segmentation more accurate. The primary key fingerprint corresponding to each file in the first segmentation information can be obtained through the data table grouping information.
[0119] Step S305: Divide files belonging to the same primary key fingerprint into the same segment to obtain the second segment information.
[0120] It should be noted that data belonging to the same primary key can be divided into the same segment. At this time, the data volume of the segment may be greater than or less than 2GB, thus obtaining the second segment information. The second segment information is the data after the fingerprints of the same primary key are divided into the same segment.
[0121] Step S306: Store the grouping information of each table in the data table and the primary key fingerprint of each file into the first index file using the second segmentation information to obtain the first index information.
[0122] In practical implementation, the grouping information of each table in the data table and the primary key fingerprint of each file can be stored in the first index file through the second segment information, thereby obtaining the first index information. By saving the grouping information of each table to the first index file of the corresponding table, and storing the primary key fingerprint of each file consecutively, the obtained first index information includes the segment where each primary key fingerprint is located and its specific position in the segment, i.e., the offset.
[0123] This embodiment obtains the data volume of a data table; sets segmentation rules for a first index file based on the data volume; segments the files in the data table according to the segmentation rules to obtain first segmentation information; obtains the primary key fingerprints corresponding to each file in the first segmentation information based on the grouping information of the data table; divides files belonging to the same primary key fingerprint into the same segment to obtain second segmentation information; stores the grouping information of each table and the primary key fingerprints of each file in the first index file using the second segmentation information to obtain first index information. By segmenting the files in the data table and grouping the primary key fingerprints of each file together to form an index of primary key fingerprints, file indexing can be quickly performed using the first index information when retrieving data.
[0124] refer to Figure 5 , Figure 5 This is a flowchart illustrating the fourth embodiment of the data retrieval method of the present invention.
[0125] Based on the first embodiment described above, step S40 of the data retrieval method in this embodiment specifically includes:
[0126] Step S401: Obtain the file information and primary key fingerprint data corresponding to the data information of the data table based on the first index information.
[0127] It should be noted that after the first index information is established or updated, the file information of each file in the data packet and the primary key fingerprint data can be obtained.
[0128] Step S402: Obtain the file number of each file based on the file information.
[0129] It should be noted that the file information includes the file number of the file, which is the file identifier of each file.
[0130] Step S403: Obtain the segment number and offset of the primary key fingerprint data.
[0131] The primary key fingerprint data includes the segment number of the primary key and the position of the primary key within that segment, i.e., the offset. Therefore, the segment number and offset of the primary key fingerprint data can be obtained. The segment number of the primary key fingerprint data is the segment identifier of each segment included in each file, and the offset is the specific position of the data corresponding to each primary key fingerprint within the segment of the first index file.
[0132] Step S404: Obtain the byte length of the primary key fingerprint of the file in the data table, and use the byte length of the primary key fingerprint as the index length.
[0133] It should be noted that the byte length of the primary key fingerprint of a file in the data table can be obtained directly from the primary key fingerprint information, and the byte length of the primary key fingerprint is the index length.
[0134] Step S405: Calculate the fingerprint value of the primary key to obtain the primary key range.
[0135] It should be noted that the primary key range can be the fingerprint range of each primary key fingerprint. Therefore, the primary key range can be obtained by calculating the primary key using hash algorithms with different principles. For example, by using three different 32-bit hash algorithms, three primary key fingerprints can be generated to form their respective fingerprint ranges, that is, the maximum and minimum fingerprint values of each hash algorithm.
[0136] Step S406: Summarize the file number, segment number, offset, index length, and primary key range to obtain the second index information.
[0137] In practice, the file number, segment number, offset, index length, and primary key range can be summarized to obtain the second index information.
[0138] It should be noted that when creating an index, the number of concurrent threads for establishing the first and second index information can be adjusted according to the number of data tables to be processed, thereby speeding up the index creation process and enabling smooth performance upgrades after resource expansion.
[0139] This embodiment obtains file information and primary key fingerprint data corresponding to the data information in the data table based on the first index information; obtains the file number of each file according to the file information; obtains the segment number and offset of the primary key fingerprint data; obtains the byte length of the primary key fingerprint of the file in the data table, and uses the byte length of the primary key fingerprint as the index length; calculates the fingerprint value of the primary key to obtain the primary key range; summarizes the file number, segment number, offset, index length, and primary key range to obtain the second index information; obtains the information in the data table through the first index information, thereby summarizing the file number, segment number, offset, index length, and primary key range to construct the second index information. By establishing a file-level index structure, the accuracy of retrieval is improved.
[0140] In some embodiments, before obtaining the primary key information and data table information of the data table, the method further includes: scanning the original compressed data file to obtain data file information; obtaining the file number, file name, start time and end time of file storage based on the data file information; and establishing a file index list using the file number, file name, start time and end time.
[0141] It should be noted that the uploaded raw compressed data files can be scanned using an index creation program. For the compressed files, an online read-compression method is used to directly process each data entry, obtaining data file information. This information includes the file number, file name, start time, and end time. Therefore, the file number, file name, and the start and end times of file storage can be obtained from this information, thus building a file index list. This facilitates initial data retrieval using the file index list, further improving retrieval efficiency.
[0142] Accordingly, obtaining the primary key information and data table information of the data table includes: parsing the data file information to obtain the primary key information and data table information of the data table.
[0143] It should be noted that after the file index list is established, more precise retrieval methods can be further developed. Therefore, the data file information can be parsed to obtain the primary key information and data table information of the data table.
[0144] like Figure 6 As shown, Figure 6The diagram illustrates the principle of creating a data index. It describes how to obtain and scan the original compressed data file to retrieve file information. A file index list is then built based on the file number, filename, start time, and end time. After the file index list is established, the data file information is parsed to obtain the primary key information and table information. The files in each table are then split into multiple sub-files. For example, the files in Table 1 are split into multiple sub-files, namely file 1, file 2, ..., file n. Each sub-file has a corresponding primary key. By generating a fingerprint of the primary key, the primary key fingerprints of each sub-file are obtained, namely key-value fingerprint 1, key-value fingerprint 2, ..., key-value fingerprint n. By storing data table information containing multiple sets of primary key fingerprints in a first index file, and segmenting the files in the data table according to segmentation rules to obtain several segments of data, and further dividing them using the primary key fingerprints to obtain primary key fingerprint data blocks, i.e., key-value fingerprint data blocks, the first index information is obtained. The first index information is a primary key fingerprint feature-level index structure. The file number of each file in the data table is obtained through the first index information, and the segment number, offset, index length, and primary key range are obtained based on the primary key fingerprint data. The second index information is then established using the above data, which is a file feature-level index structure. In the above index creation process, steps one through four can be processed in a distributed and parallel manner. The process of generating the file index list and primary key fingerprints can form concurrent threads based on the number of data files scanned. The summary results of the first two steps will be categorized according to the table. The creation process of the first and second index information can use an appropriate number of concurrent threads based on the amount of table data to be processed. Through the above distributed processing, the processing speed of index creation can be greatly accelerated, and the performance can be smoothly upgraded after resource expansion.
[0145] refer to Figure 7 , Figure 7 This is a flowchart illustrating the fifth embodiment of the data retrieval method of the present invention.
[0146] Based on the first embodiment described above, step S50 of the data retrieval method in this embodiment specifically includes:
[0147] S501: Obtain search information and a list of primary keys to be searched.
[0148] It should be noted that when a user needs to perform a data query, the system can obtain the user's search information and the list of primary keys to be searched. By receiving merged query requests from different primary key fingerprints from different tables, the system can maximize the use of data retrieval to complete the reading in one go, thereby improving the overall processing efficiency.
[0149] S502: Preprocess the list of primary keys to be retrieved to obtain multiple fingerprints of primary keys to be retrieved.
[0150] Understandably, preprocessing the primary key list to be retrieved includes validating and standardizing the primary key list, as well as calculating the primary key using three different types of 32-bit hash algorithms to obtain three valid primary key fingerprints for subsequent retrieval.
[0151] S503: Search the second index information by the range of multiple primary key fingerprints to be searched to obtain the initial index segment.
[0152] In practice, when a search is required, the search can be performed first in the second index information to obtain the initial index segments.
[0153] Optionally, the search can be performed by calculating the range of the three primary key fingerprints to be retrieved, identifying the index segments that the primary key fingerprints can hit, and locating the index segments by using the range of the three key fingerprints. This can significantly narrow down the search scope and improve search efficiency. The initial index segment includes: file number, segment number, offset, index length, and the primary key fingerprint it contains.
[0154] S504: The initial index segments are filtered using the file index list and the retrieval information to obtain filtered index segments.
[0155] In practice, by using a pre-established list of file indexes to filter the initial index segments, the initial index segments can be filtered based on information such as the time, source, and name specified by the user, thereby further narrowing the search scope and obtaining filtered index segments.
[0156] S505: The primary key fingerprint corresponding to the filtered index segment is obtained by searching the first index information using the offset and index length in the filtered index segment.
[0157] It should be noted that after obtaining the filtered index segments, the first index information can be used for further retrieval. By obtaining the offset and index length in the filtered index segments, the primary key fingerprint corresponding to the filtered index segment can be quickly found from the first index information.
[0158] S506: Compare the primary key fingerprint corresponding to the filtered index segment with the primary key fingerprint to be retrieved, and use the file corresponding to the primary key fingerprint that matches the filter index segment as the initial data information.
[0159] It should be understood that by comparing the primary key fingerprint corresponding to the filtered index segment with the primary key fingerprint to be retrieved, a primary key fingerprint that matches the primary key fingerprint to be retrieved is obtained, and the file corresponding to this primary key fingerprint is used as the initial data information.
[0160] S507: Based on the initial data information, search the original compressed data file to obtain the target data information.
[0161] It should be noted that once the initial data information is obtained, the target data information can be obtained by searching the original compressed data file using the initial data information. For example, the initial data information can be grouped, and the target data information can be obtained by searching the original compressed data file based on the grouped data information.
[0162] Further, the step of retrieving target data information from the original compressed data file based on the initial data information includes: obtaining a data table of the initial data information; obtaining the order of each primary key fingerprint in the data table to obtain the relative position of the primary key fingerprint; establishing a correspondence between the primary key fingerprint and the relative position; grouping the initial data information based on a preset grouping structure and the correspondence to obtain multiple groups of data information; and retrieving the target data information from the original compressed data file based on the multiple groups of data information.
[0163] It should be noted that the initial data information can be retrieved from the data table, and the order of each primary key fingerprint in the data table can be obtained to determine the relative position of the primary key fingerprint. The relative position indicates which position of the primary key fingerprint appears in the same data file within the table. Since determining the data table information for each record is much faster than parsing the complete data content, this relative position information can be used to quickly locate the required record in a large data file. A correspondence between primary key fingerprints and their relative positions can then be established. This correspondence, along with a pre-defined grouping structure, can then be used to group the initial data information into multiple sets of data.
[0164] The preset grouping structure is a three-level storage structure of file → table → primary key fingerprint list. That is, each table in each data file is a list of primary key fingerprints involved in this search, thereby grouping the initial data information to obtain multiple sets of data information.
[0165] In practice, once multiple sets of data information are obtained, the target data information can be obtained by searching through these multiple sets of data information in the original compressed file.
[0166] Further, the step of retrieving target data information from the original compressed data file based on the multiple sets of data information includes: obtaining relative position and file information based on the multiple sets of data information; retrieving reference data information from the original compressed data file based on the relative position and the file information; comparing the primary key fingerprint of the reference data information with the primary key fingerprint to be retrieved, and removing the reference data information corresponding to the primary key fingerprint that is inconsistent with the primary key fingerprint to be retrieved, thereby obtaining the target data information.
[0167] Once multiple sets of data information are obtained, they can be searched within the original compressed data file. This allows for quick location of the corresponding data record based on file information and relative position, thus obtaining the target data information.
[0168] The retrieval data in this solution can be of various types, such as JSON, CSV, TSV, etc. This embodiment does not limit this; this embodiment uses JSON format data as an example for illustration.
[0169] like Figure 8 As shown, Figure 8 This diagram illustrates the data retrieval principle. By obtaining a list of primary keys to be retrieved, a first-level index-based retrieval (file feature-level retrieval) is performed. All primary keys are verified and standardized. Then, the primary keys are calculated using one of the three pre-defined 32-bit hash algorithms to form three valid primary key fingerprints for subsequent retrieval. The range of these three fingerprints is used to identify the index segments that the primary key can be matched with. Figure 8The first-level index segment is located by segmenting the index using three primary key fingerprint ranges, which can significantly narrow the search scope and improve search efficiency. Index segment information includes: file number, segment number, offset, index length, and the included primary key, and is filtered using a file index list. By specifying the time, source, and name in the user search, the obtained index segments can be filtered to further narrow the search scope. The search is then further refined using a second-level data table index. That is, using the filtered index segments, offsets, and index lengths, all primary key fingerprints of the current index segment can be quickly found from the second-level data index and compared with the primary key fingerprint to be searched, thus obtaining the specific data information that has been hit. The data information hit by the second-level index is grouped. The grouping structure uses a three-level storage structure of file → table → key-value fingerprint list, that is, a list of key-value fingerprints involved in each table of each data file in this search. Crucially, in the third-level storage structure, a pairing relationship between relative position and key-value fingerprint must be established. Here, relative position refers to the position of the key value in all the data of this table within the same data file. Since determining the data table information for each record is significantly faster than parsing the complete data content during data retrieval, this relative position information can be used to quickly locate the required record in a large data file. Using the resulting grouping information, a search is performed within the corresponding original compressed data file, quickly locating the relevant data record through file information and relative position. Furthermore, because hash algorithms can experience collisions, a successful primary key fingerprint has a very small chance of being a false hit. Therefore, a secondary check can be performed on the primary key of the retrieved data record; if it does not match the primary key fingerprint to be retrieved, it is discarded, thus avoiding this false hit problem.
[0170] By establishing an indexing mechanism based on compressed JSON data files, storage space usage is reduced, thus saving a significant amount of storage space. Furthermore, index creation and data retrieval are handled entirely by the application, eliminating the need for dedicated DBAs for maintenance and management.
[0171] The relative position and file information are obtained by using multiple sets of data. Based on the relative position and file information, a search is performed in the original compressed data file to obtain reference data information. Since hash algorithms may have collisions, the obtained consistent primary key fingerprints may be false hits. Therefore, the reference data information needs to be screened again. The primary key fingerprints of the reference data information are verified a second time. The reference data information corresponding to the primary key fingerprints that are inconsistent with the primary key fingerprint to be searched is removed, thus obtaining the removed data information, i.e., the target data information, thereby improving the accuracy of the search.
[0172] This embodiment acquires retrieval information and a list of primary keys to be retrieved; preprocesses the list of primary keys to be retrieved to obtain multiple primary key fingerprints; searches the second index information using the range of the multiple primary key fingerprints to obtain initial index segments; filters the initial index segments using a file index list and the retrieval information to obtain filtered index segments; searches the first index information using the offset and index length of the filtered index segments to obtain the primary key fingerprints corresponding to the filtered index segments; compares the primary key fingerprints corresponding to the filtered index segments with the primary key fingerprints to be retrieved, and uses the files corresponding to the matching primary key fingerprints as initial data information; and searches the original compressed data files based on the initial data information to obtain the target data information, thereby accelerating the data retrieval process and improving the retrieval accuracy.
[0173] Reference Figure 9 , Figure 9 This is a structural block diagram of the first embodiment of the data retrieval device of the present invention.
[0174] like Figure 9 As shown, the data retrieval device proposed in this embodiment of the invention includes:
[0175] Module 10 is used to retrieve the primary key information and data table information of the data table in response to the index creation command.
[0176] The decomposition module 20 is used to decompose the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data.
[0177] The update module 30 is used to update the data table grouping information to the first index file according to the preset partitioning rules to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment.
[0178] The aggregation module 40 is used to aggregate the file information and the corresponding primary key fingerprint data corresponding to the data information of the data table based on the first index information to obtain the second index information.
[0179] The retrieval module 50 is used to perform data retrieval based on the first index information and the second index information.
[0180] In this embodiment, in response to an index creation command, the primary key information and data table information of the data table are obtained. Based on the primary key information and the data table information, the data table is decomposed to obtain data table grouping information including multiple sets of primary key fingerprint data. The data table grouping information is updated to a first index file using a preset partitioning rule to obtain first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of that segment. Based on the first index information, the file information corresponding to the data information of the data table and the corresponding primary key fingerprint data are summarized to obtain second index information. Data retrieval is performed based on the first index information and the second index information. By segmenting the data table, a two-level index is created based on the position of the primary key fingerprint within the segment, greatly improving index processing efficiency. The primary key fingerprint indexing mechanism meets business retrieval needs, improving data retrieval efficiency and accuracy.
[0181] In one embodiment, the decomposition module 20 is further configured to obtain file information corresponding to the data information of each data table based on the data table information; to segment the file information to obtain multiple sub-file information; to generate a fingerprint based on each primary key in the multiple sub-file information to obtain a primary key fingerprint, wherein each primary key in the primary key information corresponds to each of the sub-file information; and to group and store the primary key fingerprints belonging to the same file and the primary key fingerprints belonging to the same table to obtain data table grouping information including multiple sets of primary key fingerprint data.
[0182] In one embodiment, the update module 30 is further configured to: obtain the data volume of the data table; set segmentation rules for the first index file based on the data volume; segment the files in the data table according to the segmentation rules to obtain first segmentation information; obtain the primary key fingerprint corresponding to each file in the first segmentation information based on the data table grouping information; divide files belonging to the same primary key fingerprint into the same segment to obtain second segmentation information; and store the grouping information of each table and the primary key fingerprint of each file in the data table into the first index file using the second segmentation information to obtain first index information.
[0183] In one embodiment, the summarization module 40 is further configured to: obtain file information and primary key fingerprint data corresponding to the data information of the data table based on the first index information; obtain the file number of each file according to the file information; obtain the segment number and offset of the primary key fingerprint data; obtain the byte length of the primary key fingerprint of the file in the data table, and use the byte length of the primary key fingerprint as the index length; perform fingerprint value calculation on the primary key to obtain the primary key range; and summarize the file number, the segment number, the offset, the index length, and the primary key range to obtain the second index information.
[0184] In one embodiment, the acquisition module 10 is further configured to scan the original compressed data file to obtain data file information; obtain the file number, file name, start time and end time of file storage based on the data file information; and establish a file index list using the file number, file name, start time and end time. Correspondingly, the acquisition of the primary key information and data table information of the data table includes: parsing the data file information to obtain the primary key information and data table information of the data table.
[0185] In one embodiment, the retrieval module 50 is further configured to: acquire retrieval information and a list of primary keys to be retrieved; preprocess the list of primary keys to be retrieved to obtain multiple primary key fingerprints to be retrieved; search the second index information using the range of the multiple primary key fingerprints to be retrieved to obtain an initial index segment; filter the initial index segment using a file index list and the retrieval information to obtain a filtered index segment; search the first index information using the offset and index length in the filtered index segment to obtain the primary key fingerprint corresponding to the filtered index segment; compare the primary key fingerprint corresponding to the filtered index segment with the primary key fingerprint to be retrieved, and use the file corresponding to the primary key fingerprint that matches the filter as initial data information; and search the original compressed data file based on the initial data information to obtain target data information.
[0186] In one embodiment, the retrieval module 50 is further configured to: obtain a data table of the initial data information; obtain the order of each primary key fingerprint in the data table to obtain the relative position of the primary key fingerprint; establish a correspondence between the primary key fingerprint and the relative position; group the initial data information based on a preset grouping structure and the correspondence to obtain multiple groups of data information; and perform a retrieval in the original compressed data file based on the multiple groups of data information to obtain target data information.
[0187] In one embodiment, the retrieval module 50 is further configured to obtain relative position and file information based on multiple sets of data information; perform a retrieval in the original compressed data file based on the relative position and the file information to obtain reference data information; compare the primary key fingerprint of the reference data information with the primary key fingerprint to be retrieved, and remove the reference data information corresponding to the primary key fingerprint that is inconsistent with the primary key fingerprint to be retrieved to obtain target data information.
[0188] Furthermore, to achieve the above objectives, the present invention also proposes a data retrieval device, which includes: a memory, a processor, and a data retrieval program stored in the memory and executable on the processor, the data retrieval program being configured to implement the steps of the data retrieval method described above.
[0189] Since this data retrieval device adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be described in detail here.
[0190] Furthermore, embodiments of the present invention also propose a storage medium storing a data retrieval program, which, when executed by a processor, implements the steps of the data retrieval method described above.
[0191] Since this storage medium adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.
[0192] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0193] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0194] In addition, for technical details not described in detail in this embodiment, please refer to the data retrieval method provided in any embodiment of the present invention, which will not be repeated here.
[0195] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0196] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0198] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A data retrieval method, characterized by, The data retrieval method includes: In response to an index creation command, retrieve the primary key information and table information of the data table; Based on the primary key information and the data table information, the data table is decomposed to obtain data table grouping information including multiple sets of primary key fingerprint data; The data table grouping information is updated to the first index file by a preset partitioning rule to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment. Based on the first index information, the file information and the corresponding primary key fingerprint data corresponding to the data information of the data table are summarized to obtain the second index information, wherein the first index information is a primary key fingerprint level index and the second index information is a file level index; Data retrieval is performed based on the first index information and the second index information; The data retrieval based on the first index information and the second index information includes: Retrieve search information and the list of primary keys to be searched; The list of primary keys to be retrieved is preprocessed to obtain multiple fingerprints of primary keys to be retrieved; The initial index segments are obtained by searching the range of multiple primary key fingerprints to be retrieved in the second index information. The initial index segments are filtered using the file index list and the retrieval information to obtain filtered index segments; The primary key fingerprint corresponding to the filtered index segment is obtained by retrieving the first index information using the offset and index length in the filtered index segment. Compare the primary key fingerprint corresponding to the filtered index segment with the primary key fingerprint to be retrieved, and use the file corresponding to the primary key fingerprint that matches the primary key fingerprint as the initial data information. Based on the initial data information, the target data information is obtained by searching the original compressed data file.
2. The data retrieval method of claim 1, wherein, The data table is decomposed based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data, including: Based on the data table information, obtain the file information corresponding to the data information of each data table; The file information is split into multiple sub-file information; A fingerprint is generated based on each primary key in the multiple sub-file information to obtain a primary key fingerprint, wherein each primary key in the primary key information corresponds to each of the sub-file information. The primary key fingerprints belonging to the same file and the primary key fingerprints belonging to the same table are grouped and stored to obtain data table grouping information containing multiple groups of primary key fingerprint data.
3. The data retrieval method of claim 1, wherein, The step of updating the data table grouping information to the first index file according to a preset partitioning rule to obtain the first index information includes: Get the amount of data in the data table; The segmentation rules for the first index file are set based on the amount of data. The files in the data table are segmented according to the segmentation rules to obtain the first segmentation information; Based on the grouping information in the data table, obtain the primary key fingerprint corresponding to each file in the first segment information; Files belonging to the same primary key fingerprint are divided into the same segment to obtain the second segment information; The grouping information of each table in the data table and the primary key fingerprint of each file are stored in the first index file using the second segment information to obtain the first index information.
4. The data retrieval method of claim 1, wherein, The step of summarizing the file information and corresponding primary key fingerprint data corresponding to the data information of the data table based on the first index information to obtain the second index information includes: Based on the first index information, obtain the file information and primary key fingerprint data corresponding to the data information of the data table; The file number of each file is obtained based on the file information; Obtain the segment number and offset of the primary key fingerprint data; Obtain the byte length of the primary key fingerprint of the file in the data table, and use the byte length of the primary key fingerprint as the index length; Calculate the fingerprint value of the primary key to obtain the primary key range; The file number, segment number, offset, index length, and primary key range are summarized to obtain the second index information.
5. The data retrieval method of any one of claims 1 to 4, wherein, Before obtaining the primary key information and data table information of the data table, the following steps are also included: Scan the original compressed data file to obtain data file information; The file number, file name, start time, and end time of file storage are obtained based on the data file information. A file index list is created using the file number, the file name, the start time, and the end time. Accordingly, obtaining the primary key information and data table information of the data table includes: The data file information is parsed to obtain the primary key information and data table information of the data table.
6. The data retrieval method of claim 1, wherein, The step of retrieving target data information from the original compressed data file based on the initial data information includes: The data table for obtaining the initial data information; Obtain the order of each primary key fingerprint in the data table to get the relative position of the primary key fingerprint; Establish the correspondence between the primary key fingerprint and the relative position; The initial data information is grouped based on a preset grouping structure and the corresponding relationship to obtain multiple groups of data information. Based on the multiple sets of data information, the target data information is obtained by searching the original compressed data file.
7. The data retrieval method of claim 6, wherein, The process of retrieving target data information from the original compressed data file based on the multiple sets of data information includes: The relative location and file information are obtained based on the multiple sets of data information described above; Based on the relative position and the file information, a search is performed in the original compressed data file to obtain reference data information; The primary key fingerprint of the reference data information is compared with the primary key fingerprint to be retrieved. The reference data information corresponding to the primary key fingerprint that is inconsistent with the primary key fingerprint to be retrieved is removed to obtain the target data information.
8. A data retrieval apparatus, characterized by comprising: The data retrieval device includes: The retrieval module is used to retrieve the primary key information and data table information of the data table in response to the index creation command; The decomposition module is used to decompose the data table based on the primary key information and the data table information to obtain data table grouping information including multiple sets of primary key fingerprint data; The update module is used to update the data table grouping information to the first index file according to the preset partitioning rules to obtain the first index information, wherein the first index information includes the segment where each primary key fingerprint is located and the offset of the segment. The aggregation module is used to aggregate the file information and the corresponding primary key fingerprint data of the data information of the data table based on the first index information to obtain the second index information, wherein the first index information is a primary key fingerprint level index and the second index information is a file level index; The retrieval module is used to perform data retrieval based on the first index information and the second index information; The retrieval module is further configured to: acquire retrieval information and a list of primary keys to be retrieved; preprocess the list of primary keys to be retrieved to obtain multiple primary key fingerprints to be retrieved; search the second index information using the range of the multiple primary key fingerprints to be retrieved to obtain initial index segments; filter the initial index segments using a file index list and the retrieval information to obtain filtered index segments; search the first index information using the offset and index length of the filtered index segments to obtain the primary key fingerprints corresponding to the filtered index segments; compare the primary key fingerprints corresponding to the filtered index segments with the primary key fingerprints to be retrieved, and use the files corresponding to the primary key fingerprints that match as initial data information; and search the original compressed data files based on the initial data information to obtain target data information.
9. A data retrieval apparatus, characterized by comprising: The data retrieval device includes: a memory, a processor, and a data retrieval program stored in the memory and executable on the processor, the data retrieval program being configured to implement the data retrieval method as described in any one of claims 1 to 7.
10. A storage medium, characterized by The storage medium stores a data retrieval program, which, when executed by a processor, implements the data retrieval method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data indexing method and device and electronic equipment
CN113157689A
Data query method, nonvolatile storage medium and electronic equipment
CN113312313A