Data Reading Method and Device, Electronic Device, Storage Medium

By constructing a table-level composite index with preset Bloom filters and double-layer sparse indexes in the OLAP scenario, the problem of reading performance reduction caused by network interaction between primary and secondary nodes is solved, and more efficient data reading is achieved.

CN119357188BActive Publication Date: 2025-06-24本原数据(北京)信息技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411429436.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-06-24
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

In the OLAP scenario, unnecessary network interaction between the master and spare nodes leads to a decrease in the read performance of the database nodes, affecting the efficiency of data reading.

Method used

By building a table-level composite index with preset Bloom filters and preset double-layer sparse indexes, the index pages and data pages of candidate areas are divided, and data reading is performed based on these structures to avoid unnecessary network interactions.

Benefits of technology

It effectively avoids unnecessary network interaction between the master and spare nodes, improves the read performance of the database nodes, and thus improves the efficiency of data reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357188B_ABST
    Figure CN119357188B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a data reading method, apparatus, electronic device, and storage medium, belonging to the technical field of databases. The method includes: determining the number of first pages occupied by a corresponding table-level composite index based on the area size of the candidate area, the preset number of index fields, and the preset index type in the initial data table; dividing the pages included in the candidate area based on the number of first pages to determine the index pages and data pages of the number of first pages; constructing a table-level composite index based on the preset Bloom filter and the preset double-layer sparse index at the district level, and storing the table-level composite index into the index pages; storing the field data of all candidate pages in the candidate area into the data pages, and constructing a target area structure corresponding to the candidate area based on the stored index pages and data pages; constructing a target data table based on multiple target area structures of the same initial data table, and performing data reading based on the target data table. The embodiment of the present application can improve the data reading efficiency between database nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technology, and in particular, to a data reading method and apparatus, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of emerging technologies and services such as big data and the Internet of Things, the amount of data generated is growing at a high speed. When users query data with complex requirements from a large amount of data, it involves not only querying or manipulating one or several records in a relational table, but also performing data analysis and information synthesis on tens of millions of records in multiple tables. Moreover, with the sharp increase in the amount of data, the large amount of data also brings pressure on storage costs. To solve the problem that the storage capacity in a traditional high-availability (HA) deployed database cluster doubles compared to a single machine, related technologies have proposed building a resource pooling architecture by combining segment-page storage to solve the problem of too many file handles in page storage, and it can be well compatible with other database storage layer modules based on page storage currently.

[0003] Currently, under the resource pooling architecture of related technologies, the primary node has read and write permissions, and the standby node only has read permissions. All nodes share a copy of the data. Moreover, in the cluster of database nodes, the primary node manages the status of all resources, and the primary and standby nodes interact through the network. Taking a read request of a standby node as an example, before reading a certain page, the standby node will first apply for the S lock of the page. After judging by the holding situation, the primary node will let the holding node transfer the page to the requesting node through the network and let the requesting node also hold the S lock. This method, in the on-line analytical processing (OLAP) scenario, while bringing real-time consistency of the standby machine, will introduce a large number of network interactions between nodes, which is likely to reduce the read performance of the cluster. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a data reading method and apparatus, an electronic device, and a storage medium, which can avoid unnecessary network interactions between the primary and standby nodes in the OLAP scenario, improve the read performance of database nodes, and thus improve the efficiency of data reading.

[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a data reading method, and the method includes:

[0006] Obtain the region size, the number of preset index fields, and the preset index type in the candidate region of the initial data table, where the initial data table includes a plurality of the candidate regions, and the candidate region includes a plurality of candidate pages;

[0007] Determine the number of first pages occupied by the table-level composite index corresponding to the candidate area based on the area size, the number of preset index fields, and the preset index type;

[0008] Based on the number of first pages, perform page division on the pages included in the candidate area to determine the index pages and data pages of the candidate area, where the number of index pages is the number of first pages;

[0009] Construct the table-level composite index based on a preset Bloom filter and a preset two-layer sparse index, and store the table-level composite index in the index pages; the preset Bloom filter is a Bloom filter constructed in advance at the area level, and the preset two-layer sparse index includes a first sparse index at the area level and a second sparse index at the page level. The first sparse index is used to indicate the data index fields of the candidate area, and the second sparse index is used to indicate the data index fields of each candidate page in the candidate area;

[0010] Store the field data of all candidate pages in the candidate area in the data pages, and construct the target area structure corresponding to the candidate area based on the stored index pages and data pages;

[0011] Construct a target data table based on multiple target area structures of the same initial data table, and perform data reading based on the target data table.

[0012] In some embodiments, the performing data reading based on the target data table includes:

[0013] Receive a read request, where the read request includes the target data table and a query predicate condition;

[0014] Determine a target area list based on the target data table, where the target area list includes multiple target areas;

[0015] Select a to-be-filtered area from multiple target areas, and obtain the target index page corresponding to the to-be-filtered area, where the target index page stores the table-level composite index corresponding to the to-be-filtered area;

[0016] Perform pre-filtering on the query predicate condition based on the table-level composite index corresponding to the to-be-filtered area to obtain a composite filtering result;

[0017] If the composite filtering result indicates that the to-be-filtered area contains data that meets the query predicate condition, perform page query based on the composite filtering result, the query predicate condition, and the data page corresponding to the to-be-filtered area to determine the read page;

[0018] Perform data reading by scanning the read page.

[0019] In some embodiments, pre-filtering the query predicate condition based on the table-level composite index corresponding to the area to be filtered to obtain a composite filtering result includes:

[0020] Performing area-level filtering on the query predicate condition based on the preset Bloom filter corresponding to the area to be filtered and the first sparse index at the district level to obtain an area-level filtering result;

[0021] If the area-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, performing page-level sparse index detection on the table-level composite index corresponding to the area to be filtered based on the query predicate condition to obtain a page-level sparse index detection result;

[0022] If the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered contains the second sparse index to be filtered, performing page-level filtering on the query predicate condition based on the second sparse index to obtain the page-level filtering result;

[0023] Determining the composite filtering result based on the page-level filtering results corresponding to all the second sparse indexes to be filtered.

[0024] In some embodiments, before the step of if the area-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, performing page-level sparse index detection on the table-level composite index corresponding to the area to be filtered based on the query predicate condition to obtain a page-level sparse index detection result, the method further includes:

[0025] If the area-level filtering result indicates that the area to be filtered does not contain data that meets the query predicate condition, updating the area to be filtered based on multiple target areas;

[0026] Performing pre-filtering on the query predicate condition according to the table-level composite index corresponding to the updated area to be filtered to obtain a composite filtering result until pre-filtering on all the target areas is completed.

[0027] In some embodiments, performing a page query based on the composite filtering result, the query predicate condition, and the data page corresponding to the area to be filtered to determine the read page includes:

[0028] If the composite filtering result indicates that data that meets the query predicate condition is matched in the second sparse index of the area to be filtered and the corresponding memory contains a page that meets the query predicate condition, determining the read page from the data page based on the matched second sparse index; or,

[0029] If the composite filtering result indicates that data meeting the conditions of the query predicate is matched in the second sparse index in the area to be filtered, and the corresponding memory does not contain pages meeting the conditions of the query predicate, a data page request is sent to the corresponding distributed memory service module, and the read page returned by the distributed memory service module is received.

[0030] In some embodiments, obtaining the target index page corresponding to the area to be filtered includes:

[0031] If it is detected that the corresponding memory contains multiple index pages of the target area, obtain the target index page corresponding to the area to be filtered; or,

[0032] If it is detected that the corresponding memory does not contain multiple index pages of the target area, an index page request is sent to the corresponding distributed memory service module, and the target index page corresponding to the area to be filtered returned by the distributed memory service module is received.

[0033] In some embodiments, the preset Bloom filter in the candidate area is constructed in the following manner:

[0034] Based on the area size, the number of fields in the table-level composite index, the storage size per unit page, and the preset optimization parameter, perform bit number calculation to obtain the target bit number of the preset Bloom filter;

[0035] Based on the shared bit array with the target bit number and the preset hash function, construct the preset Bloom filter corresponding to the candidate area.

[0036] To achieve the above object, a second aspect of the embodiments of the present application proposes a data reading device, and the device includes:

[0037] An acquisition module, configured to acquire the area size, the preset index field number, and the preset index type of the candidate area in the initial data table, where the initial data table includes multiple candidate areas, and the candidate area includes multiple candidate pages;

[0038] A page number determination module, configured to determine the number of the first pages occupied by the table-level composite index corresponding to the candidate area based on the area size, the preset index field number, and the preset index type;

[0039] A page division module, configured to perform page division on the pages included in the candidate area based on the number of the first pages, and determine the index page and the data page of the candidate area, where the number of the index pages is the number of the first pages;

[0040] A composite index construction module, configured to construct the table-level composite index based on a preset Bloom filter and a preset two-layer sparse index, and store the table-level composite index into the index page; the preset Bloom filter is a pre-constructed Bloom filter based on districts, and the preset two-layer sparse index includes a first sparse index based on districts and a second sparse index based on pages. The first sparse index is used to indicate the data index fields of the candidate districts, and the second sparse index is used to indicate the data index fields of each candidate page in the candidate districts;

[0041] A storage module, configured to store the field data of all the candidate pages in the candidate districts into the data page, and construct a target district structure corresponding to the candidate districts based on the stored index page and data page;

[0042] A data reading module, configured to construct a target data table based on multiple target district structures of the same initial data table, and perform data reading based on the target data table.

[0043] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0044] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed, the method described in the first aspect above is implemented.

[0045] A data reading method and device, an electronic device, and a storage medium proposed in an embodiment of the present application. First, obtain the region size, the number of preset index fields, and the preset index type in the candidate region of the initial data table. The initial data table includes multiple candidate regions, and each candidate region includes multiple candidate pages. Further, based on the region size, the number of preset index fields, and the preset index type, determine the number of first pages occupied by the table-level composite index corresponding to the candidate region. Further, based on the number of first pages, perform page division on the pages included in the candidate region to determine the index pages and data pages of the candidate region, and the number of index pages is the number of first pages. Further, construct a table-level composite index based on a preset Bloom filter and a preset two-layer sparse index, and store the table-level composite index in the index pages. The preset Bloom filter is a Bloom filter constructed in advance at the region level, and the preset two-layer sparse index includes a first sparse index at the region level and a second sparse index at the page level. The first sparse index is used to indicate the data index fields of the candidate region, and the second sparse index is used to indicate the data index fields of each candidate page in the candidate region. Further, store the field data of all candidate pages in the candidate region in the data pages, and construct a target region structure corresponding to the candidate region based on the stored index pages and data pages. Further, construct a target data table based on the target region structures of the same initial data table, and perform data reading based on the target data table. The embodiment of the present application proposes an innovative way to construct a target data table, that is, index pages and data pages are divided in the target region structure of the target data table, and the index pages store the table-level composite index constructed based on the preset Bloom filter and the preset two-layer sparse index, which can play a hierarchical screening effect in subsequent data reading and avoid reading invalid pages. Therefore, compared with the related art, the present application can effectively avoid unnecessary network interactions between the master and standby nodes in the OLAP scenario, improve the read performance of the database nodes, and thus improve the efficiency of data reading. Description of the Drawings

[0046] Figure 1 is a schematic diagram of an architecture of a resource pooling architecture provided by the related art;

[0047] Figure 2 is a schematic diagram of an organization form based on segment-page storage provided by an embodiment of the present application;

[0048] Figure 3 is a first flowchart of the data reading method provided by an embodiment of the present application;

[0049] Figure 4 is a schematic diagram of a structure of a candidate region provided by an embodiment of the present application;

[0050] Figure 5 is a schematic diagram of a structure for dividing pages in the candidate region structure provided by an embodiment of the present application;

[0051] Figure 6 It is a schematic structural diagram of storing a corresponding table-level composite structure in an index page of a candidate area provided by an embodiment of the present application;

[0052] Figure 7A It is a specific structural diagram of storing a table-level composite index provided by an embodiment of the present application;

[0053] Figure 7B It is another specific structural diagram of storing a table-level composite index provided by an embodiment of the present application;

[0054] Figure 8 It is a schematic diagram of a tuple storage structure of a table-level composite index provided by an embodiment of the present application;

[0055] Figure 9 It is Figure 3 a flowchart of step S340 in

[0056] Figure 10 It is a filtering schematic diagram of a preset Bloom filter provided by an embodiment of the present application;

[0057] Figure 11 It is Figure 3 a flowchart of step S360 in

[0058] Figure 12 It is Figure 11 a flowchart of step S1140 in

[0059] Figure 13 It is a specific flowchart of a data reading method provided by an embodiment of the present application;

[0060] Figure 14 It is a schematic structural diagram of a data reading device provided by an embodiment of the present application;

[0061] Figure 15 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0062] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. The terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0065] First, several nouns involved in this application are analyzed as follows:

[0066] Page-based storage: It is a storage method that logically organizes, allocates, and manages storage space in units of pages, and data is stored in pages in a fixed format. Under the page-based storage mechanism, the data of each table independently corresponds to a logically large file (maximum support 32TB), and this logical file is divided into multiple physical files according to a fixed size (default 1GB) and stored in the corresponding directory. A table is divided into multiple partition tables and split into multiple sub-tables, and the number of underlying files required is relatively large.

[0067] Segmented page-based storage: Logically, segmented page-based storage is organized, allocated, and managed in the form of segments, extents, and pages. A data table in the database corresponds to a logical segment, and all the data of this table exists on this segment. Each segment will mount multiple extents, and each extent is a continuous physical page. An extent contains multiple pages. Extents are not necessarily continuous, but the pages within the same extent are continuous. In physical files, extents of the same type are stored in the same physical file, and the maximum size of a single physical file is 1GB. Different physical files store data with different forknums.

[0068] Database node: It refers to an independent entity in a database system that is responsible for storing and managing data. A database node can be a physical server, a virtual server, or other computing resources.

[0069] Scan operator: It is an operator used for traversing the results of tables, result sets, linked list subqueries, etc., and each time it obtains a tuple as the input for the upper-level node. Scan operators include sequential scan (SeqScan), index scan (IndexScan), bitmap scan (BitmapHeapScan), etc.

[0070] Bloom Filter: A probabilistic data structure with high space efficiency, used to determine whether an element is in a set. A Bloom Filter usually consists of a very long binary vector (bit array) and a series of random mapping functions (hash functions). When an element is added to the Bloom Filter, hash values are calculated through hash functions and the corresponding bits are set to 1. When querying, the given element can be hashed again. If all corresponding bits are 1, the element may exist; otherwise, the element must not be in the set.

[0071] Shared Lock (i.e., S Lock): A lock that allows multiple transactions to read a resource simultaneously but not modify it. The S Lock can ensure that the resource can be shared and read among multiple transactions, but prevents modification, thus avoiding data inconsistency. Among them, multiple transactions can hold the S Lock simultaneously, allowing concurrent reads.

[0072] Predicate Condition: A conditional expression used to filter data in programming languages and database queries. Predicate conditions are usually used to determine whether the value of a certain attribute or field meets specific conditions, thereby deciding whether to select or operate on the corresponding data records.

[0073] Distributed Memory Service (DMS): A dynamic library integrated inside the database. It transmits page content through the Transmission Control Protocol (TCP) / Remote Direct Memory Access (TCP / RDMA) network, fuses the primary and standby memories, provides memory pooling capabilities, and thus realizes the real-time consistent read function for the standby machine. Memory pooling realizes the real-time exchange of primary and standby pages through the DMS component of the distributed memory service, providing real-time consistency capabilities for the standby machine. That is, after the host transaction is committed, it can be immediately read on the standby machine without the phenomenon of delayed reading (the transaction isolation level is Read-Committed).

[0074] Transaction Visibility: Refers to which modifications of other transactions can be seen by the read operations on the database within a transaction. This is mainly affected by the transaction isolation level, and different isolation levels determine different data visibility between transactions.

[0075] With the rapid development of emerging technologies and businesses such as big data and the Internet of Things, the amount of data generated is growing at a high speed. When users query data with complex requirements from massive data, it involves not only querying or manipulating one or a few records in a relational table, but also performing data analysis and information integration on tens of millions of records in multiple tables. Moreover, with the sharp increase in the amount of data, the massive data also brings pressure on storage costs. Based on the above problems, researchers in the database field have proposed On-Line Analytical Processing (OLAP) technology. OLAP is specifically designed to support complex analysis operations, focusing on decision support for decision-makers and senior management personnel, and can quickly and flexibly perform complex query processing on large amounts of data according to the requirements of analysts. In the OLAP scenario, its data often has the following characteristics: First, its data volume is extremely large, often at the terabyte (TB) or petabyte (PB) level. Reading operations on this level of data volume pose challenges to the scanning performance of the database; Second, the possibility of its data being changed after being written is relatively low; Third, the actual distribution of its data in the storage medium often shows aggregation on some fields.

[0076] With the sharp increase in the amount of data, the massive data also brings pressure on storage costs. To solve the problem that the storage capacity in a traditional High Availability (HA) deployed database cluster doubles compared to a single machine, related technologies have proposed a way to introduce the resource pooling feature into the database. As Figure 1 shown, Figure 1 is a schematic diagram of the resource pooling architecture provided by related technologies. This resource pooling feature provides a new form of HA deployment, which provides the ability for all nodes in the cluster to share a single storage. While reducing storage costs, it brings excellent features such as real-time consistent reads for standby machines and low Recovery Time Objective (RTO) times.

[0077] Among them, segment-page storage is a storage architecture different from traditional page storage. It can solve the problem of too many file handles in page storage and can be well compatible with other database storage layer modules based on page storage currently. As Figure 2 shown, Figure 2It is a schematic diagram of an organizational form based on segment-page storage provided by an embodiment of the present application. Among them, the possible values of the size of an Extent are 64K (1K = 1024 bytes), 1M (1M = 1024K), 8M, 64M, etc., and the size of a data page is fixed at 8K. Therefore, by combining the segment-page storage method to build a resource pooling architecture, the problem of too many file handles in page storage can be solved, and it can be well compatible with other database storage layer modules based on page storage currently.

[0078] Currently, as Figure 1 shown, under the resource pooling architecture of the related technology, the master node has read and write permissions, and the standby node only has read permissions. All nodes share a copy of the data. In a cluster of database nodes, there are multiple database nodes in a cluster. Resources refer to data pages and regular locks, etc. And the lock and page resources are managed by the DMS module. Among them, the master node manages the status of all resources, and the master and standby nodes interact through the network. Taking a read request of the standby node as an example, the standby node will first apply for the S lock of a certain page before reading the page. For this reason, the standby node will request the master node to hold the S lock of the page, and the master node will respond according to the current status of the page. If no node in the current cluster holds any lock on the page, then it will let the requesting node read from the storage and hold the S lock; if there is a node in the current cluster writing to the page, then the requesting node needs to wait; if there is another node in the cluster holding the S lock, then the master node will let the holding node transfer the page to the requesting node through the network and let the requesting node also hold the S lock. However, this method will introduce a large number of network interactions between nodes while bringing real-time consistency of the standby machine in the On-Line Analytical Processing (OLAP) scenario, which is likely to reduce the read performance of the cluster.

[0079] Based on this, the embodiments of the present application provide a data reading method, device, electronic device, and storage medium, which can avoid unnecessary network interactions between the master and standby nodes in the OLAP scenario, improve the read performance of the database nodes, and thus improve the efficiency of data reading.

[0080] The data reading method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; the software can be an application implementing the data reading method, etc., but is not limited to the above forms.

[0081] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0082] Please refer to Figure 3 , Figure 3 which is an optional flowchart of the data reading method provided by the embodiments of the present application. Figure 3 The method in Figure 3 may specifically include but is not limited to steps S310 to S360. The following will introduce these six steps in detail with reference to

[0083] Step S310, obtain the area size, the number of preset index fields, and the preset index type of the candidate area in the initial data table;

[0084] Step S320, determine the number of the first pages occupied by the table-level composite index corresponding to the candidate area based on the area size, the number of preset index fields, and the preset index type;

[0085] Step S330: Based on the first number of pages, divide the pages included in the candidate area to determine the index page and data pages of the candidate area. The number of index pages is the first number of pages.

[0086] Step S340: Construct a table-level composite index based on a preset Bloom filter and a preset two-layer sparse index, and store the table-level composite index in the index page.

[0087] Step S350: Store the field data of all candidate pages in the candidate area into the data pages, and construct a target area structure corresponding to the candidate area based on the stored index pages and data pages.

[0088] Step S360: Construct a target data table based on the target area structures of multiple initial data tables of the same initial data table, and perform data reading based on the target data table.

[0089] In steps S310 to S360 of some embodiments, the data reading method provided by the embodiments of the present application can function in both traditional clusters using segment-page storage and resource-pooled architecture clusters. That is, when facing a large amount of read operations in a traditional cluster using segment-page storage, this method can filter multiple data pages at once through a composite index at the Extent level, avoiding invalid disk I / O. In addition, the embodiments of the present application can be better applied to tables using segment-page storage, thereby greatly improving the data read performance.

[0090] In step S310 of some embodiments, the initial data table refers to a data table initially set up based on a segment-page storage structure. The storage of this initial data table is organized, stored, allocated, and managed in the form of segments, extents, and pages. An initial data table corresponds to a logically segment, and all data of this table exists on this segment. This initial data table includes multiple candidate areas (that is, each segment will mount multiple extents), and a candidate area includes multiple candidate pages (that is, each extent will contain multiple pages).

[0091] Among them, the extent size of the candidate area refers to the space occupied by the candidate area in the initial data table, which can be the number of rows, the number of pages, or other metrics. The preset number of index fields refers to the number of fields used to construct the table-level composite index. By setting the preset number of index fields and the corresponding preset index fields, it can be known which fields will be used to construct the index to accelerate the retrieval operation, and these fields should be selected according to the query usage frequency and data distribution. The preset index type refers to a preset index type, and the index type can include Bloom index and sparse index. In addition, the present application does not specially limit the space used by the index page.

[0092] In step S320 of some embodiments, further, the number of first pages occupied by the table-level composite index corresponding to the candidate extent may be determined based on the extent size, the number of preset index fields, and the preset index type.

[0093] In step S330 of some embodiments, since the spatial size of the candidate extent is set, the number of pages it contains is also fixed. After determining the number of first pages occupied by the table-level composite index corresponding to the candidate extent, the pages included in the candidate extent may be page-divided based on the number of first pages to determine the index pages and data pages of the candidate extent. The number of index pages is the number of first pages, and the remaining pages in the candidate extent are data pages for storing table data. That is, the candidate extent can be constructed based on the index pages and data pages. As Figure 4 shown, Figure 4 FIG. is a schematic structural diagram of a candidate extent provided by an embodiment of the present application. Among them, for the candidate extents extent0 and extent1 included in the initial database, their corresponding storage space sizes are both 64K. Based on the extent size, the number of preset index fields, and the preset index type, it can be determined that the number of first pages occupied by the table-level composite index corresponding to extent0 is the number of pages included in area D1, and the number of first pages occupied by the table-level composite index corresponding to extent1 is the number of pages included in area D9. It can be understood that area D1 in extent0 is used to store the table-level composite index corresponding to extent0, and areas D2 to D8 are used to store the real data corresponding to extent0; area D9 in extent1 is used to store the table-level composite index corresponding to extent1, and areas D10 to D16 are used to store the real data corresponding to extent1.

[0094] In step S340 of some embodiments, the preset Bloom filter is a pre-constructed Bloom filter based on the extent level. The preset double-layer sparse index includes a first sparse index based on the extent level and a second sparse index based on the page level. The first sparse index is used to indicate the data index fields of the candidate extent, and the second sparse index is used to indicate the data index fields of each candidate page in the candidate extent. Among them, the index pages divided from the candidate extent are used to store the table-level composite index constructed based on the preset Bloom filter and the preset double-layer sparse index.

[0095] It should be noted that the sparse index is an indexing technology that only stores the key information of part of the data, rather than all the data. This can save storage space and improve query efficiency in some cases. The preset double-layer sparse index created in the present application is a sparse index based on Extent-Page double-layer filtering, which means that the index is first filtered at the Extent level and then further refined to the Page level. This hierarchical filtering can more effectively locate the page containing the required data.

[0096] It should be noted that the index page for storing the table-level composite index can include two parts: a preset Bloom filter page (i.e., the page for storing the preset Bloom filter at the Extent level) and a sparse index storage page (i.e., the page for storing the preset double-layer sparse index). Among them, in the index page, the sparse index storage page is stored first, and the preset Bloom filter page is stored later, and the head of the index page can record the offset of the preset Bloom filter at the Extent level in the index page. Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of dividing pages in the candidate area structure provided by an embodiment of the present application. Among them, in the table file storing the initial data table, there can be multiple Segments. One Segment is equivalent to an initial data table. One Segment can include multiple Extents (i.e., candidate areas), and one Extent can include multiple pages. At this time, in combination with the above steps S310 to S330, the first several pages of the Extent header can be used as index pages for storing the corresponding table-level composite index, and the remaining pages can be used as data pages for storing the real data corresponding to the candidate area.

[0097] It should be noted that, please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of storing the corresponding table-level composite structure in the index page of the candidate area provided by an embodiment of the present application. Among them, in the index page, for the storage of the preset double-layer sparse index, a column-first storage method can be adopted (but this method is not limited, only for example), that is, the involved sparse index can be stored in a continuous storage space on the index page. First, in the first index page, the data index fields of the first sparse index based on the district level can be stored, that is, the maximum value (max) and the minimum value (min) of the data in the established index fields are saved. As Figure 6 shown, the min / max of this extent contains 4 data index fields, column1 to column4, and at this time, the data index field corresponding to column1 represents the min and max of all page pages statistically from the perspective of the extent. In the subsequent index pages, the min / max of the candidate pages page0 to page n (n represents the total number of pages included in the extent - 1) included in this extent can be stored, and at this time, the data index fields of page 0 to page n in the same column all belong to the sub-fields of the data index field corresponding to this column. In this embodiment, the sparse indexes of one column can be placed together, or the structure can be adjusted according to actual needs, which will not be elaborated here.

[0098] Exemplarily, please refer toFigure 7A and Figure 7B , Figure 7A is Figure 4 a schematic diagram of the specific structure of area D1 in extent0 for storing the table-level composite index corresponding to extent0. Figure 7B is Figure 4 a schematic diagram of the specific structure of area D1 in extent1 for storing the table-level composite index corresponding to extent1. For Figure 7A this extent, the min / max of this extent contains 4 data index fields, colum1 to colum4, and this extent contains candidate pages page 1 to page 7. For the min (i.e., the minimum value of the field) of the data index field of colum1 in the extent is -1, and the max (i.e., the maximum value of the field) is 57. At this time, the min (i.e., the minimum value of the field) of the data index field of colum1 in page1 is -1, and the max (i.e., the maximum value of the field) is 10. The min (i.e., the minimum value of the field) of the data index field of colum1 in page2 is 11, and the max (i.e., the maximum value of the field) is 20, and so on. And the min (i.e., the minimum value of the field) of the data index field of colum1 in page7 is 48, and the max (i.e., the maximum value of the field) is 57. Thus, the value range of the data index field of each candidate page in the same colum is a subset of the value range of the data index field of the same colum in the extent to which it belongs. Based on this, it is possible to preliminarily determine whether the extent contains the data to be read based on the value range of the data index field at the page where the extent is located.

[0099] It should be noted that the page for storing the preset double-layer word count index may include multiple tuples. Then, the first of all sparse index tuples is the sparse index tuple at the Extent level, followed by the sparse index tuples at each page level. Please refer to Figure 8 , Figure 8 which is a schematic diagram of the tuple storage structure of the table-level composite index provided by an embodiment of the present application. Among them, in the tuple structure of the preset double-layer sparse index, one tuple can store the index information of a candidate area or a candidate page. For example, col1min and col1max store the numerical value range of the index field at the Extent level, and col2min and col2max store the numerical value range of the index field at the page level in the Extent, and so on.

[0100] Please refer to Figure 9 , Figure 9 which is an alternative flowchart of step S340 provided by an embodiment of the present application. In some embodiments, this step S340 may specifically include steps S910 to S930. The following combines Figure 9These two steps are introduced in detail.

[0101] Step S910: Calculate the number of bits based on the extent size, the number of fields in the table-level composite index, the unit page storage size, and a preset optimization parameter to obtain the target number of bits of a preset Bloom filter.

[0102] Step S920: Construct a preset Bloom filter corresponding to the candidate extent based on the shared bit array of the target number of bits and a preset hash function.

[0103] In steps S910 and S920 of the above embodiments, the preset Bloom filter constructed in this application can be an adaptive multi-key Bloom filter that is aware of the number of fields at the Extent level and the Extent size, and can be better used for efficient existence judgment on fields with an enumerated value type.

[0104] In the embodiments of this application, all fields used to establish the preset Bloom filter can share a bit array (bitarray), that is, a shared bit array. And the target number of bits of this shared bit array can be jointly determined by the size of the Extent where it is located and the number of fields used to create the index. Therefore, the process of determining the target number of bits in the embodiments of this application can be calculated by the following formula:

[0105] A = K * n * (m / 8KB)

[0106] Where A represents the number of bits used by the shared bit array, that is, the target number of bits; n is the number of fields used to create the table-level composite index, that is, the number of fields in the table-level composite index; m is the extent size of the candidate extent corresponding to the preset Bloom filter, K is a preset optimization parameter, which can be a hyperparameter set according to actual needs and is not limited; 8KB represents the unit page storage size.

[0107] It should be noted that the larger the number of bits used by the shared bit array, the smaller the probability of making mistakes, but if the number of bits is too large, many problems will also be caused. Therefore, based on the extent size, the number of fields in the table-level composite index, the unit page storage size, and the preset optimization parameter, this application can determine a target number of bits that better meets the requirements and needs of the candidate extent, and better improve the read performance of the database node, thereby improving the efficiency of data reading.

[0108] In some embodiments, please refer to Figure 10 , Figure 10It is a filtering schematic diagram of a preset Bloom filter provided by an embodiment of the present application. Among them, after determining the target number of bits, a shared bit array with a length of the target number of bits can be created first, and all bits are initialized to 0. When inserting an element, q different hash functions can be used to calculate the hash values corresponding to the attribute values of the element (such as attribute1: value1 and attribute2: value2 in the figure), and then map these hash values to the corresponding positions in the shared bit array, and set the bits at these positions to 1. When querying whether an element is in the set, k hash functions are also used to calculate the hash value of the element, and then check the positions corresponding to these hash values in the shared bit array. If the bits at all positions are 1, then the element may be in the set; if the bit at any one position is 0, then the element is definitely not in the set.

[0109] In step S350 of some embodiments, further, the field data of all candidate pages in the candidate area can be stored in the data page, and a target area structure corresponding to the candidate area can be constructed based on the stored index page and data page.

[0110] The embodiment of the present application optimizes the segment-page storage in the predicate pre-filtering framework based on the composite index under the segment-page storage, divides the pages in each Extent into index pages and data pages, where the index pages are used to store the composite index, and the data pages are used to store the real data. And, the user can, according to the data characteristics when creating a table, specify the fields and index types for establishing indexes when creating a table, select a certain proportion of pages in each Extent as index pages, and the rest are used as data pages.

[0111] In step S360 of some embodiments, after constructing multiple target areas based on the target area structure construction method of the present application, a target data table can be constructed based on the target area structures of multiple same initial data tables, and data reading can be performed based on the target data table.

[0112] Please refer to Figure 11 , Figure 11 is an optional flowchart of step S360 provided by an embodiment of the present application. In some embodiments, the data reading based on the target data table in step S360 may specifically include steps S1110 to S1160. The following combines Figure 11 to introduce these six steps in detail.

[0113] Step S1110, receive a read request;

[0114] Step S1120, determine a target area list based on the target data table;

[0115] Step S1130: Select a zone to be filtered from multiple target zones, and obtain the target index page corresponding to the zone to be filtered;

[0116] Step S1140: Based on the table-level composite index corresponding to the zone to be filtered, pre-filter the query predicate conditions to obtain a composite filtering result;

[0117] Step S1150: If the composite filtering result indicates that the zone to be filtered contains data that meets the query predicate conditions, perform a page query based on the composite filtering result, the query predicate conditions, and the data page corresponding to the zone to be filtered to determine the page to be read;

[0118] Step S1160: Read data by scanning the page to be read.

[0119] In step S1110 of some embodiments, the read request refers to a request for a target node (from a client or other system) to apply to other nodes for querying and reading data. The read request can be an identifier, condition, SQL statement, etc. of a specific data item. The target node can be the primary node or the standby node in the same node cluster. The read request includes the target data table and the query predicate conditions. The target data table refers to the data table where the data to be read is located, and the target data table can be constructed based on the target zone structure of the present application. The query predicate conditions refer to the conditions for filtering out the data to be read in the target data table, and are used to filter the data stored in the target data table. The query predicate conditions can also be specifically limited to the pages in the target data table to be read, without specific limitation. For example, the SQL statement corresponding to the read request is "select * from tbl where colum1 < -80 and colum4 > 750". The target of this read request is to return all data from the tbl table that simultaneously meet the conditions that the value of column1 is less than -80 and the value of column4 is greater than 750. If no record meets both conditions, the query result will be an empty set. At this time, the target data table of this read request is the tb1 table, and the query predicate conditions are "colum1 < -80 and colum4 > 750".

[0120] In step S1120 of some embodiments, based on the known target data table, all target zones included in the target data table can be known, so that a target zone list can be obtained, that is, the target zone list includes multiple target zones.

[0121] In step S1130 of some embodiments, further, a to-be-filtered area is selected from multiple target areas. This is because each target area represents a part of the data stored in the target data table, and these areas need to be checked to find qualified data. Therefore, each target area in the target data table needs to be used as the to-be-filtered area in turn for filtering to achieve a complete read of the data containing qualified data in the target data table. When judging the to-be-filtered area, first obtain the target index page corresponding to the to-be-filtered area. At this time, the target index page refers to the page storing the table-level composite index corresponding to the to-be-filtered area proposed in the above embodiments.

[0122] In some embodiments, obtaining the target index page corresponding to the to-be-filtered area may specifically include:

[0123] If it is detected that the corresponding memory contains index pages of multiple target areas, obtain the target index page corresponding to the to-be-filtered area; or,

[0124] If it is detected that the corresponding memory does not contain index pages of multiple target areas, send an index page request to the corresponding distributed memory service module, and receive the target index page corresponding to the to-be-filtered area returned by the distributed memory service module.

[0125] It can be understood that when other nodes receive a read request, they will first check whether the memory shared by the nodes of the database already contains index pages of multiple target areas related to the target query, that is, check whether the page corresponding to the to-be-filtered area is hit in the memory. If it is included, directly obtain the target index page corresponding to the to-be-filtered area from the memory. In this way, the data reading efficiency can be improved because memory access is much faster than retrieving data from a disk or a remote service.

[0126] If the index page corresponding to the to-be-filtered area is not found in the memory and currently under the resource pooling architecture, then a request can be sent to the DMS module corresponding to the node, which also points to the distributed feature of the data storage architecture. In other cases, data can be directly read from the disk. During this process, the system will send an index page request to the distributed memory service module, and this request includes information about the target area to be queried. Once the distributed memory service module processes the request and returns the target index page corresponding to the to-be-filtered area, these target index pages can be used to perform subsequent read operations.

[0127] In step S1140 of some embodiments, after determining the target index page of the to-be-filtered area, the query predicate condition can be pre-filtered based on the table-level composite index corresponding to the to-be-filtered area stored in the target index page to obtain a composite filtering result.

[0128] Please refer to Figure 12 , Figure 12It is an alternative flowchart of step S1140 provided by an embodiment of the present application. In some embodiments, this step S1140 may specifically include steps S1210 to S1240. The following combines Figure 12 to introduce these four steps in detail.

[0129] Step S1210: Perform extent-level filtering on the query predicate condition based on the preset Bloom filter corresponding to the area to be filtered and the first sparse index at the extent level, to obtain an extent-level filtering result;

[0130] Step S1220: If the extent-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, perform page-level sparse index detection on the table-level composite index corresponding to the area to be filtered based on the query predicate condition, to obtain a page-level sparse index detection result;

[0131] Step S1230: If the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered contains the second sparse index to be filtered, perform page-level filtering on the query predicate condition based on the second sparse index, to obtain a page-level filtering result;

[0132] Step S1240: Determine a composite filtering result based on the page-level filtering results corresponding to all the second sparse indexes to be filtered.

[0133] In step S1210 of some embodiments, during specific filtering, the query predicate condition can be first filtered at the extent level based on the preset Bloom filter corresponding to the area to be filtered and the first sparse index at the extent level, to obtain an extent-level filtering result. That is to say, first perform index filtering at the extent level. When establishing the table-level composite index in the present application, the preset Bloom filter and the sparse index have no dependency relationship with each other and can be established independently. Therefore, for a field in the table for which a table-level composite index is established, when using the table-level composite index for extent-level filtering, if a preset Bloom filter at the extent level is established for this field, the preset Bloom filter is used for filtering; if a first sparse index at the extent level is established, the first sparse index at the extent level is used for filtering. There is no limitation on the specific filtering order and which one to specifically use for filtering. Further, if either of them results in a negative, it is determined that the area to be filtered does not contain data that meets the query predicate condition.

[0134] In step S1220 of some embodiments, if the extent-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, then further perform page-level sparse index detection on the target page in the table-level composite index corresponding to the area to be filtered based on the query predicate condition, to determine whether there is an available second sparse index at the page level, so as to obtain a page-level sparse index detection result.

[0135] In some embodiments, before the page-level sparse index detection is performed on the table-level composite index corresponding to the area to be filtered based on the query predicate condition when the area-level filtering result in step S1220 indicates that the area to be filtered contains data that meets the query predicate condition, the data reading method of the present application may further include:

[0136] If the area-level filtering result indicates that the area to be filtered does not contain data that meets the query predicate condition, update the area to be filtered based on multiple target areas;

[0137] Perform pre-filtering on the query predicate condition according to the table-level composite index corresponding to the updated area to be filtered to obtain a composite filtering result until pre-filtering is completed for all target areas.

[0138] It can be understood that if the area-level filtering result indicates that the area to be filtered does not contain data that meets the query predicate condition, it means that the current area to be filtered does not have the required data, and the next target area of the target data table can be judged, so the area to be filtered can be updated based on multiple target areas. Further, perform pre-filtering on the query predicate condition according to the table-level composite index corresponding to the updated area to be filtered, that is, repeat steps S1140 and subsequent steps to obtain a composite filtering result until pre-filtering is completed for all target areas.

[0139] It can be understood that if the area-level filtering result at the extent level is yes and there is a sparse index on this field in the table-level composite index, use the page-level sparse index to filter the data when reading each page. If the result is no, there is no required data on that page, otherwise perform a scan operation on that data page. In a non-resource pooling architecture, the read and write processes of the index page are the same as those of the data page. Among them, the table-level composite index is updated while writing data. When responding to a read request, first determine the list of areas involved in the read request and request the index page for each processed area. If the filtering result is no, all subsequent pages in that area are skipped. If the area-level filtering result is yes, use the corresponding sparse index of the data page to filter when the scan operator processes each data page.

[0140] In step S1230 of some embodiments, if the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered contains a second sparse index to be filtered at the page level, perform page-level filtering on the query predicate condition based on the second sparse index to obtain a page-level filtering result. The page-level filtering result at this time may be empty, that is, there is no data that meets the requirements in the corresponding page, or it may be a specific range of matching data fields, that is, the corresponding page contains data that meets the requirements, so as to facilitate subsequent specific data reading on the data page.

[0141] It should be noted that if the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered does not contain the second sparse index based on the page level to be filtered, the target pages in the area to be filtered will be directly scanned.

[0142] In step S1240 of some embodiments, after performing page-level filtering on all target pages in the area to be filtered, the composite filtering result can be determined based on the page-level filtering results corresponding to all the second sparse indexes to be filtered and the results of directly scanning the pages. At this time, the composite filtering result includes the data field ranges in the area to be filtered that all meet the requirements of the query predicate conditions.

[0143] In step S1150 of some embodiments, if the composite filtering result indicates that there is data in the area to be filtered that meets the query predicate conditions, a page query is performed based on the composite filtering result, the query predicate conditions, and the data pages corresponding to the area to be filtered to determine the pages to be read. Among them, the data on the scanned pages can be directly read. The pages to be read at this time are the pages corresponding to the data field ranges that meet the requirements of the query predicate conditions. In this way, it can be ensured that only the necessary data pages are scanned, further improving the data reading efficiency.

[0144] In some embodiments, the page query based on the composite filtering result, the query predicate conditions, and the data pages corresponding to the area to be filtered in step S1150 to determine the pages to be read may specifically include:

[0145] If the composite filtering result indicates that data that meets the query predicate conditions is matched in the second sparse index in the area to be filtered, and the corresponding memory contains the pages that meet the query predicate conditions, the pages to be read are determined from the data pages based on the matched second sparse index; or,

[0146] If the composite filtering result indicates that no data that meets the query predicate conditions is matched in the second sparse index in the area to be filtered, and the corresponding memory also does not contain the pages that meet the query predicate conditions, a data page request is sent to the corresponding distributed memory service module, and the pages to be read returned by the distributed memory service module are received.

[0147] It can be understood that if the composite filtering result indicates that data that meets the query predicate conditions is matched in the second sparse index in the area to be filtered, and the memory corresponding to the node contains the pages that meet the query predicate conditions, it means that the data pages stored in the memory can hit the corresponding data. Otherwise, a data page request needs to be sent to the distributed memory service module corresponding to the node to obtain the pages to be read.

[0148] In step S1160 of some embodiments, further, the requested data is obtained by scanning and reading the determined read page. This step can extract the required information from the physical data page and return it to the user or application. Among them, data reading may involve decoding, conversion and formatting the data into a form that is easy for the user to understand, for example, returning JSON, XML or other formats, which are not specifically limited here.

[0149] It should be noted that, after the present application completes data reading of a page of a to-be-filtered area by scanning, the to-be-filtered area will be further updated until all target areas of the target data table are judged and complete data is read.

[0150] Thus, the present invention proposes an efficient composite index to filter predicate conditions (i.e., query condition data), wherein the adaptive multi-key Bloom filter based on the number of fields at the Extent level and the Extent size perception, and the preset sparse index of the Extent-Page double-layer filtering are both low-cost and efficient, and each filtering operation of these two indexes has a time complexity of O(1), which means that the time overhead of each filtering operation is extremely short, and each successful filtering may save several to thousands of data pages of disk reading and network overhead, which can greatly speed up the read operation. At the same time, the composite index at the Extent level in this application does not need to introduce the Extent level lock, but only needs to rely on the original data page level lock. In this way, because the S lock at the data page level is obtained, it does not affect the modification of other pages at the Extent level. And under the resource pooling architecture, the composite index will not be affected by the transaction visibility during use, and there is no need to specially consider the version visibility problem when using it.

[0151] It should be noted that the composite index in this application can filter fields of any data type and can be established on any field without any special restrictions when used. The function of this application only relies on segment page storage, has no other dependencies, and does not conflict with existing functions. This function is directly enabled through parameters in the table creation SQL statement when used. In summary, the present invention has good ease of use.

[0152] For example, assume that the statement corresponding to the read request is "select * from tbl where colum1<-80 and colum4>750". Figure 4 , Figure 7A and Figure 7B, the target area in the target data table tbl includes extent0 and extent1. When performing pre-filtering, first read the index page of extent0. From the filtering in the first sparse index at the extent level of extent0 by the condition colum1 < -80, it can be seen that there is no data in extent0 that meets the requirements, so the data page scan of this extent can be directly skipped. Further, read the index page of extent1. From the first sparse index at the extent level of extent1, it can be known that extent1 contains data that meets the requirements. At this time, the second sparse index at the page level corresponding to page10 to page16 in extent1 can be compared in turn. It can be known that the data that meets the conditions is included in page10 and page16, that is, the data field range of colum1 in page10 is [-90, -70], including the data of "colum1 < -80", and the data field range of colum4 in page16 is [701, 810], including the data of "colum4 > 750". Further, the data in page10 and page16 can be scanned and read, while other pages are skipped.

[0153] It should be noted that the embodiments of the present application are used to perform quick existence judgments on fields of all value types. The present application establishes a preset two-layer sparse index on the fields specified by the user. The sparse index maintains the min and max values of the field at the Extent and Page levels. The min takes the minimum value in all versions, and the max takes the maximum value in all versions, which can avoid the problem of version visibility. When filtering, first use the first sparse index at the Extent level for filtering. If the filtering result is negative, all the pages included in the Extent are skipped. If the filtering result is positive, when actually reading each data page, use the second sparse index at the Page level for filtering. This method has no restrictions on the field value type and is relatively friendly to range type predicate conditions.

[0154] As Figure 13 shown, Figure 13 is a specific flowchart of the data reading method provided by the embodiments of the present application. Among them, taking an example where a standby node sends a read request to the master node once, the data reading method may specifically include steps S1301 to step S1310.

[0155] Step S1301, start;

[0156] Step S1302, receive a read request;

[0157] Step S1303, determine the list of target areas involved (equivalent to determining the list of target areas based on the target data table in the read request);

[0158] Step S1304, select a zone to be filtered from multiple target zones, and request the target index page corresponding to the zone to be filtered;

[0159] Step S1305, determine whether the index page is hit in the memory (equivalent to detecting that the index page of multiple target zones is included in the corresponding memory). If not, execute Step S1306; if so, execute Step S1307;

[0160] Step S1306, request the index page through the distributed memory service module;

[0161] Step S1307, determine whether the extent-level index filtering based on the zone to be filtered contains data that meets the requirements (equivalent to determining the zone-level filtering result of filtering the query predicate condition based on the preset Bloom filter corresponding to the zone to be filtered and the first sparse index at the zone level). If so, that is, it contains, execute Step S1308; if not, execute Step S1313;

[0162] Step S1308, determine whether there is an available page-level sparse index (equivalent to determining whether the table-level composite index corresponding to the zone to be filtered contains the second sparse index to be filtered). If so, that is, it contains, execute Step S1309; if not, that is, it does not contain, execute Step S1310;

[0163] Step S1309, determine whether the page-level sparse index filters out data that meets the requirements. If so, that is, it contains, execute Step S1310; if not, that is, it does not contain, execute Step S1309;

[0164] Step S1310, determine whether the data page is hit in the memory (equivalent to the corresponding memory contains the data page that meets the query predicate condition). If so, that is, it contains, execute Step S1312; if not, that is, it does not contain, execute Step S1311;

[0165] Step S1311, request the data page through the distributed memory service module;

[0166] Step S1312, scan the hit page;

[0167] Step S1313, the page of the current zone to be filtered or the processing of the zone to be filtered ends, and process the page or zone to be filtered of the next zone to be filtered.

[0168] The data reading method provided by the embodiments of the present application optimizes the segment-page storage, that is, constructs a predicate pre-filtering framework based on a composite index under the segment-page storage. By constructing an adaptive multi-key Bloom filter that perceives the number of fields at the Extent level and the Extent size for page filtering, the sparse index of the fields used in the predicate condition can be quickly located in memory. Moreover, the size of the bitarray used in the Bloom filter can be determined according to the number of fields specified by the user for building the Bloom filter and the Extent size where they are located, making the constructed preset Bloom filter more in line with the user's requirements and effectively improving the efficiency of data screening. Therefore, the present application can avoid unnecessary network interactions between the primary and standby nodes in the OLAP scenario, improve the read performance of the database nodes, and thus improve the efficiency of data reading.

[0169] Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of the data reading device provided by the embodiments of the present application. This device can implement the data reading method of the above embodiments. Specifically, the device may include:

[0170] An acquisition module 1410, configured to acquire the extent size, the preset number of index fields, and the preset index type of the candidate area in the initial data table. The initial data table includes multiple candidate areas, and each candidate area includes multiple candidate pages;

[0171] A page number determination module 1420, configured to determine the number of the first pages occupied by the table-level composite index corresponding to the candidate area based on the extent size, the preset number of index fields, and the preset index type;

[0172] A page division module 1430, configured to divide the pages included in the candidate area based on the number of the first pages to determine the index pages and data pages of the candidate area. The number of index pages is the number of the first pages;

[0173] A composite index construction module 1440, configured to construct a table-level composite index based on a preset Bloom filter and a preset two-layer sparse index, and store the table-level composite index into the index pages. The preset Bloom filter is a Bloom filter based on the extent level. The preset two-layer sparse index includes a first sparse index based on the extent level and a second sparse index based on the page level. The first sparse index is used to indicate the data index fields of the candidate area, and the second sparse index is used to indicate the data index fields of each candidate page in the candidate area;

[0174] A storage module 1450, configured to store the field data of all candidate pages in the candidate area into the data pages, and construct a target area structure corresponding to the candidate area based on the stored index pages and data pages;

[0175] A data reading module 1460 is configured to construct a target data table based on multiple target area structures of the same initial data table and perform data reading based on the target data table.

[0176] It should be noted that the data reading device in the embodiments of the present application is used to implement the data reading method in the above embodiments. The data reading device in the embodiments of the present application corresponds to the aforementioned data reading method. For the specific processing process, please refer to the aforementioned data reading method and will not be elaborated here.

[0177] The embodiments of the present application further provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the data reading method in the embodiments of the present application. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0178] Please refer to Figure 15 , Figure 15 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0179] A processor 1510, which can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0180] A memory 1520, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1520 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1520 and are called by the processor 1510 to execute the data reading method in the embodiments of the present application;

[0181] An input / output interface 1530, which is configured to implement information input and output;

[0182] A communication interface 1540, which is configured to implement communication interaction between this device and other devices. It can implement communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0183] The bus 1550 transmits information between various components of the device (such as the processor 1510, the memory 1520, the input / output interface 1530, and the communication interface 1540);

[0184] Among them, the processor 1510, the memory 1520, the input / output interface 1530, and the communication interface 1540 achieve communication connections with each other inside the device through the bus 1550.

[0185] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program for causing a computer to execute the data reading method in the above embodiment.

[0186] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0188] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine some steps, or different steps.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0191] In the description of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0192] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0193] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0194] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0196] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0197] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A data reading method, characterized in that: The method comprises: Obtaining a region size, a preset number of index fields, and a preset index type of a candidate region in an initial data table, wherein the initial data table includes a plurality of candidate regions, and the candidate region includes a plurality of candidate pages; Determine the first number of pages occupied by the table-level composite index corresponding to the candidate zone based on the zone size, the preset number of index fields, and the preset index type; Dividing the pages included in the candidate area into pages based on the first number of pages to determine index pages and data pages of the candidate area, where the number of index pages is the first number of pages; The table-level composite index is constructed based on a preset Bloom filter and a preset double-layer sparse index, and the table-level composite index is stored in the index page; the preset Bloom filter is a pre-constructed Bloom filter based on the zone level, the preset double-layer sparse index includes a first sparse index based on the zone level and a second sparse index based on the page level, the first sparse index is used to indicate the data index field of the candidate zone, and the second sparse index is used to indicate the data index field of each candidate page in the candidate zone; storing the field data of all the candidate pages in the candidate area into the data page, and constructing a target area structure corresponding to the candidate area based on the stored index page and the data page; A target data table is constructed based on multiple target area structures of the same initial data table, and data is read based on the target data table; the data reading based on the target data table includes: receiving a read request, the read request includes the target data table and a query predicate condition; determining a target area list based on the target data table, the target area list includes multiple target areas; selecting a to-be-filtered area from the multiple target areas, and obtaining a target index page corresponding to the to-be-filtered area, the target index page storing the table-level composite index corresponding to the to-be-filtered area; pre-filtering the query predicate condition based on the table-level composite index corresponding to the to-be-filtered area to obtain a composite filtering result; if the composite filtering result indicates that the to-be-filtered area contains data that meets the query predicate condition, performing a page query based on the composite filtering result, the query predicate condition and the data page corresponding to the to-be-filtered area to determine a read page; and reading data by scanning the read page; Among them, the pre-filtering of the query predicate condition based on the table-level composite index corresponding to the area to be filtered to obtain a composite filtering result includes: performing zone-level filtering on the query predicate condition based on the preset Bloom filter corresponding to the area to be filtered and the first sparse index based on the zone level to obtain a zone-level filtering result; if the zone-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, performing page-level sparse index detection on the table-level composite index corresponding to the area to be filtered based on the query predicate condition to obtain a page-level sparse index detection result; if the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered contains the second sparse index to be filtered, performing page-level filtering on the query predicate condition based on the second sparse index to obtain the page-level filtering result; and determining the composite filtering result based on the page-level filtering results corresponding to all the second sparse indexes to be filtered.

2. The method according to claim 1, characterized in that: If the zone-level filtering result indicates that the to-be-filtered zone contains data that meets the query predicate condition, before performing page-level sparse index detection on the table-level composite index corresponding to the to-be-filtered zone based on the query predicate condition and obtaining the page-level sparse index detection result, the method further includes: If the zone-level filtering result indicates that the zone to be filtered does not contain data that meets the query predicate condition, updating the zone to be filtered based on the plurality of target zones; The query predicate condition is pre-filtered according to the updated table-level composite index corresponding to the to-be-filtered area to obtain a composite filtering result until all the target areas are pre-filtered.

3. The method according to claim 1, characterized in that Performing a page query based on the composite filtering result, the query predicate condition, and the data page corresponding to the to-be-filtered area to determine a page to be read includes: If the composite filtering result indicates that data that meets the query predicate condition is matched in the second sparse index of the to-be-filtered area, and the corresponding memory contains pages that meet the query predicate condition, the read page is determined from the data pages based on the matched second sparse index; or, If the composite filtering result indicates that data that meets the query predicate condition is matched in the second sparse index of the area to be filtered, and the corresponding memory does not contain pages that meet the query predicate condition, a data page request is sent to the corresponding distributed memory service module, and the read page returned by the distributed memory service module is received.

4. The method according to claim 1, characterized in that The step of obtaining a target index page corresponding to the area to be filtered includes: If it is detected that the corresponding memory contains multiple index pages of the target area, the target index page corresponding to the area to be filtered is obtained; or If it is detected that the corresponding memory does not contain multiple index pages of the target area, an index page request is sent to the corresponding distributed memory service module, and the target index page corresponding to the to-be-filtered area returned by the distributed memory service module is received.

5. The method according to any one of claims 1 to 4, characterized in that: The preset Bloom filter in the candidate area is constructed in the following way: Performing bit calculation based on the zone size, the number of fields of the table-level composite index, the unit page storage size, and preset optimization parameters to obtain a target bit number of the preset Bloom filter; The preset Bloom filter corresponding to the candidate area is constructed based on the shared bit array of the target number of bits and a preset hash function.

6. A data reading device, characterized in that: The device comprises: An acquisition module, used to acquire a region size, a preset number of index fields, and a preset index type of a candidate region in an initial data table, wherein the initial data table includes a plurality of candidate regions, and the candidate region includes a plurality of candidate pages; A page number determination module, configured to determine the first number of pages occupied by the table-level composite index corresponding to the candidate zone based on the zone size, the preset number of index fields and the preset index type; a page division module, configured to divide the pages included in the candidate area into pages based on the first number of pages, and determine index pages and data pages of the candidate area, wherein the number of the index pages is the first number of pages; A composite index construction module, configured to construct the table-level composite index based on a preset Bloom filter and a preset double-layer sparse index, and store the table-level composite index in the index page; the preset Bloom filter is a pre-constructed zone-level Bloom filter, the preset double-layer sparse index includes a first zone-level sparse index and a second page-level sparse index, the first sparse index is used to indicate a data index field of the candidate zone, and the second sparse index is used to indicate a data index field of each candidate page in the candidate zone; A storage module, used for storing the field data of all the candidate pages in the candidate area into the data page, and constructing a target area structure corresponding to the candidate area based on the stored index page and the data page; A data reading module, for constructing a target data table based on multiple target area structures of the same initial data table, and reading data based on the target data table; the data reading based on the target data table includes: receiving a read request, the read request includes the target data table and a query predicate condition; determining a target area list based on the target data table, the target area list includes multiple target areas; selecting a to-be-filtered area from the multiple target areas, and obtaining a target index page corresponding to the to-be-filtered area, the target index page storing the table-level composite index corresponding to the to-be-filtered area; pre-filtering the query predicate condition based on the table-level composite index corresponding to the to-be-filtered area to obtain a composite filtering result; if the composite filtering result indicates that the to-be-filtered area contains data that meets the query predicate condition, performing a page query based on the composite filtering result, the query predicate condition and the data page corresponding to the to-be-filtered area to determine a read page; scanning the read page Data reading is performed; wherein, the query predicate condition is pre-filtered based on the table-level composite index corresponding to the area to be filtered to obtain a composite filtering result, including: performing area-level filtering on the query predicate condition based on the preset Bloom filter corresponding to the area to be filtered and the first sparse index based on the area level to obtain an area-level filtering result; if the area-level filtering result indicates that the area to be filtered contains data that meets the query predicate condition, performing page-level sparse index detection on the table-level composite index corresponding to the area to be filtered based on the query predicate condition to obtain a page-level sparse index detection result; if the page-level sparse index detection result indicates that the table-level composite index corresponding to the area to be filtered contains the second sparse index to be filtered, performing page-level filtering on the query predicate condition based on the second sparse index to obtain the page-level filtering result; and determining the composite filtering result based on the page-level filtering results corresponding to all the second sparse indexes to be filtered.

7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Data query method, device and equipment based on Hudi and storage medium

    CN113094340A

  • Intelligent contract state data ad-hoc query method and device, equipment and storage medium

    CN118035314A