Parallel scanning method, apparatus, device, storage medium and computer program
By determining the target physical block based on data query conditions in the database and allocating it to multiple processing units for parallel scanning, the problem of low database query efficiency is solved, and the query speed and user experience are improved.
Patent Information
- Application Number
- PCT/CN2024/129021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, database query efficiency is low, thread waiting time is long, processing efficiency is reduced, computing resources are seriously wasted, and user experience is affected.
By determining the target physical block from multiple physical blocks based on data query conditions and assigning them to multiple processing units for parallel scanning, serial acquisition tasks are avoided and query efficiency is improved.
It effectively avoids the waste of computing resources of processing units, improves the efficiency and user experience of database querying data, and improves the speed and accuracy of data query.
Smart Images

Figure CN2024129021_03072025_PF_FP_ABST
Abstract
Description
Parallel scanning method, device, equipment, storage medium and computer program
[0001] This application claims priority to Chinese patent application No. 202311864350.8 filed on December 29, 2023, entitled “Parallel scanning method, device, equipment, storage medium and computer program”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of database technology, and in particular to a parallel scanning method, apparatus, device, storage medium and computer program. Background Art
[0003] With the development of computer hardware technology, symmetric multi-processing (SMP) technology is increasingly being applied to databases to improve database query performance. SMP technology utilizes a computer's multi-core CPU architecture to implement multi-threaded parallel computing, fully utilizing CPU resources to improve query performance. This multi-threaded computing can exchange data and information within the same process using various inter-thread synchronization techniques, reducing computing resource usage and increasing the effective utilization of system computing resources.
[0004] In related art, when a user queries a database for target data, the database can organize multiple physical blocks storing the data into a task queue. Multiple threads can then sequentially retrieve the next physical block to be scanned from the task queue and scan the physical blocks based on the query criteria, thereby determining the query result for the target data based on the scan results of the multiple threads. The multiple threads serially retrieve the next physical block to be scanned and scan the retrieved physical storage blocks in parallel, meaning the multiple threads serially retrieve tasks and scan in parallel.
[0005] However, since the process of multiple threads acquiring tasks is a serial process, this will result in longer waiting times for the multiple threads, reducing processing efficiency. In addition, when there are many threads and the amount of data in the physical block is small, the multiple threads need to frequently acquire tasks from the task queue, which will cause the threads to be in a waiting state for most of the time, greatly reducing data query efficiency and affecting the user experience.
[0006] Summary of the Invention
[0007] This application provides a parallel scanning method, apparatus, device, storage medium, and computer program that can solve the problem of low data query efficiency in related technologies. The technical solution is as follows:
[0008] In a first aspect, a parallel scanning method is provided, which is applied to a first data node included in a database, wherein the first data node includes multiple physical blocks, and the method includes: determining at least one target physical block from the multiple physical blocks based on a data query condition, wherein the target physical block refers to a physical block related to the data to be queried, and the data query condition is a condition satisfied by the data to be queried; based on the total number of the at least one target physical block and the number of parallel processing units, the at least one target physical block is allocated to multiple processing units, wherein the multiple processing units refer to processing units that perform data query using a parallel scanning method, and the number of parallel processing units is the number of the multiple processing units; the target physical blocks allocated to each of the multiple processing units are scanned by the multiple processing units, and data satisfying the data query condition is determined based on the scanning results of the multiple processing units.
[0009] The present application can determine at least one target physical block from multiple physical blocks based on data query conditions. Since the target physical block can refer to a physical block related to the data to be queried, this can effectively avoid wasting the computing resources of the processing unit and improve the efficiency of database data query. Moreover, since the present application can also allocate the at least one target physical block to multiple processing units based on the total number of at least one target physical block and the number of parallel processing units, compared with related technologies that require serial acquisition tasks, the present application does not require serial acquisition tasks. Each processing unit can scan the target physical block assigned to it, thereby greatly improving data query efficiency and enhancing user experience.
[0010] In practical applications, the first data node stores a target base table, which refers to a base table containing the data to be queried. The first data node may also store at least one index directory corresponding to the target base table. In this case, the data query condition may instruct the first data node to query data from the target base table, or may instruct the first data node to query data from one of the at least one index directories. When different data query conditions indicate different content, the implementation methods for determining at least one target physical block from multiple physical blocks based on the data query condition vary, and each of these methods will be described below.
[0011] When a data query condition instructs the first data node to query data from a target base table, the data query condition includes a constraint condition for at least one query field and a base table identifier, the base table identifier indicates the target base table, and the multiple physical blocks are used to store the target base table. In this case, the first data node can directly determine the multiple physical blocks as the at least one target physical block.
[0012] In the case where a data query condition instructs the first data node to query data from an index directory of a target base table, the data query condition includes a constraint condition for at least one query field and an index directory identifier, the index directory identifier indicating a target index directory, the target index directory being an index directory constructed based on data stored by the first data node (also referred to as a target base table), the target index directory being one of at least one index directory corresponding to the target base table, and the multiple physical blocks being used to store the target index directory. In this case, if a sort field in the target index directory exists in the at least one query field, then based on the constraint condition for the sort field and the target index directory, at least one target physical block is determined from the multiple physical blocks, the index stored in the target physical block satisfying the constraint condition for the sort field.
[0013] Since the target index directory is constructed based on the sorting field (also known as the index field), if the sorting field in the target index directory exists in the query field, the first data node can directly filter out the indexes that meet the constraints of the sorting field from the target index directory based on the constraints of the sorting field. In this way, the indexes that do not meet the constraints of the sorting field can be filtered, effectively avoiding the subsequent waste of processing unit computing resources and improving the efficiency of database data query.
[0014] Optionally, if the sorting field in the target index directory does not exist in the at least one query field, the multiple physical blocks are determined as the at least one target physical block.
[0015] If the sorting field in the target index directory does not exist in the query field, the first data node cannot filter the multiple physical blocks according to the target index directory, and therefore can directly determine the multiple physical blocks as the at least one physical block.
[0016] It should be noted that, in actual applications, the data query condition may not include query field constraints. In this case, the data query instruction instructs the first data node to perform a full table query on the target base table indicated by the base table identifier, or the target index directory indicated by the index directory identifier. If the data query condition includes a base table identifier, the first data node uses the multiple physical blocks where the target base table indicated by the base table identifier is located as the at least one target physical block. If the data query condition includes an index directory identifier, the first data node uses the multiple physical blocks where the target index directory indicated by the index directory identifier is located as the at least one target physical block.
[0017] In actual applications, the database also includes a coordination node. Before determining at least one target physical block from the multiple physical blocks based on the data query conditions, the first data node can also receive a data query instruction sent by the coordination node. The data query instruction is determined by the coordination node based on the data query request sent by the user, and the data query instruction includes the above-mentioned data query conditions.
[0018] Optionally, when a user needs to query data, he or she can send a data query request to the coordination node, and the data query request indicates the base table where the user's data to be queried is located, as well as the constraints of the user's data to be queried. In this case, the coordination node in the database determines the distribution of the base table where the user's data to be queried is located in the data nodes based on the base table where the user's data to be queried is located, and based on the distribution and the constraints of the user's data to be queried, determines at least one second data node that needs to be queried from the data nodes included in the database, and determines the data query instructions corresponding to the at least one second data node. For each of the at least one second data node, the coordination node sends the data query instruction corresponding to the second data node to the second data node to instruct the second data node to perform the data query and whether to use parallel scanning to perform the data query.
[0019] Optionally, the at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform data query using parallel scanning. In this case, the first data node can receive the data query instruction sent by the coordination node.
[0020] Optionally, the multiple processing units may be multiple processes in the first data node, or may be multiple threads in the first data node.
[0021] Optionally, the total number of the at least one target physical block is divided by the number of parallel processing units to obtain a candidate number. If the candidate number is an integer, the at least one target physical block is allocated to the multiple processing units according to the candidate number, so that the number of target physical blocks allocated to each processing unit is the candidate number.
[0022] Optionally, if the candidate number is not an integer, the candidate number is rounded up to obtain a target number, and based on the target number and the number of parallel processing units, the at least one target physical block is allocated to the multiple processing units so that the number of target physical blocks allocated to each processing unit is no greater than the target number.
[0023] The at least one target physical block is allocated to N processing units among the multiple processing units according to the target number, so that the number of target physical blocks allocated to the N processing units is the target number, N is equal to the number of parallel processing units minus 1, if the value after N multiplied by the target number is less than the total number of the at least one target physical block, then according to the first difference, the at least one target physical block is allocated to the remaining processing units, and the remaining processing units are the processing units other than the N processing units among the multiple processing units, and the first difference is the difference between the value after N multiplied by the target number and the total number of the at least one target physical blocks; if the value after N multiplied by the target number is not less than the total number of the at least one target physical blocks, no target physical block is allocated to the remaining processing units.
[0024] If the value obtained by multiplying N by the target number is less than the total number of the at least one target physical blocks, it means that the total number of physical blocks allocated to the N processing units is less than the total number of the at least one target physical blocks, there are still unallocated physical blocks in the at least one target physical block, and the number of the unallocated physical blocks is the first difference. Therefore, the target physical blocks with the first difference can be allocated to the remaining processing units. If the value obtained by multiplying N by the target number is not less than the total number of the at least one target physical blocks, it means that the total number of physical blocks allocated to the N processing units is not less than the total number of the at least one target physical blocks, there are no unallocated physical blocks in the at least one target physical block, and therefore, the target physical blocks may not be allocated to the remaining processing units.
[0025] Based on the above description, when the candidate number is not an integer, the candidate number can be rounded up. In practical applications, the candidate number can also be rounded down to obtain the target number. At this time, there is at least one thread among the multiple threads to which the number of target physical blocks allocated is not greater than the target number.
[0026] Optionally, the data query instruction further includes the number of parallel processing units. In this case, the first data node can determine the number of parallel processing units.
[0027] Since the first data node can query data from the target base table or from the target index directory, in different cases, multiple processing units are used to scan the target physical blocks assigned to each of them, and the method of determining the data that meets the data query conditions based on the scanning results of the multiple processing units is different.
[0028] In the case where the first data node queries data from the target base table, the multiple processing units can scan the target physical blocks allocated to them respectively, and determine data that meets the data query condition from the at least one target physical block.
[0029] In practical applications, the leaf nodes of the target index directory may store data identifiers corresponding to the indexes within the second value range of the leaf node, or the leaf nodes may store specific data corresponding to the indexes within the second value range of the leaf node, and the data identifiers can indicate the data corresponding to the indexes. For example, a primary key index directory usually stores specific data corresponding to the indexes in its leaf nodes.
[0030] When the first data node queries data from the target index directory, if the contents stored in the leaf nodes of the target index directory are different, the multiple processing units can scan the target physical blocks assigned to each of them, and the methods of determining the data that meets the data query conditions from the at least one target physical block are different, which will be introduced below.
[0031] If specific data corresponding to the index is stored in the leaf node of the target index directory, multiple processing units can scan the target physical blocks assigned to them and determine data that meets the data query condition from the at least one target physical block.
[0032] If the leaf node in the target index directory stores the data identifier corresponding to the index, multiple processing units can scan the target physical blocks assigned to them. For each of the multiple processing units, the processing unit can determine the data identifier corresponding to at least one index in the at least one target physical block that meets the data query conditions, and perform a table retrieval based on the at least one data identifier to obtain the data that meets the data query conditions.
[0033] Optionally, after determining the data that meets the data query condition based on the scanning results of the multiple processing units, the first data node may further determine a data query result based on the data that meets the data query condition, and then send the data query result to the coordination node.
[0034] In a second aspect, a parallel scanning device is provided, which has the function of implementing the parallel scanning method described in the first aspect. The parallel scanning device includes at least one module, which is used to implement the parallel scanning method described in the first aspect.
[0035] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program for executing the parallel scanning method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the parallel scanning method described in the first aspect.
[0036] Optionally, the computer device may further include a communication bus, which is used to establish a connection between the processor and the memory.
[0037] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the parallel scanning method described in the first aspect.
[0038] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is caused to perform the steps of the parallel scanning method described in the first aspect. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer is caused to perform the steps of the parallel scanning method described in the first aspect.
[0039] The technical effects obtained in the above-mentioned second, third, fourth and fifth aspects are similar to those obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0041] FIG2 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application;
[0042] FIG3 is a flow chart of a parallel scanning method provided in an embodiment of the present application;
[0043] FIG4 is a schematic diagram of multiple threads provided in an embodiment of the present application;
[0044] FIG5 is a flow chart of another parallel scanning method provided in an embodiment of the present application;
[0045] FIG6 is a schematic diagram of a target physical block provided in an embodiment of the present application;
[0046] FIG7 is a schematic diagram of a first thread scanning range provided in an embodiment of the present application;
[0047] FIG8 is a schematic diagram of a second thread scanning range provided in an embodiment of the present application;
[0048] FIG9 is a schematic structural diagram of a parallel scanning device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0050] For ease of understanding, before explaining in detail the parallel scanning method provided in the embodiment of the present application, the nouns, application scenarios and implementation environment involved in the embodiment of the present application are first introduced.
[0051] First, the nouns involved in the embodiments of the present application are introduced.
[0052] Symmetric multi-processing (SMP) is a technology that uses a computer's multi-core CPU architecture to implement multi-threaded parallel computing, fully utilizing central processing unit (CPU) resources to improve database query performance.
[0053] Distributed database: A logically unified database formed by connecting multiple physically dispersed database units through a computer network. Each connected database unit is called a data site or data node.
[0054] Centralized database: refers to a database in which the data is stored on a physical device and the data processing is also completed on the physical device.
[0055] Base table: that is, table. A table is an object used to store data in a database. It is a collection of structured data and the foundation of the entire database system. A table is a database object that contains all the data in the database. A table is defined as a collection of columns.
[0056] Index (index directory): a data structure that sorts the values of one or more fields (one or more columns) in a base table to assist in quickly querying and updating data in the database table.
[0057] NULL value: A null value. It is a special marker used in Structured Query Language. It is an identifier for unknown or missing attributes and is used to indicate uncertain values in the database.
[0058] Primary key: A table often has a column or combination of columns whose values uniquely identify each row in the table. Such a column or columns is called the primary key of the table and is used to enforce the entity integrity of the table.
[0059] Point query: refers to querying a specific value, for example, querying the data of a specific identity document (ID).
[0060] Range query: refers to querying values within a certain range, for example, querying data within a certain time range.
[0061] Next, we'll use another example to illustrate point and range queries. Consider a student profile table containing two fields: student ID and class. In this case, if a user requests data for a student with an ID of 500, the query is a point query. If a user requests data for students with IDs greater than 500 and less than 1000, the query is a range query.
[0062] Next, the application scenarios involved in the embodiments of this application are introduced.
[0063] A database is a technology for storing and managing data and is an indispensable part of computer application systems. It provides a way to organize, store, and manage large amounts of data, and supports operations such as querying, updating, and deleting data. Databases are a core component of modern information systems and are widely used in various fields, such as finance, healthcare, education, and e-commerce. With the development of computer hardware technology, symmetric multi-processing (SMP) technology is increasingly being used in databases to improve database query performance. SMP technology is a technology that uses a computer's multi-core CPU architecture to implement multi-threaded parallel computing to fully utilize CPU resources to improve query performance. This multi-threading can use various inter-thread synchronization technologies within the same process to exchange data and information, thereby reducing the occupation of computing resources and increasing the effective utilization of system computing resources.
[0064] In related art, when a user queries a database for target data, the database can organize multiple physical blocks storing the data into a task queue. Multiple threads can then sequentially retrieve the next physical block to be scanned from the task queue and scan the physical blocks based on the query criteria, thereby determining the query result for the target data based on the scan results of the multiple threads. The multiple threads serially retrieve the next physical block to be scanned and scan the retrieved physical storage blocks in parallel, meaning the multiple threads serially retrieve tasks and scan in parallel.
[0065] For example, a database stores multiple physical storage blocks, numbered 1-4, containing data. These four physical blocks are listed in the task queue in the order 1-4, and two threads (thread 1 and thread 2) are currently scanning these multiple physical blocks. Thread 1 retrieves the next physical block to be scanned from the task queue, which is physical block 1. After thread 1 completes the task, it scans physical block 1. After thread 1 completes the task, thread 2 retrieves the next physical block to be scanned from the task queue, physical block 2, and scans physical block 2. Thread 1's scan of physical block 1 and thread 2's scan of physical block 2 are performed in parallel. If thread 1 completes scanning physical block 1 first, thread 1 proceeds to retrieve the next physical block to be scanned, physical block 3, from the task queue. If thread 2 completes scanning physical block 2 first, thread 2 must wait until thread 1 completes scanning physical block 1 and retrieves physical block 3 before it can retrieve the next physical block to be scanned, physical block 4, from the task queue.
[0066] However, since the process of the multiple threads acquiring tasks is a serial process, this will result in a long waiting time for the multiple threads, reducing processing efficiency. Moreover, when the number of threads is large and the amount of data in the physical blocks is small, the multiple threads need to frequently acquire tasks from the task queue, which will cause the threads to be in a waiting state most of the time, greatly reducing data query efficiency and affecting the user experience. In addition, since not all physical blocks in the multiple physical blocks store the target data, that is, there are physical blocks that do not store the target data and physical blocks that store the target data in the multiple physical blocks, the thread is only effective when scanning the physical blocks that store the target data. However, the thread can only know whether the physical block stores the target data when scanning the physical block. If the thread acquires the next physical block to be scanned that does not store the target data, the thread needs to scan the physical block to know that the target data is not stored in the physical block, and then acquire the next task from the task queue. In this case, the thread's scanning of the physical block that does not store the target data is ineffective, resulting in a waste of the thread's computing resources, greatly reducing the efficiency of database query data.
[0067] Based on this, an embodiment of the present application provides a parallel scanning method, which can determine at least one target physical block from multiple physical blocks based on data query conditions. Since the target physical block can refer to a physical block related to the data to be queried, this can effectively avoid wasting the computing resources of the processing unit and improve the efficiency of database data query. Moreover, since the embodiment of the present application can also allocate the at least one target physical block to multiple processing units based on the total number of at least one target physical block and the number of parallel processing units, compared with related technologies that require serial acquisition tasks, the embodiment of the present application does not require serial acquisition tasks. Each processing unit can scan the target physical block assigned to it, thereby greatly improving data query efficiency and enhancing user experience.
[0068] The execution subject of the parallel scanning method provided in the embodiment of the present application is the first data node included in the database. The first data node is a data node that uses parallel scanning to perform data query. The database can be a distributed database or a centralized database, and the embodiment of the present application does not limit this.
[0069] Next, the implementation environment involved in the embodiments of this application is introduced.
[0070] Based on the above description, the database provided in the embodiment of the present application can be a distributed database or a centralized database, which will be introduced below.
[0071] In the case where the database in the embodiment of the present application is a distributed database, please refer to Figure 1, which is a schematic diagram of an implementation environment provided by the embodiment of the present application. The implementation environment includes a coordination node and multiple data nodes (four data nodes are exemplarily represented in Figure 1), the data nodes are used to store data in the database, and the coordination node is used to manage the multiple data nodes.
[0072] It should be noted that the coordinating node can be any one of the multiple data nodes, or a node other than the multiple data nodes. In the case where the coordinating node is any one of the multiple data nodes, the coordinating node can be used to store data in the database and manage its own data and the data of other data nodes.
[0073] Next, the coordination node and the multiple data nodes will be introduced as different nodes.
[0074] When a user needs to query certain data, the user can send a data query request to the coordination node. Based on the data query request sent by the user, the coordination node determines at least one second data node that needs to be queried from multiple data nodes, and determines the data query instructions corresponding to each of the at least one second data node. For each of the at least one second data node, the coordination node sends the data query instruction corresponding to the second data node to the second data node, instructing the second data node to perform the data query and the method of performing the data query. The second node can receive the data query instruction sent by the coordination node, and then perform a data query based on the data query instruction to obtain a data query result, and then the second data node can send the data query result to the coordination node. The coordination node can receive the data query results sent by each second data node, and then determine the target data query result based on the data query result, and return the target data query result to the user.
[0075] It should be noted that, when the coordinating node is any one of the multiple data nodes, and the at least one second data node includes the coordinating node, the coordinating node can instruct itself to perform a data query to obtain a data query result.
[0076] In some embodiments, the at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform data query in a parallel scanning manner, and the data query instruction includes a data query condition. In this case, for each first data node in the at least one first data node, after receiving the data query instruction sent by the coordination node, the first data node can determine at least one target physical block from the multiple physical blocks based on the data query condition, and then allocate the at least one target physical block to multiple processing units based on the total number of the at least one target physical block and the number of parallel processing units. The multiple processing units refer to processing units that perform data query in a parallel scanning manner, scan the target physical blocks allocated to each of the multiple processing units, determine the data that meets the data query condition based on the scanning results of the multiple processing units, and then determine the data query result based on the data that meets the data query condition.
[0077] It should also be noted that when the coordinating node is any one of multiple data nodes and the at least one first data node includes a coordinating node, the coordinating node can determine its own data query conditions, and based on the data query conditions, determine at least one target physical block from the multiple physical blocks, and finally obtain the data query result.
[0078] In the case where the database in the embodiments of the present application is a centralized database, the centralized database includes only one data node. In this case, the data node also functions as the coordination node. That is, when a user needs to query certain data, the user can send a data query request to the data node. Based on the data query request sent by the user, the data node instructs itself to perform a data query to obtain the target data query result and returns the target data query result to the user.
[0079] In some embodiments, when the data node instructs itself to use a parallel scanning method to perform data query, the data node is also referred to as a first data node. The first data node can determine its own data query conditions, and based on the data query conditions, determine at least one target physical block from the multiple physical blocks, and then based on the total number of at least one target physical block and the number of parallel processing units, allocate the at least one target physical block to multiple processing units. The multiple processing units refer to processing units that use a parallel scanning method to perform data query. The target physical blocks assigned to each of them are scanned by the multiple processing units, and the data that meets the data query conditions is determined based on the scanning results of the multiple processing units. Then, based on the data that meets the data query conditions, the target data query result is determined.
[0080] Those skilled in the art should understand that the above-mentioned coordination nodes and data nodes are only examples. Other existing or future coordination nodes and data nodes that are applicable to the embodiments of the present application should also be included in the scope of protection of the embodiments of the present application and are included here by reference.
[0081] It should be noted that the application scenarios and implementation environments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0082] Please refer to Figure 2, which is a schematic diagram illustrating the structure of a computer device according to an embodiment of the present application. The computer device can serve as part or all of the first data node described above. The computer device includes at least one processor 201, a communication bus 202, a memory 203, and at least one communication interface 204.
[0083] The processor 201 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solution of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0084] The communication bus 202 is used to transmit information between the above components. The communication bus 202 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0085] The memory 203 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compact disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 203 may exist independently and be connected to the processor 201 via the communication bus 202. The memory 203 may also be integrated with the processor 201.
[0086] The communication interface 204 uses any device, such as a transceiver, for communicating with other devices or communication networks. The communication interface 204 includes a wired communication interface and may also include a wireless communication interface. For example, the wired communication interface may be an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0087] In a specific implementation, as an embodiment, the processor 201 may include one or more CPUs, such as CPU0 and CPU1 shown in FIG. 2 .
[0088] In a specific implementation, as an embodiment, a computer device may include multiple processors, such as processor 201 and processor 205 shown in FIG2 . Each of these processors may be a single-core processor or a multi-core processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0089] In a specific implementation, as an embodiment, the computer device may further include an output device 206 and an input device 207. The output device 206 communicates with the processor 201 and can display information in a variety of ways. For example, the output device 206 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 207 communicates with the processor 201 and can receive user input in a variety of ways. For example, the input device 207 can be a mouse, a keyboard, a touch screen device, or a sensor device.
[0090] In some embodiments, the memory 203 is used to store program code 210 for executing the solution of the present application, and the processor 201 can execute the program code 210 stored in the memory 203. The program code 210 may include one or more software modules, and the computer device can implement the parallel scanning method provided in the embodiment of Figure 3 below through the processor 201 and the program code 210 in the memory 203.
[0091] FIG3 is a flowchart of a parallel scanning method provided by an embodiment of the present application, wherein the method is applied to a first data node included in a database, and the first data node includes multiple physical blocks. Referring to FIG3 , the method includes the following steps.
[0092] Step 301: Determine at least one target physical block from a plurality of physical blocks based on a data query condition, where the target physical block refers to a physical block related to the data to be queried. The data query condition is a condition satisfied by the data to be queried.
[0093] In practical applications, the first data node stores a target base table, which refers to a base table containing the data to be queried. The first data node may also store at least one index directory corresponding to the target base table. In this case, the data query condition may instruct the first data node to query data from the target base table, or may instruct the first data node to query data from one of the at least one index directories. When different data query conditions indicate different content, the implementation methods for determining at least one target physical block from multiple physical blocks based on the data query condition vary, and each of these methods will be described below.
[0094] When a data query condition instructs the first data node to query data from a target base table, the data query condition includes a constraint condition for at least one query field and a base table identifier, the base table identifier indicates the target base table, and the multiple physical blocks are used to store the target base table. In this case, the first data node can directly determine the multiple physical blocks as the at least one target physical block.
[0095] In the case where a data query condition instructs the first data node to query data from an index directory of a target base table, the data query condition includes a constraint condition for at least one query field and an index directory identifier, the index directory identifier indicating a target index directory, the target index directory being an index directory constructed based on data stored by the first data node (also referred to as a target base table), the target index directory being one of at least one index directory corresponding to the target base table, and the multiple physical blocks being used to store the target index directory. In this case, if a sort field in the target index directory exists in the at least one query field, then based on the constraint condition for the sort field and the target index directory, at least one target physical block is determined from the multiple physical blocks, the index stored in the target physical block satisfying the constraint condition for the sort field.
[0096] Since the target index directory is constructed based on the sorting field (also known as the index field), if the sorting field in the target index directory exists in the query field, the first data node can directly filter out the indexes that meet the constraints of the sorting field from the target index directory based on the constraints of the sorting field. In this way, the indexes that do not meet the constraints of the sorting field can be filtered, effectively avoiding the subsequent waste of processing unit computing resources and improving the efficiency of database data query.
[0097] In some embodiments, the target index directory includes multiple child nodes, each child node also includes at least one leaf node, the multiple child nodes respectively correspond to multiple first value ranges, and the multiple first value ranges do not overlap. For each child node, the at least one leaf node included in the child node respectively corresponds to at least one second value range, the at least one second value range does not overlap, the union of the at least one second value range is the same as the first value range of the child node, and the first value range and the second value range are the value ranges of the index (also known as the sort field) in the target index directory.
[0098] If the constraint condition of the sorting field is a point query, based on the constraint value in the constraint condition and the first value range corresponding to multiple child nodes of the target index directory, the first child node corresponding to the constraint value is determined from the multiple child nodes, and based on the second value range of at least one leaf node included in the first child node, the leaf node corresponding to the constraint value is determined from the at least one leaf node, and the physical block where the leaf node corresponding to the constraint value is located is the target physical block.
[0099] In some embodiments, the second value ranges of the above-mentioned leaf nodes are different, and the leaf nodes are arranged in order from large to small according to the maximum value (or minimum value) in the value range, or from small to large.
[0100] If the constraint condition of the sorting field is a range query, based on the maximum constraint value and the minimum constraint value in the constraint condition, from the first value range corresponding to multiple child nodes of the target index directory, determine the second child node corresponding to the maximum constraint value and the third child node corresponding to the minimum constraint value; based on the second value range of at least one leaf node included in the second child node, determine the leaf node corresponding to the maximum constraint value as the ending leaf node; based on the second value range of at least one leaf node included in the third child node, determine the leaf node corresponding to the minimum constraint value as the starting leaf node; and use the physical blocks where the starting leaf node, the ending leaf node, and all leaf nodes between the starting leaf node and the ending leaf node are located as the at least one target physical block.
[0101] In some embodiments, if the sorting field in the target index directory does not exist in the at least one query field, the multiple physical blocks are determined as the at least one target physical block.
[0102] If the sorting field in the target index directory does not exist in the query field, the first data node cannot filter the multiple physical blocks according to the target index directory, and therefore can directly determine the multiple physical blocks as the at least one physical block.
[0103] It should be noted that the above description is based on an example in which the data query condition includes a constraint condition for at least one query field. In actual applications, the data query condition may not include a constraint condition for the query field. In this case, the data query instruction instructs the first data node to perform a full table query on the target base table indicated by the base table identifier, or the target index directory indicated by the index directory identifier. If the data query condition includes a base table identifier, the first data node uses the multiple physical blocks where the target base table indicated by the base table identifier is located as the at least one target physical block. If the data query condition includes an index directory identifier, the first data node uses the multiple physical blocks where the target index directory indicated by the index directory identifier is located as the at least one target physical block.
[0104] In actual applications, the database also includes a coordination node. Before determining at least one target physical block from the multiple physical blocks based on the data query conditions, the first data node can also receive a data query instruction sent by the coordination node. The data query instruction is determined by the coordination node based on the data query request sent by the user, and the data query instruction includes the above-mentioned data query conditions.
[0105] Optionally, when a user needs to query data, he or she can send a data query request to the coordination node, and the data query request indicates the base table where the user's data to be queried is located, as well as the constraints of the user's data to be queried. In this case, the coordination node in the database determines the distribution of the base table where the user's data to be queried is located in the data nodes based on the base table where the user's data to be queried is located, and based on the distribution and the constraints of the user's data to be queried, determines at least one second data node that needs to be queried from the data nodes included in the database, and determines the data query instructions corresponding to the at least one second data node. For each of the at least one second data node, the coordination node sends the data query instruction corresponding to the second data node to the second data node to instruct the second data node to perform the data query and whether to use parallel scanning to perform the data query.
[0106] In some embodiments, the at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform data query using a parallel scanning method. In this case, the first data node can receive the data query instruction sent by the coordination node.
[0107] It should be noted that the above is that the coordination node determines at least one second data node based on the data query request sent by the user, and generates a corresponding data query instruction for the at least one second data node. However, in actual applications, the coordination node can also determine at least one second data node based on the data query request sent by the user, and send the data query request to each second data node. Based on the data query request sent by the coordination node, each second data node determines its own data query conditions and whether to use parallel scanning to perform data query. The embodiment of the present application does not limit this.
[0108] Step 302: Based on the total number of at least one target physical block and the number of parallel processing units, the at least one target physical block is allocated to multiple processing units, where the multiple processing units refer to processing units that use a parallel scanning method to perform data query, and the number of parallel processing units is the number of the multiple processing units.
[0109] It should be noted that the multiple processing units in the embodiment of the present application can be multiple processes in the first data node, or multiple threads in the first data node, and the embodiment of the present application is not limited to this.
[0110] In some embodiments, the total number of the at least one target physical block is divided by the number of parallel processing units to obtain a candidate number. If the candidate number is an integer, the at least one target physical block is allocated to the multiple processing units according to the candidate number, so that the number of target physical blocks allocated to each processing unit is the candidate number.
[0111] Optionally, if the candidate number is not an integer, the candidate number is rounded up to obtain a target number, and based on the target number and the number of parallel processing units, the at least one target physical block is allocated to the multiple processing units so that the number of target physical blocks allocated to each processing unit is no greater than the target number.
[0112] The at least one target physical block is allocated to N processing units among the multiple processing units according to the target number, so that the number of target physical blocks allocated to the N processing units is the target number, N is equal to the number of parallel processing units minus 1, if the value after N multiplied by the target number is less than the total number of the at least one target physical block, then according to the first difference, the at least one target physical block is allocated to the remaining processing units, and the remaining processing units are the processing units other than the N processing units among the multiple processing units, and the first difference is the difference between the value after N multiplied by the target number and the total number of the at least one target physical blocks; if the value after N multiplied by the target number is not less than the total number of the at least one target physical blocks, no target physical block is allocated to the remaining processing units.
[0113] If the value obtained by multiplying N by the target number is less than the total number of the at least one target physical blocks, it means that the total number of physical blocks allocated to the N processing units is less than the total number of the at least one target physical blocks, there are still unallocated physical blocks in the at least one target physical block, and the number of the unallocated physical blocks is the first difference. Therefore, the target physical blocks with the first difference can be allocated to the remaining processing units. If the value obtained by multiplying N by the target number is not less than the total number of the at least one target physical blocks, it means that the total number of physical blocks allocated to the N processing units is not less than the total number of the at least one target physical blocks, there are no unallocated physical blocks in the at least one target physical block, and therefore, the target physical blocks may not be allocated to the remaining processing units.
[0114] Based on the above description, when the number of candidates is not an integer, the number of candidates can be rounded up. In actual applications, the number of candidates can also be rounded down to obtain the target number. At this time, there is at least one thread among the multiple threads to which the number of target physical blocks allocated is not greater than the target number. This embodiment of the present application does not limit this.
[0115] In some embodiments, before allocating the at least one target physical block to a plurality of processing units based on the total number of the at least one target physical block and the number of parallel processing units, the first data node needs to determine the number of parallel processing units.
[0116] In some embodiments, the data query instruction further includes the number of parallel processing units. In this case, the first data node can determine the number of parallel processing units. Of course, in actual applications, the first data node can also store the number of parallel processing units. In this case, the first data node can also directly determine the number of parallel processing units.
[0117] Based on the above description, the number of parallel processing units may be included in the data query instruction and sent by the coordinating node to the first data node. Of course, in actual applications, the number of parallel processing units may also not be included in the data query instruction, but rather the coordinating node may separately send the number of parallel processing units to the first data node before the first node allocates the at least one target physical block to multiple processing units based on the total number of the at least one target physical block and the number of parallel processing units. This embodiment of the present application is not limited to this.
[0118] Step 303: Scan the target physical blocks assigned to each of the multiple processing units, and determine data that meets the data query condition based on the scanning results of the multiple processing units.
[0119] Based on the above description, since the first data node can query data from the target base table or from the target index directory, in different cases, multiple processing units are used to scan the target physical blocks assigned to each of them, and the method of determining the data that meets the data query conditions based on the scanning results of the multiple processing units is different.
[0120] In the case where the first data node queries data from the target base table, the multiple processing units can scan the target physical blocks allocated to them respectively, and determine data that meets the data query condition from the at least one target physical block.
[0121] In practical applications, the leaf nodes of the target index directory may store data identifiers corresponding to the indexes within the second value range of the leaf node, or the leaf nodes may store specific data corresponding to the indexes within the second value range of the leaf node, and the data identifiers can indicate the data corresponding to the indexes. For example, a primary key index directory usually stores specific data corresponding to the indexes in its leaf nodes.
[0122] When the first data node queries data from the target index directory, if the contents stored in the leaf nodes of the target index directory are different, the multiple processing units can scan the target physical blocks assigned to each of them, and the methods of determining the data that meets the data query conditions from the at least one target physical block are different, which will be introduced below.
[0123] If specific data corresponding to the index is stored in the leaf node of the target index directory, multiple processing units can scan the target physical blocks assigned to them and determine data that meets the data query condition from the at least one target physical block.
[0124] If the leaf node in the target index directory stores the data identifier corresponding to the index, multiple processing units can scan the target physical blocks assigned to them. For each of the multiple processing units, the processing unit can determine the data identifier corresponding to at least one index in the at least one target physical block that meets the data query conditions, and perform a table retrieval based on the at least one data identifier to obtain the data that meets the data query conditions.
[0125] In some embodiments, the data identifier is the physical address of the data, and the first data node supports multiple processing units to perform read operations on the same physical block at the same time. In this case, the multiple processing units can perform a table lookup based on the physical addresses of at least one data corresponding to the multiple processing units to obtain data that meets the data query conditions.
[0126] In other embodiments, the data identifier is the physical address of the data, and the first data node does not support multiple processing units performing read operations on the same physical block at the same time. In this case, the first data node stores the starting physical address and the ending physical address of each physical block. At this time, the first data node can summarize the physical addresses of at least one data corresponding to the multiple processing units respectively, and determine the multiple second physical blocks that need to be read and the total number of the multiple second physical blocks based on the physical addresses of the multiple data and the starting physical address and the ending physical address of each physical block. Then, based on the total number of the multiple second physical blocks and the number of parallel processing units, the at least one second physical block is allocated to multiple processing units, and the multiple processing units scan the second physical blocks allocated to each of them to obtain data that meets the data query conditions.
[0127] The implementation method of allocating the at least one second physical block to multiple processing units based on the total number of the multiple second physical blocks and the number of parallel processing units is similar to the implementation method of allocating the at least one target physical block to multiple processing units based on the total number of at least one target physical block and the number of parallel processing units. For the detailed implementation process, please refer to the relevant content above and will not be repeated here.
[0128] In actual applications, the data identifier can also be the primary key value of the data. In this case, the database also stores the primary key index directory of the target base table, and the primary key value is used to obtain the data corresponding to the primary key value from the leaf node of the primary key index directory.
[0129] Optionally, after determining the data that meets the data query condition based on the scanning results of the multiple processing units, the first data node may further determine a data query result based on the data that meets the data query condition, and then send the data query result to the coordination node.
[0130] In some embodiments, after the first data node determines the data that meets the data query conditions, it also needs to process the data that meets the data query conditions, such as summing, etc. In this case, the first data node can also process the data that meets the data query conditions according to the actual situation and the relevant algorithm to obtain the data query results.
[0131] In actual applications, before using multiple processing units to scan the target physical blocks assigned to each of them, the first data node needs to start the multiple processing units, the number of which is the same as the number of the parallel processing units. After starting the multiple processing units, the multiple processing units can scan the target physical blocks assigned to each of them.
[0132] It should be noted that the timing for starting the multiple processing units can be before executing step 301 or before executing step 303, and this embodiment of the application does not limit this. If the timing for starting the multiple processing units is before executing step 301, any one processing unit can be selected from the multiple processing units to execute the above steps 301 and 302.
[0133] For ease of description, the processing unit among the multiple processing units that executes steps 301 and 302 is subsequently referred to as a first processing unit. After the first processing unit allocates the at least one target physical block to the multiple processing units, the first processing unit can store the allocation result of the at least one target physical block into a shared memory, so that the multiple processing units can read their corresponding target physical blocks from the shared memory.
[0134] It should also be noted that after the multiple processing units obtain their own scanning results, the multiple processing units can send their own scanning results to a second processing unit. The second processing unit can be any one of the multiple processing units, or a processing unit other than the multiple processing units. The second processing unit (also called the summary processing unit) is used to summarize the data of the multiple processing units.
[0135] For example, when the multiple processing units are multiple threads, please refer to Figure 4. The multiple threads are threads 1-3, the target physical blocks assigned to thread 1 are physical blocks 1-3, the target physical blocks assigned to thread 2 are physical blocks 4-6, and the target physical blocks assigned to thread 3 are physical blocks 7-9. After the scanning is completed, threads 1-3 can send their own scanning results to the summary thread (that is, the above-mentioned summary processing unit) so that the summary thread can summarize the data of the multiple threads.
[0136] Next, the parallel scanning method provided in the embodiment of the present application will be introduced again with reference to FIG5 .
[0137] When a user needs to query certain data, the user can send a data query request to the coordinating node. Based on the data query request sent by the user, the coordinating node determines at least one second data node from a plurality of data nodes on which the data query needs to be performed, and determines a data query instruction corresponding to each of the at least one second data node. The at least one second data node includes at least one first data node. The data query instruction corresponding to the first data node instructs the first data node to perform the data query using a parallel scanning method. The data query instruction includes a data query condition. In this case, for each of the at least one first data node, after receiving the data query instruction sent by the coordinating node, the first data node can determine at least one target physical block from the plurality of physical blocks based on the data query condition, and then assign the at least one target physical block to a plurality of processing units based on the total number of the at least one target physical block and the number of parallel processing units. The plurality of processing units refers to processing units that perform the data query using a parallel scanning method. The plurality of processing units scan the assigned target physical blocks, and after the scanning is completed, determine data that meets the data query condition based on the scanning results of the plurality of processing units, and then determine a data query result based on the data that meets the data query condition.
[0138] The following will specifically illustrate the parallel scanning method provided in the embodiment of the present application by taking examples.
[0139] The embodiment of the present application can be applied in the database corresponding to the school's student registration system. In the student registration system, there is an archive table (portfolio) as shown in Table 1 below, and the table attributes of the portfolio are shown in Table 2 below. In Table 1, the portfolio includes four fields, namely, the studentid field, the docid field, the name field, and the classid field. As can be seen in Table 2, the data type of the studentid field is integer (integer), the data type of the docid field is a string of up to 10 characters, the data type of the name field is a variable-length string with a length of up to 50 characters, the data type of the classid field is an integer, and the primary key of the archive table is studentid and classid. In other words, the primary key index directory (portfolio_idx) of the archive table is a joint index of studentid and classid, wherein the studentid is the sorting field of the primary key index directory.
[0140] Table 1
[0141] Table 2
[0142] Next, the process of performing data query on the database corresponding to the above-mentioned student registration system will be introduced through the following examples 1-5, taking the multiple processing units as multiple threads as an example.
[0143] Example 1: A user wishes to query the class ID of a student with ID 1000. In this case, the user sends the data query request "select classid from portfolio where studentid = 1000" to the coordination node. After receiving the data query request, the coordination node determines the data query instructions corresponding to at least one second data node for the data query. The at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform the data query using a parallel scan. The data query instructions corresponding to the first data node are shown below.
[0144] Streaming(type:LOCAL GATHER dop:1 / 2)
[0145] ->Index Only Scan using portfolio_idx on portfolio
[0146] Index Cond:(studentid=1000)
[0147] This data query instructs the first data node to start two threads to scan in parallel, using the index directory named "portfolio_idx" to retrieve all records with a studentID equal to 1000 from the table named "portfolio". The query is processed in a streaming manner with a parallelism of 1 / 2. This parallelism of 1 / 2 means that any one of the two threads is selected as the second thread to perform data aggregation.
[0148] The query field in the data query instruction is studentid, the constraint condition of the query field is student=1000, the index directory identifier is portfolio_idx, and the number of parallel threads is 2.
[0149] Please refer to Figure 6. The first thread (i.e., the first processing unit mentioned above) determines that the target physical block is physical block 6 based on the constraints of the sort field and the target index directory. The total number of target physical blocks is divided by the number of parallel threads, that is, 1 divided by 2, to obtain a candidate number of 0.5. Since the candidate number is not an integer, the candidate number is rounded up to obtain a target number of 1. The number of target physical blocks allocated to the first thread is 1. Since the value after multiplying 1 by 1 is not less than the total number of the at least one target physical block, the target physical block is not allocated to the second thread. Therefore, the first thread can put the allocation result of the at least one target physical block into the shared memory, so that the two threads can read their corresponding target physical blocks from the shared memory. The allocation result is that the physical block allocated to the first thread is physical block 6. The second thread is not allocated a physical block. The first thread scans physical block 6 and obtains a piece of data with studentid=1000, which is 23. The data is sent to the second thread. The second thread summarizes the scanning results of the first and second threads and returns 23 to the coordination node. The coordination node returns 23 to the user.
[0150] Example 2: A user wants to query the number of students whose student IDs are between 23 and 1500. In this case, the user can send a data query request select count(*)from portfolio where studentid>=23AND studentid<=1500 to the coordination node. After receiving the data query request, the coordination node can query at least one second data node and determine the data query instructions corresponding to the at least one second data node. The at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform data query using a parallel scan method. The data query instruction corresponding to the first data node is shown below.
[0151] This data query instructs the first data node to start two threads to scan in parallel, using the index directory named "portfolio_idx" to retrieve the total number of records with a studentID greater than or equal to 23 and less than or equal to 1500 from the table named "portfolio". The query is processed in a streaming manner with a parallelism of 1 / 2. This parallelism of 1 / 2 means that any one of the two threads is selected as the second thread to perform data aggregation.
[0152] The query field in the data query instruction is studentid, the constraint condition of the query field is (studentid>=23) AND (studentid<=1500), the index directory identifier is portfolio_idx, and the number of parallel threads is 2.
[0153] Based on the constraints of the sort field and the target index directory, the first thread determines that the target physical blocks are physical blocks 1-9. The total number of target physical blocks is divided by the number of parallel threads, that is, 9 divided by 2, to obtain a candidate number of 4.5. Since the candidate number is not an integer, the candidate number is rounded down to obtain a target number of 4. The number of target physical blocks allocated to the first thread is 4. Since the value after multiplying 1 by 4 is less than the total number of the at least one target physical blocks, the at least one target physical block is allocated to the second thread according to the first difference, and the number of target physical blocks allocated to the second thread is 5. The first thread can put the allocation result of the at least one target physical block into the shared memory so that the two threads can read their corresponding target physical blocks from the shared memory. The allocation result is that the physical blocks allocated to the first thread are physical blocks 1-4. The physical blocks allocated to the second thread are physical blocks 5-9.
[0154] Please refer to Figure 7. The first thread starts scanning from the tuple with studentid>=23 in physical block 1 to the end of block number 4, and counts the tuples that meet the conditions. Please refer to Figure 8. The second thread starts scanning from the starting position of physical block 5 to the last tuple with studentid<=1500 in physical block 9, and counts the tuples that meet the conditions. The second thread summarizes the scanning results of each thread to obtain the number of students with student ID between 23 and 1500, which is 1478. The second thread returns 1478 to the coordination node, and the coordination node returns 1478 to the user.
[0155] Example 3: A user wants to query the number of students in classes with IDs 19, 23, 45, 80, 100, 150, 151, 170, 187, 195, and 199. In this case, the user can send the data query request select / *+indexonlyscan(portfolio portfolio_idx)* / count(*)from portfolio where classid in(19, 23, 45, 80, 100, 150, 151, 170, 187, 195, 199) to the coordination node. After receiving the data query request, the coordination node can identify at least one second data node for data query and determine the data query instructions corresponding to each of the at least one second data node. The at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform the data query using a parallel scan. The data query instructions corresponding to the first data node are shown below.
[0156] This data query instructs the first data node to start two threads to scan in parallel, using the index directory named "portfolio_idx" to retrieve the total number of students with classid equal to 19, 23, 45, 80, 100, 150, 151, 170, 187, 195, and 199 from the table named "portfolio". The query is then processed in a streaming manner with a parallelism of 1 / 2. This parallelism of 1 / 2 means that any one of the two threads is selected as the second thread to perform data aggregation.
[0157] The query field in the data query instruction is classid, the constraint condition of the query field is classid=ANY{19, 23, 45, 80, 100, 150, 151, 170, 187, 195, 199}, the index directory identifier is portfolio_idx, and the number of parallel threads is 2.
[0158] Since the sort field in the target index directory does not exist in the query field, the multiple physical blocks are determined to be the at least one physical block, that is, the at least one target physical block is physical blocks 1-100. Referring to Figure 000, the physical blocks assigned to the first thread are physical blocks 1-50. The physical blocks assigned to the second thread are physical blocks 51-100. Finally, the number of students with class IDs 19, 23, 45, 80, 100, 150, 151, 170, 187, 195, and 199 is 1089. 1089 is returned to the coordination node, and the coordination node returns 1089 to the user.
[0159] Example 4: The user wishes to query the number of students whose student id column is not null. In this case, the user can send a data query request, select count(*) from portfolio where studentid is not null, to the coordination node. After receiving the data query request, the coordination node can query at least one second data node and determine the data query instructions corresponding to each of the at least one second data node. The at least one second data node includes at least one first data node, and the data query instruction corresponding to the first data node instructs the first data node to perform data query using a parallel scan. The data query instruction corresponding to the first data node is shown below.
[0160] This data query instructs the first data node to start two threads to scan in parallel, and retrieve the total number of students whose studentid is not NULL from the table named "portfolio" using the index directory named "portfolio_idx". This is processed in a streaming manner with a parallelism of 1 / 2. This parallelism of 1 / 2 means that any one of the two threads is selected as the second thread to perform data aggregation.
[0161] The query field in this data query instruction is classid, the constraint condition of the query field is studentid IS NOT NULL, the index directory identifier is portfolio_idx, and the number of parallel threads is 2.
[0162] The first data node determines that the at least one target physical block is physical block 1-100, and finally obtains the number of students with non-empty studentID as 20,000. The second thread returns 20,000 to the coordination node, and the coordination node returns 20,000 to the user.
[0163] Example 5: A user wishes to query the number of students whose student ID equals their class ID. In this case, the user can send a data query request, "select count(*)from portfolio as t0 where exists(select 1from portfolio as t where t0.studentid=t.classid)," to the coordination node. Upon receiving the data query request, the coordination node can determine at least one second data node for data query and the corresponding data query instructions for each of the at least one second data node. The at least one second data node includes at least one first data node, and the corresponding data query instruction for each first data node instructs the first data node to perform data query using a parallel scan. The data query instruction for each first data node is shown below.
[0164] This data query instructs the first data node to start two threads to scan in parallel, using the index directory named "portfolio_idx" to retrieve the total number of students where t0.studentid = t.classid from the table named "portfolio". This query is performed in a streaming manner with a parallelism of 1 / 2. This parallelism of 1 / 2 means that any one of the two threads is selected as the second thread to perform data aggregation.
[0165] The query fields in this data query instruction are classid and studentid, the query field constraints are t0.studentid = t.classid, the index directory identifier is portfolio_idx, and the number of parallel threads is 2. The at least one target physical block is physical blocks 1-100. The first data node determines that the final number of students with studentid = classid is 200. The second thread returns 200 to the coordination node, which then returns 200 to the user.
[0166] It should be noted that the above examples do not constitute a limitation on the usage scenarios of the parallel scanning method provided in the embodiments of the present application. The technical solutions provided in the embodiments of the present application are also applicable to other similar technical problems.
[0167] The embodiment of the present application can determine at least one target physical block from multiple physical blocks based on the data query conditions. Since the target physical block can refer to a physical block related to the data to be queried, this can effectively avoid wasting thread computing resources and improve the efficiency of database data query. Moreover, since the embodiment of the present application can also allocate the at least one target physical block to multiple threads based on the total number of at least one target physical block and the number of parallel threads, compared with the related technology that requires serial acquisition tasks, the embodiment of the present application does not require serial acquisition tasks. Each thread can scan the target physical block assigned to it, thereby greatly improving data query efficiency and improving user experience. Moreover, since the target index directory is constructed based on the sorting field (also known as the index field), if the sorting field in the target index directory exists in the query field, the first data node can directly filter out the index that meets the constraint conditions of the sorting field from the target index directory based on the constraint conditions of the sorting field. In this way, the index that does not meet the constraint conditions of the sorting field can be filtered, effectively avoiding the waste of thread computing resources and improving the efficiency of database data query.
[0168] FIG9 is a schematic diagram of the structure of a parallel scanning device provided in an embodiment of the present application. The parallel scanning device can be implemented by software, hardware, or a combination of both to form part or all of the first data node described above, where the first data node includes multiple physical blocks. Referring to FIG9 , the device includes a first determination module 901, an allocation module 902, and a second determination module 903.
[0169] The first determination module 901 is configured to determine at least one target physical block from a plurality of physical blocks based on a data query condition, where the target physical block is a physical block associated with the data to be queried. The data query condition is a condition satisfied by the data to be queried. The detailed implementation process is described in the corresponding sections of the above embodiments and will not be repeated here.
[0170] Allocation module 902 is configured to allocate the at least one target physical block to multiple processing units based on the total number of the at least one target physical block and the number of parallel processing units. The multiple processing units are processing units that perform data queries using a parallel scanning method, and the number of parallel processing units is the number of the multiple processing units. The detailed implementation process is described in the corresponding embodiments above and will not be repeated here.
[0171] The second determining module 903 is configured to scan the target physical blocks assigned to each of the multiple processing units and determine the data that meets the data query condition based on the scanning results of the multiple processing units. The detailed implementation process is referred to the corresponding content of the above embodiments and will not be repeated here.
[0172] Optionally, the data query condition includes a constraint condition of at least one query field and an index directory identifier, the index directory identifier indicates a target index directory, the target index directory is an index directory constructed based on data stored in the first data node, and the multiple physical blocks are used to store the target index directory;
[0173] The first determining module 901 is specifically configured to:
[0174] If a sorting field in the target index directory exists in at least one query field, then based on the constraints of the sorting field and the target index directory, at least one target physical block is determined from multiple physical blocks, and the index stored in the target physical block meets the constraints of the sorting field.
[0175] Optionally, the first determining module 901 is specifically configured to:
[0176] If the sorting field in the target index directory does not exist in at least one query field, multiple physical blocks are determined as at least one target physical block.
[0177] Optionally, the allocation module 902 is specifically configured to:
[0178] Dividing the total number of the at least one target physical block by the number of parallel processing units to obtain a candidate number;
[0179] If the candidate number is an integer, at least one target physical block is allocated to a plurality of processing units according to the candidate number, so that the number of target physical blocks allocated to each processing unit is the candidate number.
[0180] Optionally, the allocation module 902 is specifically configured to:
[0181] If the candidate number is not an integer, the candidate number is rounded up to obtain the target number;
[0182] Based on the target number and the number of parallel processing units, at least one target physical block is allocated to the plurality of processing units so that the number of target physical blocks allocated to each processing unit is not greater than the target number.
[0183] Optionally, the database further includes a coordination node, and the device further includes:
[0184] The receiving module is configured to receive a data query instruction sent by the coordination node. The data query instruction is determined by the coordination node based on a data query request sent by a user. The data query instruction includes a data query condition.
[0185] Optionally, the data query instruction also includes the number of parallel processing units.
[0186] Optionally, the device further comprises:
[0187] A third determining module, configured to determine a data query result based on data that meets the data query condition;
[0188] The sending module is used to send the data query results to the coordination node.
[0189] Optionally, the multiple processing units are multiple processes in the first data node, or multiple threads in the first data node.
[0190] In an embodiment of the present application, since the target physical block can refer to a physical block related to the data to be queried, this can effectively avoid wasting thread computing resources and improve the efficiency of database data query. Moreover, since the embodiment of the present application can also allocate the at least one target physical block to multiple threads based on the total number of at least one target physical block and the number of parallel threads, compared with the related technology that requires serial acquisition tasks, the embodiment of the present application does not require serial acquisition tasks. Each thread can scan the target physical block assigned to it, thereby greatly improving data query efficiency and improving user experience. Moreover, since the target index directory is constructed based on the sorting field (also known as the index field), if the sorting field in the target index directory exists in the query field, the first data node can directly filter out the index that meets the constraint conditions of the sorting field from the target index directory based on the constraint conditions of the sorting field. In this way, the index that does not meet the constraint conditions of the sorting field can be filtered, effectively avoiding the waste of thread computing resources and improving the efficiency of database data query.
[0191] It should be noted that the parallel scanning device provided in the above embodiment is merely an example of the division of the functional modules described above when performing parallel scanning. In actual applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the parallel scanning device provided in the above embodiment and the parallel scanning method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0192] An embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the parallel scanning method described in the above embodiment.
[0193] The present application also provides a computer program product comprising instructions that, when executed on a computer, cause the computer to perform the steps of the parallel scanning method described in the above embodiment. Alternatively, the present application also provides a computer program that, when executed on a computer, causes the computer to perform the steps of the parallel scanning method described in the above embodiment.
[0194] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0195] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0196] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the data query conditions involved in the embodiments of this application are all obtained with full authorization.
[0197] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A parallel scanning method, characterized in that, Applied to the first data node included in the database, the first data node includes a plurality of physical blocks, and the method includes: Determine at least one target physical block from the plurality of physical blocks based on a data query condition, where the target physical block refers to a physical block related to the data to be queried, and the data query condition is the condition satisfied by the data to be queried; Based on the total number of the at least one target physical block and the number of parallel processing units, allocate the at least one target physical block to a plurality of processing units, where the plurality of processing units refer to processing units that perform data query in a parallel scan manner, and the number of parallel processing units is the number of the plurality of processing units; Scan the target physical blocks allocated to each of the plurality of processing units through the plurality of processing units, and determine the data that satisfies the data query condition based on the scan results of the plurality of processing units.
2. The method according to claim 1, characterized in that, The data query condition includes constraint conditions of at least one query field and an index directory identifier, where the index directory identifier indicates a target index directory, and the target index directory is an index directory constructed based on the data stored in the first data node, and the plurality of physical blocks are used to store the target index directory; The determining at least one target physical block from the plurality of physical blocks based on the data query condition includes: If there is a sorting field in the target index directory among the at least one query field, determine the at least one target physical block from the plurality of physical blocks based on the constraint condition of the sorting field and the target index directory, where the index stored in the target physical block satisfies the constraint condition of the sorting field.
3. The method according to claim 2, wherein The method further includes: If there is no sorting field in the target index directory among the at least one query field, determine the plurality of physical blocks as the at least one target physical block.
4. The method according to any one of claims 1 to 3, characterized in that, The allocating the at least one target physical block to a plurality of processing units based on the total number of the at least one target physical block and the number of parallel processing units includes: Divide the total number of the at least one target physical block by the number of parallel processing units to obtain a candidate number; If the candidate number is an integer, allocate the at least one target physical block to the plurality of processing units according to the candidate number, so that the number of target physical blocks allocated to each processing unit is the candidate number.
5. The method according to claim 4, characterized in that The method further includes: If the candidate number is not an integer, round up the candidate number to obtain a target number; Based on the target number and the number of parallel processing units, allocate the at least one target physical block to the plurality of processing units, so that the number of target physical blocks allocated to each processing unit is not greater than the target number.
6. The method according to any one of claims 1-5, characterized in that The database further includes a coordination node. Before determining at least one target physical block from the plurality of physical blocks based on the data query condition, the method further includes: Receive a data query instruction sent by the coordination node, where the data query instruction is determined by the coordination node based on a data query request sent by a user, and the data query instruction includes the data query condition.
7. The method according to claim 6, wherein The data query instruction further includes the number of parallel processing units.
8. The method according to claim 6, wherein After determining the data that meets the data query condition based on the scanning results of the multiple processing units, the method further includes: Determining a data query result based on the data that meets the data query condition; Sending the data query result to the coordination node.
9. The method according to any one of claims 1-8, characterized in that, The multiple processing units are multiple processes in the first data node or multiple threads in the first data node.
10. A parallel scanning device, characterized in that, Applied to a first data node included in a database, the first data node includes multiple physical blocks, and the apparatus includes: A first determination module, configured to determine at least one target physical block from the multiple physical blocks based on a data query condition, where the target physical block refers to a physical block related to the data to be queried, and the data query condition is a condition satisfied by the data to be queried; An allocation module, configured to allocate the at least one target physical block to multiple processing units based on the total number of the at least one target physical block and the number of parallel processing units, where the multiple processing units refer to processing units that perform data query in a parallel scanning manner, and the number of parallel processing units is the number of the multiple processing units; A second determination module, configured to scan the target physical blocks allocated to each of them by the multiple processing units, and determine the data that meets the data query condition based on the scanning results of the multiple processing units.
11. A computer device, characterized in that, The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is configured to execute the computer program stored in the memory to implement the steps of the method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions run on the computer, the computer is caused to execute the steps of the method according to any one of claims 1-9.
13. A computer program product containing instructions, characterized in that, When the instructions run on the computer, the computer is caused to execute the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Parallel scanning method and device, equipment, storage medium and computer program
CN120234369A
Distributed database system, method for building index therein and query method
CN102375853A
Data processing method, device and system and computer storage medium
CN113419824A
Data processing system and method, and storage medium
CN116795875A
Data query method, electronic equipment and storage medium
CN118503311A