Data query method, device and equipment and computer readable storage medium
Patent Information
- Application Number
- CN202311330674.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-13
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-10-13
AI Technical Summary
多个并行线程同时访问数据,如果访问模式不合理,可能会引发数据访问冲突,包括锁竞争、数据竞争等问题,导致数据查询效率下降或结果不正确等问题
[0039] The data query method, apparatus, device, and computer-readable storage medium provided in this disclosure utilize a cursor variable in shared memory to indicate the data retrieval progress in the current database. Multiple worker processes alternately drive the cursor variable, eliminating the need for data exchange between worker processes. This greatly simplifies the progress synchronization process between multiple worker processes, avoids introducing additional communication overhead, and effectively improves the efficiency and accuracy of data query.
Smart Images

Figure CN117520366B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a data query method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] Data scanning of the main database table is one of the most fundamental operations in relational databases. After the database receives a user's query request, searches the query path, generates an execution plan, and scans and reads the data from the specified database table, it further performs various computationally intensive tasks such as filtering and projection, finally returning the data that meets the query requirements to the user. While this process involves complex computationally intensive tasks, the time overhead from the I / O-intensive task of scanning and reading database table data is the most significant.
[0003] With the significant improvement in hardware parallelism and concurrency capabilities, the idea of accelerating data querying by using multiple processes or threads to perform parallel scanning of databases has been widely adopted in recent years.
[0004] However, parallel table scanning involves cooperation and communication between multiple parallel threads. Threads need to exchange data and results, which introduces additional communication overhead. Multiple parallel threads accessing data simultaneously, if the access pattern is unreasonable, may cause data access conflicts, including lock contention and data races, leading to decreased data query efficiency or incorrect results. Summary of the Invention
[0005] To address the aforementioned technical problems, this disclosure provides a data query method, apparatus, device, and computer-readable storage medium to improve the efficiency and accuracy of data querying.
[0006] In a first aspect, embodiments of this disclosure provide a data query method, including:
[0007] Generate an expression information tree from the expression tree;
[0008] Copy the expression information tree to shared memory;
[0009] Multiple worker processes sequentially adjust the cursor variable in the shared memory that indicates the progress of database data acquisition, and retrieve the target data corresponding to the adjustment operation from the database to the DPU. The worker processes run in the DPU.
[0010] Each of the worker processes accesses the shared memory, performs a data query on the target data based on the expression information tree, and obtains the data query result.
[0011] In some embodiments, generating an expression information tree from an expression tree includes:
[0012] Determine the computational information required for each node in the expression tree and the access sequence number of the node;
[0013] An expression information tree is generated based on the expression tree, the calculation information corresponding to each node, and the access sequence number.
[0014] In some embodiments, the computational information includes at least one or more of the following:
[0015] The operands of the node, the operation type of the node, and the constants that the node operation needs to be converted into DPU computing format.
[0016] In some embodiments, copying the expression information tree to shared memory includes:
[0017] Based on the access sequence number of each node in the expression information tree and the data contained in the node, the expression information tree is saved in shared memory as an array.
[0018] In some embodiments, the worker process includes a main thread and worker threads. The step of sequentially adjusting the cursor variable in the shared memory that indicates the progress of database data acquisition through multiple worker processes, and retrieving the target data corresponding to the adjustment operation from the database to the DPU, includes:
[0019] For each worker process, the starting block number indicated by the current cursor variable is read through the main thread;
[0020] Advance the cursor variable by a preset number of blocks so that the cursor variable points to the target block number;
[0021] Read the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and obtain the target data into the working thread.
[0022] In some embodiments, accessing the shared memory through each of the worker processes, performing a data query on the target data according to the expression information tree, and obtaining the data query result includes:
[0023] For each worker process, the expression information tree in the form of an array in shared memory is accessed according to the access sequence number to obtain the information in the expression information tree;
[0024] The target data is queried based on the information in the expression information tree to obtain the data query results.
[0025] In some embodiments, the method includes:
[0026] Perform memory alignment operations on the data in the shared memory.
[0027] Secondly, embodiments of this disclosure provide a data query device, comprising:
[0028] The generation module is used to generate an expression information tree based on the expression tree;
[0029] A copy module is used to copy the expression information tree to shared memory;
[0030] The acquisition module is used to sequentially adjust the cursor variable in the shared memory that indicates the progress of database data acquisition through multiple worker processes, and to acquire the target data corresponding to the adjustment operation from the database to the DPU. The worker processes run in the DPU.
[0031] The query module is used to access the shared memory through each of the worker processes, perform data queries on the target data according to the expression information tree, and obtain data query results.
[0032] Thirdly, embodiments of this disclosure provide an electronic device, including:
[0033] Memory;
[0034] Processor; and
[0035] Computer programs;
[0036] The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.
[0037] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described in the first aspect.
[0038] Fifthly, embodiments of this disclosure also provide a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the data query method described above.
[0039] The data query method, apparatus, device, and computer-readable storage medium provided in this disclosure utilize a cursor variable in shared memory to indicate the data retrieval progress in the current database. Multiple worker processes alternately drive the cursor variable, eliminating the need for data exchange between worker processes. This greatly simplifies the progress synchronization process between multiple worker processes, avoids introducing additional communication overhead, and effectively improves the efficiency and accuracy of data query. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0041] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 A flowchart of a data query method provided in this embodiment of the disclosure;
[0043] Figure 2 A schematic diagram illustrating an application scenario provided by an embodiment of this disclosure;
[0044] Figure 3 A flowchart of a data query method provided in another embodiment of this disclosure;
[0045] Figure 4 This is a schematic diagram of an expression information tree node mapping provided in an embodiment of the present disclosure;
[0046] Figure 5 This is a schematic diagram of multi-process collaboration provided in an embodiment of the present disclosure;
[0047] Figure 6 This is a schematic diagram of the structure of the data query device provided in the embodiments of this disclosure;
[0048] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0049] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0050] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0051] Data scanning of the master table is one of the most fundamental operations in relational databases. After receiving a user's query request, the database searches for a query path, generates an execution plan, scans and reads data from the specified database tables, and then performs various computationally intensive tasks such as filtering and projection. Finally, it returns the data that meets the query requirements to the user. The query path refers to the path data takes from storage media (such as a hard drive) to the query processor when a query operation is executed in a computer system or database, including all steps of data retrieval, transmission, and processing. The execution plan refers to the sequence of steps and operations taken when executing a query statement in the database management system. The database management system generates an optimal execution plan based on information such as the query statement and table structure to efficiently retrieve and process data. The execution plan includes details such as the order of query operations, join methods, and data access methods.
[0052] In the above process, the time overhead caused by the I / O-intensive task of scanning and reading database table data is the most significant. I / O (Input / Output) refers to the process of data transmission in a computer system. Input refers to introducing external data or signals into the computer system, while output refers to transmitting data from the computer system to the external environment.
[0053] Currently, the method used to accelerate data queries, relying on the multi-core capabilities of processors, is based on data partitioning. Data partitioning divides the dataset to be scanned into multiple subsets, each of which is processed concurrently using a certain number of processor cores to achieve some degree of acceleration. However, data partitioning itself has certain limitations. First, it introduces communication overhead. Parallel table scanning involves cooperation and communication between multiple parallel threads. Threads need to exchange data and results, which introduces communication overhead. Secondly, in parallel table scanning, the data table is divided into multiple partitions, but sometimes the partitioning is uneven, causing some parallel threads to be overloaded while others are underloaded. This leads to uneven system performance, with some threads potentially becoming bottlenecks and failing to fully leverage the advantages of parallel computing. Furthermore, existing parallel scanning techniques have scalability limitations when processing large-scale data. For example, they cannot effectively utilize a large number of processor cores, thus failing to meet the demands of high concurrency and large-scale data processing.
[0054] To address the aforementioned problems, this disclosure provides a data query method, which will be described below with reference to specific embodiments.
[0055] Figure 1 A flowchart illustrating a data query method provided in this embodiment. This method can be applied to... Figure 2The application scenario shown includes a data query system, specifically comprising a master process 21, shared memory 22, and worker processes 23. The master process 21 and shared memory 22 are located in the host, while the worker processes are executed by the DPU, which communicates with the host. It is understood that the data query method provided in this embodiment can also be applied to other scenarios.
[0056] The following is combined Figure 2 The application scenarios shown are for Figure 1 The data query method shown is described below, and the specific steps of this method are as follows:
[0057] S101. Generate an expression information tree based on the expression tree.
[0058] The user's data query request includes the expression required for the current query task. The expression is the condition carried in the query task, used to filter and select data in the database that meets the requirements. The expression is generally expressed in a tree structure, called an expression tree.
[0059] The expression information tree is closely aligned with the expression tree structure. In addition to storing the expression information of the corresponding expression node in the expression tree, each node in the expression information tree also stores information on how the expression node should be calculated, i.e., the calculation information of the expression node, such as the operands of the node, the operation type of the node, and the constants that the node operation needs to be converted into the DPU calculation format, etc.
[0060] It should be noted that the expression information tree described above is similar in form to an expression tree, but the information stored in each node differs from that stored in the corresponding node in the expression tree. Any complex expression can be constructed into an expression information tree in this way with a linear (single-pass) time complexity.
[0061] After obtaining the expression tree, the leading process generates the corresponding expression information tree based on the expression tree.
[0062] S102. Copy the expression information tree to shared memory.
[0063] Based on the aforementioned expression information tree, the controlling process takes advantage of the opportunity provided by the operating system in PostgreSQL to create a shared memory region in kernel space, and copies the expression information tree into this shared memory. PostgreSQL, also known as Postgres, is an open-source relational database.
[0064] Understandably, the leading process will also copy some relevant information and data about this copy operation and the expression information tree.
[0065] S103. The cursor variable indicating the progress of database data acquisition in the shared memory is adjusted sequentially through multiple working processes, and the target data corresponding to the adjustment operation is retrieved from the database and sent to the DPU.
[0066] A cursor variable exists in memory; it is an atomic variable used to indicate the progress of data reading in the current database (such as PostgreSQL mentioned above). In a database, data exists in the form of data blocks. Reading data from the database means sequentially reading one or more data blocks. The cursor variable represents the data block currently being read. Atomic variables ensure that operations on shared variables are not interfered with by operations from other threads, thus avoiding problems such as race conditions and deadlocks.
[0067] During parallel scanning, multiple worker processes read the database sequentially. When reading data, each worker process first obtains the current indication of the cursor variable, determines the data block to be read, and advances the cursor variable to the value corresponding to the data block where the current read ends, according to the preset data volume. This allows the cursor variable to indicate the data block where the current read ends, so that the next worker process can continue reading subsequent data blocks from this data block.
[0068] Specifically, the working process runs in a Data Processing Unit (DPU), which is a hardware device or computing unit used to perform data processing tasks. In this embodiment, DPU refers to a big data computing DPU card based on the KPU (Kernel Processing Unit) architecture, which is widely used in various business scenarios to meet the requirements of various scenarios for data query acceleration and is compatible with Spark and PostgreSQL scenarios.
[0069] After the workflow adjusts the cursor variable, it retrieves the target data corresponding to the adjustment operation from the database and stores it in the DPU. The target data can be one or more data blocks, the specific number of which is determined by a preset data read volume. Furthermore, the target data refers to the data blocks between the data block indicated by the cursor variable before the adjustment and the data block indicated by the cursor variable after the adjustment.
[0070] S104. Access shared memory via DPU, perform data query on the target data according to the expression information tree, and obtain data query results.
[0071] The DPU can access the expression information tree and related information by accessing shared memory. Specifically, each worker process on the DPU can access the shared memory and map the expression information tree and related information therein into its own memory space for data calculation and processing.
[0072] After each worker process on the DPU obtains its assigned target data, it scans the target data according to the expression information tree and related information to complete the data query operation, thereby obtaining the data query results that meet the data query requirements corresponding to the expression information tree in the target data. This achieves the goal of parallel scanning of data in the database among the various worker processes.
[0073] Furthermore, each worker process returns its data query results to the host after obtaining them, and continues to retrieve the next portion of target data to be scanned, until all data in the database that needs to be scanned has been retrieved and the scan is complete. Finally, the master process collects the data query results from each worker process, aggregates them into the final query result, and sends it back to the user.
[0074] This embodiment generates an expression information tree based on an expression tree; copies the expression information tree to shared memory; sequentially adjusts a cursor variable in the shared memory indicating the progress of database data retrieval through multiple worker processes, and retrieves the target data corresponding to the adjustment operation from the database to the DPU; the DPU accesses the shared memory, performs a data query on the target data based on the expression information tree, and obtains the data query result. The cursor variable in the shared memory indicates the current progress of data retrieval in the database. Multiple worker processes alternately drive the cursor variable, eliminating the need for data exchange between worker processes, greatly simplifying the progress synchronization process between multiple worker processes, and avoiding the introduction of additional communication overhead, effectively improving the efficiency and accuracy of data query.
[0075] Figure 3 A flowchart of a data query method provided in another embodiment of this disclosure is shown below. Figure 3 As shown, the method includes the following steps:
[0076] S301. Determine the computational information required for each node in the expression tree and the access sequence number of the node.
[0077] Specifically, for each node, the computation information includes at least one or more of the following: the operands of the node, the operation type of the node, and constants of the node operation that need to be converted into DPU computation format.
[0078] In some embodiments, the access sequence number can be a DFN (Depth First Number). DFN is a specific access sequence constructed when performing depth-first traversal in tree structures (such as trees, graphs, etc.). It is widely used in graph theory (such as Tarjan's algorithm for strongly connected components in directed graphs and biconnected components in undirected graphs) and can be used to quickly locate and compare nodes in trees or graphs.
[0079] S302. Generate an expression information tree based on the expression tree, the calculation information corresponding to each node, and the access sequence number.
[0080] The leading process generates an expression information tree based on the expression tree, the computation information corresponding to each node, and the access sequence number, flattens it, and calculates the space it needs to occupy.
[0081] Figure 4 This is a schematic diagram of an expression information tree node mapping provided in an embodiment of this disclosure. For example... Figure 4 As shown, tree structure 41 is an expression tree, with each node containing a corresponding expression; tree structure 42 is an expression information tree generated from the expression tree. Each node, in addition to the expression of the corresponding node in the expression tree, also includes the calculation information required for each node and the access sequence number of that node. Only the access sequence number of each node is shown in the figure. For example, for the root node `bool`, its access sequence number is 1. Figure 4 In the expression information tree 42, it is shown as bool_1, and so on.
[0082] In some embodiments, the access sequence number of each node in the expression information tree is determined based on the traversal order of each node when traversing the expression tree. For example, as... Figure 4 As shown, the expression tree 41 is traversed in the order of preorder traversal. The index of each node in the expression information tree is the order in which that node is traversed in the preorder traversal.
[0083] S303. Based on the access sequence number of each node in the expression information tree and the data contained in the node, save the expression information tree in the form of an array to shared memory.
[0084] The leading process further flattens the expression information tree and copies it as an array to shared memory. As mentioned above, the expression information tree is a tree structure, which, after flattening, forms an array. Figure 4 As shown, in shared memory, the two-dimensional array in shared memory is the array representation of the expression information tree 42. Each element corresponds to a node in the expression information tree, including the access sequence number of that node and the data within that node.
[0085] In some embodiments, the data copied to shared memory includes three parts: the first part is the header information of this copy operation, such as the memory starting offset of this copy, the constants involved in this data query, and the relative offsets of these constants with respect to the starting offset; the second part is an array representation of the flattened expression information tree; and the third part is the constants that have been converted into DPU format.
[0086] It is understandable that the expression information tree mentioned above includes constants that need to be converted into DPU calculation format for a certain node operation. The constants that have been converted into DPU format and stored in the shared memory here are the DPU format constants after format conversion according to the expression information tree.
[0087] Multiple worker processes started by PostgreSQL successively obtain the starting offset of the shared memory mapped to their own private space. Each worker process then establishes an address mapping between the expression tree and the expression information tree in the shared memory.
[0088] S304. For each working process, read the starting block number indicated by the current cursor variable.
[0089] S305. Advance the cursor variable by a preset number of blocks so that the cursor variable points to the target block number.
[0090] S306. Read the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and obtain the target data into the working thread.
[0091] During parallel scanning, multiple worker processes sequentially read from the database. Each worker process consists of a main thread and worker threads. When reading data, the main thread first obtains the current cursor variable's indication, determines the starting block number indicated by the current cursor variable, and advances the cursor variable to the target block number corresponding to the data block where the current read ended, according to a preset number of blocks. This allows the next worker process to continue reading subsequent data blocks from the data block corresponding to the target block number.
[0092] Figure 5 This is a schematic diagram illustrating a multi-process collaboration as provided in an embodiment of this disclosure. Figure 5 As shown, process 51 moves the cursor variable from position A to position B and reads the target data between positions A and B; process 52 moves the cursor variable from position B to position C and reads the target data between positions B and C; process 53 moves the cursor variable from position C to position D and reads the target data between positions C and D, and so on, with each process alternating until the data file is completely read. EOF in the operating system indicates that there is no more data to read from the data source.
[0093] S307. For each working process, access the array-like expression information tree in shared memory according to the access sequence number to obtain information from the expression information tree.
[0094] Worker processes randomly access the array-based expression information tree in shared memory by accessing the sequence number (e.g., DFN number). Each created worker process can access the array-based expression information tree in shared memory. Since the array-based expression information tree is pre-computed and does not need to be modified, it is in a read-only state, and concurrent access by multiple processes is safe.
[0095] like Figure 4 As shown, taking any one of the worker processes as an example, the worker process randomly accesses the two-dimensional array in shared memory through the access sequence number, regardless of the access order of each element in the two-dimensional array. The worker process can obtain the relationship between each node in the expression information tree based on the access sequence number, and thus obtain all the information in the expression information tree.
[0096] S308. Perform a data query on the target data based on the information in the expression information tree to obtain the data query result.
[0097] Each worker process scans the target data according to the expression information tree and its related information to complete the data query operation, thereby obtaining the data query results in the target data that meet the data query requirements corresponding to the expression information tree. The purpose of parallel scanning of data in the database is achieved in each worker process.
[0098] Finally, the lead process collects the data query results from each worker process, aggregates them into the final query result, and feeds it back to the user.
[0099] This embodiment of the disclosure determines the computational information required for each node in the expression tree and the access sequence number of the node; generates an expression information tree based on the expression tree, the computational information corresponding to each node, and the access sequence number; saves the expression information tree in array form to shared memory based on the access sequence number of each node and the data contained in the node; for each working process, the main thread reads the starting block number indicated by the current cursor variable; advances the cursor variable by a preset number of blocks so that the cursor variable indicates the target block number; reads the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and obtains the target data to the... In the worker threads, for each worker process, an array-like expression information tree in shared memory is accessed according to the access sequence number to obtain information from the expression information tree. Based on the information in the expression information tree, the target data is queried to obtain the query results. By constructing the expression information tree and simplifying the complex tree-like expression into a linear structure that facilitates fast access and computation based on the access sequence number of each node, the expression is passed to other worker processes via shared memory. This fully utilizes the heterogeneous resources of the CPU and DPU, avoiding process and thread switching overhead while performing expression computation at a low cost. This significantly improves the throughput of table data scanning and drastically reduces the latency for users to obtain query results.
[0100] Based on the above embodiments, the method includes: performing memory alignment operations on the data in the shared memory.
[0101] Specifically, the data copied to shared memory consists of three parts. The first part is the header information for this copy operation, such as the starting offset of the memory for this copy, the constants involved in this data query, and the relative offsets of these constants relative to the starting offset. The second part is an array representation of the flattened expression information tree. The third part consists of constants converted to DPU format. These three parts are stored in a compact form on a contiguous space in shared memory. Since each data entry has a different length, zeros are padded between the data entries to align them and ensure that each data entry has the same length. In some embodiments, the length of each data entry can be padded to 2. n Bytes, where n is an integer greater than 0. It's understandable that the length of the padded data is greater than the length of the original data.
[0102] This disclosure embodiment accelerates processor addressing and computation by aligning data in shared memory, thereby further improving the efficiency of data retrieval.
[0103] Figure 6This is a schematic diagram of the structure of a data query device provided in an embodiment of this disclosure. The data query device can be a data query system as described in the above embodiments, or it can be a component or part of that data query system. The data query device provided in this embodiment can execute the processing flow provided in the data query method embodiments, such as... Figure 6 As shown, the data query device 60 includes: a generation module 61, a copy module 62, an acquisition module 63, and a query module 64; wherein, the generation module 61 is used to generate an expression information tree based on the expression tree; the copy module 62 is used to copy the expression information tree to shared memory; the acquisition module 63 is used to sequentially adjust the cursor variable in the shared memory that indicates the progress of database data acquisition through multiple worker processes, and retrieve the target data corresponding to the adjustment operation from the database to the DPU, wherein the worker processes run in the DPU; the query module 64 is used to access the shared memory through each of the worker processes, perform data query on the target data according to the expression information tree, and obtain the data query result.
[0104] Optionally, the generation module 61 includes a first determining unit 611 and a generation unit 612; the first determining unit 611 is used to determine the calculation information required for each node in the expression tree and the access sequence number of the node; the generation unit 612 is used to generate an expression information tree based on the expression tree and the calculation information and access sequence number corresponding to each node.
[0105] Optionally, the computation information includes at least one or more of the following: the operands of the node, the operation type of the node, and constants of the node operation that need to be converted into DPU computation format.
[0106] Optionally, the copy module 62 is further configured to save the expression information tree in the form of an array to shared memory based on the access sequence number of each node in the expression information tree and the data contained in the node.
[0107] Optionally, the working process includes a main thread and worker threads. The acquisition module 63 includes a reading unit 631, a pushing unit 632, and a first acquisition unit 633. The reading unit 631 is used to read the starting block number indicated by the current cursor variable through the main thread for each worker process. The pushing unit 632 is used to push the cursor variable forward a preset number of blocks so that the cursor variable indicates the target block number. The first acquisition unit 633 is used to read the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and acquire the target data into the worker thread.
[0108] Optionally, the query module 64 includes a second acquisition unit 641 and a query unit 642; the second acquisition unit 641 is used to access the array-like expression information tree in shared memory for each working process according to the access sequence number to obtain information in the expression information tree; the query unit 642 is used to perform data query on the target data according to the information in the expression information tree to obtain data query results.
[0109] Optionally, the data query device 60 further includes an alignment module 65 for performing memory alignment operations on the data in the shared memory.
[0110] Figure 6 The data query device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0111] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device can be a data query system as described in the above embodiments. The electronic device provided in this embodiment can execute the processing flow provided in the data query method embodiments, such as… Figure 7 As shown, the electronic device 70 includes: a memory 71, a processor 72, a computer program, and a communication interface 73; wherein the computer program is stored in the memory 71 and configured to be executed by the processor 72 using the data query method described above.
[0112] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the data query method described in the above embodiments.
[0113] Furthermore, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the data query method described above.
[0114] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0115] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data query method, characterized in that, The method includes: Generate an expression information tree from the expression tree; Copy the expression information tree to shared memory; Multiple worker processes sequentially adjust the cursor variable in the shared memory that indicates the progress of database data acquisition, and retrieve the target data corresponding to the adjustment operation from the database to the DPU. The worker processes run in the DPU. Each of the worker processes accesses the shared memory and performs a data query on the target data according to the expression information tree to obtain the data query result. The working process includes a main thread and worker threads. The step of sequentially adjusting the cursor variable in shared memory that indicates the progress of database data retrieval through multiple worker processes, and retrieving the target data corresponding to the adjustment operation from the database to the DPU, includes: For each worker process, the starting block number indicated by the current cursor variable is read through the main thread; Advance the cursor variable by a preset number of blocks so that the cursor variable points to the target block number; Read the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and obtain the target data into the working thread.
2. The method according to claim 1, characterized in that, The step of generating an expression information tree based on the expression tree includes: Determine the computational information required for each node in the expression tree and the access sequence number of the node; An expression information tree is generated based on the expression tree, the calculation information corresponding to each node, and the access sequence number.
3. The method according to claim 2, characterized in that, The computational information includes at least one or more of the following: The operands of the node, the operation type of the node, and the constants that the node operation needs to be converted into DPU computing format.
4. The method according to claim 2, characterized in that, The step of copying the expression information tree to shared memory includes: Based on the access sequence number of each node in the expression information tree and the data contained in the node, the expression information tree is saved in shared memory as an array.
5. The method according to claim 4, characterized in that, The process of accessing the shared memory through each of the worker processes, querying the target data according to the expression information tree, and obtaining the data query result includes: For each worker process, the expression information tree in the form of an array in shared memory is accessed according to the access sequence number to obtain the information in the expression information tree; The target data is queried based on the information in the expression information tree to obtain the data query results.
6. The method according to claim 4, characterized in that, The method includes: Perform memory alignment operations on the data in the shared memory.
7. A data query device, characterized in that, include: The generation module is used to generate an expression information tree based on the expression tree; A copy module is used to copy the expression information tree to shared memory; The acquisition module is used to sequentially adjust the cursor variable in the shared memory that indicates the progress of database data acquisition through multiple worker processes, and to acquire the target data corresponding to the adjustment operation from the database to the DPU. The worker processes run in the DPU. The query module is used to access the shared memory through each of the worker processes, perform data query on the target data according to the expression information tree, and obtain the data query results; The working process includes a main thread and a worker thread, and the acquisition module includes a reading unit, a pushing unit, and a first acquisition unit; The read unit is used to read the starting block number indicated by the current cursor variable for each worker process via the main thread; The advance unit is used to advance the cursor variable by a preset number of blocks so that the cursor variable indicates the target block number; The first acquisition unit is used to read the data of each data block between the data block corresponding to the starting block number and the data block corresponding to the target block number, and acquire the target data into the working thread.
8. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Method for realing sharing internal stored data base and internal stored data base system
CN1740978A