Data processing method and device, storage medium and computer program product
By adding target computing nodes to the execution plan of the database system for batch processing, the cumbersome computing and resource overhead problems caused by item-by-item calculation methods are solved, and more efficient computing task execution is achieved.
Patent Information
- Application Number
- CN202410033809.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
In existing database systems, as computing tasks tend to be complex and diversified, the calculation method one by one leads to cumbersome computing process and high computing resource overhead.
Add target computing nodes between the computing nodes that execute the plan, implement batch processing of tuples, and perform batch operations on multiple tuples through the target computing node.
The calculation process is simplified, the computing resource overhead is reduced, and the applicable scenarios of computing tasks are expanded, especially in high-performance scenarios such as machine learning, confidential computing and heterogeneous computing.
Smart Images

Figure CN120296045A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing, and in particular, to a data processing method, device, storage medium, and computer program product. Background Art
[0002] In a database system, an execution plan is a data structure generated based on a query statement. It describes the operation flow taken by the database system to execute a query. It is composed of multiple computing nodes composed of operators. Taking the database system based on the volcano model as an example, different computing nodes will be logically organized into a tree structure, and each computing node is responsible for performing different processing operations. Each computing node accepts the output of its child node's computing node as its own input, and then passes its own output upward to its parent node's computing node. In a database system, this data transfer is realized through tuples, which refers to a record in a database table. A computing node accepts a tuple as input at a time, performs calculations within the computing node, and generates a tuple as output.
[0003] From the above description, it can be seen that the computing nodes are calculated one by one. As the computing tasks of modern database systems tend to be more complex and diversified, such as in various databases that support machine learning, confidential computing, and heterogeneous computing with high performance, which involve a large number of tuple calculations, if the one-by-one calculation method is still used, the calculation process will be cumbersome, and it may also bring more computing resource overhead and poor computing effect. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, device, storage medium and computer program product to solve the problem that when the computing tasks of the database system are performed in a one-by-one calculation manner in the prior art, the calculation process is cumbersome and may bring more computing resource overhead.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, comprising:
[0006] Get the query statement;
[0007] Generate an execution plan for the query statement, and add a target computing node between a first computing node and a second computing node in the execution plan; wherein the first computing node is a computing node that performs a target computing operation; the target computing node is used as a parent node of the first computing node; and the second computing node is used as a parent node of the target computing node;
[0008] executing the execution plan to provide the acquired tuple as an output result to the target computing node at the first computing node;
[0009] The target computing node stores the tuples provided by the first computing node, performs batch processing on the stored multiple tuples according to the target computing operation, and provides multiple output results obtained by the batch processing to the second computing node respectively.
[0010] In a second aspect, a computing device is provided in an embodiment of the present application, comprising a storage component and a processing component; the storage component stores one or more computer program instructions, the computer program instructions are called and executed by the processing component, and the processing component executes the one or more computer program instructions to implement the data processing method as described in the first aspect.
[0011] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, wherein the computer program is executed by a computer to implement the data processing method as described in the first aspect.
[0012] In a fourth aspect, an embodiment of the present application provides a computer program product, which stores a computer program, and when the computer program is executed by a computer, implements the data processing method described in the first aspect.
[0013] In an embodiment of the present application, a query statement can be obtained, an execution plan for the query statement can be generated, and a target computing node can be added between a first computing node and a second computing node of the execution plan, wherein the first computing node is a computing node that performs a target computing operation, the target computing node is used as a parent node of the first computing node, and the second computing node is used as a parent node of the target computing node. The execution plan is then executed to provide the obtained tuple as an output result to the target computing node at the first computing node, save the tuple provided by the first computing node at the target computing node, batch process the saved multiple tuples according to the target computing operation, and provide the multiple output results obtained by the batch processing to the second computing node respectively. By adding the target computing node as the parent node of the computing node that executes the target operation in the execution plan, the tuple obtained by the first computing node can be output to the target computing node for storage, and multiple tuples can be batch processed at the target computing node, and multiple output results can be provided to the second computing node, that is, the parent node of the target computing node. Compared with the traditional solution in which the computing node accepts a tuple as input for calculation and generates a tuple as output, the one-by-one calculation method realizes batch execution, simplifies the calculation process, and can better handle computing tasks in scenarios involving a large number of tuple calculations, thereby expanding the applicable scenarios for the execution of computing tasks.
[0014] These and other aspects of the present application will become more apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 The flowchart of an embodiment of a data processing method provided by the present application is shown;
[0017] Figure 2 The conversion schematic diagram of a custom type in an actual application of the embodiment of the present application is shown;
[0018] Figure 3 The structural schematic diagram of an execution plan in an actual application of the embodiment of the present application is shown;
[0019] Figure 4 The scenario interaction schematic diagram in an actual application of the embodiment of the present application is shown;
[0020] Figure 5 The structural schematic diagram of an embodiment of a data processing device provided by the present application is shown;
[0021] Figure 6 The structural schematic diagram of an embodiment of a computing device provided by the present application is shown. Detailed implementation manners
[0022] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.
[0023] In some processes described in the specification, claims and the above drawings of the present application, multiple operations that appear in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are different types.
[0024] In a database system, an execution plan is a data structure generated based on a query statement, which describes the operation process that the database system takes to execute the query. It consists of multiple computing nodes composed of operators. Taking a database system based on the volcano model as an example, different computing nodes are logically organized into a tree structure, and each computing node is responsible for performing different processing operations. Each computing node accepts the output of its child computing nodes as its own input, and then passes its output upward to the computing nodes of its parent node. In a database system, this data transfer is achieved through tuples, which are also a record in a database table. A computing node accepts one tuple as input at a time for calculation within the computing node and generates one tuple as output.
[0025] As can be seen from the above description, computing nodes perform calculations one by one. With the increasing complexity and diversification of computing tasks in modern database systems, such as in various databases that support machine learning, confidential computing, and heterogeneous computing with high performance in task scenarios involving a large number of tuple calculations, if the one-by-one calculation method is still used, it will lead to a cumbersome calculation process, and at the same time may bring more computing resource overhead and poor calculation results.
[0026] To solve the above technical problems, the inventor proposed the technical solution of this application. In the embodiment of this application, a query statement is obtained; an execution plan for the query statement is generated, and a target computing node is added between the first computing node and the second computing node in the execution plan; wherein, the first computing node is a computing node that executes a target operation; the target computing node is used as the parent node of the first computing node; the second computing node is used as the parent node of the target computing node; the execution plan is executed to provide the tuple obtained by the first computing node as an output result to the target computing node; the tuple provided by the first computing node is saved in the target computing node, and the multiple saved tuples are batch-processed according to the target operation, and the multiple output results obtained from the batch processing are respectively provided to the second computing node.
[0027] In the embodiment of this application, by adding a target computing node as the parent node above the computing node that executes the target operation in the execution plan, the tuple obtained by the first computing node can be output to the target computing node for saving, and multiple tuples are batch-processed in the target computing node, and multiple output results are provided to the second computing node, that is, the parent node of the target computing node. Compared with the traditional method in which a computing node accepts one tuple as input for calculation and generates one tuple as output for one-by-one calculation, batch execution is realized, the operation process is simplified, and it can better handle the computing tasks in scenarios involving a large number of tuple calculations, expanding the applicable scenarios for executing computing tasks.
[0028] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0029] Figure 1 The flowchart of an embodiment of a data processing method provided by the present application is shown. The technical solution of this embodiment can be executed by a database engine. In practical applications, the technical solution of the embodiment of the present application can be applied to a database processing system composed of a user end and a server end. The database engine can be deployed in the server end. The method can include the following steps:
[0030] 101: Obtain a query statement.
[0031] In this embodiment, the data processing method can be applied to the server end. Optionally, the query statement can be sent by the user end, and the server end can receive the query statement sent by the user end.
[0032] 102: Generate an execution plan for the query statement, and add a target computing node between the first computing node and the second computing node of the execution plan.
[0033] Among them, the first computing node is the computing node that executes the target operation, and the second computing node is the original parent node of the first computing node. After adding the target computing node, the target computing node is used as the parent node of the first computing node, and the second computing node is used as the parent node of the target computing node.
[0034] The execution plan can be implemented as various data structures such as a tree structure and a graph structure. Taking a database system based on the volcano model as an example, an execution plan tree for the query statement can be generated. The execution plan tree can include multiple computing nodes. For the convenience of description, the computing node that executes the target operation can be called the first computing node, and the original parent node of the first computing node can be called the second computing node. Among them, the target operation can include, for example, numerical operation, comparison operation, logical operation, etc.
[0035] Furthermore, a target computing node can be added between the first computing node and the second computing node. At this time, the added target computing node can be used as the parent node of the first computing node, and the second computing node can be used as the parent node of the target computing node.
[0036] Optionally, there may be multiple first computing nodes in the above execution plan. For each first computing node, a corresponding target computing node may be added above the first computing node as the parent node of the first computing node.
[0037] 103: Execute the execution plan to provide the obtained tuples as output results to the target computing node at the first computing node.
[0038] Execute the above execution plan, use the first computing node to obtain tuples, and provide the obtained tuples as output results to the target computing node. That is to say, different from the traditional solution where the first computing node needs to perform target operation on the data in the tuples, in the solution of the embodiment of the present application, the target operation is not actually performed in the first computing node, and all the obtained tuples will pass the verification and be provided to its upper layer node, that is, only the obtained tuples are provided to its upper layer node.
[0039] 104: Save the tuples provided by the first computing node at the target computing node, batch process the saved multiple tuples according to the target operation, and provide the multiple output results obtained from the batch processing to the second computing node respectively.
[0040] The target computing node may first save the tuples provided by the first computing node.
[0041] And batch process the saved multiple tuples according to the target operation, and provide the multiple output results obtained to the second computing node.
[0042] Optionally, the target computing node may process the saved multiple tuples according to the target operation when the tuple save capacity meets the processing requirements. Among them, the processing requirements may be set according to the actual application scenario. For example, it may be set that when the tuple save capacity meets the preset quantity, the saved multiple tuples are batch processed, or when the tuple save capacity reaches the storage threshold, the saved multiple tuples are batch processed, or when it is determined that the first computing node has provided all the obtained tuples and no longer provides tuples, the saved multiple tuples are batch processed, and so on.
[0043] In a practical application, the technical solution of the embodiment of the present application can be applied to scenarios where it is necessary to call external hardware to perform target computing operations. For example, in order to achieve computing scenarios such as encryption processing or acceleration processing, calling external hardware can be, for example, calling a graphics processing unit (GPU) for computing, calling an external field programmable gate array (FPGA) for computing, etc. Therefore, optionally, in the target computing node, external hardware can be called to batch process the data objects in the multiple tuples stored according to the target computing operation.
[0044] Specifically, the target computing node batch processes the stored multiple tuples according to the target computing operation, and provides the obtained multiple output results to the second computing node, which will be described in the subsequent embodiments.
[0045] According to the above output results, the query result corresponding to the execution plan can be obtained. Optionally, the server can also feed back the query result corresponding to the execution plan to the user.
[0046] In this embodiment, a query statement can be obtained, an execution plan for the query statement can be generated, and a target computing node can be added between a first computing node and a second computing node of the execution plan, wherein the first computing node is a computing node that performs a target computing operation, the target computing node is used as a parent node of the first computing node, and the second computing node is used as a parent node of the target computing node. The execution plan is then executed to provide the obtained tuple as an output result to the target computing node at the first computing node, save the tuple provided by the first computing node at the target computing node, batch process the saved multiple tuples according to the target computing operation, and provide the multiple output results obtained by the batch processing to the second computing node respectively. By adding the target computing node as the parent node of the computing node that executes the target operation in the execution plan, the tuple obtained by the first computing node can be output to the target computing node for storage, and multiple tuples can be batch processed at the target computing node, and multiple output results can be provided to the second computing node, that is, the parent node of the target computing node. Compared with the traditional solution in which the computing node accepts a tuple as input for calculation and generates a tuple as output, the one-by-one calculation method realizes batch execution, simplifies the calculation process, and can better handle computing tasks in scenarios involving a large number of tuple calculations, thereby expanding the applicable scenarios for the execution of computing tasks.
[0047] In practical applications, an execution plan may include multiple computing nodes. During the execution of a computing task, not every computing node will perform the target operation. Therefore, the computing nodes that perform the target operation, i.e., the first computing nodes, can be determined therefrom, and then the target computing nodes can be added on top of them. Therefore, in some embodiments, generating the execution plan of the query statement and adding a target computing node between the first computing node and the second computing node in the execution plan may include: generating the execution plan of the query statement, and when the query statement involves a target operation that meets a preset condition, adding a target computing node between the first computing node and the second computing node in the execution plan.
[0048] Specifically, the preset condition may include one or more of the following implementation manners:
[0049] An operation on a specified data type;
[0050] An operation that meets a predetermined operation requirement;
[0051] An operation that calls external hardware;
[0052] And, an operation for generating data of a custom type according to a pre-configured custom type and the operation requirements of the custom type.
[0053] Among them, the specified data type may include, for example, a numeric type, a string type, etc. Taking the specified data type as a numeric type as an example, at least one computing node for performing an operation on the numeric type can be determined from multiple computing nodes in the execution plan as the first computing node, and a corresponding target computing node can be added on top of it as the parent node.
[0054] The predetermined operation requirement may include, for example, a comparison operation, a logical operation, etc. Taking the predetermined operation requirement as a comparison operation as an example, at least one computing node for performing a comparison operation can be determined from multiple computing nodes in the execution plan as the first computing node, and a corresponding target computing node can be added on top of it as the parent node.
[0055] An operation that calls external hardware may include, for example, calling a GPU for operation, calling an FPGA for operation, etc. At least one computing node that calls external hardware can be determined from multiple computing nodes in the execution plan as the first computing node, and a corresponding target computing node can be added on top of it as the parent node.
[0056] Custom types can be distinguished from primitive data types such as numeric types and string types and can be configured according to actual application scenarios. The operation requirements for custom types can include generating custom type data from operation operations on primitive data types and generating custom type data from operation operations on custom types. Thus, from multiple computing nodes of an execution plan, at least one computing node for performing the operation operation that generates custom type data according to the pre-configured custom type and the operation requirements of the custom type can be determined as the first computing node, and a corresponding target computing node can be added above it as the parent node.
[0057] By setting the above preset conditions, the determination of the first computing node in the execution plan can be achieved, facilitating the addition of a target computing node above the first computing node. The target computing node performs batch operations on multiple tuples provided by the corresponding first computing node to meet the batch processing requirements.
[0058] The following uses the example that the target operation operation includes generating custom type data according to the pre-configured custom type and the operation requirements of the custom type to illustrate this data processing method.
[0059] Each primitive data type can define its corresponding custom type. For example, the primitive numeric type can correspond to a custom numeric type, the primitive string type can correspond to a custom string type, the primitive boolean type can correspond to a custom boolean type, and so on.
[0060] The operation requirements for custom types can include generating custom type data from operation operations on primitive data types and generating custom type data from operation operations on custom types. For example, it can be pre-configured to generate custom numeric type data from numeric operation operations on the primitive numeric type, generate custom boolean type data from comparison operations on the custom boolean type, and so on.
[0061] In some embodiments, in the case of involving the above target operation operation, the operation result of the first computing node can be set to verified. For example, if the target operation operation includes generating custom boolean type data from comparison operations on primitive data types and also generating custom boolean type data from comparison operations on custom boolean type data, the operation result of the first computing node can be set to "true". That is to say, for any comparison operation on primitive data type data and any comparison operation on custom data type data, the obtained comparison result is "true". Thus, the first computing node can output all the data participating in the target operation operation and provide it to the target computing node.
[0062] Moreover, in the target computing node, the data in the saved tuples that participates in the target operation can be batch-processed according to the target operation, and multiple operation results of a custom type can be generated. For example, the target operation includes generating custom boolean type data for a comparison operation on the original data type, and also generating custom boolean type data for a comparison operation on the custom boolean type data. In the target computing node, a comparison operation can be performed on the data participating in the target operation to generate operation results of the custom boolean type.
[0063] Also, considering that in practical applications, the output results provided by the target computing node to the second computing node are usually of the original type, therefore, in the target computing node, the operation results of the custom type can also be converted into the corresponding original data type.
[0064] Therefore, the method for the first computing node to provide the obtained tuple as an output result to the target computing node may include:
[0065] In the first computing node, when the target operation is performed on the obtained tuple, if it is determined that the operation result passes the verification, the data participating in the target operation is saved as a data object in the tuple, and the tuple is provided as an output result to the target computing node.
[0066] Furthermore, the method for batch-processing the saved multiple tuples according to the target operation may include:
[0067] In the target computing node, the data objects in the saved multiple tuples are batch-processed according to the target operation to generate multiple operation results of a custom type;
[0068] The multiple operation results are respectively converted into the corresponding original data types;
[0069] According to the multiple operation results, multiple output results to be provided to the second computing node are determined.
[0070] For example, after the multiple operation results generated by the target computing node are converted into the corresponding original data types, including multiple "true" operation results and multiple "false" operation results, then multiple output results that meet the query request to be provided to the second computing node can be determined. For example, the tuples corresponding to the multiple "true" operation results are determined as the multiple output results to be provided to the second computing node.
[0071] In this embodiment, when the target operation is performed on the obtained tuple in the first computing node and the operation result is determined to pass the verification, it is realized that the first computing node no longer actually performs the target operation, but saves the data participating in the target operation into the tuple and provides it to the target computing node for batch processing in the target computing node, thus realizing the batch processing of the tuple.
[0072] To achieve batch processing, a target computing node is added between the first computing node and the second computing node. The first computing node directly provides the obtained tuple as the output result to the target computing node, and the target computing node performs batch processing on multiple tuples according to the target operation. It can be understood that in this process, the target operation on the data in the tuple is delayed and executed in batch in the target computing node, while the first computing node no longer performs the target operation. To avoid modifying the structure of the first computing node in the execution plan, the modification method of the query statement can be adopted. Therefore, in some embodiments, the method of generating an execution plan for the query statement and adding a target computing node between the first computing node and the second computing node in the execution plan may include:
[0073] Detect that the query statement involves a target operation;
[0074] Rewrite the query statement to apply a target execution function outside the operation expression corresponding to the target operation;
[0075] Parse the query statement to generate an execution plan, and when the query statement includes the target execution function, add a target computing node between the first computing node and the second computing node in the execution plan.
[0076] Specifically, when it is detected that the query statement involves a target operation, such as when the query statement includes an operation expression corresponding to the target operation, the query statement can be rewritten to apply a target execution function outside the operation expression corresponding to the target operation.
[0077] Taking the query statement "SELECT * FROM t WHERE a > encrypt(3)" as an example, this query statement can be implemented for a fully encrypted database (a database solution that uses ciphertext processing technology to achieve data security protection, which can rely on encryption technologies such as Trusted Execution Environment (TEE) and Fully Homomorphic Encryption (FHE)). This query statement can mean retrieving all tuples from table t where the data content in column a is greater than the encrypted value of 3, that is, comparing the data in column a of the tuples in table t with the encrypted data "encrypt(3)". The target operation involved is a comparison operation, and the operation expression corresponding to this target operation is a > encrypt(3). Then the target execution function can be represented as, for example, the force function. This target execution function can be called by the computing node in the execution plan to delay the target operation on the data in the tuple from the first computing node to the target computing node for execution, thus achieving batch processing.
[0078] After that, the rewritten query statement can be parsed to generate an execution plan. The rewritten query statement can be implemented as "SELECT * FROM t WHERE force(a > encrypt(3))". Also, when the query statement includes the target execution function, a target computing node is added between the first computing node and the second computing node in the execution plan.
[0079] By wrapping the operation expression corresponding to the target operation in the query statement with the target execution function to rewrite the query statement, generating an execution plan using the rewritten query statement, and adding a target computing node above the first computing node, it is possible to, without modifying the structure of the first computing node, directly provide the retrieved tuples as the output result to the target computing node, that is, delay the target operation on the data in the tuple to the target computing node for batch execution, thereby achieving batch processing.
[0080] On this basis, still taking the target operation as an operation that generates custom - type data according to a pre - configured custom type and the operation requirements of the custom type as an example, this data processing method will be described.
[0081] In some embodiments, the above-mentioned target execution function can be used to set the operation result of the first computing node as verified when it comes to the above-mentioned target operation. For example, the target operation includes generating custom Boolean type data for comparison operations on the original data type, and also generating custom Boolean type data for comparison operations on the custom Boolean type data. The operation result of the first computing node can be set to "true", that is, for comparison operations on any original data type data and for comparison operations on any custom data type data, the obtained comparison result is "true". Thus, the first computing node can output all the data participating in the target operation and provide it to the target computing node.
[0082] Moreover, the target execution function can also be used to batch process the data participating in the target operation in the saved tuples according to the target operation in the target computing node and generate multiple operation results of the custom type. For example, the target operation includes generating custom Boolean type data for comparison operations on the original data type, and also generating custom Boolean type data for comparison operations on the custom Boolean type data. In the target computing node, comparison operations can be performed on the data participating in the target operation to generate operation results of the custom Boolean type.
[0083] And, considering that in practical applications, the output result provided by the target computing node to the second computing node is usually of the original type data, therefore, the target execution function can also be used to convert the operation results of the custom type into the corresponding original data type in the target computing node.
[0084] Therefore, the method for the first computing node to provide the obtained tuple as the output result to the target computing node can include:
[0085] In the first computing node, call the target execution function to determine that the operation result is verified when performing the target operation on the obtained tuple, save the data participating in the target operation as a data object into the tuple, and provide the tuple as the output result to the target computing node.
[0086] Furthermore, the method for batch processing the saved multiple tuples according to the target operation can include:
[0087] In the target computing node, call the target execution function to batch process the data objects in the saved multiple tuples according to the target operation to generate multiple operation results of the custom type;
[0088] Convert the multiple operation results into the corresponding original data types respectively;
[0089] Determine multiple output results to be provided to a second computing node based on multiple operation results.
[0090] For example, after the multiple operation results generated by the target computing node are converted into corresponding original data types, including multiple "true" operation results and multiple "false" operation results, multiple output results that meet the query request can be determined to be provided to the second computing node. For example, determine the tuples corresponding to the multiple "true" operation results as the multiple output results provided to the second computing node.
[0091] In some embodiments, after determining the multiple output results, the above method may further include:
[0092] Store the multiple output results in the processing order for the second computing node to sequentially obtain the multiple output results.
[0093] Specifically, the method of storing the multiple output results in the processing order may include:
[0094] Put the multiple output results into the storage queue in sequence according to the order in which the corresponding tuples are provided to the target computing node.
[0095] By storing the output results in the form of a queue, the sequential acquisition and ejection of tuples by the target computing node are realized, and the calculation correctness can be ensured without modifying the original execution logic of the database.
[0096] Optionally, a circular buffer can be used to simultaneously ensure the relative order between output results and reduce data movement. Alternatively, a general data structure such as a linked list can also be used to maintain the queue, and the present application does not limit this.
[0097] The technical solution of the embodiment of the present application can be applied to the computing scenario for a fully secret database in a practical application. The fully secret database is a database solution that uses ciphertext processing technology to achieve data security protection, which can be implemented by relying on encryption technologies such as Trusted Execution Environment (TEE) and Fully Homomophic Encryption (FHE). It is necessary to call external hardware to perform data processing in a trusted execution environment. In a fully secret database system, the data content exists in the form of ciphertext, and the ciphertext processing occurs immediately and is calculated immediately, and the ciphertext is transmitted in the database as an intermediate form. Taking a fully secret database based on TEE as an example, the tuple contains secret data. In the traditional way, when performing calculations, it is usually firstly used to implement the secure call of the computing instance running in the TEE using untrusted conventional software, and then the secret data is decrypted in the computing instance, and the calculation is performed in plain text, and finally the calculation result is output in the form of ciphertext. This way of implementing calculations one by one requires secure calls, encryption and decryption of the computing instance in the TEE for each calculation, and the TEE is provided by external hardware, and calling external hardware is likely to bring about a large computing resource overhead. Therefore, the batch processing calculation solution proposed in this application will significantly optimize the above-mentioned fully encrypted database system.
[0098] In a fully encrypted database, since the data content exists in ciphertext form, it is usually necessary to define the ciphertext type and the operation type corresponding to the ciphertext type. Taking the numeric type as an example, in a plaintext database, a plaintext numeric type can be defined, such as the int4 type, which can represent a 4-byte plaintext numeric type. Correspondingly, in a fully encrypted database, a ciphertext numeric type can be defined, such as the enc_int4 type, which can represent an encrypted 4-byte numeric type.
[0099] When processing plain text, digital operations may include addition, subtraction, multiplication, and division, and comparison operations may include greater, lesser, and equal operations. Digital operations on plain text digital type data will generate digital type data, while comparison operations will generate Boolean type data. Correspondingly, when processing cipher text, digital operations on cipher text digital type data will also generate cipher text digital type data, while comparison operations can generate plain text Boolean type data or cipher text Boolean type data, which can be set according to the actual application scenario.
[0100] Corresponding to the above-mentioned original numeric types, the custom types may include custom numeric types, such as delay_int4, delay_enc_int4, etc., and custom Boolean types, such as delay_bool.
[0101] The operation requirements of the custom type may include generating custom digital type data for digital operation operations on the original digital type, generating custom boolean type data for comparison operations on the original digital type, generating custom digital type data for digital operation operations on the custom digital type, and generating custom boolean type data for comparison operations on the custom boolean type, etc. For details, please refer to Figure 2 the schematic diagram of an embodiment of a custom type and its operations shown in
[0102] Taking the data participating in the target operation as digital type data and the target operation including digital operation operations and / or comparison operations for generating custom type data according to the pre-configured custom type and the operation requirements of the custom type as an example, in the target computing node, the target execution function is called to batch process the data objects in the saved multiple tuples according to the target operation to generate multiple operation results of the custom type, which may include:
[0103] In the target computing node, when the target operation includes a digital operation operation for generating custom type data according to the pre-configured custom type and the operation requirements of the custom type, the data objects in the saved multiple tuples are batch processed according to the digital operation operation to generate the operation results of the custom digital type;
[0104] In the case that the target operation includes a comparison operation for generating custom type data according to the pre-configured custom type and the operation requirements of the custom type, the data objects in the saved multiple tuples are batch processed according to the comparison operation to generate the operation results of the custom boolean type;
[0105] In the case that the target operation includes a digital operation operation and a comparison operation for generating custom type data according to the pre-configured custom type and the operation requirements of the custom type, the data objects in the saved multiple tuples are batch processed by sequentially executing the digital operation operation and the comparison operation according to the operation sequence. In the case that the last operation operation is a digital operation operation, the operation results of the custom digital type are generated, and in the case that the last operation operation is a comparison operation, the operation results of the custom boolean type are generated.
[0106] Furthermore, converting the multiple operation results into the corresponding original data types respectively may include:
[0107] In the case where the operation result includes the operation result of a custom digital type, convert the operation result of the custom digital type into the operation result of the original digital type; and, in the case where the operation result includes the operation result of a custom Boolean type, convert the operation result of the custom Boolean type into the operation result of the original Boolean type.
[0108] The following takes the data participating in the target operation operation including ciphertext digital type data, and the target operation operation including generating a comparison operation of custom type data according to a pre-configured custom type and the operation requirements of the custom type as an example to illustrate this data processing method.
[0109] For example, still taking the query statement as SELECT * FROM t WHERE a > encrypt(3) as an example, this query statement means to obtain all tuples from table t where the data content in column a is greater than encrypted 3. Taking the general volcano model of relational databases as an example, the execution plan generated according to this query statement can include two computing nodes, namely the SCAN node and the Output node. Among them, the SCAN node can obtain a tuple from the storage medium each time it executes, and complete the judgment on whether the data content in column a of the tuple is greater than encrypted 3. The Output node, as the parent node of the SCAN node, will continuously call the SCAN node to obtain the tuples output by the SCAN node until the SCAN node no longer outputs tuples.
[0110] In the embodiment of this application, it is detected that the above query statement involves a target operation operation. The operation expression corresponding to this target operation operation is a > encrypt(3). The computing node for executing this target operation operation is the SCAN node. Therefore, the SCAN node can be used as the first computing node, and the Output node can be used as the second computing node. Rewrite this query statement, enclose the operation expression with the force function, and the rewritten query statement is implemented as SELECT * FROM t WHERE force(a > encrypt(3)). Parse the rewritten query statement and generate an execution plan, and add a target computing node between the SCAN node and the Output node, which can be represented by the BUFFER node. For the sake of understanding, Figure 3 shows a schematic diagram of an embodiment of an execution plan.
[0111] As Figure 3As shown, the SCAN node can obtain a tuple from the storage medium each time it executes, and call the force function. When performing the judgment and verification of the WHERE statement, that is, when performing the target operation of a > encrypt(3) on the obtained tuple, the operation result of each tuple is set to "true", that is, it is determined that the data content in column a of each tuple is greater than the encrypted value of 3, and all the operation results pass the verification. Then, the data participating in this target operation is saved as a data object in the tuple, and the tuple is provided as the output result to the BUFFER node.
[0112] As the parent node of the SCAN node, the BUFFER node can save multiple tuples provided by the SCAN node. When the tuple storage capacity meets the processing requirements, the force function is called to batch process the data objects in the multiple saved tuples according to the target operation of a > encrypt(3), generating multiple operation results of a custom boolean type. Specifically, the decryption component can be called to decrypt the data object in the tuple, process the decrypted data object, and then encrypt the processed operation data to obtain the operation result. Among them, the decryption component can be implemented as internal hardware or external hardware, and the present application does not limit this. As Figure 3 shown, when the BUFFER node batch processes the data objects in the N saved tuples, N operation results of a custom boolean type can be generated, where N is an integer greater than 1. Then, the BUFFER node can also convert the N operation results into the original boolean type respectively.
[0113] After generating multiple operation results, the BUFFER node can also determine the multiple output results to be provided to the Output node, that is, the second computing node. Specifically, the tuples with the operation result of "true" are determined as the output results provided to the Output node. Then, the multiple output results can be sequentially placed in the storage queue according to the order of the corresponding tuples provided to this BUFFER node.
[0114] As the parent node of the BUFFER node, the Output node can call the BUFFER node to obtain tuples from the storage queue.
[0115] In the embodiments of the present application, by adding a target computing node to the execution plan, the operation order of the traditional execution plan is changed. When the target computing node needs to return tuples to the second computing node, that is, its parent node, it is necessary to continuously obtain the first computing node, that is, its child node, save the obtained multiple tuples, and then perform batch processing. Thus, the actual target operation on the data in the tuples is transferred from the first computing node to the target computing node for batch execution, realizing local caching and batch computing of tuples, and can generally expand various existing databases to support batch execution task scenarios such as machine learning, confidential computing, and heterogeneous computing with high performance. At the same time, since multiple tuples are saved in the order provided to the target computing node, the execution order of the data in the tuples remains unchanged. Moreover, the solution of this embodiment belongs to a plug-in expansion of the original computing solution of the database, without modifying the original nodes and computing engines of the database, and is easy to iterate and upgrade with the product. When the database performance is improved, it will not affect its stability.
[0116] As can be seen from the above description, the query statement can be sent by the user through the user terminal, such as Figure 4 shown, which shows the schematic diagram of scenario interaction in an actual application of the embodiments of the present application.
[0117] The user can send a query request to the server 402 through the user terminal 401, and the query request may include a query statement.
[0118] Among them, the user terminal can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a small program, a lightweight application program) or a cloud application, etc. The user terminal can be deployed in an electronic device and needs to rely on the device or some apps in the device to run, etc.
[0119] The server can include servers that provide various services, such as a server that processes the query request sent by the user terminal. It should be noted that the server can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0120] A database engine is deployed in the server, which can obtain and parse the query statement, generate an execution plan for the query statement, and add a target computing node between the first computing node and the second computing node of the execution plan. After that, the execution plan is executed to provide the obtained tuple as an output result to the target computing node at the first computing node, and save the tuple provided by the first computing node at the target computing node, and batch process the saved multiple tuples according to the target operation, and provide the multiple output results obtained by the batch processing to the second computing node respectively. The second computing node can continue to process the multiple output results and provide the output results obtained by the processing to its parent node, etc. After the execution plan is executed, the query result corresponding to the query statement can be obtained. After that, the server 402 can feed back the query result corresponding to the execution plan to the user end 401.
[0121] By adding the target computing node as the parent node of the computing node that executes the target operation in the execution plan, the tuple obtained by the first computing node can be output to the target computing node for storage, and when the tuple storage capacity of the target operation node meets the processing requirements, multiple tuples can be batch processed, and multiple output results can be provided to the second computing node, that is, the parent node of the target computing node. Compared with the traditional solution in which the computing node accepts a tuple as input for calculation and generates a tuple as output, the one-by-one calculation method realizes batch execution, simplifies the calculation process, and can better handle computing tasks in scenarios involving a large number of tuple calculations, thereby expanding the applicable scenarios for the execution of computing tasks.
[0122] Figure 5 The structure diagram of an embodiment of a data processing device provided by the present application is shown, and the device may include the following modules:
[0123] The acquisition module 501 can be used to acquire a query statement;
[0124] The generation module 502 may be used to generate an execution plan for the query statement, and add a target computing node between a first computing node and a second computing node of the execution plan; wherein the first computing node is a computing node that performs a target computing operation; the target computing node is used as a parent node of the first computing node; and the second computing node is used as a parent node of the target computing node;
[0125] The execution module 503 can be used to execute the execution plan to provide the obtained tuples as output results to the target computing node at the first computing node; and save the tuples provided by the first computing node at the target computing node, batch process the saved multiple tuples according to the target operation, and provide the multiple output results obtained from the batch processing to the second computing node respectively.
[0126] In some embodiments, the generation module 502 can be specifically used to generate an execution plan for the query statement, and add a target computing node between the first computing node and the second computing node in the execution plan when the query statement involves a target operation that meets preset conditions.
[0127] In some embodiments, the generation module 502 can be specifically used to detect that the query statement involves a target operation; rewrite the query statement to enclose a target execution function outside the operation expression corresponding to the target operation; parse the query statement to generate an execution plan, and add a target computing node between the first computing node and the second computing node in the execution plan when the query statement includes the target execution function.
[0128] In some embodiments, the target operation includes an operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type;
[0129] The execution module 503 can be specifically used to, in the first computing node, when the target operation is executed for the obtained tuples, determine that the operation result is verified, save the data participating in the target operation as a data object to the tuples, and provide the tuples as output results to the target computing node; in the target computing node, batch process the data objects in the saved multiple tuples according to the target operation to generate multiple operation results of the custom type; convert the multiple operation results into corresponding original data types respectively; and determine multiple output results for providing to the second computing node according to the multiple operation results.
[0130] In some embodiments, the device may further include a storage module for storing the multiple output results in the processing order for the second computing node to obtain the multiple output results in sequence.
[0131] In some embodiments, the storage module can be specifically used to sequentially put the multiple output results into a storage queue in the order in which the corresponding tuples are provided to the target computing node.
[0132] In some embodiments, the data participating in the target operation includes original digital type data, and the target operation includes a digital operation and / or a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type;
[0133] The execution module 503 may specifically be configured to, in the target computing node, when the target operation includes a digital operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the saved multiple tuples according to the digital operation to generate an operation result of the custom digital type; when the target operation includes a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the saved multiple tuples according to the comparison operation to generate an operation result of the custom boolean type; when the target operation includes a digital operation and a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the saved multiple tuples by sequentially performing the digital operation and the comparison operation according to the operation sequence. When the last operation is a digital operation, an operation result of the custom digital type is generated, and when the last operation is a comparison operation, an operation result of the custom boolean type is generated; and when the operation result includes an operation result of the custom digital type, convert the operation result of the custom digital type into an operation result of the original digital type; when the operation result includes an operation result of the custom boolean type, convert the operation result of the custom boolean type into an operation result of the original boolean type.
[0134] In some embodiments, the database corresponding to the query statement is a ciphertext database, and the data participating in the target operation includes ciphertext data type data;
[0135] The execution module 503 may specifically be configured to call external hardware to decrypt the data objects in the saved multiple tuples, batch process the decrypted data objects according to the target operation to generate multiple operation data, and encrypt the multiple operation data to obtain multiple operation results of the custom type.
[0136] In some embodiments, the acquisition module 501 may specifically be configured to receive a query statement sent by the client;
[0137] The apparatus may further include:
[0138] A sending module, which can be used to feedback the query result corresponding to the execution plan to the client.
[0139] Figure 6 FIG. 4 shows a schematic structural diagram of an embodiment of a computing device provided by the present application. The device may include a storage component 601 and a processing component 602;
[0140] The storage component 601 stores one or more computer program instructions, and one or more computer program instructions are called and executed by the processing component 602 to implement Figure 1 The data processing method shown.
[0141] In practical applications, the computing device may be implemented as a server in the system architecture shown in Figure 4 FIG. 14.
[0142] The processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0143] The storage component is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0144] Of course, the above computing device may also include other components, such as input / output interfaces, communication components, etc.
[0145] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.
[0146] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by the computer, it can implement Figure 1 The data processing method shown. The computer-readable medium may be included in the computing device described in the above embodiment; or it may exist separately and not be assembled into the computing device.
[0147] A computer-readable storage medium may be, for example, but not limited to, a system, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above, and the like.
[0148] The embodiments of the present application also provide a computer program product storing a computer program, and when the computer program is executed by a computer, it can implement Figure 1 the data processing method shown.
[0149] In such an embodiment, the computer program can be downloaded and installed from a network and / or installed from a removable medium. When the computer program is executed by a processor, it executes various functions defined in the system of the present application.
[0150] It should be noted that the above computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can also be implemented as a single server or a single terminal device.
[0151] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain a query statement; Generate an execution plan for the query statement, and add a target computing node between a first computing node and a second computing node in the execution plan; wherein, the first computing node is a computing node that executes a target operation; the target computing node is used as the parent node of the first computing node; the second computing node is used as the parent node of the target computing node; Execute the execution plan to provide the obtained tuples as output results from the first computing node to the target computing node; Save the tuples provided by the first computing node at the target computing node, batch process the saved multiple tuples according to the target operation, and provide the multiple output results obtained from the batch process to the second computing node respectively.
2. The method according to claim 1, characterized in that The generating the execution plan for the query statement and adding a target computing node between the first computing node and the second computing node in the execution plan includes: Generate the execution plan for the query statement, and add a target computing node between the first computing node and the second computing node in the execution plan when the query statement involves a target operation that meets a preset condition.
3. The method according to claim 2, wherein The preset condition includes one or more of the following implementation manners: An operation for a specified data type; An operation that meets a predetermined operation requirement; An operation that calls external hardware; And, an operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type.
4. The method according to claim 1, wherein The generating the execution plan for the query statement and adding a target computing node between the first computing node and the second computing node in the execution plan includes: Detect that the query statement involves a target operation; Rewrite the query statement to enclose a target execution function outside the operation expression corresponding to the target operation; Parse the query statement to generate an execution plan, and add a target computing node between the first computing node and the second computing node in the execution plan when the query statement includes the target execution function.
5. The method according to claim 1, wherein The target operation includes an operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type; The providing the obtained tuples as output results from the first computing node to the target computing node includes: In the first computing node, when the target operation is executed for the obtained tuples, determine that the operation result verification passes, save the data participating in the target operation as a data object in the tuples, and provide the tuples as output results to the target computing node; The batch processing the saved multiple tuples according to the target operation includes: In the target computing node, batch process the data objects in the saved multiple tuples according to the target operation to generate multiple operation results of the custom type; Convert the multiple operation results into corresponding original data types respectively; Determine a plurality of output results to be provided to the second computing node according to the plurality of operation results.
6. The method according to claim 1, wherein The method further includes: Store the plurality of output results in the processing order for the second computing node to sequentially obtain the plurality of output results.
7. The method according to claim 6, characterized in that, The storing the plurality of output results in the processing order includes: Put the plurality of output results into a storage queue in the order corresponding to the tuples provided to the target computing node; the storage queue is a circular buffer or a linked list structure.
8. The method according to claim 1, wherein The batch processing of the data objects in the plurality of saved tuples according to the target operation includes: In the target computing node, call external hardware to batch process the data objects in the plurality of saved tuples according to the target operation.
9. The method according to claim 5, characterized in that, The data participating in the target operation includes original digital type data, and the target operation includes a digital operation and / or a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type; The batch processing of the data objects in the plurality of saved tuples according to the target operation in the target computing node to generate a plurality of operation results of the custom type includes: In the target computing node, when the target operation includes a digital operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the plurality of saved tuples according to the digital operation to generate operation results of the custom digital type; When the target operation includes a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the plurality of saved tuples according to the comparison operation to generate operation results of the custom boolean type; When the target operation includes a digital operation and a comparison operation for generating custom type data according to a pre-configured custom type and the operation requirements of the custom type, batch process the data objects in the plurality of saved tuples by sequentially executing the digital operation and the comparison operation according to the operation order. When the last operation is a digital operation, generate operation results of the custom digital type, and when the last operation is a comparison operation, generate operation results of the custom boolean type; The converting the plurality of operation results into corresponding original data types respectively includes: When the operation results include operation results of the custom digital type, convert the operation results of the custom digital type into operation results of the original digital type; When the operation results include operation results of the custom boolean type, convert the operation results of the custom boolean type into operation results of the original boolean type.
10. The method according to claim 5, wherein The database corresponding to the query statement is a fully homomorphic encryption database, and the data participating in the target operation includes ciphertext data type data; Batch-processing the data objects in the saved multiple tuples according to the target operation to generate multiple calculation results of the custom type includes: Invoking a decryption component to decrypt the data objects in the saved multiple tuples, batch-processing the decrypted data objects according to the target operation to generate multiple calculation data, and encrypting the multiple calculation data to obtain multiple calculation results of the custom type.
11. The method according to claim 1, wherein The obtaining of the query statement includes: Receiving a query statement sent by a client; The method further includes: Feeding back the query result corresponding to the execution plan to the client.
12. A computing device, characterized in that, Including a storage component and a processing component; the storage component stores one or more computer program instructions, the computer program instructions are called and executed by the processing component, and the processing component executes the one or more computer program instructions to implement the data processing method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, Storing a computer program, the computer program is executed by a computer to implement the data processing method according to any one of claims 1 to 11.
14. A computer program product, characterized in that, Storing a computer program, when the computer program is executed by a computer, it implements the data processing method according to any one of claims 1 to 11.