Data table processing method and device, equipment and computer program product
By extending the attribute item and aligning data columns in a distributed environment, the problem of inefficient processing of multi-partition data tables is solved, and efficient and secure data processing is achieved.
Patent Information
- Application Number
- CN202410027884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In a distributed computing environment, it is difficult for the prior art to efficiently process different data tables stored in multiple partitions, resulting in inefficient data processing and risk of privacy leakage.
By acquiring the first data table and the second data table, expanding operations based on their attribute items, adding data columns to align data, and finally performing association processing, forming an association data table to improve data processing quality and efficiency.
It realizes efficient correlation operations on data tables in multiple distributed partitions, improves data processing quality and efficiency, reduces traffic, and ensures data security.
Smart Images

Figure CN120277110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method, apparatus, device, and computer program product for processing data tables. Background Art
[0002] With the advent of the big data era, the distributed computing method, as an effective solution to the insufficient computing power for big data, has gradually become popular. Specifically, distributed computing usually requires splitting data and computing tasks into multiple partitions for processing to improve the quality and efficiency of data processing.
[0003] However, since different data is stored in different partitions, when a data processing task involves data in multiple partitions, it is often necessary to perform data processing operations on each involved partition in sequence to obtain a data processing result. This greatly increases the data processing process and reduces the data processing efficiency at the same time. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, device, and computer program product for processing data tables, which can perform an association operation on different data tables stored in multiple distributed partitions to obtain an associated data table, and then can perform corresponding data processing operations based on the associated data table, thereby improving the quality and efficiency of data processing.
[0005] In a first aspect, an embodiment of the present invention provides a method for processing a data table, including:
[0006] Obtain a first data table and a second data table, where both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items;
[0007] Expand the second data table based on the attribute items in the first data table to obtain a second expanded table;
[0008] Expand the first data table based on the attribute items in the second data table to obtain a first expanded table;
[0009] Add a data column to the first expanded table to obtain an adjusted expanded table, where the parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table;
[0010] Associate the second expanded table with the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0011] In a second aspect, an embodiment of the present invention provides a data table processing apparatus, including:
[0012] An acquisition module, configured to acquire a first data table and a second data table, where the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items;
[0013] An extension module, configured to extend the second data table based on the attribute items in the first data table to obtain a second extended table;
[0014] The extension module is further configured to extend the first data table based on the attribute items in the second data table to obtain a first extended table;
[0015] A processing module, configured to add a data column to the first extended table to obtain an adjusted extended table, where the parameter values included in the data column are used to reorder the first extended table so that the data in the first extended table is aligned with the data in the second extended table;
[0016] The processing module is further configured to associate the second extended table with the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0017] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the processing method of the data table in the first aspect above is implemented.
[0018] In a fourth aspect, an embodiment of the present invention provides a computer storage medium for storing a computer program, and when the computer program is executed by a computer, the processing method of the data table in the first aspect above is implemented.
[0019] In a fifth aspect, an embodiment of the present invention provides a computer program product, including: a computer program, when the computer program is executed by a processor of an electronic device, causing the processor to execute the steps in the processing method of the data table in the first aspect above.
[0020] The data table processing method, device, equipment and computer program product provided by this embodiment obtain a first data table and a second data table, expand the second data table based on the attribute items in the first data table to obtain a second expanded table; expand the first data table based on the attribute items in the second data table to obtain a first expanded table; then add a data column to the first expanded table to obtain an adjusted expanded table, and associate the second expanded table with the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table, effectively realizing the association operation on different data tables stored in multiple distributed partitions, and then corresponding data processing operations can be performed based on the obtained associated data table, which greatly improves the quality and efficiency of data processing and ensures the practicability of this method, facilitating market promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 Schematic diagram of the principle of a data table processing method provided by an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of the flow of a data table processing method provided by an embodiment of the present invention;
[0024] Figure 3 Schematic diagram of the flow of expanding the second connection table to obtain a second expanded table provided by an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of the flow of calculating the prefix sum of the expansion times to obtain the calculated parameter provided by an embodiment of the present invention;
[0026] Figure 5 Schematic diagram of the structure of a data table processing device provided by an embodiment of the present invention;
[0027] Figure 6 For Figure 5 Schematic diagram of the structure of an electronic device corresponding to the data table processing device shown in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0030] It should be understood that the term "and / or" used herein is only a kind of association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0031] Depending on the context, the words "if", "when" as used herein may be interpreted as "when" or "when...", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0032] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such commodity or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including said element.
[0033] In addition, the step timings in the following method embodiments are only examples and are not strictly limited.
[0034] Term definition:
[0035] Trusted Execution Environment (TEE): It is an isolated execution environment with security functions where data can be securely computed without the worry of being stolen.
[0036] TEE-based encrypted database: A common encrypted database solution. Users encrypt the data and upload it to the database, and upload the key to the TEE. During computation, the data in the database is securely decrypted in the TEE for operation, and the operation result is re-encrypted before leaving the TEE and then sent to the user. Since the data exists in encrypted form except inside the TEE during the data flow process, the data security is protected to a certain extent.
[0037] Data Shuffle: An operation to redistribute data in a certain way according to data processing requirements in a distributed environment. The purpose is to facilitate subsequent data processing, and the implementation method is generally that each partition sends the results to other partitions as needed.
[0038] Data aggregation operation (reduceByKey): Used to combine values with the same key together.
[0039] Oblivious algorithm: Often used to describe an algorithm. In a single-machine scenario, an oblivious algorithm means that the reads or writes of memory locations by the algorithm are completely independent of the input of the algorithm and only related to some public information (such as the length of the input or output data). In a distributed scenario, an oblivious algorithm means that when data is exchanged between partitions, the amount of data sent is independent of the input, and whether it is required that the reads or writes of memory locations by each partition during local computation are completely independent of the input of the algorithm depends on the context. In this proposal, there is no oblivious requirement for the algorithms of local computation of partitions.
[0040] Dummy data: Meaningless elements used as padding bits in an oblivious algorithm. The purpose is to pad the amount of data to a specified size to prevent the actual size of the data volume from leaking privacy information.
[0041] In order to understand the specific implementation process of the technical solution in this embodiment, the related technologies are briefly described as follows:
[0042] In the field of encrypted databases based on the Trusted Execution Environment (TEE), data is securely protected through encryption. During operations, it usually needs to be securely decrypted within the trusted area before calculation. In the era of big data, it is difficult for a single machine to complete the calculation operations of a large amount of data, and the newly added encryption and decryption operations undoubtedly further increase the load of the calculation operations. Therefore, the distributed computing method, as an effective solution to address the insufficient computing power for big data calculation, has gradually become popular. Distributed computing usually requires splitting data and computing tasks across multiple partitions for processing to improve the quality and efficiency of data processing.
[0043] However, even though the data has been encrypted and protected by the encrypted database, when each partition performs data calculation under the distributed environment, there are still serious privacy leakage problems in the amount of data communicated and transmitted between each other. For example, when the distributed open-source processing system executes the data aggregation operator (reduceByKey), in the first step, the values with the same key are concentrated into the same partition through the data redistribution operator (shuffle) for subsequent operations. By observation, during the process of data redistribution by shuffle, the amount of data received by each partition can roughly estimate the data distribution of the original data on the key. Therefore, different from non-encrypted databases, these operators need to be redesigned into oblivious algorithms so that the amount of data communicated and sent between partitions during the operation process is independent of the data input, thereby achieving all-round protection of the data.
[0044] Common operators for database queries can include: filter operator, group by aggregate operator, project operator, order by operator, etc. It is not very difficult to transform them into oblivious algorithms. Specifically:
[0045] Related technology 1 implemented a distributed database query system that meets the oblivious requirements. This distributed database query system can support sorting, filtering, and aggregation operations. However, this query system does not support general natural join operators and only supports natural join operators of the primary key-foreign key type, lacking generality.
[0046] Related technology 2 implemented an efficient distributed database operator calculation scheme. This calculation scheme optimized the filtering and aggregation operations, removed the dependence on the sorting operation, and also supported the operation of general natural joins. However, this scheme has the following disadvantages in terms of security and efficiency:
[0047] In terms of security, the algorithm relies on two publicly available values, α and β, which represent the maximum repetition degrees of the two tables participating in the natural join on the keys of the corresponding join conditions respectively. If these two values themselves are also private information, then this algorithm will not be applicable, otherwise it will violate the security requirements;
[0048] In terms of efficiency, the algorithm will fill each partition with some dummy data up to a total of M / p + αβ, where M is the number of rows in the output result of the natural join operator and p is the total number of partitions. An algorithm that achieves computational balance should have exactly M / p data in each partition, meaning that the operator provided by this scheme fills up to αβ data to meet the requirements of the oblivious algorithm. Before and after filling, the algorithm has two rounds of complete communication, and the total communication volume is approximately N + 2M + 2αβp. In the worst case, αβ = N 2 , where N is the number of rows in the two tables participating in the natural join. Therefore, this scheme is actually very inefficient in many scenarios; moreover, for common natural join operators, it is not an obvious thing to transform them into oblivious algorithms. If the algorithm is not designed properly, it will lead to a significant increase in communication volume, a decrease in computational efficiency, and even deviate from the original intention of distributed computing.
[0049] To solve the above technical problems, this embodiment provides a method, device, and equipment for processing data tables. Referring to the attached Figure 1 As shown, the execution subject of the data table processing method provided in this embodiment can be a data table processing device. It should be noted that this data table processing device can be implemented as a terminal device, a personal computer, a tablet computer, a local server, or a cloud server. At this time, when the data table processing device is implemented as a cloud server, the data table processing method can be executed in the cloud. Several computing nodes (cloud servers) can be deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, multiple computing nodes can be organized to provide a certain service. Of course, a single computing node can also provide one or more services. The way the cloud provides this service can be to provide an external service interface, and users can call this service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).
[0050] The processing device of the data table is communicatively connected to the client. Among them, the client is used for the user to perform applications so as to enable the processing operation of the data table. The above-mentioned client can be any computing device with certain data transmission capabilities. Specifically, the client can be a mobile phone, a personal computer (PC), a tablet computer, a set application program, etc. In addition, the basic structure of the client can include: at least one processor. The number of processors depends on the configuration and type of the client. The client can also include a memory, which can be volatile, for example: Random Access Memory (RAM), or non-volatile, for example: Read-Only Memory (ROM), flash memory, etc., or can also include both types at the same time. Usually, an operating system (OS), one or more application programs, and program data can be stored in the memory. In addition to the processing unit and the memory, the client also includes some basic configurations, such as a network card chip, an IO bus, a display component, and some peripheral devices, etc. Optionally, some peripheral devices can include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated here.
[0051] The processing device of the data table refers to a device that can provide the processing operation of the data table in a network virtual environment, usually referring to a device that uses the network for information planning and the processing operation of the data table. Physically, the processing device of the data table can be any device that can provide computing services, respond to the processing request of the data table, and can perform the processing operation of the data table based on the processing request of the data table. For example: it can be a cluster server, a conventional server, a cloud server, a cloud host, a virtual center, etc. The composition of the processing device of the data table mainly includes a processor, a hard disk, a memory, a system bus, etc., which is similar to a general computer architecture.
[0052] In the above-mentioned embodiment of the present invention, the client is network-connected to the processing device of the data table, and this network connection can be a wireless or wired network connection. If the client can be communicatively connected to the processing device of the data table, the network mode of the mobile network can be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, 6G, etc.
[0053] In an embodiment of the present application, a client is configured to generate or obtain a processing request for a data table. Specifically, the client can display a human-computer interaction interface, obtain an execution operation input by a user in the human-computer interaction interface, and generate or obtain a processing request for the data table based on the execution operation. In order to implement the processing operation of the data table, the processing request for the data table can be sent to a data table processing device.
[0054] The data table processing device is configured to obtain the processing request for the data table sent by the client, and then can determine a first data table and a second data table corresponding to the processing request for the data table. Among them, the first data table and the second data table are respectively complete data tables, and both are stored in multiple distributed partitions, that is, the multiple data sub-tables constituting the first data table are evenly distributed and stored in each partition, and the multiple data sub-tables constituting the second data table are evenly distributed and stored in each partition. Moreover, for the first data table and the second data table, the first data table and the second data table may include the same attribute items, and the same attribute items may include one category, two categories or more categories, etc.
[0055] Since the first data table and the second data table are stored in multiple distributed partitions, in order to improve the quality and efficiency of data processing and minimize the data communication volume as much as possible, after obtaining the first data table and the second data table corresponding to the processing request for the data table, an expansion operation can be performed on the second data table based on the attribute items in the first data table to obtain a second extended table. Similarly, an expansion operation can be performed on the first data table based on the attribute items in the second data table to obtain a first extended table.
[0056] After performing the expansion operation on the first data table and the second data table, the data in the obtained first extended table and second extended table is not aligned. In order to stably process the first data table and the second data table, a data column can be added to the first extended table, so that an adjusted extended table aligned with the data in the second extended table can be obtained. After obtaining the second extended table and the adjusted extended table, an association operation can be performed on the second extended table and the adjusted extended table, so that an association data table corresponding to the first data table and the second data table can be obtained. Then, corresponding data processing operations can be performed based on the association data table. In this way, when data processing operations need to be performed based on the first data table and the second data table, the data processing operations can be directly performed based on the association data table, thereby effectively improving the quality and efficiency of data processing and being conducive to market promotion and application.
[0057] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step timing in the following method embodiments is only an example, not strictly limited.
[0058] Figure 2 Schematic flowchart of a method for processing a data table provided by an embodiment of the present invention; refer to the appendix Figure 2 As shown, this embodiment provides a method for processing a data table. The execution subject of this method can be a data table processing device. It can be understood that this data table processing device can be implemented as software, or a combination of software and hardware. Specifically, when the data table processing device is implemented as hardware, it can specifically be various electronic devices capable of implementing data table processing operations, including but not limited to tablet computers, personal computers (PCs), servers, and the like. When the data table processing device is implemented as software, it can be installed in the above-mentioned exemplified electronic devices. Based on the above data table processing device, the method for processing a data table in this embodiment can include the following steps:
[0059] Step S201: Obtain a first data table and a second data table. Both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items.
[0060] Step S202: Expand the second data table based on the attribute items in the first data table to obtain a second expanded table.
[0061] Step S203: Expand the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0062] Step S204: Add a data column to the first expanded table to obtain an adjusted expanded table. The parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table.
[0063] Step S205: Associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0064] The specific implementation principles and implementation effects of the above steps will be described in detail below:
[0065] Step S201: Obtain a first data table and a second data table. Both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items.
[0066] Among them, when the user has a data table processing requirement, the data table processing device can obtain a first data table and a second data table. In some instances, the first data table and the second data table can be obtained through human-computer interaction operations. At this time, obtaining the first data table and the second data table can include: displaying a human-computer interaction interface for implementing data table processing operations; displaying multiple data tables in the human-computer interaction interface; obtaining the execution operations input by the user for the multiple data tables, and obtaining the first data table and the second data table based on the execution operations. For example, when the data table processing device stores multiple data tables including: data table a, data table b, data table c, data table d, and data table e, when the user has a data table processing requirement for the above data table b and data table c, the user can input a selection operation for the above data table b and data table c, so as to obtain the to-be-processed data table b and data table c.
[0067] In other instances, the first data table and the second data table can not only be obtained through human-computer interaction operations, but also through analysis and processing. At this time, obtaining the first data table and the second data table can include: obtaining a client or a preset device communicatively connected to the data table processing device; actively or passively obtaining a data table processing request through the client or the preset device; determining the first data table and the second data table corresponding to the data table processing request, thereby effectively ensuring the accuracy and reliability of obtaining the first data table and the second data table.
[0068] It should be noted that the first data table and the second data table can be complete data tables respectively. The above first data table and second data table are both stored in multiple distributed partitions. Among them, a partition can be a computing unit under distribution, which can be a machine, such as: a server; or, a partition can also be a core of a machine processor. Each partition has independent computing, storage, and network resources, and the partitions cooperate to complete computing tasks through network communication.
[0069] In addition, for the first data table and the second data table, the first data table and the second data table include the same attribute items. For example, the first data table is R(A,B). The above first data table can be composed of the data included in sub-table R1, sub-table R2, and sub-table R3. Sub-table R1 can be stored in partition a, sub-table R2 can be stored in partition b, and sub-table R3 can be stored in partition c. That is, the first data table R(A,B) can be distributed and stored in the above-mentioned multiple partitions. Similarly, the second data table is S(B,C). The second data table can be composed of the data included in sub-table S1, sub-table S2, and sub-table S3. Sub-table S1 can be stored in partition a, sub-table S2 can be stored in partition b, and sub-table S3 can be stored in partition c. And, the above first data table R(A,B) and second data table S(B,C) can include the same attribute item B. It can be understood that the same attribute items existing between the first data table and the second data table can be not only of one type. Those skilled in the art can flexibly adjust the number of the same attribute items existing between the first data table and the second data table according to requirements.
[0070] When there are the same attribute items between the first data table and the second data table, the first data table and the second data table can be associated and processed. When there are no same attribute items between the first data table and the second data table, it is impossible to perform the operation of associating and processing the first data table and the second data table.
[0071] Step S202: Expand the second data table based on the attribute items in the first data table to obtain a second expanded table.
[0072] After obtaining the first data table and the second data table, in order to implement the association operation of the data tables and ensure the data security in the first data table and the second data table, the second data table can be expanded based on the attribute items in the first data table. Among them, the attribute items in the first data table are the common attribute items included in the first data table and the second data table, so that a second expanded table can be obtained.
[0073] In some examples, the expansion operation can be implemented through a preset machine learning model, neural network model, or preset algorithm. At this time, expanding the second data table based on the attribute items in the first data table to obtain a second expanded table can include: obtaining a pre-trained machine learning model or neural network model; inputting the attribute items in the first data table and the second data table into the machine learning model or neural network model to obtain the second expanded table output by the machine learning model or neural network model.
[0074] In other instances, the extension operation can be implemented not only through a machine learning model or a neural network model, but also based on the repetition degree of the attribute items in the first data table. At this time, the second data table is extended based on the attribute items in the first data table, and obtaining the second extended table may include: performing grouped aggregation processing on the first data table based on the attribute items in the first data table to obtain a first aggregated table, where the first aggregated table includes the repetition degrees of each attribute value located under the attribute item; performing a natural join on the first aggregated table and the second data table to obtain a second joined table; and extending the second joined table to obtain the second extended table.
[0075] Specifically, after obtaining the first data table and the second data table, grouped aggregation processing can be performed on the first data table based on the attribute items in the first data table to obtain a first aggregated table, and the first aggregated table may include the repetition degrees of each attribute value located under the attribute item.
[0076] For example, the first data table R(A, B) can be {(a1, b1), (a1, b2), (a2, b2), (a1, b3), (a2, b4)}, and the second data table S(B, C) can be {(b1, c1), (b1, c2), (b2, c2), (b2, c3), (b3, c4)}. As can be seen from the above, there is the same attribute item B between the first data table and the second data table, that is, the attribute item B in the first data table is the value corresponding to the b column. Then, by performing a statistical operation on the repetition degree of the attribute items in the first data table, the repetition degrees of each attribute value in the above attribute items can be obtained, that is, the repetition degree of b1 is 1, the repetition degree of b2 is 2, the repetition degree of b3 is 1, and the repetition degree of b4 is 1. Based on the repetition degrees of the attribute values in the first data table above, grouped aggregation processing is performed on the first data table to obtain a first aggregated table R(B, D), and the first aggregated table R(B, D) can be {(b1, 1), (b2, 2), (b3, 1), (b4, 1)}.
[0077] After obtaining the first aggregated table and the second data table, a natural join can be performed on the first aggregated table and the second data table to obtain a second joined table. In some examples, the natural join operation can be implemented through a preset machine learning model or neural network model. For example, when the first aggregated table R(B,D) is {(b1,1),(b2,2),(b3,1),(b4,1)} and the second data table S(B,C) is {(b1,c1),(b1,c2),(b2,c2),(b2,c3),(b3,c4)}, by performing a natural join operation on the first aggregated table and the second data table, the second joined table S2(B,C,D) can be obtained. The second joined table S2(B,C,D) can be {(b1,c1,1),(b1,c2,1),(b2,c2,2),(b2,c3,2),(b3,c4,1)}. After obtaining the second joined table, an expansion operation can be performed on the second joined table, thus effectively implementing the expansion operation on the second data table, and the second expanded table can be stably obtained.
[0078] Step S203: Expand the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0079] After obtaining the first data table and the second data table, in order to implement the association operation of the data tables and ensure the data security in the first data table and the second data table, an expansion operation can be performed on the first data table based on the attribute items in the second data table. Among them, the attribute items in the second data table are the common attribute items included in the first data table and the second data table, so that a first expanded table can be obtained.
[0080] In some examples, expanding the first data table based on the attribute items in the second data table to obtain a first expanded table may include: performing grouped aggregation processing on the second data table based on the attribute items in the second data table to obtain a second aggregated table, where the second aggregated table includes the repetition degrees of each attribute value located under the attribute item; performing a natural join on the second aggregated table and the first data table to obtain a first joined table; and expanding the first joined table to obtain a first expanded table.
[0081] It should be noted that the specific implementation method, implementation effect, and implementation principle of the expansion operation on the first data table in this embodiment are similar to those of the expansion operation on the second data table in step S202 above. For specific reference, please refer to the above description, and details will not be repeated here.
[0082] Step S204: Add a data column to the first expanded table to obtain an adjusted expanded table. The parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table.
[0083] For the first data table and the second data table, after performing an expansion operation on the first data table and the second data table, the data in the first extended table obtained is scrambled with the data in the second extended table. At this time, the data in the first extended table and the data in the second extended table are often not aligned. Therefore, in order to enable the processing operation of the data table, after obtaining the first extended table, a data column can be added to the first extended table to obtain an adjusted extended table. The parameter values included in the data column are used to perform a reordering operation on the first extended table so that the data in the first extended table is aligned with the data in the second extended table.
[0084] In some examples, adding a data column to the first extended table to obtain an adjusted extended table may include: adding a data column with an initial value of 1 to the first extended table to obtain an intermediate extended table; comparing two attribute values of two adjacent groups of data in the intermediate extended table; when the two attribute values corresponding to two adjacent groups of data are the same, performing a summation process on the data columns in the two adjacent groups of data to obtain a summation data column, and determining the two attribute values and the summation data column to be the data for constructing the adjusted extended table; when the two attribute values corresponding to two adjacent groups of data are different, determining the latter data in the two adjacent groups of data to be the data for constructing the adjusted extended table.
[0085] For example, when the first extended table is By adding a data column with an initial value of 1 to the first extended table, an intermediate extended table can be obtained. The intermediate extended table can be Then, two attribute values of two adjacent groups of data in the intermediate extended table can be compared, that is, comparing the attribute values "a" and "b" in the data {a1, b1, 1} and the data {a2, b2, 1}. When a1 = a2 and b1 = b2, a summation process can be performed on the data columns in the two adjacent groups of data, and thus a summation data column can be obtained. The summation data column is {a1, b1, 2}. Then, the two attribute values and the summation data column can be determined to be the data for constructing the adjusted extended table, which can ensure the accuracy and reliability of obtaining the adjusted extended table.
[0086] Correspondingly, when the two attribute values corresponding to two adjacent groups of data are different, that is, a1 ≠ a2 or b1 ≠ b2, the latter data in the two adjacent groups of data can be determined to be the data for constructing the adjusted extended table. That is, the data for constructing the adjusted extended table at this time is {a2, b2, 1}, which effectively ensures the accuracy and reliability of obtaining the adjusted extended table.
[0087] It should be noted that in order to align the data in the first extended table and the second extended table, not only can data columns be added to the first extended table, but also data columns can be added to the second extended table, so as to obtain an adjusted extended table corresponding to the second extended table. At this time, step S204 can be changed to "add a data column to the second extended table to obtain an adjusted extended table, and the parameter values included in the data column are used to reorder the second extended table so that the data in the first extended table and the second extended table are aligned".
[0088] Step S205: Associate the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0089] Since the data in the adjusted extended table is aligned with the data in the second data table, in order to accurately process the first data table and the second data table, after obtaining the adjusted extended table and the second extended table, an association operation can be performed on the second extended table and the adjusted extended table, so as to obtain an associated data table corresponding to the first data table and the second data table.
[0090] In some instances, the association operation can be implemented by a pre-trained machine learning model or neural network model. At this time, associating the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table may include: obtaining a pre-trained machine learning model or neural network model; inputting the second extended table and the adjusted extended table into the machine learning model or neural network model to obtain the associated data table output by the machine learning model or neural network model corresponding to the first data table and the second data table. This associated data table is the result obtained after performing a natural join operation on the first data table and the second data table, effectively ensuring the accuracy and reliability of determining the associated data table. For the obtained associated data table, the associated data table can be stored in multiple distributed partitions, which can ensure the quality and effect of storing the associated data table.
[0091] In other instances, not only can the association operation of data tables be implemented by a pre-trained machine learning model or neural network model, but also the association operation of data tables can be implemented using preset rules. At this time, associating the first extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table includes: obtaining a preset rule for sorting the first extended table and the adjusted extended table; using the preset rule to sort the first extended table and the adjusted extended table respectively to obtain a first sorted table and a second sorted table; associating the first sorted table and the second sorted table to obtain an associated data table corresponding to the first data table and the second data table.
[0092] After obtaining the first extended table, a preset rule for sorting the first extended table can be obtained first. The preset rule can be a lexicographical order or other sequence stored in a preset area or a preset device. Similarly, after obtaining the adjusted extended table, a preset rule for sorting the adjusted extended table can be obtained first. The preset rule can be a lexicographical order or other sequence pre-stored in a preset area or a preset device. It should be noted that the preset rule for sorting the first extended table and the preset rule for sorting the adjusted extended table can be obtained synchronously or asynchronously.
[0093] After obtaining the preset rules for sorting the first extended table and the adjusted extended table, the first extended table and the adjusted extended table can be sorted respectively using the preset rules, so that a first sorted table and a second sorted table can be obtained. After obtaining the first sorted table and the second sorted table, an association operation can be performed on the first sorted table and the second sorted table, so that an associated data table corresponding to the first data table and the second data table can be obtained. For example, if the i-th data in the first data table is (b, c) and the i-th data in the second data table is (a, b1, f), where b = b1, when performing an association operation on the first data table and the second data table, an associated data table can be obtained, and the i-th item data in the associated data table is (a, b, c), thus effectively realizing the association operation of the data tables.
[0094] In some other examples, after obtaining the associated data table corresponding to the first data table and the second data table, to improve the practicability of the method, a data query operation can be implemented based on the obtained associated data table. At this time, the method in this embodiment may further include: obtaining a data query request; determining the associated data table corresponding to the data query request; and obtaining a data query result based on the associated data table.
[0095] Specifically, when the user has a data query requirement, the processing device of the data table can be made to obtain the data query request. In some examples, the data query request can be obtained through a human-computer interaction operation, or the data query request can be obtained actively or passively through a preset device. After obtaining the data query request, the data query request can be analyzed and processed to determine the associated data table corresponding to the data query request. After obtaining the associated data table, a corresponding data query operation can be performed based on the associated data table, so that a data query result can be obtained, thus effectively realizing the data query operation based on the associated data table and further improving the practicability of the method.
[0096] The data table processing method provided in this embodiment obtains a first data table and a second data table, expands the second data table based on the attribute items in the first data table to obtain a second extended table, expands the first data table based on the attribute items in the second data table to obtain a first extended table, then adds a data column to the first extended table to obtain an adjusted extended table, and associates the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table, effectively realizing the association operation on different data tables stored in multiple distributed partitions. Then, corresponding data processing operations can be performed based on the obtained associated data table, greatly improving the quality and efficiency of data processing and ensuring the practicability of this method, which is conducive to market promotion and application.
[0097] Figure 3 It is a schematic flowchart for expanding the second connection table provided in the embodiment of the present invention to obtain a second extended table; on the basis of the above embodiment, refer to the attached Figure 3 As shown, this embodiment does not limit the implementation manner of expanding the second connection table. In some instances, a pre-trained machine learning model or neural network model can be used to expand the second connection table, or the second connection table can be expanded based on the data output volume corresponding to each partition. At this time, expanding the second connection table provided in this embodiment to obtain a second extended table may include:
[0098] Step S301: Obtain the data output volume corresponding to each partition.
[0099] Since the expansion operation of the second connection table is related to the data output volume corresponding to each partition, in order to accurately implement the expansion operation of the second connection table, the data output volume corresponding to each partition can be obtained. In some instances, the data output volume corresponding to each partition may be data pre-configured by the user and stored in a preset area or a preset device. At this time, the data output volume corresponding to each partition can be obtained by accessing the preset area or the preset device.
[0100] Alternatively, in some other examples, the data output volume corresponding to each partition can not only be pre-configured by the user, but also be determined based on the number of partitions and the total output data volume. At this time, obtaining the data output volume corresponding to each partition may include: obtaining the total output data volume M and the number of partitions p; based on the total output data volume M and the number of partitions p, determining the data output volume m = M / p corresponding to each partition. For example, when the total output data volume M is 18 and the number of partitions p is 3, the data output volume corresponding to each partition can be determined to be 18 / 3 = 6. At this time, multiple partitions perform corresponding data processing operations on the data using an equalization algorithm.
[0101] Step S302: Determine the duplication degree of each attribute value as the expansion times for expanding each data in the second connection table.
[0102] After performing a natural join operation on the first aggregation table and the second data table, a second connection table can be obtained. The second connection table may include the duplication degree of each attribute value, and then the duplication degree of each attribute value can be determined as the expansion times for expanding each data in the second connection table. For example, when the second connection table S2(B, C, D) is {(b1, c1, 1), (b1, c2, 1), (b2, c2, 2), (b2, c3, 2), (b3, c4, 1)}, the "1" in (b1, c1, 1) can be determined as the expansion times for expanding the data "(b1, c1)", that is, the data "(b1, c1)" needs to be expanded 1 time. Similarly, the "1" in (b1, c2, 1) is determined as the expansion times for expanding the data "(b1, c2)", the "2" in (b2, c2, 2) is determined as the expansion times for expanding the data "(b2, c2)", the "2" in (b2, c3, 2) is determined as the expansion times for expanding the data "(b2, c3)", and the "1" in (b3, c4, 1) is determined as the expansion times for expanding the data "(b3, c4)", effectively ensuring the accuracy and reliability of determining the expansion times for expanding each data in the second connection table.
[0103] Step S303: Based on the expansion times and the data output volume, determine the target partition corresponding to each data.
[0104] Since each data in the data table is located in different distributed partitions, when performing an expansion operation, it is necessary to send each data to the corresponding target partition. Therefore, in order to be able to perform an expansion operation on the second joined table, after obtaining the expansion times and the data output volume, the expansion times and the data output volume can be analyzed and processed to determine the target partition corresponding to each data. In some instances, the target partition can be determined by analyzing and processing the expansion times and the data output volume through a pre-trained machine learning model or neural network model. At this time, based on the expansion times and the data output volume, determining the target partition corresponding to each data can include: obtaining a pre-trained machine learning model or neural network model; inputting the expansion times, the data output volume, and the second joined table into the machine learning model or neural network model to obtain the target partition corresponding to each data output by the machine learning model or neural network model.
[0105] In other instances, not only can the target partition corresponding to each data be determined through a pre-trained machine learning model or neural network model, but the prefix sum algorithm can also be used to analyze and process the expansion times and the data output volume to determine the target partition corresponding to each data. At this time, based on the expansion times and the data output volume, determining the target partition corresponding to each data in the second joined table can include: calculating the prefix sum of the expansion times to obtain the calculated parameter; determining the 0 value and the calculated parameter as the processed parameter; obtaining the quotient of the processed parameter divided by the data output volume; determining the target partition corresponding to the data based on the sum value of the quotient and 1.
[0106] For example, when the expansion times are (1, 3, 1, 1, 5, 2, 1, 1, 3), after calculating the prefix sum of the expansion times, the calculated parameter obtained can be (1, 4, 5, 6, 11, 13, 14, 15). Then, the 0 value and the calculated parameter can be determined as the processed parameter, and the processed parameter is (0, 1, 4, 5, 6, 11, 13, 14, 15). When the number of partitions is 3 and the output data volume of each partition is 6, the quotient of each of the above processed parameters divided by the data output volume can be obtained, that is, the quotient of 0 / 6 is 0, the quotient of 1 / 6 is 0, the quotient of 4 / 6 is 0, the quotient of 5 / 6 is 0, the quotient of 6 / 6 is 1, the quotient of 11 / 6 is 1, the quotient of 13 / 6 is 2, the quotient of 14 / 6 is 2, and the quotient of 15 / 6 is 2.
[0107] After obtaining the quotient values corresponding to the above-mentioned processed parameters, the sum value of the quotient value and 1 can be obtained, so as to obtain a sum value sequence (1, 1, 1, 1, 2, 2, 3, 3, 3) corresponding to multiple calculated parameters. From the above, it can be determined that the target partition corresponding to the data corresponding to the sum value "1" is the partition with the preset identifier "1", the target partition corresponding to the data corresponding to the sum value "2" is the partition with the preset identifier "2", and the target partition corresponding to the data corresponding to the sum value "3" is the partition with the preset identifier "3", thus effectively ensuring the accuracy and reliability of determining the target partition corresponding to the data.
[0108] It should be noted that the target partition corresponding to the data can be determined not only by the sum value of the quotient value and 1, but also by first obtaining the difference between the processed parameter and 1, then obtaining the ratio of the difference to the data output volume, and performing a ceiling operation on the ratio value. Finally, the target partition corresponding to the data can be determined based on the parameter after the ceiling operation, which can also ensure the accuracy and reliability of determining the target partition.
[0109] Step S304: Expand each data based on the data output volume to obtain expanded data.
[0110] In order to implement the data expansion operation, after obtaining the data output volume, each data can be expanded based on the data output volume to obtain expanded data. In some instances, the expansion operation can be implemented by a pre-trained machine learning model or neural network model. At this time, expanding each data based on the data output volume to obtain expanded data may include: obtaining a pre-trained machine learning model or neural network model; inputting the data output volume and each data into the machine learning model or neural network model to obtain the expanded data output by the machine learning model or neural network model.
[0111] In other instances, the data expansion operation can be implemented not only by a pre-trained machine learning model or neural network model, but also based on the total number of digits of the data output. At this time, expanding each data based on the data output volume to obtain expanded data may include: determining the total number of digits of the data output based on the data output volume; expanding each data based on the total number of digits of the data output to obtain expanded data.
[0112] Among them, since the data expansion operation is related to the total number of bits of data output, in order to accurately implement the data expansion operation, after obtaining the data output amount, the total number of bits of data output can be determined based on the data output amount. In some instances, determining the total number of bits of data output based on the data output amount may include: obtaining a pre-configured mapping relationship between the data output amount and the total number of bits of data output, and determining the total number of bits of data output based on the data output amount and the mapping relationship.
[0113] In other instances, not only can the total number of bits of data output be determined through a preset mapping relationship, but also through a preset formula. At this time, determining the total number of bits of data output based on the data output amount may include: performing a square root processing operation on the data output amount to obtain a processed value; determining the sum value between the data output amount and the processed value, and performing a rounding operation on the sum value to obtain the total number of bits of data output. For example, when the data output amount is 6, then through the calculation, the total number of bits of data output = 9 can be obtained, which effectively ensures the accuracy and reliability of determining the total number of bits of data output.
[0114] After obtaining the total number of bits of output data, each data can be expanded based on the total number of bits of data output to obtain expanded data. In some instances, not only can each data be expanded through a pre-trained machine learning model or neural network model, but also the data expansion operation can be achieved by combining preset padding bits. At this time, expanding each data based on the total number of bits of data output to obtain expanded data may include: obtaining all the data to be sent to the same target partition; detecting whether the sum of the data bits of all the data meets the total number of bits of data output; if not, performing a data expansion operation based on the sum of the data bits of all the data and the preset padding bits to obtain expanded data, and the number of bits of the expanded data meets the total number of bits of data output.
[0115] For example, when the number of partitions is 3, the data corresponding to each of the above data are as follows: The first partition corresponds to {(a, 0), (c, 4), (f, 11), (h, 14)}, the second partition corresponds to {(b, 1), (e, 6), (g, 13)}, and the third partition corresponds to {(d, 5), (i, 15)}. The processed parameter corresponding to data a in the first partition above is 0. Based on the processed parameter and the data output volume, the first partition can be determined as the target partition corresponding to data a. Similarly, through analysis and processing, it can be determined that: the target partition corresponding to data c is the first partition, the target partition corresponding to data f is the second partition, the target partition corresponding to data h is the third partition, the target partition corresponding to data b is the first partition, the target partition corresponding to data e is the second partition, the target partition corresponding to data g is the third partition, the target partition corresponding to data d is the first partition, and the target partition corresponding to data i is the third partition.
[0116] After statistics, the target partitions of data a and data c in the first partition are both the first partition. At this time, since the data bit lengths of data a and data c (2 bits) cannot meet the total data output bit length (9 bits), data expansion operations can be performed based on the total data bit length of all data and the preset padding bits to obtain extended data, that is, expanding data (a, c) into data (a, c, x, x, x, x, x, x, x), where the above "x" is the preset placeholder data (or called the preset padding bit), thus effectively realizing the data expansion operation.
[0117] Similarly, the target partition of data f in the first partition is the second partition. At this time, since data f cannot meet the total data output bit length (9 bits), data expansion operations can be performed based on the total data bit length of all data and the preset padding bits to obtain extended data, that is, expanding data (f) into data (f, x, x, x, x, x, x, x), where the above "x" is the preset padding bit. The target partition of data h in the first partition is the third partition, and then data (h) can be expanded into data (h, x, x, x, x, x, x, x), thus effectively realizing the data expansion operation.
[0118] The target partition corresponding to the data b in the second partition is the first partition, and then the data (b) can be expanded to data (b, x, x, x, x, x, x, x, x). The target partition corresponding to the data e in the second partition is the second partition, and then the data (e) can be expanded to data (e, x, x, x, x, x, x, x, x). The target partition corresponding to the data g in the second partition is the third partition, and then the data (g) can be expanded to (g, x, x, x, x, x, x, x, x). The target partition corresponding to the data d in the third partition is the first partition, and then the data (d) can be expanded to (d, x, x, x, x, x, x, x, x). The target partition corresponding to the data i in the third partition is the third partition, and then the data (e) can be expanded to (e, x, x, x, x, x, x, x, x). In this way, the expansion operation on each data is effectively realized, so that the expanded data can be stably obtained.
[0119] Step S305: Send the expanded data to the target partition and rewrite the expanded data to obtain the expanded data for constructing the second extended table.
[0120] After obtaining the expanded data and determining the target partition corresponding to the expanded data, the expanded data can be sent to the target partition. Since the data received in the target partition includes a lot of preset padding bits, in order to ensure the quality and effect of data expansion, after sending the expanded data to the target partition, a rewrite operation can be performed on the expanded data, so that the expanded data for constructing the second extended table can be obtained. In some instances, a pre-trained machine learning model or neural network model can be used to rewrite the expanded data, so that the second extended table can be stably obtained;
[0121] Alternatively, in some other instances, not only can the rewrite operation on the expanded data be realized through a pre-trained machine learning model or neural network model, but also the rewrite operation can be realized by deleting the redundant preset padding bits. At this time, rewriting the expanded data to obtain the expanded data for constructing the second extended table may include: obtaining the remainder corresponding to the division operation between the processed parameter and the data output amount; determining the position where the data is stored in the target partition based on the remainder; deleting the redundant preset padding bits based on the position and the data output amount to obtain the sorted data; rewriting the preset padding bits included in the sorted data to the previous adjacent data to obtain the target data for constructing the second extended table.
[0122] For example, when the processed parameters are (0, 1, 4, 5, 6, 11, 13, 14, 15), the data corresponding to the above processed parameters includes: the first partition corresponds to {(a, 0), (c, 4), (f, 11), (h, 14)}, the second partition corresponds to {(b, 1), (e, 6), (g, 13)}, the third partition corresponds to {(d, 5), (i, 15)}. When the output data volume is 6, the remainder corresponding to the division operation between the processed parameters and the data output volume can be obtained. Specifically, the remainders of the processed parameters corresponding to the above data a, b, c, d, e, f, g, h are: 0, 1, 4, 5, 1, 5, 1, 2, 3. The above remainders are used to represent the positions where the data is stored in the target partition, and the target partition corresponding to the data can be determined by the quotient value corresponding to the division operation between the processed parameters and the data output volume.
[0123] After analysis and processing, it can be known that the above data a is in the 0th position in the target partition (i.e., the first partition), data b is in the 1st position in the target partition (i.e., the first partition), data c is in the 4th position in the target partition (i.e., the first partition), data d is in the 5th position in the target partition (i.e., the first partition). Then, based on the above positions and the data output volume, the redundant preset padding bits can be deleted, and the sorted data in the target partition (the first partition) can be obtained as: a, b, x, x, c, d, where the above x is the remaining preset padding bit. Similarly, the sorted data in the target partition (the second partition) can be obtained as: e, x, x, x, x, f, where the above x is the remaining preset padding bit; the sorted data in the target partition (the third partition) can be obtained as: x, g, h, i, x, x, where the above x is the remaining preset padding bit.
[0124] After obtaining the sorted data, the preset padding bits included in the sorted data can be rewritten as the previous adjacent data. That is, change "x, x" in the first partition to "b, b" to obtain the first target data "a, b, b, b, c, d" in the first partition. Similarly, change "x" in the second partition to "e" to obtain the second target data in the second partition as "e, e, e, e, e, f". For the first "x" in the third partition, it is adjacent to the last data "f" in the second partition. Then, the first "x" in the third partition can be rewritten as "f", and the fifth and sixth "x" in the third partition can be rewritten as "i". After the above rewriting operations, the third target data in the third partition can be obtained as "f, g, h, i, i, i". The above first target data, second target data, and third target data are all the target data that make up the second extended table, effectively ensuring the stability and reliability of obtaining the second extended table.
[0125] In this embodiment, by obtaining the data output volume corresponding to each partition, the repetition degree of each attribute value is determined as the expansion times for performing expansion operations on each data in the second connection table. Then, based on the expansion times and the data output volume, the target partition corresponding to each data is determined, and each data is expanded based on the data output volume to obtain expanded data. After that, the expanded data is sent to the target partition, and the expanded data is rewritten, so that the expanded data for constituting the second expansion table can be stably obtained, thus effectively realizing the stable expansion operation of the second connection table.
[0126] Figure 4 It is a schematic flow chart for calculating the prefix sum of the expansion times and obtaining the calculated parameters provided by the embodiment of the present invention; on the basis of the above embodiment, refer to the attached Figure 4 As shown, this embodiment provides an implementation manner for calculating the prefix sum of the expansion times. Specifically, the calculation of the prefix sum of the expansion times in this embodiment to obtain the calculated parameters may include:
[0127] Step S401: Obtain the current partition corresponding to the expansion times.
[0128] After obtaining the expansion times, the current partition corresponding to the expansion times can be determined based on the data corresponding to each expansion time. In some instances, the current partition corresponding to each data can be determined as the current partition corresponding to the expansion times. For example, when the expansion times are (1, 3, 1, 1, 5, 2, 1, 1, 3), it can be determined that the current partition corresponding to the above expansion times (1, 3, 1) is the first partition, the current partition corresponding to the above expansion times (1, 5, 2) is the second partition, and the current partition corresponding to the above expansion times (1, 1, 3) is the third partition.
[0129] Step S402: Perform prefix sum calculation processing on the expansion times in the same partition to obtain the first calculation processing sequence corresponding to each partition. The first calculation processing sequence includes multiple processing results.
[0130] After obtaining the current partition corresponding to the expansion times, prefix sum calculation processing operations can be performed on the expansion times in the same partition, so that the first calculation processing sequences in each partition can be obtained. The first calculation processing sequence includes multiple processing results. In some instances, a pre-trained machine learning model or neural network model can be used to perform prefix sum calculation processing on the expansion times in the same partition to obtain the first calculation processing sequence corresponding to each partition.
[0131] For example, when storing the expansion times (1, 3, 1) in the first partition, after performing the prefix sum calculation processing operation on the above expansion times, a first calculation processing sequence corresponding to the first partition can be obtained, and this first calculation processing sequence can be: 1, 4, 5. Similarly, when storing the expansion times (1, 5, 2) in the second partition, after performing the prefix sum calculation processing operation on the above expansion times, a first calculation processing sequence corresponding to the second partition can be obtained, and this first calculation processing sequence can be: 1, 6, 8; when storing the expansion times (1, 1, 3) in the third partition, after performing the prefix sum calculation processing operation on the above expansion times, a first calculation processing sequence corresponding to the third partition can be obtained, and this first calculation processing sequence can be: 1, 2, 5, which effectively ensures the accuracy and reliability of determining the first calculation processing sequences corresponding to each partition.
[0132] Step S403: Send the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition to obtain a processing result sequence stored in the preset partition.
[0133] After obtaining the first calculation processing sequences corresponding to each partition, the last processing result in the first calculation processing sequence corresponding to each partition can be sent to a preset partition. Among them, the preset partition can be the first partition, the second partition, or the Nth partition, etc. configured in multiple partitions. Taking the first partition as an example of the preset partition, after obtaining the last processing result in the first calculation processing sequence corresponding to each partition and sending it to the first partition, a processing result sequence stored in the preset partition can be obtained.
[0134] For example, when the first partition is the preset partition, if the first calculation processing sequence corresponding to the first partition is: 1, 4, 5; the first calculation processing sequence corresponding to the second partition is: 1, 6, 8; the first calculation processing sequence corresponding to the third partition is: 1, 2, 5, then the last processing results (5, 8, 5) in the first calculation processing sequences corresponding to the above partitions can be sent to the first partition, so as to obtain the processing result sequence stored in the first partition: 5, 8, 5.
[0135] Step S404: Perform prefix sum calculation on the processing result sequence to obtain calculated parameters.
[0136] After obtaining the processed result sequence, prefix sum calculation processing can be performed on the processed result sequence, so that the calculated parameter can be obtained. In some examples, performing prefix sum calculation on the processed result sequence to obtain the calculated parameter may include: performing prefix sum calculation on the processed result sequence to obtain a second calculation processing sequence, where the second calculation processing sequence includes multiple calculation processing results; determining the target partition corresponding to each calculation processing result in the second calculation processing sequence; sending the calculation processing result to the target partition, and summing the calculation processing result and the expansion times stored in the same target partition to obtain the calculated parameter.
[0137] Among them, determining the target partition corresponding to each calculation processing result in the second calculation processing sequence may include: sorting the multiple calculation processing results included in the second calculation processing sequence based on the partition identifier corresponding to each calculation processing result to obtain result sorting information; determining the sum value of the current sorting position where each calculation processing result is located and 1 as the target partition corresponding to the calculation processing result.
[0138] For example, when the processed result sequence is: 5, 8, 5, prefix sum calculation can be performed on the processed result sequence to obtain a second calculation processing sequence, and the second calculation processing sequence can be: 5, 13, 18. Then, the second calculation processing sequence can be sorted based on the identifier of the source partition corresponding to the processing result in each of the above processed result sequences. Since 5 corresponds to the first partition, 13 corresponds to the second partition, and 18 corresponds to the third partition, after sorting the second calculation processing sequence, result sorting information can be obtained, that is, 5, 13, 18, that is, the result sorting information is exactly the same as the second calculation processing sequence.
[0139] Then, the target partition corresponding to the above 5 can be determined as the second partition, the target partition corresponding to 13 can be determined as the third partition, and the above 5 can be sent to the second partition, and 13 can be sent to the third partition. Then, the original data stored in the second partition and the data "5" can be summed to obtain the calculated parameter stored in each partition. For example: 1, 3, 1 are stored in the first partition; 6, 8, 7 are stored in the second partition; 14, 14, 16 are stored in the third partition. In this way, the prefix sum calculation operation on the processed result sequence is stably realized, and the accuracy and reliability of obtaining the calculated parameter are ensured.
[0140] In this embodiment, by obtaining the current partition corresponding to the number of expansion times, calculating the prefix sum of the expansion times in the same partition, obtaining the first calculation processing sequence corresponding to each partition, and then sending the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition, obtaining the processing result sequence stored in the preset partition, and calculating the prefix sum of the processing result sequence, the calculated parameters can be stably obtained. Then, the processing operation of the data table can be implemented based on the calculated parameters, further ensuring the practicability of the method.
[0141] In specific applications, taking the number of distributed partitions as three as an example, this application embodiment provides an oblivious efficient algorithm for the natural join operator in a distributed scenario. The oblivious efficient algorithm for the natural join operator relies on the following basic operators to implement: the primary key-foreign key natural join operator, the sorting operator, the grouping aggregation operator, the prefix sum operator, and the expansion operator. Among them, the primary key-foreign key natural join, the sorting operator, and the grouping aggregation operator can directly follow the implementation methods in related technologies. To facilitate the understanding of the implementation principle and implementation effect of the oblivious efficient algorithm in this application embodiment, the implementation principles of the prefix sum operator and the expansion operator are described below:
[0142] Taking the input parameters as {x1, x2,..., x N} and the output parameters as {x1, x1 + x2,..., x1 + x2 +... + x N} as an example to illustrate the implementation principle of the prefix sum operator. Among them, the above "+" operation is not limited to arithmetic addition and can be any binary operation that satisfies the associative law. For example: arithmetic multiplication, logical exclusive OR, taking the larger value, taking the smaller value, taking the latter (if the latter is not the preset placeholder data dummy) or the former (if the latter is the preset placeholder data dummy), etc. Specifically, the implementation process of the prefix sum operator can include the following steps:
[0143] Step 1: Each partition (the first partition, the second partition, and the third partition) calculates the prefix sum of the local stored data and sends the last number of the calculation result to the first partition (or the preset partition).
[0144] For example, the number of partitions p = 3, which are the first partition (corresponding to label 1), the second partition (corresponding to label 2), and the third partition (corresponding to label 3) respectively. The total amount of data N = 9, and the total amount of data stored in each partition n = N / p = 3. The N input numbers are 1, 2, 3, 1, 2, 1, 3, 4, 2 in sequence. Before performing the prefix sum operator, the starting data corresponding to each partition can be: in the first partition, 1, 2, 3 are stored; in the second partition, 1, 2, 1 are stored; in the third partition, 3, 4, 2 are stored; after performing the prefix sum calculation on the stored data in each of the above partitions locally, the following data can be obtained: the first partition - 1, 3, 6; the second partition - 1, 3, 4; the third partition - 3, 7, 9. Then, the data 6, 4, 9 after the prefix sum calculation for the above three partitions can be sent to the first partition, so that the first partition can obtain 6, 4, 9.
[0145] Step 2: The first partition can arrange the p received numbers together in the order of the partition numbers and perform the prefix sum calculation operation, and then send the i-th number to the partition labeled i + 1, where i = 1, 2,..., p - 1.
[0146] Among them, after the first partition obtains 6, 4, 9, it can arrange the p received numbers together, and then perform the prefix sum calculation on the above data to obtain 6, 10, 19. Then, the above 6 can be sent to the (i + 1)-th partition (i.e., the second partition), and the above 10 can be sent to the (i + 1)-th partition (i.e., the third partition).
[0147] Step 3: After each partition receives the corresponding data, the received data can be added to the value after the prefix sum calculation within each partition.
[0148] Since the first partition does not receive any data, the data in the first partition is 1, 3, 6; while the second partition receives the data 6, and the values after the prefix sum calculation in the second partition are 1, 3, 4. After addition and summation, 7, 9, 10 can be obtained and stored in the second partition; the third partition receives the data 10, and the values after the prefix sum calculation in the third partition are 3, 7, 9. After addition and summation, 13, 17, 19 can be obtained and stored in the third partition, thus effectively realizing the prefix sum operation of the data.
[0149] It should be noted that the communication volumes corresponding to the above prefix sum calculation operations are: the communication volume of the above step 1 is p - 1, the communication volume of step 2 is p - 1, and the total communication volume is 2p - 2.
[0150] On the other hand, with the input parameters {x1, x2,..., x N}, N positive integers {d1, d2,..., d with a sum of M N}, as the number of expansion times, and the output parameters are {x1,..., x1, x2…, x2,...x N ,...x N} is used as an example to illustrate the implementation principle of the expansion operator. Among them, x i appears continuously d i times, i = 1, 2,...N. Among them, n = N / p, m = M / p. It should be noted that if N or M is not a multiple of p, dummy data of preset placeholder data is added to make N and M both multiples of p. Specifically, the implementation process of the expansion operator can include the following steps:
[0151] Step 11: Call the prefix sum operator for the sequence {0, d1, d2,..., d N-1}. The addition in the above prefix sum operator is ordinary integer addition, and the result is recorded as {e1, e2,...e N}.
[0152] For example, the number of partitions p = 3, the total amount of data N = 9, the total amount of data in each partition n = N / p = 3, the total amount of output data M = 18, and the amount of output data in each partition m = M / p = 6. In addition, for the sake of illustration, the constant parameter used to control the failure rate can be set to 1. When the N numbers in the input are a, b, c, d, e, f, g, h, i in sequence, and the corresponding positive integers are 1, 3, 1, 1, 5, 2, 1, 1, 3 in sequence, after calling the prefix sum operator for {0, 1, 3, 1, 1, 5, 2, 1, 1}, the result {0, 1, 4, 5, 6, 11, 13, 14, 15} can be obtained.
[0153] Step 12: Rewrite the data {x1, x2,...x N} into the structure of (x1, e1), (x2, e2), …(x N , e N ). Independently and uniformly randomly select one partition from p partitions as the target partition for each data, and then send all the numbers to their respective designated target partitions.
[0154] After the above rewriting operation, the local data stored in each partition will be shuffled. Suppose the data in each partition after shuffling can be:
[0155] First partition: (a, 0), (c, 4), (f, 11), (h, 14);
[0156] Second partition: (b, 1), (e, 6), (g, 13);
[0157] Third partition: (d, 5), (i, 15).
[0158] Step 13: For the data received by each partition, rewrite the data (x1, e1), (x2, e2), … (x N , e N ) as (x1, r1), (x2, r2), … (x N , r N ).
[0159] Assume its data content is (y1, f1), (y2, f2), … (y K , f K ). For (y i , f i ), assume f i = mq i + r i , that is, the quotient of f i divided by m is q i , and the remainder is r i . Rewrite the data content as (x i , r i ), and assign the target partition for this data as the (q i + 1)-th. For each partition j, if the total number of data sent to partition j is less than pieces, where c is a constant parameter pre-configured to control the failure rate. If the total number of data sent to partition j is less than 9, then the preset placeholder data dummy can be supplemented until the total number of data is equal to this value, and then all the data is sent to the specified partition.
[0160] If the total number of data sent to partition j is greater than this value, the entire algorithm fails. At this time, the processing operation of the data table can be restarted, or a prompt message can be generated to remind the user to configure the constant parameter for controlling the failure rate based on the prompt message.
[0161] After the above process, the amount of data sent to each partition will be filled to exactly 9. After sending the data to the corresponding partition, the data of each partition can be respectively (x represents the preset placeholder data dummy):
[0162] First partition: (a, 0), (c, 4), x, x, x, x, x, x, x, (b, 1), x, x, x, x, x, x, x, x, (d, 5), x, x, x, x, x, x, x, x;
[0163] Second partition: (f, 5), x, x, x, x, x, x, x, x, (e, 0), x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x;
[0164] The third partition: (h, 2), x, x, x, x, x, x, x, x, (g, 1), x, x, x, x, x, x, x, x, (i, 3), x, x, x, x, x, x, x, x.
[0165] In addition, for the above constant parameter c, which is used to control the failure probability of the algorithm, through theoretical analysis, it can be proved that the failure probability of the algorithm is less than In an actual scenario, the user can select or configure an appropriate parameter c according to specific application requirements and the required failure rate of the application.
[0166] Step 14: Among the data received in each partition, assuming that the data that is not the preset placeholder data dummy is (z1, g1), …, (z i , g i ), place the data z i at the (g i + 1)-th position, and fill the remaining positions with dummy. The total amount of data is m. Discard the extra dummy data. Finally, call the prefix sum algorithm on the data. Here, the addition in the prefix sum algorithm is defined as x + y = x (if y is dummy data) or y (if y is not dummy data).
[0167] After the above processing operations, the data can be moved to the corresponding positions and the extra dummy can be removed. The results are as follows:
[0168] The first partition: a, b, x, x, c, d;
[0169] The second partition: e, x, x, x, x, f;
[0170] The third partition: x, g, h, i, x, x.
[0171] Then call the prefix sum algorithm again. All dummy data will be rewritten as the non-dummy number adjacent to it in front. Note that the global prefix sum operator is called here. Therefore, the first x in the third partition will recognize the non-dummy number adjacent to or closest to it in front as f, and then the "x" in the third partition can be rewritten as f, so that the processed result can be obtained, which is specifically as follows:
[0172] The first partition: a, b, b, b, c, d;
[0173] The second partition: e, e, e, e, e, f;
[0174] The third partition: f, g, h, i, i, i.
[0175] It should be noted that the communication volumes corresponding to the above extension operations are as follows: the communication volume of step 11 above is 2p - 2, the communication volume of step 12 is N, the communication volume of step 13 is approximately M (which can be ignored), and the communication volume of step 14 is 2p - 2. Thus, the total communication volume can be obtained as approximately N + M.
[0176] Based on the above prefix sum operator and extension operator, taking two data tables of R(A, B) and S(B, C) as input and outputting a data table of T(A, B, C) with M rows of data as an example, where each of the tables R(A, B) and S(B, C) has N rows of data. After the natural join operation, the obtained output result has M rows. Specifically, the implementation process of the oblivious efficient algorithm for the natural join operator can include the following steps:
[0177] Step 111: Invoke the grouping aggregation operator on the table R(A, B) to calculate the repetition degree of each b in R(A, B), and record the result as the table R1(B, D), where the column B is the primary key in the table R1, and D represents the repetition degree of the primary key B in R.
[0178] For example, when N = 5, the data of the table R(A, B) is {(a1, b1), (a1, b2), (a2, b2), (a1, b3), (a2, b4)}, and the data of the table S(B, C) is {(b1, c1), (b1, c2), (b2, c2), (b2, c3), (b3, c4)}. Then the correct natural join result T(A, B, C) of R and S should be {(a1, b1, c1), (a1, b1, c2), (a1, b2, c2), (a2, b2, c2), (a1, b2, b3), (a2, b2, c3), (a1, b3, c4)}, that is, M = 7. For the above table R(A, B), calculate the repetition degree of each b in R(A, B), and the obtained R1(B, D) is {(b1, 1), (b2, 2), (b3, 1), (b4, 1)}.
[0179] Step 112: Invoke the primary key - foreign key natural join operator on the tables R1 and S, and record the result as the table S2(B, C, D).
[0180] Specifically, perform a natural join operation on the tables R1 and S through the primary key - foreign key natural join operator to obtain the table S2(B, C, D). The above table S2(B, C, D) can be {(b1, c1, 1), (b1, c2, 1), (b2, c2, 2), (b2, c3, 2), (b3, c4, 1)}.
[0181] Step 113: Invoke the extension operator on the table S2, the data to be extended is the column combination (B, C), and the number of extension times is the column D. Record the extension result as the table S3(B, C).
[0182] Specifically, after obtaining Table S2, the extension operator can be used to perform an extension process on Table S2, and the obtained extended result Table S3(B, C) is {(b1, c1), (b1, c2), (b2, c2), (b2, c2), (b2, c3), (b2, c3), (b3, c4)}.
[0183] Step 114: Call the sorting operator to sort Table S3 in lexicographical order according to columns (B, C).
[0184] Among them, the sorting operator can receive N elements (or a table with N rows) as input, rearrange the N elements (or N rows) according to a certain size relationship, and then output. When these elements are tuples containing multiple components, sometimes multiple components (columns) are specified for the size relationship, which is called lexicographical sorting by multiple columns. That is, when comparing two tuples, the comparison is first made according to the specified first column. If the values of the tuples in this column are the same, then the second column is used for comparison, and so on. Through calculation, it can be known that the communication volume of the above sorting operation is approximately 3N.
[0185] Step 115: Call the grouping and aggregation operator on Table S to calculate the repetition degree of each b in S(B, C), and record the result as Table S1(B, E). Among them, column B is the primary key in Table S1, and E represents the repetition degree of the primary key B in S.
[0186] Among them, the grouping and aggregation operator corresponds to the "group by" query statement in database queries. This operator can receive a table as input, group the rows of the table according to the specified column combination, and perform aggregation operations on another specified column within the same group, such as summation, maximum value calculation, average value calculation, etc. Through calculation, it can be known that the communication volume of the above grouping and aggregation operation is approximately N.
[0187] Specifically, after obtaining Table S, the grouping and aggregation operator can be used to analyze Table S to obtain the repetition degree corresponding to each value b in the data column B of Table S, so that the result S1(B, E) can be obtained as {(b1, 2), (b2, 2), (b3, 1)}.
[0188] Step 116: Call the primary key-foreign key natural join operator on Table S1 and R, and record the result as R2(A, B, E).
[0189] Among them, the primary key-foreign key natural join operator is a special case of the natural join operator. It takes two tables, table R(A, B) and table S(B, C), as input. Each table has N rows of data, and the B column is unique in table S, that is, the degree of repetition is 1. The output is a table in the form of T(A, B, C) with at most N rows of data, and the content of the table is a set. Since the primary key-foreign key natural join operator needs to merge the two tables together and calls the sorting and prefix sum operators once, its communication volume is relatively large, specifically about 6N.
[0190] Specifically, when table S1(B, E) is {(b1, 2), (b2, 2), (b3, 1)} and table R is {(a1, b1), (a1, b2), (a2, b2), (a1, b3), (a2, b4)}, then the primary key-foreign key natural join operator can be used to perform a natural join operation on table S1 and table R, and the result R2(A, B, E) can be obtained as {(a1, b1, 2), (a1, b2, 2), (a2, b2, 2), (a1, b3, 1)}.
[0191] Step 117: Call the extension operator on table R2. The data to be extended is the column combination (A, B), and the number of extension times is column E, and the extended result is R3(A, B).
[0192] After obtaining table R2, the extension operator can be used to perform an extension operation on table R2 to obtain the extended result R3(A, B). Specifically, R3(A, B) can be {(a1, b1), (a1, b1), (a1, b2), (a1, b2), (a2, b2), (a2, b2), (a1, b3)}.
[0193] Step 118: Add column F to table R3 with an initial value of 1, and denote the result as R4(A, B, F). Call the prefix sum operator on R4. Among them, the addition is defined as: assume that the two tuples to be summed are (a1, b1, f1) and (a2, b2, f2) respectively. If a1 = a2 and b1 = b2, then their sum is (a2, b2, f1 + f2), otherwise it is (a2, b2, f2). Denote the result as R5(A, B, F).
[0194] Among them, the added F column is used to determine the preset order for adjusting table R3 to perform data alignment operation between table R3 and table S. In addition, after calling the prefix sum operator on R4, the obtained result R5(A, B, F) can be {(a1, b1, 1), (a1, b1, 2), (a1, b2, 1), (a1, b2, 2), (a2, b2, 1), (a2, b2, 2), (a1, b3, 1)}.
[0195] Step 119: Call the sorting operator to sort R5 in lexicographic order by column (B, F, A).
[0196] Step 1110: After obtaining Table S3 and Table R5, the merging operation can be performed on Table S3 and Table R5 to obtain the calculation table T(A, B, C).
[0197] For example, assume that the i-th row in Table S3 is (b, c), and the i-th row in Table R5 is (a, b′, f). After merging Table S3 and Table R5, the i-th row in the calculation table T can be obtained as (a, b, c). Specifically, the content of Table T can be: {(a1, b1, c1), (a1, b1, c2), (a1, b2, c2), (a2, b2, c2), (a1, b2, c3), (a2, b2, c3), (a1, b3, c4)}, thus effectively ensuring the quality and effect of data table processing.
[0198] It should be noted that Steps 111 - 113 and Steps 114 - 116 in the above embodiments are completely symmetric, and the parameter b = b′ in the above embodiments. In addition, the communication volumes corresponding to the above extension operations are as follows: the communication volume corresponding to Step 111 is N, the communication volume corresponding to Step 112 is 6N, the communication volume corresponding to Step 113 is N + M, the communication volume corresponding to Step 114 is 3M, the communication volume corresponding to Step 115 is N, the communication volume corresponding to Step 116 is 6N, the communication volume corresponding to Step 117 is N + M, the communication volume corresponding to Step 118 is p - 1, the communication volume corresponding to Step 119 is 3M, the communication volume corresponding to Step 11110 is 0, and the total communication volume is approximately 16N + 8M.
[0199] The technical solution provided in the embodiments of this application realizes an oblivious and efficient algorithm for the natural join operator in a distributed scenario. Specifically, by setting the prefix sum operator and the extension operator, it effectively realizes the data of each table being extended and repeated multiple times in an efficient and concise manner, and the number of times is exactly equal to its repetition degree in another table. Then, through sorting, the data is aligned, so that the natural join operation can be efficiently completed. And since the amount of data sent by each partition is the same or similar, this effectively meets the oblivious requirement, and the communication volume of the entire data table processing operation is relatively low. Furthermore, it effectively solves the problems in the related art that the oblivious algorithm is difficult to implement and has low efficiency, and there is an additional risk of information leakage in the field of encrypted data. In addition, the above implementation process also improves the quality and efficiency of data table processing, further improves the practicability of this method, and is conducive to market promotion and application.
[0200] Figure 5Schematic structural diagram of a data table processing device provided by an embodiment of the present invention; refer to the appendix Figure 5 As shown, this embodiment provides a data table processing device, and this data table processing device is used to perform the data table processing operation shown above Figure 2 Specifically, this data table processing device may include:
[0201] An acquisition module 11, configured to acquire a first data table and a second data table. The first data table and the second data table are stored in different distributed partitions, and the first data table and the second data table include the same attribute items;
[0202] An expansion module 12, configured to expand the second data table based on the attribute items in the first data table to obtain a second expanded table;
[0203] The expansion module 12 is further configured to expand the first data table based on the attribute items in the second data table to obtain a first expanded table;
[0204] A processing module 13, configured to add a data column to the first expanded table to obtain an adjusted expanded table. The parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table;
[0205] The processing module 13 is further configured to associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0206] In some instances, when the expansion module 12 expands the second data table based on the attribute items in the first data table to obtain a second expanded table, the expansion module 12 is configured to perform: perform grouped aggregation processing on the first data table based on the attribute items in the first data table to obtain a first aggregation table, where the first aggregation table includes the repetition degrees of each attribute value located under the attribute item; perform a natural join on the first aggregation table and the second data table to obtain a second join table; expand the second join table to obtain a second expanded table.
[0207] In some instances, when the expansion module 12 expands the second join table to obtain a second expanded table, the expansion module 12 is configured to perform: obtain the data output amount corresponding to each partition; determine the repetition degree of each attribute value as the expansion times for performing expansion operations on each data in the second join table; based on the expansion times and the data output amount, determine the target partition corresponding to each data; expand each data based on the data output amount to obtain expanded data; send the expanded data to the target partition and rewrite the expanded data to obtain the expanded data for constituting the second expanded table.
[0208] In some instances, when the expansion module 12 determines the target partition corresponding to each data in the second connection table based on the number of expansion times and the data output volume, the expansion module 12 is configured to perform: calculating the prefix sum of the number of expansion times to obtain a calculated parameter; determining the processed parameter by using 0 and the calculated parameter; obtaining the quotient of the processed parameter divided by the data output volume; and determining the target partition corresponding to the data based on the sum value of the quotient and 1.
[0209] In some instances, when the expansion module 12 calculates the prefix sum of the number of expansion times to obtain a calculated parameter, the expansion module 12 is configured to perform: obtaining the current partition corresponding to the number of expansion times; calculating the prefix sum of the number of expansion times in the same partition to obtain the first calculation sequence corresponding to each partition, where the first calculation sequence includes multiple processing results; sending the last processing result in the first calculation sequence corresponding to each partition to a preset partition to obtain a processing result sequence stored in the preset partition; and calculating the prefix sum of the processing result sequence to obtain the calculated parameter.
[0210] In some instances, when the expansion module 12 calculates the prefix sum of the processing result sequence to obtain a calculated parameter, the expansion module 12 is configured to perform: calculating the prefix sum of the processing result sequence to obtain a second calculation sequence, where the second calculation sequence includes multiple calculation results; determining the target partition corresponding to each calculation result in the second calculation sequence; and sending the calculation result to the target partition, and summing the calculation result stored in the same target partition and the number of expansion times to obtain the calculated parameter.
[0211] In some instances, when the expansion module 12 determines the target partition corresponding to each calculation result in the second calculation sequence, the expansion module 12 is configured to perform: sorting the multiple calculation results included in the second calculation sequence based on the partition identifier corresponding to each calculation result to obtain result sorting information; and determining the sum value of the current sorting position where each calculation result is located and 1 as the target partition corresponding to the calculation result.
[0212] In some instances, when the expansion module 12 expands each data based on the data output volume to obtain expanded data, the expansion module 12 is configured to perform: determining the total number of bits of data output based on the data output volume; and expanding each data based on the total number of bits of data output to obtain expanded data.
[0213] In some instances, when the expansion module 12 expands each piece of data based on the total number of data output bits to obtain expanded data, the expansion module 12 is used to perform: obtaining all the data that needs to be sent to the same target partition; detecting whether the sum of the data bits of all the data meets the total number of data output bits; if not, performing a data expansion operation based on the sum of the data bits of all the data and a preset padding bit to obtain expanded data, where the number of data bits of the expanded data meets the total number of data output bits.
[0214] In some instances, when the expansion module 12 rewrites the expanded data to obtain the expanded data for constructing the second expansion table, the expansion module 12 is used to perform: obtaining the remainder corresponding to the division operation between the processed parameter and the data output amount; determining the position where the data is stored in the target partition based on the remainder; deleting the redundant preset padding bits based on the position and the data output amount to obtain sorted data; rewriting the preset padding bits included in the sorted data as the previous adjacent data to obtain the target data for constructing the second expansion table.
[0215] In some instances, when the expansion module 12 expands the first data table based on the attribute items in the second data table to obtain the first expansion table, the expansion module 12 is used to perform: performing a grouped aggregation process on the second data table based on the attribute items in the second data table to obtain a second aggregation table, where the second aggregation table includes the repetition degree of each attribute value under the attribute item; performing a natural join on the second aggregation table and the first data table to obtain a first join table; expanding the first join table to obtain the first expansion table.
[0216] In some instances, when the processing module 13 adds a data column to the first expansion table to obtain an adjusted expansion table, the processing module 13 is used to perform: adding a data column with an initial value of 1 to the first expansion table to obtain an intermediate expansion table; comparing the two attribute values of two adjacent groups of data in the intermediate expansion table; when the two attribute values corresponding to two adjacent groups of data are the same, performing a summation process on the data columns in the two adjacent groups of data to obtain a summation data column, and determining the two attribute values and the summation data column as the data for constructing the adjusted expansion table; when the two attribute values corresponding to two adjacent groups of data are different, determining the latter data in the two adjacent groups of data as the data for constructing the adjusted expansion table.
[0217] In some examples, when the processing module 13 associates the second extended table with the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table, the processing module 13 is configured to perform: obtaining a preset rule for sorting the first extended table and the adjusted extended table; using the preset rule to sort the first extended table and the adjusted extended table respectively to obtain a first sorted table and a second sorted table; and associating the first sorted table with the second sorted table to obtain an associated data table corresponding to the first data table and the second data table.
[0218] Figure 5 The device shown can execute Figures 1 - 4 the method of the embodiment shown. For parts not described in detail in this embodiment, reference may be made to the relevant description of the Figures 1 - 4 embodiment shown. For the execution process and technical effects of this technical solution, refer to the description in the Figures 1 - 4 embodiment shown, which will not be elaborated here.
[0219] In a possible design, Figure 5 the structure of the processing device for the data table shown can be implemented as an electronic device, which can be various devices such as a controller, a personal computer, a partition, etc. As Figure 6 shown, the electronic device may include: a first processor 21 and a first memory 22. Among them, the first memory 22 is used to store a program for the corresponding electronic device to execute the data table processing method provided in the Figures 1 - 2 embodiment shown, and the first processor 21 is configured to execute the program stored in the first memory 22.
[0220] The program includes one or more computer instructions. When the one or more computer instructions are executed by the first processor 21, the following steps can be implemented: obtaining a first data table and a second data table, both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; extending the second data table based on the attribute items in the first data table to obtain a second extended table; extending the first data table based on the attribute items in the second data table to obtain a first extended table; adding a data column to the first extended table to obtain an adjusted extended table, and the parameter values included in the data column are used to re-sort the first extended table so that the data in the first extended table is aligned with the data in the second extended table; and associating the second extended table with the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0221] Further, the first processor 21 is also used to execute all or part of the steps in the Figures 1 - 4 embodiment shown above.
[0222] Among them, the structure of the electronic device may further include a first communication interface 23 for the electronic device to communicate with other devices or communication networks.
[0223] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by an electronic device, which includes a program for executing the data table processing method in the above Figures 1 - 4 illustrated embodiment.
[0224] Furthermore, an embodiment of the present invention provides a computer program product, including: a computer-readable storage medium storing computer instructions, when the computer instructions are executed by one or more processors, causing the one or more processors to execute the steps in the data table processing method in the above Figures 1 - 4 illustrated method embodiment.
[0225] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0226] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0227] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0228] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 or more processes and / or boxes Figure 1 or more boxes specified in one box.
[0229] These computer program instructions can also be loaded onto a computer or other programmable device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one process Figure 1 or more processes and / or boxes Figure 1 or more boxes specified in one box.
[0230] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0231] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0232] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0233] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing a data table, characterized in that, Including: Obtain a first data table and a second data table, both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; Expand the second data table based on the attribute items in the first data table to obtain a second expanded table; Expand the first data table based on the attribute items in the second data table to obtain a first expanded table; Add a data column to the first expanded table to obtain an adjusted expanded table, and the parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table; Associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
2. The method according to claim 1, wherein Expanding the second data table based on the attribute items in the first data table to obtain a second expanded table includes: Perform grouped aggregation processing on the first data table based on the attribute items in the first data table to obtain a first aggregated table, and the first aggregated table includes the repetition degrees of each attribute value under the attribute item; Perform a natural join on the first aggregated table and the second data table to obtain a second joined table; Expand the second joined table to obtain a second expanded table.
3. The method according to claim 2, wherein Expanding the second joined table to obtain a second expanded table includes: Obtain the data output amount corresponding to each partition; Determine the repetition degree of each attribute value as the expansion times for performing expansion operations on each data in the second joined table; Based on the expansion times and the data output amount, determine the target partition corresponding to each data; Expand each data based on the data output amount to obtain expanded data; Send the expanded data to the target partition and rewrite the expanded data to obtain the expanded data for forming the second expanded table.
4. The method according to claim 3, wherein Based on the expansion times and the data output amount, determining the target partition corresponding to each data in the second joined table includes: Perform a prefix sum calculation on the expansion times to obtain a calculated parameter; Determine 0 value and the calculated parameter as the processed parameter; Obtain the quotient value of the processed parameter divided by the data output amount; Based on the sum value of the quotient value and 1, determine the target partition corresponding to the data.
5. The method according to claim 4, characterized in that Performing a prefix sum calculation on the expansion times to obtain a calculated parameter includes: Obtain the current partition corresponding to the expansion times; Perform a prefix sum calculation process on the expansion times in the same partition to obtain a first calculation process sequence corresponding to each partition, and the first calculation process sequence includes multiple processing results; Send the last processing result in the first calculation process sequence corresponding to each partition to a preset partition to obtain a processing result sequence stored in the preset partition; Perform a prefix sum calculation on the processing result sequence to obtain a calculated parameter.
6. The method according to claim 5, characterized in that, Performing a prefix sum calculation on the processing result sequence to obtain a calculated parameter includes: Perform a prefix sum calculation on the processed result sequence to obtain a second calculation and processing sequence, where the second calculation and processing sequence includes multiple calculation and processing results; Determine the target partition corresponding to each calculation and processing result in the second calculation and processing sequence; Send the calculation and processing result to the target partition, and sum the calculation and processing result and the expansion times stored in the same target partition to obtain a calculated parameter.
7. The method according to claim 6, characterized in that Determine the target partition corresponding to each calculation and processing result in the second calculation and processing sequence, including: Sort the multiple calculation and processing results included in the second calculation and processing sequence based on the partition identifier corresponding to each calculation and processing result to obtain result sorting information; Determine the sum value of the current sorting position where each calculation and processing result is located and 1 as the target partition corresponding to the calculation and processing result.
8. The method according to claim 3, wherein Expand each data based on the data output volume to obtain expanded data, including: Based on the data output volume, determine the total number of data output bits; Expand each data based on the total number of data output bits to obtain expanded data.
9. The method according to claim 8, wherein Expand each data based on the total number of data output bits to obtain expanded data, including: Obtain all the data that needs to be sent to the same target partition; Detect whether the sum of the data bits of all the data satisfies the total number of data output bits; If not, perform a data expansion operation based on the sum of the data bits of all the data and a preset padding bit to obtain expanded data, where the number of data bits of the expanded data satisfies the total number of data output bits.
10. The method according to claim 3, characterized in that, Rewrite the expanded data to obtain expanded data for constructing the second expansion table, including: Obtain the remainder corresponding to the division operation between the processed parameter and the data output volume; Determine the position where the data is stored in the target partition based on the remainder; Delete the redundant preset padding bits based on the position and the data output volume to obtain sorted data; Rewrite the preset padding bits included in the sorted data as the previous adjacent data to obtain the target data for constructing the second expansion table.
11. The method according to any one of claims 1-10, characterized in that, Expand the first data table based on the attribute items in the second data table to obtain a first expansion table, including: Perform a grouping and aggregation process on the second data table based on the attribute items in the second data table to obtain a second aggregation table, where the second aggregation table includes the repetition degrees of each attribute value under the attribute item; Perform a natural join on the second aggregation table and the first data table to obtain a first join table; Expand the first join table to obtain a first expansion table.
12. The method according to any one of claims 1 to 10, characterized in that Add a data column to the first expansion table to obtain an adjusted expansion table, including: Add a data column with an initial value of 1 to the first expansion table to obtain an intermediate expansion table; Compare the two attribute values of two adjacent groups of data in the intermediate expansion table; When the two attribute values corresponding to two adjacent groups of data are the same, sum the data columns in the two adjacent groups of data to obtain a summed data column, and determine the two attribute values and the summed data column as the data for constructing the adjusted expansion table; When the two attribute values corresponding to two adjacent sets of data are different, the latter data in the two adjacent sets of data is determined as the data for constructing the adjusted extended table.
13. The method according to any one of claims 1 to 10, characterized in that Associate the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table, including Obtain a preset rule for sorting the first extended table and the adjusted extended table; Use the preset rule to sort the first extended table and the adjusted extended table respectively to obtain a first sorted table and a second sorted table; Associate the first sorted table and the second sorted table to obtain an associated data table corresponding to the first data table and the second data table.
14. A processing device for a data table, characterized in that, Including: An acquisition module, configured to acquire a first data table and a second data table, where the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; An extension module, configured to extend the second data table based on the attribute items in the first data table to obtain a second extended table; The extension module is further configured to extend the first data table based on the attribute items in the second data table to obtain a first extended table; A processing module, configured to add a data column to the first extended table to obtain an adjusted extended table, and the parameter values included in the data column are used to re-sort the first extended table so that the data in the first extended table is aligned with the data in the second extended table; The processing module is further configured to associate the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
15. An electronic device, characterized in that, Including: A memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method according to any one of the above claims 1-13 is implemented.
16. A computer program product, characterized in that, Including: A computer program, when the computer program is executed by a processor of an electronic device, causes the processor to execute the steps in the method according to any one of the above claims 1-13.