Data table processing
By extending the attribute item and aligning data columns for data tables in a distributed computing environment, the problem of inefficient multi-partition data processing is solved, and efficient and secure data table association and processing is achieved.
Patent Information
- Application Number
- PCT/IB2024/063264
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-17
AI Technical Summary
In a distributed computing environment, it is difficult for the prior art to efficiently process different data tables stored in multiple partitions, resulting in low data processing efficiency and large traffic. Especially under the requirements of natural connection operators and security, the existing solutions have insufficient universality and efficiency problems.
By obtaining the attribute items of the first data table and the second data table, they are expanded respectively, data columns are added to align the data, and correlation processing is performed to form an associated data table to facilitate subsequent data processing operations.
It improves the quality and efficiency of data processing, reduces data traffic, enhances the stability and security of data processing, and is suitable for data table processing in distributed computing environments.
Smart Images

Figure IB2024063264_17072025_PF_FP_ABST
Abstract
Description
Technical Field of Data Table Processing
[0001] This application relates to the technical field of data processing, and particularly to the processing of data tables. Background Art
[0002] With the advent of the big data era, the distributed computing method, as an effective solution to the insufficient computing power for big data, has gradually become popular. Specifically, distributed computing usually requires splitting data and computing tasks across multiple partitions for processing to improve the quality and efficiency of data processing.
[0003] However, since different data is stored in different partitions, when a data processing task involves data in multiple partitions, it is often necessary to perform data processing operations on each involved partition in sequence to obtain the data processing result. This greatly increases the data processing process and reduces the data processing efficiency. Summary of the Invention
[0004] The embodiments of this application provide a method, apparatus, device, and computer program product for processing data tables, which can perform association operations on different data tables stored in multiple distributed partitions to obtain an associated data table, and then perform corresponding data processing operations based on the associated data table, capable of improving the quality and efficiency of data processing.
[0005] In a first aspect, the embodiments of this application provide a method for processing data tables, including: obtaining a first data table and a second data table, where both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; expanding the second data table based on the attribute items in the first data table to obtain a second expanded table; expanding the first data table based on the attribute items in the second data table to obtain a first expanded table; adding a data column to the first expanded table to obtain an adjusted expanded table, where the parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table; associating the second expanded table with the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0006] Second aspect, an embodiment of the present application provides a processing device for a data table, including: an acquisition module, configured to acquire a first data table and a second data table, where the first data table and the second data table are stored in a plurality of distributed partitions, and the first data table and the second data table include the same attribute items; an expansion module, configured to expand the second data table based on the attribute items in the first data table to obtain a second expanded table; the expansion module is further configured to expand the first data table based on the attribute items in the second data table to obtain a first expanded table; a processing module, configured to add a data column to the first expanded table to obtain an adjusted expanded table, where the parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table; the processing module is further configured to associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0007] Third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the processing method of the data table in the above first aspect is implemented.
[0008] Fourth aspect, an embodiment of the present application provides a computer storage medium, used to store a computer program, and when the computer program is executed by a computer, the processing method of the data table in the above first aspect is implemented.
[0009] Fifth aspect, an embodiment of the present application provides a computer program product, including: a computer program, when the computer program is executed by a processor of an electronic device, the processor is caused to execute the steps in the processing method of the data table in the above first aspect.
[0010] The processing method, device, equipment and computer program product for the data table provided in this embodiment, by acquiring the first A first data table and a second data table, expand the second data table based on the attribute items in the first data table to obtain a second expanded table; expand the first data table based on the attribute items in the second data table to obtain a first expanded table; then add a data column to the first expanded table to obtain an adjusted expanded table, and associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table, effectively realizing the association operation of different data tables stored in multiple distributed partitions, and then corresponding data processing operations can be performed based on the obtained associated data table, which greatly improves the quality and efficiency of data processing and ensures the practicability of this method, facilitating market promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] FIG. 1 is a schematic diagram of the principle of a method for processing a data table provided by an embodiment of the present application;
[0013] FIG. 2 is a schematic flowchart of a method for processing a data table provided by an embodiment of the present application;
[0014] FIG. 3 is a schematic flowchart of expanding the second connection table to obtain a second expanded table provided by an embodiment of the present application;
[0015] FIG. 4 is a schematic flowchart of calculating the prefix sum of the expansion times to obtain a calculated parameter provided by an embodiment of the present application;
[0016] FIG. 5 is a schematic structural diagram of a data table processing device provided by an embodiment of the present application;
[0017] FIG. 6 is a schematic structural diagram of an electronic device corresponding to the data table processing device shown in FIG. 5. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0019] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the" and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0020] It should be understood that the term "and / or" used herein is only a correlative relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0021] Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "when...", "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0022] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a commodity or system including a series of elements not only includes those elements, but also includes other elements not specifically listed, or further includes elements inherent to such commodity or system. Without further limitation, an element defined by the statement "including an..." does not exclude the existence of another identical element in the commodity or system including the said element.
[0023] In addition, the step timings in the following method embodiments are only examples and not strictly limited. Term definitions
[0024] Trusted Execution Environment (TEE for short): It is an isolated execution environment with security functions, where data can be securely calculated without worrying about being stolen.
[0025] TEE-based encrypted database: A common encrypted database solution. Users encrypt data and upload it to the database, and upload the encryption key to the TEE. During calculation, the data in the database is securely decrypted in the TEE for operation, and the operation result is re-encrypted before leaving the TEE and then sent to the user. Since the data exists in encrypted form except inside the TEE during the data flow process, the data security is protected to a certain extent.
[0026] Data Shuffle: An operation that redistributes data in a certain way according to data processing requirements in a distributed environment. The purpose is to facilitate subsequent data processing, and the implementation method is generally that each partition sends the results to other partitions as needed.
[0027] Data aggregation operation (reduceByKey): Used to merge values with the same key together.
[0028] Oblivious algorithm: Often used to describe an algorithm. In a single-machine scenario, an oblivious algorithm means that the read or write of memory locations by the algorithm is completely independent of the input of the algorithm and only related to some public information (such as the length of the input or output data). In a distributed scenario, an oblivious algorithm means that when data is exchanged between partitions, the amount of data sent is independent of the input, and whether it is required that the read or write of memory locations by each partition during local calculation is completely independent of the input of the algorithm depends on the context. In this proposal, there is no oblivious requirement for the algorithms of local calculation in partitions.
[0029] Dummy data: Meaningless elements used as padding bits in an oblivious algorithm. The purpose is to fill the amount of data to a specified size to prevent the actual size of the data volume from leaking privacy information.
[0030] In order to understand the specific implementation process of the technical solution in this embodiment, the related technologies will be briefly described below.
[0031] In the field of encrypted databases based on the trusted execution environment TEE, data is protected by encryption. When performing calculations, it usually needs to be securely decrypted and then calculated in the trusted area. In the era of big data, it is difficult for a single machine to complete the calculation of a large amount of data, and the newly added encryption and decryption operations will undoubtedly further increase the load of the calculation operation. For this reason, distributed computing has gradually become popular as an effective solution to the lack of computing power for big data calculations. Distributed computing usually requires data and computing tasks to be distributed to multiple partitions for processing to improve the quality and efficiency of data processing.
[0032] However, even if the data has been encrypted and protected by an encrypted database, the amount of data transmitted and communicated between partitions in a distributed manner during data calculations still poses a serious privacy leakage problem. For example, when a distributed open source processing system executes a data aggregation operator (reduceByKey), the first step is to concentrate values with the same key into the same partition through a data redistribution operator (shuffle) for subsequent operations. By observation, during the data redistribution process of shuffle, the amount of data received by each partition can roughly estimate the data distribution of the original data on the key. Therefore, unlike non-encrypted databases, these operators need to be redesigned into oblivious algorithms so that the amount of data sent and communicated between partitions during the operation process is independent of the data input, thereby achieving all-round protection of the data.
[0033] Common operators for database queries may include: filter operator, group by aggregate operator, project operator, order by operator, etc. It is not difficult to modify the algorithm.
[0034] Related technology 1 implements a distributed database query system that meets oblivious requirements. The database query system can support sorting, filtering, and aggregation operations. However, the query system does not support general natural join operators, but only supports primary key-foreign key type natural join operators, which lacks versatility.
[0035] Related technology 2 implements an efficient distributed database operator computing solution, which optimizes filtering and aggregation operations, removes the reliance on sorting operations, and also supports general natural join operations. However, the solution has the following disadvantages in terms of security and efficiency.
[0036] In terms of security, the algorithm relies on two publicly available values, a and 0, which represent the maximum repetition degrees of the two tables participating in the natural join on the keys of the corresponding join conditions respectively. If these two values themselves are also private information, then this algorithm will not be applicable, otherwise it will violate the security requirements.
[0037] In terms of efficiency, the algorithm will fill each partition with some placeholder data (dummy) up to a total of M / p + a0, where M is the number of rows in the output result of the natural join operator and p is the total number of partitions. An algorithm that achieves computational balance should have exactly M / p data in each partition, meaning that the operator provided by this scheme fills as many as a0 data to meet the requirements of the oblivious algorithm. Before and after filling, the algorithm corresponds to two rounds of complete communication, and the total communication volume is approximately N + 2M + 2a0p. In the worst case, a|3 = N 2 , where N is the number of rows of the two tables participating in the natural join. Therefore, this scheme is actually very inefficient in many scenarios; moreover, for common natural join operators, it is not an obvious thing to transform them into oblivious algorithms. If the algorithm is not designed properly, it will lead to a significant increase in communication volume, a decrease in computational efficiency, and even deviate from the original intention of distributed computing.
[0038] To solve the above technical problems, this embodiment provides a method, device, and equipment for processing data tables. As shown in Figure 1, the execution subject of the data table processing method provided in this embodiment can be a data table processing device. It should be noted that this data table processing device can be implemented as a terminal device, a personal computer, a tablet computer, a local server, or a cloud server. At this time, when the data table processing device is implemented as a cloud server, the data table processing method can be executed in the cloud. Several computing nodes (cloud servers) can be deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, multiple computing nodes can be organized to provide a certain service. Of course, a single computing node can also provide one or more services. The way the cloud provides this service can be to provide a service interface externally, and users can call this service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).
[0039] The processing device for the data table is communicatively connected to the client. Herein, the client is used for a user to perform an application so as to be able to implement the processing operation of the data table. The above-mentioned client can be any computing device with certain data transmission capabilities. Specifically, in implementation, the client can be a mobile phone, a personal computer (PC), a tablet computer, a set application program, etc. In addition, the basic structure of the client may include: at least one processor. The number of processors depends on the configuration and type of the client. The client may also include a memory, which can be volatile, for example: Random Access Memory (RAM), or non-volatile, for example: Read-Only Memory (ROM), flash memory, etc., or may also include both types at the same time. Usually, an operating system (OS), one or more application programs, and program data, etc. are stored in the memory. In addition to the processing unit and the memory, the client also includes some basic configurations, such as a network card chip, a 10 bus, a display component, and some peripheral devices, etc. Optionally, some peripheral devices may include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated herein. (Random Access Memory, abbreviated as RAM), or non-volatile, for example: Read-Only Memory (ROM), flash memory, etc., or may also include both types at the same time. Usually, an operating system (OS), one or more application programs, and program data, etc. are stored in the memory. In addition to the processing unit and the memory, the client also includes some basic configurations, such as a network card chip, a 10 bus, a display component, and some peripheral devices, etc. Optionally, some peripheral devices may include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated herein.
[0040] The processing device for the data table refers to a device that can provide the processing operation of the data table in a network virtual environment, usually referring to a device that uses the network for information planning and the processing operation of the data table. Physically, the processing device for the data table can be any device that can provide computing services, respond to the processing request of the data table, and can perform the processing operation of the data table based on the processing request of the data table, for example: it can be a cluster server, a conventional server, a cloud server, a cloud host, a virtual center, etc. The composition of the processing device for the data table mainly includes a processor, a hard disk, a memory, a system bus, etc., which is similar to a general computer architecture.
[0041] In the above-described embodiment of the present invention, the client is connected to the data table processing device through a network, and this network connection can be a wireless or wired network connection. If the client can communicate with the data table processing device, the network mode of the mobile network can be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, 6G, etc.
[0042] In an embodiment of the present application, the client is used to generate or obtain a processing request for a data table. Specifically, the client can display a human-computer interaction interface, obtain the execution operation input by the user in the human-computer interaction interface, and generate or obtain a processing request for the data table based on the execution operation. In order to implement the processing operation of the data table, the processing request for the data table can be sent to the data table processing device.
[0043] The data table processing device is used to obtain the processing request for the data table sent by the client, and then can determine a first data table and a second data table corresponding to the processing request for the data table. Among them, the first data table and the second data table are both complete data tables and are stored in multiple distributed partitions. That is, the multiple data sub-tables constituting the first data table are evenly distributed and stored in each partition, and the multiple data sub-tables constituting the second data table are evenly distributed and stored in each partition. And, for the first data table and the second data table, the first data table and the second data table may include the same attribute items, and the same attribute items may include one category, two categories, or more categories, etc.
[0044] Since the first data table and the second data table are stored in multiple distributed partitions, in order to improve the quality and efficiency of data processing and minimize the data communication volume, after obtaining the first data table and the second data table corresponding to the processing request for the data table, an expansion operation can be performed on the second data table based on the attribute items in the first data table to obtain a second expanded table. Similarly, an expansion operation can be performed on the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0045] After performing expansion operations on the first data table and the second data table, the data in the obtained first expanded table and second expanded table is not aligned. In order to stably process the first data table and the second data table, a data column can be added to the first expanded table, so that an adjusted expanded table aligned with the data in the second expanded table can be obtained. After obtaining the second expanded table and the adjusted expanded table, an association operation can be performed on the second expanded table and the adjusted expanded table, so that an associated data table corresponding to the first data table and the second data table can be obtained. Then, corresponding data processing operations can be performed based on the associated data table. In this way, when data processing operations need to be performed based on the first data table and the second data table, the data processing operations can be directly performed based on the associated data table, thereby effectively improving the quality and efficiency of data processing and facilitating market promotion and application.
[0046] The following will, with reference to the accompanying drawings, elaborate on some embodiments of the present application. Under the condition that there is no conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. Additionally, the sequence of steps in the following method embodiments is only an example and is not strictly limited.
[0047] FIG. 2 is a schematic flowchart of a method for processing a data table provided by an embodiment of the present application; as shown in FIG. 2, this embodiment provides a method for processing a data table. The execution subject of this method can be a data table processing device. It can be understood that the data table processing device can be implemented as software, or a combination of software and hardware. Specifically, when the data table processing device is implemented as hardware, it can specifically be various electronic devices capable of performing data table processing operations, including but not limited to tablet computers, personal computers (PCs), servers, and the like. When the data table processing device is implemented as software, it can be installed in the above-mentioned exemplified electronic devices. Based on the above data table processing device, the method for processing a data table in this embodiment may include the following steps S201 to S205.
[0048] Step S201: Obtain a first data table and a second data table. Both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items.
[0049] Step S202: Expand the second data table based on the attribute items in the first data table to obtain a second expanded table.
[0050] Step S203: Expand the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0051] Step S204: Add a data column to the first extended table to obtain an adjusted extended table. The parameter values included in the data column are used to reorder the first extended table so that the data in the first extended table is aligned with the data in the second extended table. The numerical values are used to reorder the first extended table so that the data in the first extended table is aligned with the data in the second extended table.
[0052] Step S205: Associate the second extended table with the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0053] The specific implementation principles and implementation effects of the above steps will be described in detail below.
[0054] Step S201: Obtain a first data table and a second data table. Both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items.
[0055] Among them, when the user has a processing requirement for the data table, the data table processing device can obtain the first data table and the second data table. In some instances, the first data table and the second data table can be obtained through human-computer interaction operations. At this time, obtaining the first data table and the second data table may include: displaying a human-computer interaction interface for implementing data table processing operations; displaying multiple data tables in the human-computer interaction interface; obtaining the execution operations input by the user for the multiple data tables, and obtaining the first data table and the second data table based on the execution operations. For example, when the data table processing device stores multiple data tables including: data table a, data table b, data table c, data table d, data table e, when the user has a processing requirement for data table b and data table c among the above data tables, the user can input a selection operation for data table b and data table c among the above data tables, so as to obtain the data tables b and c to be processed.
[0056] In other instances, the first data table and the second data table can not only be obtained through human-computer interaction operations, but also be obtained through analysis and processing. At this time, obtaining the first data table and the second data table may include: obtaining a client or a preset device communicatively connected to the data table processing device; actively or passively obtaining a data table processing request through the client or the preset device; determining the first data table and the second data table corresponding to the data table processing request, thereby effectively ensuring the accuracy and reliability of obtaining the first data table and the second data table.
[0057] It should be noted that the first data table and the second data table can be complete data tables respectively. The above-mentioned first data table and second data table are both stored in multiple distributed partitions. Among them, a partition can be a computing unit under distribution, which can be a machine, such as a server; or a partition can also be a core of a machine processor. Each partition has independent computing, storage, and network resources, and the partitions cooperate to complete computing tasks through network communication.
[0058] In addition, for the first data table and the second data table, the first data table and the second data table include the same attribute items. For example, the first data table is R(A,B). The above-mentioned first data table can be composed of the data included in sub-table R1, sub-table R2, and sub-table R3. Sub-table R1 can be stored in partition a, sub-table R2 can be stored in partition b, and sub-table R3 can be stored in partition c. That is, the first data table R(A,B) can be distributed and stored in the above-mentioned multiple partitions; similarly, the second data table is S(B,C), and the second data table can be composed of the data included in sub-table S1, sub-table S2, and sub-table S3. Sub-table S1 can be stored in partition a, sub-table S2 can be stored in partition b, and sub-table S3 can be stored in partition c. And, the above-mentioned first data table R(A,B) and second data table S(B,C) can include the same attribute item B. It can be understood that the same attribute items existing between the first data table and the second data table can be not only of 1 type, and those skilled in the art can flexibly adjust the number of the same attribute items existing between the first data table and the second data table according to needs.
[0059] When there are the same attribute items between the first data table and the second data table, the first data table and the second data table can be associated and processed. When there are no same attribute items between the first data table and the second data table, it is impossible to realize the association and processing operation of the first data table and the second data table.
[0060] Step S202: Expand the second data table based on the attribute items in the first data table to obtain a second extended table.
[0061] After obtaining the first data table and the second data table, in order to realize the association operation of the data tables and ensure the data security in the first data table and the second data table, the second data table can be expanded based on the attribute items in the first data table. Among them, the attribute items in the first data table are the common attribute items included in the first data table and the second data table, so as to obtain a second extended table.
[0062] In some instances, the extension operation can be implemented by a preset machine learning model, a neural network model, or a preset algorithm. At this time, when extending the second data table based on the attribute items in the first data table, obtaining the second extended table may include: obtaining a pre-trained machine learning model or neural network model; inputting the attribute items in the first data table and the second data table into the machine learning model or neural network model to obtain the second extended table output by the machine learning model or neural network model.
[0063] In other instances, the extension operation can not only be implemented by a machine learning model or a neural network model, but also be implemented based on the repetition degree of the attribute items in the first data table. At this time, when extending the second data table based on the attribute items in the first data table, obtaining the second extended table may include: performing grouped aggregation processing on the first data table based on the attribute items in the first data table to obtain a first aggregated table, where the first aggregated table includes the repetition degree of each attribute value under the attribute item; performing a natural join on the first aggregated table and the second data table to obtain a second joined table; and extending the second joined table to obtain the second extended table.
[0064] Specifically, after obtaining the first data table and the second data table, grouped aggregation processing can be performed on the first data table based on the attribute items in the first data table to obtain a first aggregated table, and the first aggregated table may include the repetition degree of each attribute value under the attribute item.
[0065] For example, the first data table R(A,B) can be {(al,bl),(al,b2),(a2,b2),(al,b3),(a2,b4)}, and the second data table S(B,C) can be {(bl,cl),(bl,c2),(b2,c2),(b2,c3),(b3,c4)}. As can be seen from the above, there is the same attribute item B between the first data table and the second data table, that is, the attribute item B in the first data table is the value corresponding to the b column. Then, by performing a statistical operation on the repetition degree of the attribute items in the first data table, the repetition degree of each attribute value in the above attribute items can be obtained, that is, the repetition degree of bl is 1, the repetition degree of b2 is 2, the repetition degree of b3 is 1, and the repetition degree of b4 is 1. Grouped aggregation processing is performed on the first data table based on the repetition degree of the attribute values in the first data table to obtain the first aggregated table R(B,D), and the first aggregated table R(B,D) can be {(bl,1),(b2,2),(b3,1),(b4,1)}.
[0066] After obtaining the first aggregation table and the second data table, a natural join can be performed on the first aggregation table and the second data table to obtain a second joined table. In some examples, the natural join operation can be implemented by a preset machine learning model or a neural network model. For example, when the first aggregation table R(B, D) is {(b1, 1), (b2, 2), (b3, 1), (b4, 1)} and the second data table S(B, C) is {(b1, c1), (b1, c2), (b2, c2), (b2, c3), (b3, c4)}, by performing a natural join operation on the first aggregation table and the second data table, the second joined table S2(B, C, D) can be obtained, and the second joined table S2(B, C, D) can be {(b1, c1, 1), (b1, c2, 1), (b2, c2, 2), (b2, c3, 2), (b3, c4, 1)}. After obtaining the second joined table, an expansion operation can be performed on the second joined table, effectively implementing an expansion operation on the second data table, and thus a second expanded table can be stably obtained.
[0067] Step S203: Expand the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0068] After obtaining the first data table and the second data table, in order to implement the association operation of the data tables and ensure the data security in the first data table and the second data table, an expansion operation can be performed on the first data table based on the attribute items in the second data table, where the attribute items in the second data table are the common attribute items included in the first data table and the second data table, and thus a first expanded table can be obtained.
[0069] In some examples, expanding the first data table based on the attribute items in the second data table to obtain a first expanded table may include: performing grouped aggregation processing on the second data table based on the attribute items in the second data table to obtain a second aggregation table, where the second aggregation table includes the repetition degrees of each attribute value located under the attribute item; performing a natural join on the second aggregation table and the first data table to obtain a first joined table; and expanding the first joined table to obtain a first expanded table.
[0070] It should be noted that the specific implementation manner, implementation effect, and implementation principle of the expansion operation on the first data table in this embodiment are similar to those of the expansion operation on the second data table in step S202 above. For specific reference, please refer to the above description and will not be elaborated here.
[0071] Step S204: Add a data column to the first extended table to obtain an adjusted extended table. The parameter values included in the data column are used to reorder the first extended table so that the data in the first extended table is aligned with the data in the second extended table.
[0072] For the first data table and the second data table, after the extension operations on the first data table and the second data table, the data in the obtained first extended table and the data in the second extended table are shuffled. At this time, the data in the first extended table and the data in the second extended table are often not aligned. Therefore, in order to enable the processing operation of the data table, after obtaining the first extended table, a data column can be added to the first extended table to obtain an adjusted extended table. The parameter values included in the data column are used to perform a reordering operation on the first extended table so that the data in the first extended table is aligned with the data in the second extended table.
[0073] In some instances, adding a data column to the first extended table to obtain an adjusted extended table may include: adding a data column with an initial value of 1 to the first extended table to obtain an intermediate extended table; comparing the two attribute values of two adjacent groups of data in the intermediate extended table; when the two attribute values corresponding to two adjacent groups of data are the same, perform a summation process on the data columns in the two adjacent groups of data to obtain a summation data column, and determine the two attribute values and the summation data column as the data for constructing the adjusted extended table; when the two attribute values corresponding to two adjacent groups of data are different, determine the latter data in the two adjacent groups of data as the data for constructing the adjusted extended table. Compare the two attribute values of two adjacent groups of data in the extended table, that is, compare the attribute values "a" and "b" in the data {al, bl, 1} and the data {a2, b2, l}. When al = a2 and bl = b2, a summation process can be performed on the data columns in the two adjacent groups of data, so that a summation data column can be obtained. The summation data column is {al, bl, 2}. Then, the two attribute values and the summation data column can be determined as the data for constructing the adjusted extended table, which can ensure the accuracy and reliability of obtaining the adjusted extended table.
[0075] Correspondingly, when the two attribute values corresponding to two adjacent groups of data are different, that is, al ≠ a2 or bl ≠ b2, the latter data in the two adjacent groups of data can be determined as the data for constructing the adjusted extended table. That is, the data for constructing the adjusted extended table at this time is {a2, b2, l}, which effectively ensures the accuracy and reliability of obtaining the adjusted extended table.
[0076] It should be noted that, in order to align the data in the first extended table and the second extended table, not only can data columns be added to the first extended table, but also data columns can be added to the second extended table, so as to obtain an adjusted extended table corresponding to the second extended table. At this time, step S204 can be changed to "add a data column to the second extended table to obtain an adjusted extended table, and the parameter values included in the data column are used to reorder the second extended table so that the data in the first extended table and the second extended table are aligned".
[0077] Step S205: Associate the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0078] Since the data in the adjusted extended table is aligned with the data in the second data table, in order to accurately implement the processing operations on the first data table and the second data table, after obtaining the adjusted extended table and the second extended table, an association operation can be performed on the second extended table and the adjusted extended table, so as to obtain an associated data table corresponding to the first data table and the second data table.
[0079] In some instances, the association operation can be implemented by a pre-trained machine learning model or neural network model. At this time, associating the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table may include: obtaining a pre-trained machine learning model or neural network model; inputting the second extended table and the adjusted extended table into the machine learning model or neural network model to obtain the associated data table corresponding to the first data table and the second data table output by the machine learning model or neural network model. This associated data table is the result obtained after performing a natural join operation on the first data table and the second data table, which effectively ensures the accuracy and reliability of determining the associated data table. For the obtained associated data table, the associated data table can be stored in multiple distributed partitions, which can ensure the quality and effect of storing the associated data table.
[0080] In some other instances, not only can the association operation of the data table be implemented through a pre-trained machine learning model or neural network model, but also the association operation of the data table can be implemented by using preset rules. At this time, the first extended table and the adjusted extended table are associated to obtain an association data table corresponding to the first data table and the second data table, including: obtaining a preset rule for sorting the first extended table and the adjusted extended table; using the preset rule to sort the first extended table and the adjusted extended table respectively to obtain a first sorted table and a second sorted table; and associating the first sorted table and the second sorted table to obtain an association data table corresponding to the first data table and the second data table.
[0081] After obtaining the first extended table, a preset rule for sorting the first extended table can be obtained first. The preset rule can be a lexicographical order or other sequence stored in a preset area or a preset device. Similarly, after obtaining the adjusted extended table, a preset rule for sorting the adjusted extended table can be obtained first. The preset rule can be a lexicographical order or other sequence pre-stored in a preset area or a preset device. It should be noted that the preset rule for sorting the first extended table and the preset rule for sorting the adjusted extended table can be obtained synchronously or asynchronously.
[0082] After obtaining the preset rules for sorting the first extended table and the adjusted extended table, the preset rules can be used to sort the first extended table and the adjusted extended table respectively, so that a first sorted table and a second sorted table can be obtained. After obtaining the first sorted table and the second sorted table, an association operation can be performed on the first sorted table and the second sorted table, so that an association data table corresponding to the first data table and the second data table can be obtained. For example, when the i-th data in the first data table is (b, c) and the i-th data in the second data table is (a, bl, f), where b = bl, when performing an association operation on the first data table and the second data table, an association data table can be obtained, and the i-th item data in the association data table is (a, b, c), thus effectively implementing the association operation of the data table.
[0083] In some other instances, after obtaining the association data table corresponding to the first data table and the second data table, to improve the practicability of the method, a data query operation can be implemented based on the obtained association data table. At this time, the method in this embodiment can further include: obtaining a data query request; determining the association data table corresponding to the data query request; and obtaining a data query result based on the association data table.
[0084] Specifically, when the user has a data query requirement, the processing device of the data table can obtain a data query request. In some instances, the data query request can be obtained through a human-computer interaction operation, or the data query request can be obtained actively or passively through a preset device. After obtaining the data query request, the data query request can be analyzed and processed to determine the associated data table corresponding to the data query request. After obtaining the associated data table, a corresponding data query operation can be performed based on the associated data table, so that a data query result can be obtained. In this way, the data query operation based on the associated data table is effectively realized, and the practicability of this method is further improved.
[0085] The data table processing method provided in this embodiment obtains a first data table and a second data table, expands the second data table based on the attribute items in the first data table to obtain a second extended table; expands the first data table based on the attribute items in the second data table to obtain a first extended table; then adds a data column to the first extended table to obtain an adjusted extended table, and associates the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table, effectively realizing the association operation on different data tables stored in multiple distributed partitions. Then, corresponding data processing operations can be performed based on the obtained associated data table, which greatly improves the quality and efficiency of data processing and ensures the practicability of this method, facilitating market promotion and application.
[0086] Figure 3 is a schematic flowchart of expanding the second connection table to obtain a second extended table provided in an embodiment of the present application; on the basis of the above embodiment, as shown in Figure 3, the implementation manner of expanding the second connection table in this embodiment is not limited. In some instances, a pre-trained machine learning model or neural network model can be used to expand the second connection table, or the second connection table can be expanded based on the data output volume corresponding to each partition. At this time, expanding the second connection table to obtain a second extended table in this embodiment may include steps S301 to S305.
[0087] Step S301: Obtain the data output volume corresponding to each partition.
[0088] Since the expansion operation of the second connection table is related to the data output volume corresponding to each partition, in order to accurately implement the expansion operation on the second connection table, the data output volume corresponding to each partition can be obtained. In some instances, the data output volume corresponding to each partition can be data that is pre-configured by the user and stored in a preset area or a preset device. At this time, by accessing the preset area or the preset device, the data output volume corresponding to each partition can be obtained.
[0089] Or, in other instances, the data output volume corresponding to each partition can not only be pre-configured by the user, but also be determined based on the number of partitions and the total output data volume. At this time, obtaining the data output volume corresponding to each partition can include: obtaining the total output data volume M and the number of partitions p; based on the total output data volume M and the number of partitions p, determining the data output volume m = M / p corresponding to each partition. For example, when the total output data volume M is 18 and the number of partitions p is 3, the data output volume corresponding to each partition can be determined to be 18 / 3 = 6. At this time, the multiple partitions perform corresponding data processing operations on the data using an equalization algorithm.
[0090] Step S302: Determine the repetition degree of each attribute value as the expansion times for expanding each data in the second connection table.
[0091] After performing a natural join operation on the first aggregation table and the second data table, a second join table can be obtained. The second join table may include the duplication degrees of various attribute values. Then, the duplication degrees of various attribute values can be determined as the expansion times for performing expansion operations on each piece of data in the second join table. For example, when the second join table S2(B, C, D) is {(b1, c1, 1), (b1, c2, 1), (b2, c2, 2), (b2, c3, 2), (b3, c4, 1)}, the "1" in (b1, c1, 1) can be determined as the expansion times for performing an expansion operation on the data "(b1, c1)", that is, the data "(b1, c1)" needs to be expanded 1 time. Similarly, the "1" in (b1, c2, 1) is determined as the expansion times for performing an expansion operation on the data "(b1, c2)", the "2" in (b2, c2, 2) is determined as the expansion times for performing an expansion operation on the data "(b2, c2)", the "2" in (b2, c3, 2) is determined as the expansion times for performing an expansion operation on the data "(b2, c3)", and the "1" in (b3, c4, 1) is determined as the expansion times for performing an expansion operation on the data "(b3, c4)". This effectively ensures the accuracy and reliability of determining the expansion times for performing expansion operations on each piece of data in the second join table.
[0092] Step S303: Based on the expansion times and the data output volume, determine the target partition corresponding to each piece of data.
[0093] Since each piece of data in the data table is located in different distributed partitions, when performing an expansion operation, each piece of data needs to be sent to the corresponding target partition. Therefore, in order to be able to perform an expansion operation on the second join table, after obtaining the expansion times and the data output volume, the expansion times and the data output volume can be analyzed and processed to determine the target partition corresponding to each piece of data. In some instances, the target partition can be determined by analyzing and processing the expansion times and the data output volume through a pre-trained machine learning model or neural network model. At this time, based on the expansion times and the data output volume, determining the target partition corresponding to each piece of data may include: obtaining a pre-trained machine learning model or neural network model; inputting the expansion times, the data output volume, and the second join table into the machine learning model or neural network model to obtain the target partition corresponding to each piece of data output by the machine learning model or neural network model.
[0094] In some other instances, not only can the target partition corresponding to each data be determined through a pre-trained machine learning model or neural network model, but also the prefix sum algorithm can be used to analyze and process the expansion times and data output amounts to determine the target partition corresponding to each data. At this time, based on the expansion times and data output amounts, determining the target partition corresponding to each data in the second connection table may include: performing a prefix sum calculation on the expansion times to obtain the calculated parameters; determining the 0 value and the calculated parameters as the processed parameters; obtaining the quotient of the processed parameters divided by the data output amount; and determining the target partition corresponding to the data based on the sum value of the quotient and 1.
[0095] For example, when the expansion times are (1, 3, 1, 1, 5, 2, 1, 1, 3), after performing a prefix sum calculation on the expansion times, the calculated parameters obtained may be (1, 4, 5, 6, 11, 13, 14, 15). Then, the 0 value and the calculated parameters can be determined as the processed parameters, and the processed parameters are (0, 1, 4, 5, 6, 11, 13, 14, 15). When the number of partitions is 3 and the output data amount of each partition is 6, the quotient of each of the above processed parameters divided by the data output amount can be obtained. That is, the quotient of 0 / 6 is 0, the quotient of 1 / 6 is 0, the quotient of 4 / 6 is 0, the quotient of 5 / 6 is 0, the quotient of 6 / 6 is 1, the quotient of 11 / 6 is 1, the quotient of 13 / 6 is 2, the quotient of 14 / 6 is 2, and the quotient of 15 / 6 is 2.
[96] After obtaining the quotient corresponding to each of the above processed parameters, the sum value of the quotient and 1 can be obtained, so that a sum value sequence (1, 1, 1, 1, 2, 2, 3, 3, 3) corresponding to multiple calculated parameters can be obtained. From the above, it can be determined that the target partition corresponding to the data corresponding to the sum value "1" is the partition with the preset identifier "1", the target partition corresponding to the data corresponding to the sum value "2" is the partition with the preset identifier "2", and the target partition corresponding to the data corresponding to the sum value "3" is the partition with the preset identifier "3", thus effectively ensuring the accurate and reliable determination of the target partition corresponding to the data.
[97] It should be noted that the target partition corresponding to the data can be determined not only by the sum of the quotient value and 1, but also by first obtaining the difference between the processed parameter and 1, then obtaining the ratio of the difference to the data output volume, performing a ceiling operation on the ratio value, and finally determining the target partition corresponding to the data based on the parameter after the ceiling operation. This can also ensure the accuracy and reliability of determining the target partition.
[98] Step S304: Expand each data based on the data output volume to obtain expanded data.
[99] In order to implement the data expansion operation, after obtaining the data output volume, the data expansion operation can be performed on each data based on the data output volume to obtain expanded data. In some instances, the expansion operation can be implemented by a pre-trained machine learning model or a neural network model. At this time, expanding each data based on the data output volume to obtain expanded data may include: obtaining a pre-trained machine learning model or a neural network model; inputting the data output volume and each data into the machine learning model or the neural network model to obtain the expanded data output by the machine learning model or the neural network model.
[100] In other instances, not only can the data expansion operation be implemented by a pre-trained machine learning model or a neural network model, but also the data expansion operation can be implemented based on the total number of data output bits. At this time, expanding each data based on the data output volume to obtain expanded data may include: determining the total number of data output bits based on the data output volume; expanding each data based on the total number of data output bits to obtain expanded data.
[101] Among them, since the data expansion operation is related to the total number of data output bits, in order to accurately implement the data expansion operation, after obtaining the data output volume, the total number of data output bits can be determined based on the data output volume. In some instances, determining the total number of data output bits based on the data output volume may include: obtaining a pre-configured mapping relationship between the data output volume and the total number of data output bits, and determining the total number of data output bits based on the data output volume and the mapping relationship.
[102] In some other examples, the total number of bits of data output can be determined not only through a preset mapping relationship, but also through a preset formula. At this time, based on the data output volume, determining the total number of bits of data output may include: performing a square root processing operation on the data output volume to obtain a processed value; determining the sum value between the data output volume and the processed value, and performing a rounding operation on the sum value to obtain the total number of bits of data output. For example, when the data output volume is 6, the total number of bits of data output can be obtained through the calculation of [6 + √6] = 9, which effectively ensures the accuracy and reliability of determining the total number of bits of data output.
[103] After obtaining the total number of bits of output data, each data can be extended based on the total number of bits of data output to obtain extended data. In some examples, each data can be extended not only based on a pre-trained machine learning model or neural network model, but also by combining a preset padding bit to implement the data extension operation. At this time, extending each data based on the total number of bits of data output to obtain extended data may include: obtaining all the data to be sent to the same target partition; detecting whether the sum of the data bits of all the data meets the total number of bits of data output; if not, performing a data extension operation based on the sum of the data bits of all the data and the preset padding bit to obtain extended data, and the number of bits of the extended data meets the total number of bits of data output.
[104] For example, when the number of partitions is 3, the data corresponding to each of the above data is: the first partition corresponds to {(a, 0), (c, 4), (f, 11), (h, 14)}, the second partition corresponds to {(b, 1), (e, 6), (g, 13)}, and the third partition corresponds to {(d, 5), (i, 15)}. The processed parameter corresponding to data a in the first partition above is 0. Through the processed parameter and the data output volume, the first partition can be determined as the target partition corresponding to data a. Similarly, through analysis and processing, it can be determined that: the target partition corresponding to data c is the first partition, the target partition corresponding to data f is the second partition, the target partition corresponding to data h is the third partition, the target partition corresponding to data b is the first partition, the target partition corresponding to data e is the second partition, the target partition corresponding to data g is the third partition, the target partition corresponding to data d is the first partition, and the target partition corresponding to data i is the third partition. After analysis and processing, it can be determined that: the target partition corresponding to data c is the first partition, the target partition corresponding to data f is the second partition, the target partition corresponding to data h is the third partition, the target partition corresponding to data b is the first partition, the target partition corresponding to data e is the second partition, the target partition corresponding to data g is the third partition, the target partition corresponding to data d is the first partition, and the target partition corresponding to data i is the third partition.
[0105] After statistics, the target partitions of data a and data c in the first partition are both the first partition. At this time, since the data lengths of data a and data c (2 bits) cannot meet the total data output length (9 bits), data expansion operations can be performed based on the total data length of all data and the preset padding bits to obtain expanded data, that is, expand data (a, c) to data (a, c, x, x, x, x, x, x, x), where the above "x" is the preset placeholder data (or called the preset padding bit), thus effectively implementing the data expansion operation.
[0106] Similarly, the target partition of data f in the first partition is the second partition. At this time, since data f cannot meet the total data output length (9 bits), data expansion operations can be performed based on the total data length of all data and the preset padding bits to obtain expanded data, that is, expand data (f) to data (f, x, x, x, x, x, x, x, x), where the above "x" is the preset padding bit. The target partition of data h in the first partition is the third partition, and then data (h) can be expanded to data (h, x, x, x, x, x, x, x, x), thus effectively implementing the data expansion operation.
[0107] The target partition corresponding to data b in the second partition is the first partition, and then data (b) can be expanded to data (b, x, x, x, x, x, x, x, x). The target partition corresponding to data e in the second partition is the second partition, and then data (e) can be expanded to data (e, x, x, x, x, x, x, x, x). The target partition corresponding to data g in the second partition is the third partition, and then data (g) can be expanded to (g, x, x, x, x, x, x, x, x). The target partition corresponding to data d in the third partition is the first partition, and then data (d) can be expanded to (d, x, x, x, x, x, x, x, x). The target partition corresponding to data i in the third partition is the third partition, and then data (e) can be expanded to (e, x, x, x, x, x, x, x, x), thus effectively implementing the expansion operation for each data, and then the expanded data can be stably obtained.
[0108] Step S305: Send the expanded data to the target partition and rewrite the expanded data to obtain the expanded data for constructing the second expansion table.
[0109] After obtaining the extended data and determining the target partition corresponding to the extended data, the extended data can be sent to the target partition. Since the data received in the target partition includes many preset padding bits, in order to ensure the quality and effect of data extension, after sending the extended data to the target partition, an operation of rewriting the extended data can be performed, so that the extended data constituting the second extended table can be obtained. In some instances, a pre-trained machine learning model or neural network model can be used to perform the operation of rewriting the extended data, so that the second extended table can be stably obtained.
[0110] Alternatively, in some other instances, not only can the operation of rewriting the extended data be implemented through a pre-trained machine learning model or neural network model, but also the operation of rewriting can be implemented by deleting the redundant preset padding bits. At this time, the operation of rewriting the extended data to obtain the extended data for constituting the second extended table may include: obtaining the remainder corresponding to the division operation between the processed parameter and the data output amount; determining the position where the data is stored in the target partition based on the remainder; deleting the redundant preset padding bits based on the position and the data output amount to obtain the sorted data; rewriting the preset padding bits included in the sorted data to the previous adjacent data to obtain the target data for constituting the second extended table.
[0111] For example, when the processed parameter is (0, 1, 4, 5, 6, 11, 13, 14, 15), the data corresponding to the above processed parameter includes: the first partition corresponds to {(a, 0), (c, 4), (f, 11), (h, 14)}, the second partition corresponds to {(b, 1), (e, 6), (g, 13)}, the third partition corresponds to {(d, 5), (i, 15)}, and when the output data amount is 6, the remainder corresponding to the division operation between the processed parameter and the data output amount can be obtained. Specifically, the remainders of the processed parameters corresponding to the above data a, b, c, d, e, f, g, h are: 0, 1, 4, 5, 1, 5, 1, 2, 3. The above remainders are used to represent the position where the data is stored in the target partition, and the target partition corresponding to the data can be determined by the quotient value corresponding to the division operation between the processed parameter and the data output amount. output amount.
[0112] After analysis, it can be known that the above data a is at the 0th position in the target partition (i.e., the first partition), data b is at the 1st position in the target partition (i.e., the first partition), data c is at the 4th position in the target partition (i.e., the first partition), and data d is at the 5th position in the target partition (i.e., the first partition). Then, based on the above positions and the data output volume, redundant preset padding bits can be deleted, so that the sorted data in the target partition (the first partition) can be obtained as: a, b, x, x, c, d, where the above x is the remaining preset padding bit. Similarly, the sorted data in the target partition (the second partition) can be obtained as: e, x, x, x, x, f, where the above x is the remaining preset padding bit; the sorted data in the target partition (the third partition) can be obtained as: x, g, h, I, x, x, where the above x is the remaining preset padding bit.
[0113] After obtaining the sorted data, the preset padding bits included in the sorted data can be rewritten as the previous adjacent data. That is, "x, x" in the first partition is changed to "b, b" to obtain the first target data "a, b, b, b, c, d" in the first partition. Similarly, "x" in the second partition can be rewritten as "e" to obtain the second target data "e, e, e, e, e, f" in the second partition. For the first "x" in the third partition, it is adjacent to the last data "f" in the second partition. Therefore, the first "x" in the third partition can be rewritten as "f", and the fifth and sixth "x" in the third partition can be rewritten as "i". After the above rewriting operations, the third target data "f, g, h, i, i, i" in the third partition can be obtained. The above first target data, second target data, and third target data are all the target data for constructing the second extended table, effectively ensuring the stability and reliability of obtaining the second extended table.
[0114] In this embodiment, by obtaining the data output volume corresponding to each partition, the repetition degree of each attribute value is determined as the extension times for performing extension operations on each data in the second connection table. Then, based on the extension times and the data output volume, the target partition corresponding to each data is determined, and each data is extended based on the data output volume to obtain extended data. After that, the extended data is sent to the target partition and rewritten, so that the extended data for constructing the second extended table can be stably obtained, effectively realizing the stable extension operation of the second connection table.
[0115] Figure 4 is a schematic flowchart of calculating the prefix sum of the expansion times and obtaining the calculated parameters provided by the embodiment of the present application; on the basis of the above embodiment, referring to Figure 4, this embodiment provides an implementation manner of calculating the prefix sum of the expansion times. Specifically, the calculation of the prefix sum of the expansion times and obtaining the calculated parameters in this embodiment may include steps S401 to S404.
[0116] Step S401: Obtain the current partition corresponding to the expansion times.
[0117] After obtaining the expansion times, the current partition corresponding to the expansion times can be determined based on the data corresponding to each expansion time. In some instances, the current partition corresponding to each data can be determined as the current partition corresponding to the expansion times. For example, when the expansion times are (1, 3, 1, 1, 521, 1, 3), the current partition corresponding to the above expansion times (1, 3, 1) can be the first partition, the current partition corresponding to the above expansion times (1, 5, 2) can be the second partition, and the current partition corresponding to the above expansion times (1, 1, 3) can be the third partition.
[0118] Step S402: Perform a prefix sum calculation process on the expansion times in the same partition to obtain a first calculation process sequence corresponding to each partition, where the first calculation process sequence includes multiple processing results.
[0119] After obtaining the current partition corresponding to the expansion times, a prefix sum calculation process operation can be performed on the expansion times in the same partition, so as to obtain a first calculation process sequence in each partition. The first calculation process sequence includes multiple processing results. In some instances, a pre-trained machine learning model or neural network model can be used to perform a prefix sum calculation process on the expansion times in the same partition to obtain a first calculation process sequence corresponding to each partition.
[0120] For example, when the expansion times (1, 3, 1) are stored in the first partition, after performing a prefix sum calculation process operation on the above expansion times, a first calculation process sequence corresponding to the first partition can be obtained. This first The calculation processing sequence can be: 1, 4, 5. Similarly, when the expansion times (1, 5, 2) are stored in the second partition, after performing the prefix sum calculation processing operation on the above expansion times, the first calculation processing sequence corresponding to the second partition can be obtained, and this first calculation processing sequence can be: 1, 6, 8; when the expansion times (1, 1, 3) are stored in the third partition, after performing the prefix sum calculation processing operation on the above expansion times, the first calculation processing sequence corresponding to the third partition can be obtained, and this first calculation processing sequence can be: 1, 2, 5, which effectively ensures the accuracy and reliability of determining the first calculation processing sequence corresponding to each partition.
[0121] Step S403: Send the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition to obtain a processing result sequence stored in the preset partition.
[0122] After obtaining the first calculation processing sequence corresponding to each partition, the last processing result in the first calculation processing sequence corresponding to each partition can be sent to a preset partition. Among them, the preset partition can be the first partition, the second partition, or the Nth partition, etc. among the configured multiple partitions. Taking the first partition as an example of the preset partition, after sending the last processing result in the first calculation processing sequence corresponding to each partition to the first partition, the processing result sequence stored in the preset partition can be obtained.
[0123] For example, when the first partition is the preset partition, if the first calculation processing sequence corresponding to the first partition is: 1, 4, 5; the first calculation processing sequence corresponding to the second partition is: 1, 6, 8; the first calculation processing sequence corresponding to the third partition is: 1, 2, 5, then the last processing results (5, 8, 5) in the first calculation processing sequence corresponding to the above partitions can be sent to the first partition, and thus the processing result sequence stored in the first partition can be obtained: 5, 8, 5.
[0124] Step S404: Perform prefix sum calculation on the processing result sequence to obtain the calculated parameter.
[0125] After obtaining the processed result sequence, prefix sum calculation processing can be performed on the processed result sequence, so that the calculated parameters can be obtained. In some examples, performing prefix sum calculation on the processed result sequence and obtaining the calculated parameters may include: performing prefix sum calculation on the processed result sequence to obtain a second calculation processing sequence, where the second calculation processing sequence includes multiple calculation processing results; determining the target partition corresponding to each calculation processing result in the second calculation processing sequence; sending the calculation processing results to the target partition, and summing the calculation processing results and the expansion times stored in the same target partition to obtain the calculated parameters.
[0126] Among them, determining the target partition corresponding to each calculation processing result in the second calculation processing sequence may include: sorting the multiple calculation processing results included in the second calculation processing sequence based on the partition identifiers corresponding to each calculation processing result to obtain result sorting information; determining the sum value of the current sorting position where each calculation processing result is located and 1 as the target partition corresponding to the calculation processing result.
[0127] For example, when the processed result sequence is: 5, 8, 5, prefix sum calculation can be performed on the processed result sequence to obtain a second calculation processing sequence, and the second calculation processing sequence may be: 5, 13, 18. Then, the second calculation processing sequence can be sorted based on the identifiers of the source partitions corresponding to the processing results in the above-mentioned respective processing result sequences. Since 5 corresponds to the first partition, 13 corresponds to the second partition, and 18 corresponds to the third partition, after sorting the second calculation processing sequence, result sorting information can be obtained, that is, 5, 13, 18, that is, the result sorting information is exactly the same as the second calculation processing sequence.
[0128] Then, the target partition corresponding to the above-mentioned 5 can be determined as the second partition, the target partition corresponding to 13 can be determined as the third partition, and the above-mentioned 5 can be sent to the second partition and 13 can be sent to the third partition. Then, the sum of the original data stored in the second partition and the data "5" can be calculated, so that the calculated parameters stored in each partition can be obtained. For example: in the first partition, 1, 3, 1 are stored; in the second partition, 6, 8, 7 are stored; in the third partition, 14, 14, 16 are stored. In this way, the prefix sum calculation operation on the processed result sequence is stably realized, and the accuracy and reliability of obtaining the calculated parameters are ensured.
[0129] In this embodiment, by obtaining the current partition corresponding to the number of extensions, performing a prefix sum calculation on the number of extensions in the same partition, obtaining the first calculation processing sequence corresponding to each partition, and then sending the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition, obtaining the processing result sequence stored in the preset partition, and performing a prefix sum calculation on the processing result sequence, the calculated parameters can be stably obtained. Then, the processing operation of the data table can be implemented based on the calculated parameters, further ensuring the practicability of this method. The last processing result in the first calculation processing sequence corresponding to each partition is sent to a preset partition to obtain the processing result sequence stored in the preset partition, and a prefix sum calculation is performed on the processing result sequence, so that the calculated parameters can be stably obtained. Then, the processing operation of the data table can be implemented based on the calculated parameters, further ensuring the practicability of this method.
[0130] In specific applications, taking the number of distributed partitions as three as an example, this application embodiment provides an oblivious efficient algorithm for the natural join operator in a distributed scenario. The oblivious efficient algorithm for the natural join operator depends on the following basic operators to implement: primary key-foreign key natural join operator, sorting operator, grouping aggregation operator, prefix sum operator, and extension operator. Among them, the primary key-foreign key natural join, sorting operator, and grouping aggregation operator can directly adopt the implementation methods in related technologies. To facilitate the understanding of the implementation principle and implementation effect of the oblivious efficient algorithm in this application embodiment, the implementation principles of the prefix sum operator and the extension operator are described below:
[0131] Taking the input parameters as {X1, X2,..., XN} and the output parameters as {X1, X1 + X2, X1 + X2 +... + XN} as an example to illustrate the implementation principle of the prefix sum operator. Among them, the above "+" operation is not limited to arithmetic addition and can be any binary operation that satisfies the associative law. For example: arithmetic multiplication, logical exclusive OR, taking the larger value, taking the smaller value, taking the latter (if the latter is not the preset placeholder data dummy) or the former (if the latter is the preset placeholder data dummy), etc. Specifically, the implementation process of the prefix sum operator can include the following steps 1 to 3.
[0132] Step 1: Each partition (the first partition, the second partition, and the third partition) performs a prefix sum calculation on the local stored data and sends the last number of the calculation result to the first partition (or a preset partition).
[0133] For example, the number of partitions p = 3, namely the first partition (corresponding to the identifier 1), the second partition (corresponding to the identifier 2), and the third partition (corresponding to the identifier 3). The total amount of data N = 9, and the total amount of data stored in each partition n = N / p = 3. The N input numbers are 1, 2, 3, 1, 2, 1, 3, 4, 2 in sequence. Before performing the prefix sum operator, the starting data corresponding to each partition can be: in the first partition, 1, 2, 3 are stored; in the second partition, 1, 2, 1 are stored; in the third partition, 3, 4, 2 are stored. After performing the prefix sum calculation on the stored data in each of the above partitions locally, the following data can be obtained: in the first partition -1, 3, 6; in the second partition -1, 3, 4; in the third partition -3, 7, 9. Then, the data 6, 4, 9 after the prefix sum calculation for the above three partitions can be sent to the first partition, so that the first partition can obtain 6, 4, 9.
[0134] Step 2: The first partition can arrange the p received numbers together in the order of the partition numbers, and perform the prefix sum calculation operation, and then send the i-th number to the partition with the identifier i + 1, where i = 1, 2,... p - 1.
[0135] Among them, after the first partition obtains 6, 4, 9, the p received numbers can be arranged together, and then the above data can be subjected to the prefix sum calculation to obtain 6, 10, 19. Then, the above 6 can be sent to the (i + 1)-th partition (i.e., the second partition), and the above 10 can be sent to the (i + 1)-th partition (i.e., the third partition).
[0136] Step 3: After each partition receives the corresponding data, the received data can be added to the value after the prefix sum calculation in each partition.
[0137] Since the first partition does not receive any data, the data in the first partition is 1, 3, 6; the second partition receives the data 6, and the values after the prefix sum calculation in the second partition are 1, 3, 4. After the addition and summation, 7, 9, 10 stored in the second partition can be obtained; the third partition receives the data 10, and the values after the prefix sum calculation in the third partition are 3, 7, 9. After the addition and summation, 13, 17, 19 stored in the third partition can be obtained, thus effectively realizing the prefix sum operation of the data.
[0138] It should be noted that the communication volumes corresponding to the above prefix sum calculation operations are as follows: the communication volume of step 1 above is p - 1, the communication volume of step 2 is p - 1, and the total communication volume is 2p - 2.
[0139] On the other hand, taking the input parameters as {XI, X2,..., XN}, N positive integers with a sum of M as the expansion times, and the output parameters as {Xi,.... x-L, x2,..., X2,..., XN,..., X N} as an example for expansion is described. Among them, Xi appears continuously di times, i = 1, 2,... N. Among them, n = N / p, m = M / p. It should be noted that if N or M is not a multiple of p, supplementary preset placeholder data dummy data is used to make N and M both multiples of p. Specifically, the implementation process of the expansion operator may include the following steps 11 to 14.
[0140] Step 11: Call the prefix sum operator for the sequence {0, d1, d2,..., dN - 1}. The addition in the above prefix sum operator is ordinary integer addition, and the result is denoted as {e1, e2,..., e N}„
[0141] For example, the number of partitions p = 3, the total amount of data N = 9, the total amount of data in each partition n = N / p = 3, the total amount of output data M = 18, and the amount of output data in each partition m = M / p = 6. In addition, for the sake of illustration, the constant parameter used to control the failure rate can be set to 1. When the N input numbers are a, b, c, d, e, f, g, h, i in sequence, and the corresponding positive integers are 1, 3, 1, 1, 5, 2, 1, 1, 3 in sequence, after calling the prefix sum operator for {0, 1, 3, 1, 1, 5, 2, 1, 1}, the result {0, 1, 4, 5, 6, 11, 13, 14, 15} can be obtained.
[0142] Step 12: Rewrite the data {X1, X2,..., XN} into the structure of (X1, e1), (X2, e2),...(XN, eN). Independently and uniformly randomly select one partition from p partitions as the target partition for each data, and then send all the numbers to their respective designated target partitions.
[0143] After the above rewriting operation, the local data stored in each partition will be scrambled. Assume that the data in each partition after scrambling can be: First partition: (a,0), (c,4), (f,ll), (h,14); Second partition: (b,1), (e,6), (g,13); Third partition: (d, 5), (i,15).
[0144] Step 13: For the data received by each partition, rewrite the data (x1; e1), (x2, e2),...(x N , e N ) as (x1, r1), (x2, r2),...(x N , r N ) o
[0145] Assume that its data content is (y"i), ( y2, f2),.. (yK, fK). For (yi, fi), assume that fi = mqi + ri, that is, the quotient of fi divided by m is qi, and the remainder is n. Rewrite the data content as (Xj, n), and specify the target partition for this data as the (qi + 1)th. For each partition j, if the total number of data sent to partition j is less than [m + cVm]=9, where c is a constant parameter pre-configured to control the failure rate. If the total number of data sent to partition j is less than 9, then the preset placeholder data dummy can be supplemented until the total number of data is equal to this value, and then all the data is sent to the specified partition.
[0146] If the total number of data sent to partition j is greater than this value, the entire algorithm fails. At this time, the processing operation of the data table can be restarted, or a prompt message can be generated to remind the user to configure the constant parameter for controlling the failure rate based on the prompt message.
[0147] After the above process, the amount of data sent by each partition will be filled to exactly 9, and after sending the data to the corresponding partition, the data of each partition can be respectively (x represents the placeholder data dummy filled by default): The first partition: (a,0),(c,4),x,x,x,x,x,x,x,(b, l),x,x,x,x,x,x,x,x,(d,5),x,x,x,x,x,x,x,x; The second partition: (f,5),x,x,x,x,x,x,x,x,(e,0),x,x,x,x,x,x,x,x,x,x,x,x,x,x,x,x,x; The third partition: (h,2),x,x,x,x,x,x,x,x,(g,l),x,x,x,x,x,x,x,x,(i,3),x,x,x,x,x,x,x,x.
[0148] In addition, for the above constant parameter c, it is used to control the failure probability of the algorithm. Through theoretical analysis, it can be proved that the failure probability of the algorithm is less than e^(-c^2 / 2). In an actual scenario, the user can select or configure an appropriate parameter c according to specific application requirements and the required failure rate of the application.
[0149] Step 14: Among the data received by each partition, assume that the data that is not the default placeholder data dummy is (Z1, g1),...,(Zi, gi). Place the data Zj at the (gj + 1)-th position, fill the remaining positions with dummy, and the total amount of data is m. The excess dummy data is discarded. Finally, call the prefix sum algorithm on the data. Here, the addition in the prefix sum algorithm is defined as x + y = x (if y is the dummy data) or y (if y is not the dummy data).
[0150] After the above processing operations, the data can be moved to the corresponding positions and the excess dummy can be removed. The results are as follows: The first partition: a,b,x,x,c,d; The second partition: e,x,x,x,x,f; The third partition: x,g,h,i,x,x.
[0151] Then, call the prefix sum algorithm, and the data of all dummies will be rewritten as the non-dummy number adjacent to it in front. Note that the global prefix sum operator is called here. Therefore, the first x in the third partition will recognize that the non-dummy number adjacent to it in front or the nearest non-dummy number is f, and then the "x" in the third partition can be rewritten as f, so that the processed result can be obtained, as follows: The first partition: a, b, b, b, c, d; The second partition: e, e, e, e, e, f; The third partition: f, g, h, i, i, i o
[0152] It should be noted that the communication volumes corresponding to the above expansion operations are as follows: The communication volume of step 11 above is 2p - 2, the communication volume of step 12 is N, the communication volume of step 13 is approximately M (which can be ignored), and the communication volume of step 14 is 2p - 2. Therefore, the total communication volume can be obtained as approximately N + M
[0153] Based on the above prefix sum operator and expansion operator, and taking two data tables of R(A, B) and S(B, C) as input and the output as a data table of the form T(A, B, C) with a total of M rows as an example, where each of the tables R(A, B) and S(B, C) has N rows of data. After the natural join operation, the obtained output result has a total of M rows. Specifically, the implementation process of the oblivious efficient algorithm of the natural join operator can include the following steps 111 to 11110
[0154] Step 111: Call the grouping aggregation operator for the table R(A, B), calculate the repetition degree of each b in R(A, B), and record the result as the table R1(B, D), where the column B is the primary key in the table R1, and D represents the repetition degree of the primary key B in R
[0156] Step 112: Call the primary key-foreign key natural join operator for the tables R1 and S, and record the result as the table S2(B, C, D)
[0157] Specifically, perform a natural join operation on the tables R1 and S through the primary key-foreign key natural join operator to obtain the table S2(B, C, D). The above table S2(B, C, D) can be {(yi, shao, 1), (yi, c2, 1), (b2, c2, 2), (b2, c3, 2), (b3, c4, 1)} O
[0158] Step 113: Invoke the extension operator on Table S2. The data to be extended is the column combination (B, C), and the number of extension times is column D. Denote the extension result as Table S3(B, C).
[0159] Specifically, after obtaining Table S2, the extension operator can be used to perform extension processing on Table S2 to obtain the extended result Table S3(B, C) as {(yi, Ci), (b1, C2), (b2, C2), (b2, c2), (b2, c3), (b2, c3), (b3, c4)} o
[0160] Step 114: Invoke the sorting operator to sort Table S3 in lexicographical order according to columns (B, C).
[0161] Among them, the sorting operator can receive N elements (or a table with N rows) as input, rearrange the N elements (or N rows) according to a certain size relationship, and then output. When these elements are tuples containing multiple components, sometimes multiple components (columns) are specified for the size relationship, which is called lexicographical sorting according to multiple columns. That is, when comparing two tuples, the comparison is first performed according to the specified first column. If the values of the tuples in this column are the same, then the comparison is performed according to the second column, and so on. Through calculation, it can be known that the communication volume of the above sorting operation is approximately 3N.
[0162] Step 115: Invoke the grouping and aggregation operator on Table S to calculate the repetition degree of each b in S(B, C), and denote the result as Table S1(B, E). Among them, column B is the primary key in Table S1, and E represents the repetition degree of the primary key B in S.
[0163] Among them, the grouping and aggregation operator corresponds to the "group by" query statement in database queries. This operator can receive a table as input, group the rows of the table according to the specified column combination, and perform aggregation operations on another specified column within the same group, such as summation, maximum value, average value, etc. Through calculation, it can be known that the communication volume of the above grouping and aggregation operation is approximately N.
[0164] Specifically, after obtaining Table S, the grouping and aggregation operator can be used to analyze Table S to obtain the repetition degree corresponding to each value b in data column B of Table S, so as to obtain the result S1(B, E) as {(b1, 2), (b2, 2), (b3, 1)}.
[0165] Step 116: Invoke the primary key-foreign key natural join operator on tables S1 and R, and denote the result as R2(A, B, E).
[0166] Among them, the primary key-foreign key natural join operator is a special case of the natural join operator. It receives two tables, table R(A, B) and table S(B, C), with each table having N rows of data, and the B column is unique in table S, that is, the repeatability is 1. The output is a table in the form of T(A, B, C) with at most N rows of data, and the content of the table is a set. Since the primary key-foreign key natural join operator needs to combine the two tables together and invokes a sorting and a prefix sum operator once, its communication volume is relatively large, specifically about 6N.
[0167] Specifically, when table S1(B, E) is {(Doctor, 2), (b2, 2), (b3, 1)}, and table R is {(a1; b1), (a1, b2), (a2> b 2), (ai> b3), (a2> b 4)}, then the primary key-foreign key natural join operator can be used to perform a natural join operation on tables S1 and R, so as to obtain the result R2(A, B, E) as {Gi, Doctor, 2), (ai, b2, 2), (a2, b2, 2), (a1; b3, 1)} O
[0168] Step 117: Invoke the extension operator on table R2, with the data to be extended being the column combination (A, B) and the number of extension times being column E, to obtain the extension result as R3(A, B)«
[0169] After obtaining table R2, the extension operator can be used to perform an extension operation on table R2 to obtain the extension result R3(A, B). Specifically, R3(A, B) can be {(a], Doctor), ^, b^d, b2), (a1, b2), (a2, b2), (a2, b2), (a1, b3)} o
[0170] Step 118: Add column F to table R3 with an initial value of 1, and denote the result as R4(A, B, F). Before invoking R4, the addition is defined as: Assume the two tuples to be summed are (a1; bi. fi) and (a2, b2, f2), if If it is 2, then their sum is (a2, b2, fi + f2), otherwise it is (a2, b2, f2). Denote the result as R5(A, B, F).
[0171] Among them, the added F column is used to determine the preset order for adjusting Table R3 to align the data of Table R3 with Table S. Additionally, after calling the prefix sum operator on R4, the obtained result R5(A, B, F) can be {(ai, bi, 1), (a1, b1, 2),(a1, b2, l),(a1, b2, 2),(a2, b2, l),(a2, b2, 2),(a1, b3, 1)}.
[0172] Step 119: Call the sorting operator to sort R5 in lexicographical order by columns (B, F, A).
[0173] Step 1110: After obtaining Table S3 and Table R5, the merge operation can be performed on Table S3 and Table R5 to obtain the calculation table T(A, B, C).
[0174] For example, assume that the i-th row in Table S3 is (b, c), and the i-th row in Table R5 is (a, b,, f). After merging Table S3 and Table R5, the i-th row in the calculation table T can be obtained as (a, b, c). Specifically, the content of Table T can be: {(ai, yi, Ci), (a1; b1, c2),(ai> b2, c2),(a2< b2, c2),(ai> b2, c3),(a2, b2, c3),(a1, b3, c4)}, thus effectively ensuring the quality and effect of data table processing.
[0175] It should be noted that Steps 111 - Step 113 and Steps 114 - Step 116 in the above embodiments are completely symmetric, and the parameter b = b1 in the above embodiments. Additionally, the communication volumes corresponding to the above expansion operations are as follows: the communication volume corresponding to Step 111 is N, the communication volume corresponding to Step 112 is 6N, the communication volume corresponding to Step 113 is N + M, the communication volume corresponding to Step 114 is 3M, the communication volume corresponding to Step 115 is N, the communication volume corresponding to Step 116 is 6N, the communication volume corresponding to Step 117 is N + M, the communication volume corresponding to Step 118 is p - 1, the communication volume corresponding to Step 119 is 3M, the communication volume corresponding to Step 11110 is 0, and the total communication volume is approximately 16N + 8M.
[0176] The technical solution provided by the embodiments of this application realizes an oblivious and efficient algorithm for the natural join operator in a distributed scenario. Specifically, by setting up a prefix sum operator and an expansion operator, it effectively realizes expanding the data of each table multiple times in an efficient and concise manner, and the number of times is exactly equal to its multiplicity in another table. Then, through sorting, the data is aligned, so that the natural join operation can be efficiently completed. And because the amount of data sent by each partition is the same or similar, this effectively meets the oblivious requirement, and the communication volume of the entire data table processing operation is relatively low. Furthermore, it effectively solves the problems in the oblivious algorithms implemented in the related technologies, which are highly difficult and very inefficient, and there is also an additional risk of information leakage in the field of encrypted data. In addition, the above implementation process also improves the quality and efficiency of data table processing, further improves the practicability of this method, and is conducive to market promotion and application.
[0177] Figure 5 is a schematic structural diagram of a data table processing device provided by an embodiment of this application; as shown in the reference drawing 5, this embodiment provides a data table processing device, and this data table processing device is used to execute the data table processing operation shown in Figure 2 above. Specifically, this data table processing device may include an acquisition module 11, an expansion module 12, and a processing module 13.
[0178] The acquisition module 11 is used to acquire a first data table and a second data table. The first data table and the second data table are stored in different distributed partitions, and the first data table and the second data table include the same attribute items.
[0179] The expansion module 12 is used to expand the second data table based on the attribute items in the first data table to obtain a second expanded table.
[0180] The expansion module 12 is also used to expand the first data table based on the attribute items in the second data table to obtain a first expanded table.
[0181] The processing module 13 is used to add a data column to the first expanded table to obtain an adjusted expanded table. The parameter values included in the data column are used to re-sort the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table.
[0182] The processing module 13 is also used to associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
[0183] In some instances, when the expansion module 12 expands the second data table based on the attribute items in the first data table to obtain a second expanded table, the expansion module 12 is configured to perform: grouping and aggregating the first data table based on the attribute items in the first data table to obtain a first aggregated table, where the first aggregated table includes the repetition degrees of each attribute value under the attribute item; performing a natural join on the first aggregated table and the second data table to obtain a second joined table; and expanding the second joined table to obtain a second expanded table.
[0184] In some instances, when the expansion module 12 expands the second joined table to obtain a second expanded table, the expansion module 12 is configured to perform: obtaining the data output amount corresponding to each partition; determining the repetition degree of each attribute value as the expansion times for performing expansion operations on each data in the second joined table; based on the expansion times and the data output amount, determining the target partition corresponding to each data; expanding each data based on the data output amount to obtain expanded data; sending the expanded data to the target partition, and rewriting the expanded data to obtain the expanded data for constructing the second expanded table.
[0185] In some instances, when the expansion module 12 determines the target partition corresponding to each data in the second joined table based on the expansion times and the data output amount, the expansion module 12 is configured to perform: calculating the prefix sum of the expansion times to obtain a calculated parameter; determining the 0 value and the calculated parameter as the processed parameter; obtaining the quotient of the processed parameter divided by the data output amount; and based on the sum value of the quotient and 1, determining the target partition corresponding to the data.
[0186] In some instances, when the expansion module 12 calculates the prefix sum of the expansion times to obtain a calculated parameter, the expansion module 12 is configured to perform: obtaining the current partition corresponding to the expansion times; calculating the prefix sum of the expansion times in the same partition to obtain a first calculation processing sequence corresponding to each partition, where the first calculation processing sequence includes multiple processing results; sending the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition to obtain a processing result sequence stored in the preset partition; and calculating the prefix sum of the processing result sequence to obtain a calculated parameter. In some instances, when the expansion module 12 expands the second joined table to obtain a second expanded table, the expansion module 12 is configured to perform: obtaining the data output amount corresponding to each partition; determining the repetition degree of each attribute value as the expansion times for performing expansion operations on each data in the second joined table; based on the expansion times and the data output amount, determining the target partition corresponding to each data; expanding each data based on the data output amount to obtain expanded data; sending the expanded data to the target partition, and rewriting the expanded data to obtain the expanded data for constructing the second expanded table.
[0187] In some instances, when the extension module 12 performs a prefix sum calculation on the processing result sequence to obtain the calculated parameter, the extension module 12 is used to execute: performing a prefix sum calculation on the processing result sequence to obtain a second calculation processing sequence, where the second calculation processing sequence includes multiple calculation processing results; determining the target partition corresponding to each calculation processing result in the second calculation processing sequence; sending the calculation processing result to the target partition, and summing the calculation processing result and the extension times stored in the same target partition to obtain the calculated parameter.
[0188] In some instances, when the extension module 12 determines the target partition corresponding to each calculation processing result in the second calculation processing sequence, the extension module 12 is used to execute: sorting the multiple calculation processing results included in the second calculation processing sequence based on the partition identifier corresponding to each calculation processing result to obtain result sorting information; determining the sum value of the current sorting position where each calculation processing result is located and 1 as the target partition corresponding to the calculation processing result.
[0189] In some instances, when the extension module 12 expands each data based on the data output volume to obtain the expanded data, the extension module 12 is used to execute: determining the total number of data output bits based on the data output volume; expanding each data based on the total number of data output bits to obtain the expanded data.
[0190] In some instances, when the extension module 12 expands each data based on the total number of data output bits to obtain the expanded data, the extension module 12 is used to execute: obtaining all the data to be sent to the same target partition; detecting whether the sum of the data bits of all the data meets the total number of data output bits; if not, performing a data expansion operation based on the sum of the data bits of all the data and the preset padding bits to obtain the expanded data, where the number of data bits of the expanded data meets the total number of data output bits.
[0191] In some instances, when the extension module 12 rewrites the expanded data to obtain the expanded data for forming the second expansion table, the extension module 12 is used to execute: obtaining the remainder corresponding to the division operation between the processed parameter and the data output volume; determining the position where the data is stored in the target partition based on the remainder; deleting the redundant preset padding bits based on the position and the data output volume to obtain the sorted data; rewriting the preset padding bits included in the sorted data as the previous adjacent data to obtain the target data for forming the second expansion table.
[0192] In some instances, when the extension module 12 extends the first data table based on the attribute items in the second data table to obtain the first extended table, the extension module 12 is configured to perform: grouping and aggregating the second data table based on the attribute items in the second data table to obtain a second aggregated table, where the second aggregated table includes the repetition degrees of the respective attribute values under the attribute item; performing a natural join on the second aggregated table and the first data table to obtain a first joined table; and extending the first joined table to obtain the first extended table.
[0193] In some instances, when the processing module 13 adds a data column to the first extended table to obtain an adjusted extended table, the processing module 13 is configured to perform: adding a data column with an initial value of 1 to the first extended table to obtain an intermediate extended table; comparing two attribute values of two adjacent groups of data in the intermediate extended table; when the two attribute values corresponding to the two adjacent groups of data are the same, performing a summation process on the data columns in the two adjacent groups of data to obtain a summation data column, and determining that the two attribute values and the summation data column are used to form the data of the adjusted extended table; when the two attribute values corresponding to the two adjacent groups of data are different, determining the latter data in the two adjacent groups of data as the data used to form the adjusted extended table.
[0194] In some instances, when the processing module 13 correlates the second extended table and the adjusted extended table to obtain a correlation data table corresponding to the first data table and the second data table, the processing module 13 is configured to perform: obtaining a preset rule for sorting the first extended table and the adjusted extended table; sorting the first extended table and the adjusted extended table respectively using the preset rule to obtain a first sorted table and a second sorted table; and correlating the first sorted table and the second sorted table to obtain a correlation data table corresponding to the first data table and the second data table.
[0195] The device shown in Figure 5 can execute the method of the embodiments shown in Figures 1-4. For parts not described in detail in this embodiment, reference can be made to the relevant descriptions of the embodiments shown in Figures 1-4. The execution process and technical effects of this technical solution are as described in the embodiments shown in Figures 1-4, and will not be elaborated here.
[0196] In a possible design, the structure of the processing device for the data table shown in FIG. 5 can be implemented as an electronic device, which can be various devices such as a controller, a personal computer, a partition, etc. As shown in FIG. 6, the electronic device may include: a first processor 21 and a first memory 22. Among them, the first memory 22 is used to store a program for the corresponding electronic device to execute the data table processing method provided in the embodiments shown in FIGS. 1-2 above, and the first processor 21 is configured to execute the program stored in the first memory 22.
[0197] The program includes one or more computer instructions. When the one or more computer instructions are executed by the first processor 21, the following steps can be achieved: obtaining a first data table and a second data table, both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; expanding the second data table based on the attribute items in the first data table to obtain a second extended table; expanding the first data table based on the attribute items in the second data table to obtain a first extended table; adding a data column to the first extended table to obtain an adjusted extended table, and the parameter values included in the data column are used to reorder the first extended table so that the data in the first extended table is aligned with the data in the second extended table; associating the second extended table and the adjusted extended table to obtain an associated data table corresponding to the first data table and the second data table.
[0198] Further, the first processor 21 is also used to execute all or part of the steps in the embodiments shown in FIGS. 1-4 above.
[0199] Among them, the structure of the electronic device may further include a first communication interface 23 for the electronic device to communicate with other devices or communication networks.
[0200] In addition, an embodiment of the present application provides a computer storage medium for storing computer software instructions used by an electronic device, which includes a program involved in executing the data table processing method in the embodiments shown in FIGS. 1-4 above.
[0201] In addition, an embodiment of the present application provides a computer program product, including: a computer-readable storage medium storing computer instructions, when the computer instructions are executed by one or more processors, causing the one or more processors to execute the steps in the data table processing method in the method embodiments shown in FIGS. 1-4 above.
[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by the combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a computer product. This application can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0204] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram. device.
[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.
[0206] These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.
[0207] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0208] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0209] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
23 Claims 1. A method for processing a data table, comprising: Obtain a first data table and a second data table. Both the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; expand the second data table based on the attribute items in the first data table to obtain a second expanded table; expand the first data table based on the attribute items in the second data table to obtain a first expanded table; add a data column to the first expanded table to obtain an adjusted expanded table, and the parameter values included in the data column are used to reorder the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table; Associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
2. The method according to claim 1, wherein Expanding the second data table based on the attribute items in the first data table to obtain a second expanded table includes: performing grouped aggregation processing on the first data table based on the attribute items in the first data table to obtain a first aggregated table, where the first aggregated table includes the repetition degrees of each attribute value under the attribute item; performing a natural join on the first aggregated table and the second data table to obtain a second joined table; expanding the second joined table to obtain a second expanded table.
3. The method according to claim 2, wherein Expanding the second joined table to obtain a second expanded table includes: obtaining the data output amount corresponding to each partition; determining the repetition degree of each attribute value as the expansion times for performing expansion operations on each data in the second joined table; determining the target partition corresponding to each data based on the expansion times and the data output amount; expanding each data based on the data output amount to obtain expanded data; sending the expanded data to the target partition and rewriting the expanded data to obtain the expanded data for constituting the second expanded table.
4. The method according to claim 3, wherein Determining the target partition corresponding to each data in the second joined table based on the expansion times and the data output amount includes: performing a prefix sum calculation on the expansion times to obtain a calculated parameter; determining 0 and the calculated parameter as processed parameters; obtaining the quotient of the processed parameter divided by the data output amount; determining the target partition corresponding to the data based on the sum of the quotient and 1.
5. The method according to claim 4, wherein Perform a prefix sum calculation on the number of extensions to obtain the calculated parameter, including: obtaining the current partition corresponding to the number of extensions; performing a prefix sum calculation on the number of extensions in the same partition to obtain the first calculation processing sequence corresponding to each partition, where the first calculation processing sequence includes multiple processing results; sending the last processing result in the first calculation processing sequence corresponding to each partition to a preset partition to obtain the processing result sequence stored in the preset partition; performing a prefix sum calculation on the processing result sequence to obtain the calculated parameter.
6. The method according to claim 5, wherein Perform a prefix sum calculation on the processing result sequence to obtain the calculated parameter, including: performing a prefix sum calculation on the processing result sequence to obtain a second calculation processing sequence, where the second calculation processing sequence includes multiple calculation processing results; determining the target partition corresponding to each calculation processing result in the second calculation processing sequence; sending the calculation processing result to the target partition, and summing the calculation processing result and the number of extensions stored in the same target partition to obtain the calculated parameter.
7. The method according to claim 6, wherein Determining the target partition corresponding to each calculation processing result in the second calculation processing sequence includes: sorting the multiple calculation processing results included in the second calculation processing sequence based on the partition identifier corresponding to each calculation processing result to obtain result sorting information; determining the sum of the current sorting position where each calculation processing result is located and 1 as the target partition corresponding to the calculation processing result.
8. The method according to claim 3, wherein Expand each data based on the data output volume to obtain the expanded data, including: determining the total number of data output bits based on the data output volume; expanding each data based on the total number of data output bits to obtain the expanded data.
9. The method according to claim 8, wherein Expand each data based on the total number of data output bits to obtain the expanded data, including: obtaining all the data to be sent to the same target partition; detecting whether the sum of the data bits of all the data meets the total number of data output bits; if not, performing a data expansion operation based on the sum of the data bits of all the data and a preset padding bit to obtain the expanded data, where the number of data bits of the expanded data meets the total number of data output bits.
10. The method according to claim 3, wherein Rewrite the expanded data to obtain the expanded data for constructing the second expansion table, including: obtaining the remainder corresponding to the division operation between the processed parameter and the data output volume; determining the position where the data is stored in the target partition based on the remainder; deleting the redundant preset padding bits based on the position and the data output volume to obtain the sorted data; rewriting the preset padding bits included in the sorted data as the previous adjacent data to obtain the target data for constructing the second expansion table.
11. The method according to any one of claims 1-10, wherein Expand the first data table based on the attribute items in the second data table to obtain a first expanded table, including: performing grouped aggregation processing on the second data table based on the attribute items in the second data table to obtain a second aggregated table, where the second aggregated table includes the repetition degrees of each attribute value under the attribute item; performing a natural join on the second aggregated table and the first data table to obtain a first joined table; and expanding the first joined table to obtain a first expanded table.
12. The method according to any one of claims 1-10, wherein Add a data column to the first expanded table to obtain an adjusted expanded table, including: adding a data column with an initial value of 1 to the first expanded table to obtain an intermediate expanded table; comparing the two attribute values of two adjacent groups of data in the intermediate expanded table; when the two attribute values corresponding to two adjacent groups of data are the same, perform a summation process on the data columns in the two adjacent groups of data to obtain a summation data column, and determine the two attribute values and the summation data column to be used to form the data of the adjusted expanded table; when the two attribute values corresponding to two adjacent groups of data are different, determine the latter data in the two adjacent groups of data to be used to form the data of the adjusted expanded table.
13. The method according to any one of claims 1-10, wherein Associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table, including: obtaining a preset rule for sorting the first expanded table and the adjusted expanded table; using the preset rule to sort the first expanded table and the adjusted expanded table respectively to obtain a first sorted table and a second sorted table; Associate the first sorted table and the second sorted table to obtain an associated data table corresponding to the first data table and the second data table.
14. A processing device for a data table, comprising: An acquisition module, configured to acquire a first data table and a second data table, where the first data table and the second data table are stored in multiple distributed partitions, and the first data table and the second data table include the same attribute items; an expansion module, configured to expand the second data table based on the attribute items in the first data table to obtain a second expanded table; the expansion module is further configured to expand the first data table based on the attribute items in the second data table to obtain a first expanded table; A processing module, configured to add a data column to the first expanded table to obtain an adjusted expanded table, where the parameter values included in the data column are used to re-sort the first expanded table so that the data in the first expanded table is aligned with the data in the second expanded table; The processing module is further configured to associate the second expanded table and the adjusted expanded table to obtain an associated data table corresponding to the first data table and the second data table.
15. An electronic device, comprising: A memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method according to any one of claims 1-13 above is implemented.
16. A computer program product, comprising: A computer program, which, when executed by a processor of an electronic device, causes the processor to execute the steps in the method of any one of claims 1-13 above.
Citation Information
Patent Citations
Report generation method and device, electronic equipment and computer readable medium
CN113485781A
Data table processing method and device, equipment and storage medium
CN113672625A
Determining materialized view coverage for join transactions
US8359325B1