Data processing method, medium, device and computing device
By adding labels to the entity list and shuffle processing during the big data table association process, the problem of low correlation efficiency of big data tables is solved, the association efficiency and code maintainability are improved, and maintenance costs and data reading costs are reduced.
Patent Information
- Application Number
- CN202210258955.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-16
AI Technical Summary
In the process of big data table association, the prior art has problems with low correlation efficiency, especially when a large number of hotspot entities are included, resulting in data skew and high maintenance costs.
By adding tags to the entity list, distinguishing between hotspot entities and non-hotspot entities, and disassembling the hotspot entity primary keys during shuffle processing, the entity dimension table is associated with the entity list.
It improves the correlation efficiency of big data tables, reduces code maintenance costs, and only needs to read the entity list once, significantly reducing the data reading cost.
Smart Images

Figure CN114637748B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of big data technology. More specifically, embodiments of the present disclosure relate to a data processing method, medium, device, and computing device. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The descriptions herein are not admitted to be prior art merely because they are included in this section.
[0003] With the development of big data technology, for the vast amounts of data on the network, the vast amounts of data are usually stored on a big data computer cluster through a distributed file system and read and calculated by data practitioners in the form of relational tables. Big data tables used to store vast amounts of data include, for example, fact tables (such as entity detail tables) and dimension tables (such as entity dimension tables).
[0004] In related technologies, when associating big data tables containing a large number of hot entities, it is necessary to divide the big data table into a hot data table and a non-hot data table; connect the associated hot data tables to obtain target hot data, connect the associated non-hot data tables to obtain target non-hot data; and merge the target hot data and the target non-hot data to obtain the data after associating the associated big data tables. By associating big data tables in the above manner, there is a problem of low association efficiency. Summary of the Invention
[0005] The present disclosure provides a data processing method, medium, device, and computing device to solve the problem of low association efficiency in associating big data tables through related technologies.
[0006] In the first aspect of the embodiments of the present disclosure, a data processing method is provided, including:
[0007] In response to an execution instruction, obtain a hot entity table based on an entity detail table, where the entity detail table contains parameters of entities, and the hot entity table is used to store the primary keys of hot entities;
[0008] Add labels to the entity detail table according to the hot entity table, where the labels are used to mark whether the entities in the entity detail table are hot entities;
[0009] Split an entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table, where the entity dimension table is used to store dimension attribute information of entities;
[0010] Associate the entity dimension table with the entity detail table according to the labels, the hot entity dimension table, and the non-hot entity dimension table.
[0011] In a possible implementation, the entity dimension table is associated with the entity detail table according to the label, the hot entity dimension table, and the non-hot entity dimension table, including: establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity; establishing a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label; and associating the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship.
[0012] In a possible implementation, establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity includes: broadcasting the hot entity dimension table to the node where the entity detail table is located; and establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
[0013] In a possible implementation, establishing a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label includes: performing a shuffle process on the primary key corresponding to the hot entity label in the entity detail table to obtain the shuffled primary key; and establishing a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary key in the shuffled entity detail table.
[0014] In a possible implementation, the shuffle process includes concatenating a random value and / or other fields to the primary key.
[0015] In a possible implementation, establishing a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary key in the shuffled entity detail table includes: obtaining a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; obtaining a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; determining a target hash value in the second hash value that is the same as the first hash value; and allocating the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish the second connection relationship.
[0016] In a possible implementation, adding a label to the entity detail table according to the hot entity table includes: broadcasting the hot entity table to the node where the entity detail table is located, and adding a field corresponding to the label and the value of the field to the information of each entity in the entity detail table according to the primary key in the hot entity table and the primary key in the entity detail table, where the value includes a first preset value or a second preset value. The first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
[0017] In a possible implementation, according to the hot entity table, split the entity dimension table to obtain a hot entity dimension table and a non-hot entity dimension table, including: broadcasting the hot entity table to the node where the entity dimension table is located, extracting the dimension attributes of the hot entities and storing them in a table to obtain the hot entity dimension table, and extracting the dimension attributes of the non-hot entities and storing them in a table to obtain the non-hot entity dimension table.
[0018] In a possible implementation, obtain the hot entity table based on the entity detail table, including: sorting the entities in the entity detail table according to the parameters of the entities in the entity detail table to obtain a hot entity table corresponding to a preset number of hot entities.
[0019] In a possible implementation, the execution instruction includes identifiers of multiple entity dimension tables, and the data processing method further includes: parsing the execution instruction to obtain the identifiers of the multiple entity dimension tables; obtaining the corresponding entity dimension tables according to the identifiers; and respectively associating each obtained entity dimension table with the entity detail table.
[0020] In a second aspect, an embodiment of the present disclosure provides a data processing device, including:
[0021] An acquisition module, configured to obtain a hot entity table based on an entity detail table in response to an execution instruction, where the entity detail table includes parameters of entities, and the hot entity table is used to store the primary keys of hot entities;
[0022] An adding module, configured to add labels to the entity detail table according to the hot entity table, where the labels are used to mark whether the entities in the entity detail table are hot entities;
[0023] A splitting module, configured to split the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table, where the entity dimension table is used to store entity dimension attribute information;
[0024] An association module, configured to associate the entity dimension table with the entity detail table according to the labels, the hot entity dimension table, and the non-hot entity dimension table.
[0025] In a possible implementation, the association module is specifically configured to: establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary keys of the hot entities; establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the labels; and associate the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship.
[0026] In a possible implementation, when the association module is used to establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity, it is specifically used for: broadcasting the hot entity dimension table to the node where the entity detail table is located; establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
[0027] In a possible implementation, when the association module is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label, it is specifically used for: performing a shuffle process on the primary key corresponding to the hot entity label in the entity detail table to obtain the shuffled primary key; establishing a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary key in the shuffled entity detail table.
[0028] In a possible implementation, the shuffle process includes concatenating a random value and / or other fields to the primary key.
[0029] In a possible implementation, when the association module is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary key in the shuffled entity detail table, it is specifically used for: obtaining a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; obtaining a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; determining a target hash value in the second hash value that is the same as the first hash value; and allocating the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish a second connection relationship.
[0030] In a possible implementation, the addition module is specifically used for: broadcasting the hot entity table to the node where the entity detail table is located, and adding a field corresponding to the label and the value of the field to the information of each entity in the entity detail table according to the primary key in the hot entity table and the primary key in the entity detail table, where the value includes a first preset value or a second preset value. The first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
[0031] In a possible implementation, the splitting module is specifically used for: broadcasting the hot entity table to the node where the entity dimension table is located, taking out the dimensional attributes of the hot entity and storing them in a table to obtain the hot entity dimension table, and taking out the dimensional attributes of the non-hot entity and storing them in a table to obtain the non-hot entity dimension table.
[0032] In a possible implementation manner, the obtaining module is specifically configured to: based on the parameters of the entities in the entity detail list, sort the entities in the entity detail list according to the parameters, and obtain a hot entity table corresponding to a preset number of hot entities.
[0033] In a possible implementation manner, the execution instruction includes identifiers of multiple entity dimension tables, and the data processing device further includes a processing module, configured to: parse the execution instruction to obtain the identifiers of the multiple entity dimension tables; obtain the corresponding entity dimension tables according to the identifiers; and respectively associate each obtained entity dimension table with the entity detail list.
[0034] In a third aspect, an embodiment of the present disclosure provides a computing device, including: a processor, and a memory communicatively connected to the processor;
[0035] The memory stores computer execution instructions;
[0036] The processor executes the computer execution instructions stored in the memory to implement the data processing method as described in the first aspect of the present disclosure.
[0037] In a fourth aspect, an embodiment of the present disclosure provides a storage medium storing computer program instructions, which when executed, implement the data processing method as described in the first aspect of the present disclosure.
[0038] In a fifth aspect, an embodiment of the present disclosure provides a computer program product including a computer program, which when executed by a processor, implements the data processing method as described in the first aspect of the present disclosure.
[0039] The data processing method, medium, device, and computing device provided by the present disclosure obtain a hot entity table based on an entity detail list in response to an execution instruction; add labels to the entity detail list according to the hot entity table, where the labels are used to mark whether the entities in the entity detail list are hot entities; split the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table; and associate the entity dimension table with the entity detail list according to the labels, the hot entity dimension table, and the non-hot entity dimension table. Since the present disclosure distinguishes whether the entities in the entity detail list are hot entities by adding labels, instead of dividing the entity detail list into a hot entity detail list and a non-hot entity detail list, and associates the entity dimension table with the entity detail list according to the labels, the hot entity dimension table, and the non-hot entity dimension table, it can improve the association efficiency of large data tables, enhance the maintainability of the code, reduce the maintenance cost of the code, and only need to read the entity detail list once, which can greatly reduce the data reading cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0041] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present disclosure;
[0042] Figure 2 A flowchart of a data processing method provided for an embodiment of the present disclosure;
[0043] Figure 3 A schematic diagram of a computing node calculating hot songs provided for an embodiment of the present disclosure;
[0044] Figure 4 A schematic diagram of 100 computing nodes calculating hot songs provided for an embodiment of the present disclosure;
[0045] Figure 5 A schematic diagram of multi-table association provided for an embodiment of the present disclosure;
[0046] Figure 6 A schematic diagram of multi-table association provided for another embodiment of the present disclosure;
[0047] Figure 7 A flowchart of a data processing method provided for another embodiment of the present disclosure;
[0048] Figure 8 A schematic diagram of adding hot song tags to a song play detail list provided for an embodiment of the present disclosure;
[0049] Figure 9 A schematic diagram of obtaining a hot song dimension table provided for an embodiment of the present disclosure;
[0050] Figure 10 A schematic diagram of obtaining a non-hot song dimension table provided for an embodiment of the present disclosure;
[0051] Figure 11 A schematic diagram of establishing a first connection relationship provided for an embodiment of the present disclosure;
[0052] Figure 12 A schematic diagram of establishing a second connection relationship provided for an embodiment of the present disclosure;
[0053] Figure 13 A schematic diagram of the structure of a data processing apparatus provided for an embodiment of the present disclosure;
[0054] Figure 14 A schematic diagram of a storage medium provided for an embodiment of the present disclosure;
[0055] Figure 15 Schematic structural diagram of a computing device provided by an embodiment of the present disclosure.
[0056] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners
[0057] The principles and spirit of the present disclosure will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are only provided to enable those skilled in the art to better understand and then implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these implementation manners are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0058] Those skilled in the art know that the implementation manners of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. The data involved in the present disclosure can be data authorized by the user or fully authorized by all parties, and the implementation manners / embodiments of the present disclosure can be combined with each other.
[0059] According to an implementation manner of the present disclosure, a data processing method, medium, apparatus, and computing device are provided.
[0060] In this article, it should be understood that the terms involved:
[0061] Big data table, that is, massive data is stored on a big data cluster through a distributed file system and is provided for data practitioners to read and calculate in the form of a relational table;
[0062] Inner join, that is, two tables are associated, and records with equal associated fields are returned, and the union of the two tables is obtained;
[0063] Left join, that is, the left table is retained, and if the associated fields of the right table are equal, they are returned, and if they are not equal, null values are returned;
[0064] Shuffle is divided into two steps: Shuffle write and Shuffle read. It is a process of data redistribution. When two tables are joined, first, the two tables are read from the distributed file system. To ensure that data with the same key is processed on the same task, by calculating the hash value of the key, data with the same hash value is Shuffle written to the same partition, and the downstream stage Shuffle reads the data in this partition for corresponding calculations;
[0065] A large number of hot entities refer to the uneven distribution of the data volume of the two-table join fields. The data volume of some entities is significantly larger than that of the rest, and these entities are called hot entities. If these entities cannot be enumerated, they are called large-number hot entities;
[0066] Data skew means that due to the existence of hot entities, the processing time of tasks processing hot entities far exceeds that of the rest of the tasks;
[0067] Broadcast means that when a large table is joined with a small table, the small table (such as less than 10M) is broadcast to all nodes to ensure that data with the same key semantically is definitely processed on the same task;
[0068] Scattering means that if the hash values of the primary keys are the same, they will be shuffled to the same node. By appending a random value or other fields to the primary key, the hash values of the newly generated primary keys are evenly distributed.
[0069] In addition, the quantity of any element in the accompanying drawings is for illustration only and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0070] Next, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be elaborated in detail. Summary of the Invention
[0072] The inventor of the present invention found that when performing a large data table association on a large data table containing a large number of hot entities, in a related art, two large data tables are directly joined. If there are hot entities in the join primary key, data skew problems will occur. In another related art, when performing an association on two large data tables in the case where the hot entities are enumerable, it is necessary to scatter the primary keys of the hot entities and perform hard coding on the information corresponding to the other fields of the hot entities. Therefore, the code maintenance cost is high. In yet another related art, when performing an association on two large data tables in the case where the hot entities are not enumerable, it is necessary to divide the large data table into a hot data table and a non-hot data table; then, join the hot data tables of the associated large data tables to obtain the target hot data, join the non-hot data tables of the associated large data tables to obtain the target non-hot data; finally, merge the target hot data and the target non-hot data to obtain the data after the association of the associated large data tables. When performing a large data table association through this related art, there is a problem of low association efficiency. In addition, if multiple large data tables are associated through this related art, the code volume increases by 2 n and the number of reads of the entity detail table is also 2 n , resulting in high code maintenance costs and low association efficiency.
[0073] Based on the above problems, the present disclosure provides a data processing method, medium, device, and computing device. By adding tags to distinguish whether the entities in the entity detail table are hot entities, when performing shuffle processing, the primary keys of the hot entities in the entity detail table are shuffled, so as to associate at least one entity dimension table with the entity detail table. Therefore, the association efficiency of large data tables can be improved, the code amount increases linearly, the maintainability of the code can be enhanced, the maintenance cost of the code can be reduced, and it can be applied to scenarios where hot entities can be enumerated or hot entities cannot be enumerated.
[0074] Overview of Application Scenarios
[0075] First, refer to Figure 1 to illustrate the application scenarios of the solution provided by the present disclosure by way of example. Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. As Figure 1 shown, in this application scenario, the distributed file system includes the server 101 corresponding to the master control node and the servers 102 corresponding to multiple nodes for storing large data tables. The entity detail table and the entity dimension table are stored in the corresponding nodes on the server 102. The server 101 responds to the execution instruction, obtains the storage location information of the entity detail table and the entity dimension table on the server 102. The server 101 obtains the entity detail table and the entity dimension table from the server 102 according to the storage location information, obtains the hot entity table according to the entity detail table, and associates the entity dimension table with the entity detail table based on the hot entity table. Among them, the specific implementation process of the server 101 obtaining the hot entity table according to the entity detail table and associating the entity dimension table with the entity detail table based on the hot entity table can refer to the solutions of the following embodiments.
[0076] It should be noted that Figure 1 is only a schematic diagram of an application scenario provided by an embodiment of the present disclosure. The embodiments of the present disclosure do not limit Figure 1 the devices included in Figure 1 nor the positional relationship between the devices in
[0077] Exemplary Method
[0078] Next, in combination with the Figure 1 application scenario, refer to Figure 2 to describe the data processing method according to an exemplary embodiment of the present disclosure. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0079] First, introduce the data processing method through specific embodiments.
[0080] Figure 2 The flowchart of the data processing method provided by an embodiment of the present disclosure. The method of the embodiment of the present disclosure can be applied to a computing device, which can be a server or a server cluster, etc. As Figure 2 shown, the method of the embodiment of the present disclosure includes:
[0081] S201. In response to an execution instruction, obtain a hot entity table based on an entity detail list.
[0082] Among them, the entity detail list contains parameters of entities, and the hot entity table is used to save the primary keys of hot entities.
[0083] In the embodiment of the present disclosure, the execution instruction is, for example, an execution instruction of an offline scheduling task. The execution instruction can be input by a user to the computing device executing the embodiment of the method, or sent by other devices to the computing device executing the embodiment of the method. Exemplarily, the hot entity is, for example, a hot song, and the entity detail list is, for example, a song play detail list. The data volume of the song play detail list is tens of billions per day. The song play detail list contains parameters of songs. Specifically, the parameters of songs are, for example, song identification information (identity document, id). The song play detail list may also contain other information, such as the artist id of the singer of the song, the city id of the city where the song is played, and the user id of the user who plays the song, etc. The hot entity table is, for example, used to save the primary key of the hot song. Among them, the primary key of the hot song is, for example, the song id of the hot song. In this step, in response to the execution instruction, the hot entity table can be obtained based on the song play detail list.
[0084] Further, optionally, obtaining the hot entity table based on the entity detail list may include: sorting the entities in the entity detail list according to the parameters of the entities in the entity detail list, and obtaining a hot entity table corresponding to a preset number of hot entities.
[0085] Exemplarily, the parameter of the entity in the entity detail list is, for example, the entity id, and the preset number is, for example, 10,000. The data volume of the entity id granularity can be calculated from the entity detail list, and the data volume can be sorted in descending order to obtain the first preset number (for example, represented by topN) of entity ids and write them into a table to obtain the hot entity table. It can be understood that the entities corresponding to the first preset number of entity ids are hot entities. Taking the entity detail list as the song play detail list as an example, if you want to get the top 10,000 songs in terms of play volume, refer to Figure 1The song play list is stored in a distributed file system, with a data volume on the scale of tens of billions. If it is stored on 1000 nodes, each node stores tens of millions of play data, and each node can count the play volume of each song. If 100 nodes are used as computing nodes, each computing node can calculate the play volume of each song on 10 corresponding nodes among the 1000 nodes. Figure 3 It is a schematic diagram of a computing node calculating hot songs provided by an embodiment of the present disclosure. As Figure 3 shown, the first computing node among the 100 computing nodes is computing node 1. Taking computing node 1 as an example, computing node 1 can calculate the play volume of each song on nodes 1 to 10 among the 1000 nodes, and can obtain the songs with the top 10,000 (i.e., the first 10,000 songs) of the play volume corresponding to nodes 1 to 10 respectively. Then, merge the songs with the top 10,000 of the play volume corresponding to these 10 nodes respectively, and sort the play volume of each song in descending order, and the hot song id with the top 10,000 of the play volume corresponding to computing node 1 can be obtained. By analogy, the hot song id with the top 10,000 of the play volume corresponding to each computing node among the 100 computing nodes can be obtained. Specifically, Figure 4 It is a schematic diagram of 100 computing nodes calculating hot songs provided by an embodiment of the present disclosure. As Figure 4 shown, each of the 100 computing nodes (i.e., computing node 1 to computing node 100) can calculate and obtain the songs with the top 10,000 of the play volume corresponding to 10 corresponding nodes among the 1000 nodes, and merge the calculation results of these 100 computing nodes, that is, merge the songs with the top 10,000 of the play volume corresponding to these 100 computing nodes respectively, and sort the play volume of each song in descending order, and the final hot song id with the top 10,000 of the play volume can be obtained and written into a table, and a hot song id dimension table (i.e., a hot entity table) can be obtained.
[0086] It should be noted that the entity detail table is read once in this step, and there is no need to read the entity detail table in the subsequent steps. That is, the embodiment of the present disclosure only needs to read the entity detail table once, which can greatly reduce the cost of data reading.
[0087] S202. Add labels to the entity detail table according to the hot entity table, where the labels are used to mark whether the entities in the entity detail table are hot entities.
[0088] In this step, after obtaining the hot entity table, labels can be added to the entity detail table according to the hot entity table to distinguish whether the entities in the entity detail table are hot entities through the labels. For how to add labels to the entity detail table according to the hot entity table specifically, reference can be made to the subsequent embodiments, which will not be elaborated here.
[0089] S203. Split the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table.
[0090] Among them, the entity dimension table is used to store the dimension attribute information of entities.
[0091] Exemplarily, the entity dimension table is, for example, a song dimension table. The daily data volume of the song dimension table is in hundreds of millions. The song dimension table is used to store the dimension attribute information of songs. The dimension attribute information of songs is, for example, song name, album name to which the song belongs, song style, etc. In this step, after obtaining the hot entity table, the entity dimension table can be split according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table. Among them, the hot entity dimension table is used to store the dimension attribute information of hot entities, and the non-hot entity dimension table is used to store the dimension attribute information of non-hot entities. For how to split the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table specifically, reference can be made to the subsequent embodiments and will not be elaborated here.
[0092] It should be noted that the present disclosure does not limit the execution order of steps S202 and S203. Step S202 can be executed first, and then step S203; or, step S203 can be executed first, and then step S202.
[0093] S204. Associate the entity dimension table with the entity detail table according to the tags, the hot entity dimension table, and the non-hot entity dimension table.
[0094] In this step, after obtaining the hot entity dimension table and the non-hot entity dimension table, the entity dimension table can be associated with the entity detail table according to the tags, the hot entity dimension table, and the non-hot entity dimension table. For how to associate the entity dimension table with the entity detail table according to the tags, the hot entity dimension table, and the non-hot entity dimension table specifically, reference can be made to the subsequent embodiments and will not be elaborated here. After associating the entity dimension table with the entity detail table, the dimension attribute information of hot entities or the dimension attribute information of non-hot entities can be obtained for data analysis.
[0095] The data processing method provided by the embodiments of the present disclosure obtains a hot entity table based on an entity detail list in response to an execution instruction; adds tags to the entity detail list according to the hot entity table, where the tags are used to mark whether the entities in the entity detail list are hot entities; splits the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table; and associates the entity dimension table with the entity detail list according to the tags, the hot entity dimension table, and the non-hot entity dimension table. Since the embodiments of the present disclosure distinguish whether the entities in the entity detail list are hot entities by adding tags, instead of dividing the entity detail list into a hot entity detail list and a non-hot entity detail list, and associate the entity dimension table with the entity detail list according to the tags, the hot entity dimension table, and the non-hot entity dimension table, the association efficiency of large data tables can be improved, the maintainability of the code can be enhanced, the maintenance cost of the code can be reduced, and the entity detail list only needs to be read once, which can greatly reduce the data reading cost.
[0096] Based on the above embodiments, considering that multiple entity dimension tables need to be associated with the entity detail list, correspondingly, the execution instruction includes the identifiers of multiple entity dimension tables. Therefore, the data processing method provided by the embodiments of the present disclosure may further include: parsing the execution instruction to obtain the identifiers of multiple entity dimension tables; obtaining the corresponding entity dimension tables according to the identifiers; and respectively associating each obtained entity dimension table with the entity detail list.
[0097] Exemplarily, the identifier of the entity dimension table is, for example, the table name or table ID of the entity dimension table. By parsing the execution instruction, the identifiers of multiple entity dimension tables can be obtained, and the corresponding entity dimension tables can be obtained according to the identifiers. By repeatedly executing steps S201 to S204, each obtained entity dimension table can be respectively associated with the entity detail list. Specifically, Figure 5 is a schematic diagram of multi-table association provided by an embodiment of the present disclosure. As Figure 5 shown, taking the entity dimension table B1 as an example, by executing steps S201 to S204, the entity dimension table B1 can be associated with the entity detail list A, where the primary key 1 (key1) is the primary key of the entity dimension table B1, and can also be understood as the connection key between the entity dimension table B1 and the entity detail list A. By analogy, by repeatedly executing steps S201 to S204, the entity dimension tables B2 to Bn can be respectively associated with the entity detail list A. Figure 6 is a schematic diagram of multi-table association provided by another embodiment of the present disclosure. As Figure 6As shown in the figure, the song play detail table is an entity detail table, and the song dimension table, artist dimension table, city dimension table, and user dimension table are all entity dimension tables. Among them, taking the song dimension table as an example, the primary key (PK) of the song dimension table is the foreign key (FK) of the song play detail table. By repeatedly executing steps S201 to S204, the song dimension table, artist dimension table, city dimension table, and user dimension table can be respectively associated with the song play detail table. Through the above method, the problem of associating multiple large data tables and the existence of a large number of hot entities is solved, the code amount increases linearly, the association efficiency of large data tables is improved, and the code maintenance cost is reduced.
[0098] Figure 7 The flowchart of the data processing method provided by another embodiment of the present disclosure. On the basis of the above embodiment, based on the application scenario as Figure 1 shown, the embodiments of the present disclosure further illustrate the data processing method. As Figure 7 shown, the method of the embodiments of the present disclosure may include:
[0099] S701. In response to an execution instruction, obtain a hot entity table based on the entity detail table.
[0100] Among them, the entity detail table contains the primary key of the entity, and the hot entity table is used to save the primary key of the hot entity.
[0101] For the specific description of this step, reference may be made to the relevant description of S201 in the embodiment shown in Figure 2 and details are not described herein again.
[0102] In the embodiments of the present disclosure, Figure 2 step S202 in may further include the following S702 step:
[0103] S702. Broadcast the hot entity table to the node where the entity detail table is located, and according to the primary key in the hot entity table and the primary key in the entity detail table, add a field corresponding to the label and the value of the field to the information of each entity in the entity detail table.
[0104] Among them, the value includes a first preset value or a second preset value. The first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
[0105] Exemplarily, the first preset value is, for example, 1, and the second preset value is, for example, 2. In this step, after obtaining the hot entity table, the hot entity table can be broadcast to the nodes where the entity detail table is located. According to the primary key in the hot entity table and the primary key in the entity detail table, fields corresponding to the tags and the values of the fields are added to the information of each entity in the entity detail table. For example, after adding tags to the information of each entity in the entity detail table, the tag of the hot entity in the entity detail table is 1, and the tag of the non-hot entity in the entity detail table is 2. Exemplarily, Figure 8 is a schematic diagram of adding hot song tags to the song play detail table provided by an embodiment of the present disclosure, as Figure 8 shown. Referring to the example in step S201, based on the song play detail table, a hot song id dimension table can be obtained. The hot song id dimension table is broadcast to each node where the song play detail table is located (i.e., Figure 8 the nodes 1 to 1000 in
[0106] In the embodiments of the present disclosure, Figure 2 step S203 in
[0107] S703 can further include the following step S703:
[0108] Exemplarily, Figure 9 is a schematic diagram of obtaining a hot song dimension table provided by an embodiment of the present disclosure, as Figure 9 shown. Based on the above embodiment, the entity dimension table is, for example, a song dimension table. Assume that the song dimension table is stored on the 200 nodes of the distributed file system. The hot song id dimension table is broadcast to the 200 nodes where the song dimension table is located (i.e., Figure 9 the nodes 1 to 200 in
[0109] Exemplarily, Figure 10 is a schematic diagram of obtaining a non-hot song dimension table provided by an embodiment of the present disclosure, as Figure 10 shown. Based on the above embodiment, since the hot song id dimension table has been broadcast to the 200 nodes where the song dimension table is located (i.e., Figure 10Nodes 1 to 200), so the non-hot song IDs in the song dimension table can be determined based on the hot song IDs stored in the hot song ID dimension table, and then the dimension attributes of the non-hot songs can be retrieved and stored in a table to obtain the non-hot song dimension table (which can also be called the non-hot song dimension attribute table) on each node.
[0110] In the embodiments of the present disclosure, Figure 2 Step S204 in may further include the following three steps of S704 to S706:
[0111] S704. Establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
[0112] In this step, after obtaining the hot entity dimension table, a first connection relationship between the entity detail table and the hot entity dimension table can be established according to the primary key of the hot entity in the hot entity dimension table.
[0113] Further, optionally, establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity includes: broadcasting the hot entity dimension table to the nodes where the entity detail table is located; establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
[0114] Exemplarily, Figure 11 is a schematic diagram of establishing the first connection relationship provided by an embodiment of the present disclosure. As Figure 11 shown, based on the above embodiment, the hot entity dimension table is, for example, the hot song dimension table. The hot song dimension table is broadcast to each node where the song play detail table is located (i.e., Figure 11 Nodes 1 to 1000); a first connection relationship between the song play detail table and the hot song dimension table is established according to the primary key of the hot song in the hot song dimension table, so that the song play detail table can obtain the dimension attributes of the hot song by querying.
[0115] S705. Establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label.
[0116] In this step, after obtaining the non-hot entity dimension table, a second connection relationship between the non-hot entity dimension table and the entity detail table can be established in a Shuffle connection manner according to the label in the entity detail table.
[0117] Further, optionally, according to the tags, a second connection relationship is established between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner, including: shuffling the primary keys corresponding to the hot entity tags in the entity detail table to obtain the shuffled primary keys; establishing a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table. The shuffling process includes concatenating random values and / or other fields to the primary keys.
[0118] Exemplarily, Figure 12 FIG. is a schematic diagram of establishing a second connection relationship provided by an embodiment of the present disclosure. As Figure 12 shown, based on the above embodiment, the entity detail table is, for example, a song play detail table, and the non-hot entity dimension table is, for example, a non-hot song dimension table. A random value and other fields are concatenated to the primary keys corresponding to the hot entity tags in the song play detail table, or a random value is concatenated to the primary keys corresponding to the hot entity tags in the song play detail table, or other fields are concatenated to the primary keys corresponding to the hot entity tags in the song play detail table to obtain the shuffled primary keys corresponding to the hot entity tags. The primary keys in the shuffled entity detail table include the primary keys corresponding to the hot entity tags and the primary keys corresponding to the non-hot entity tags. Therefore, a second connection relationship can be established between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the primary keys in the shuffled entity detail table.
[0119] By shuffling the primary keys corresponding to the hot entity tags in the entity detail table, it can be ensured that the data connection with the non-hot entities in the entity dimension table is not established, and the problem of data skew can be avoided.
[0120] It should be noted that step S704 is preferentially executed, and then step S705 is executed.
[0121] Further, optionally, establishing a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table includes: obtaining a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; obtaining a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; determining a target hash value in the second hash values that is the same as the first hash value; and allocating the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish a second connection relationship.
[0122] Exemplarily, the first hash value corresponding to the primary key of the first entity in the non-hot entity dimension table can be obtained through a preset algorithm, and the second hash value corresponding to the primary key of the second entity in the entity detail table can be obtained. Exemplarily, the first hash value includes, for example, a1b1c1 and a3b3c3, and the second hash value includes, for example, a1b1c1, a2b2c2, a3b3c3, and a4b4c4. Then, the target hash values in the second hash value that are the same as the first hash value can be determined as a1b1c1 and a3b3c3. Therefore, the first entity in the non-hot entity dimension table corresponding to the first hash value a1b1c1 and the second entity in the entity detail table corresponding to the target hash value a1b1c1 can be assigned to the same node, and the first entity in the non-hot entity dimension table corresponding to the first hash value a3b3c3 and the second entity in the entity detail table corresponding to the target hash value a3b3c3 can be assigned to the same node, thereby establishing the second connection relationship.
[0123] S706. According to the first connection relationship and the second connection relationship, associate the entity dimension table with the entity detail table.
[0124] In this step, after obtaining the first connection relationship and the second connection relationship, the entity dimension table can be associated with the entity detail table according to the first connection relationship and the second connection relationship. After associating the entity dimension table with the entity detail table, the dimension attribute information of the hot entity or the dimension attribute information of the non-hot entity can be obtained for data analysis.
[0125] The data processing method provided by the embodiments of the present disclosure obtains a hot entity table based on an entity detail table in response to an execution instruction; broadcasts the hot entity table to the node where the entity detail table is located, and adds fields corresponding to tags and the values of the fields to the information of each entity in the entity detail table according to the primary keys in the hot entity table and the primary keys in the entity detail table; broadcasts the hot entity table to the node where the entity dimension table is located, extracts the dimension attributes of the hot entities and stores them in a table to obtain a hot entity dimension table, and extracts the dimension attributes of the non-hot entities and stores them in a table to obtain a non-hot entity dimension table; establishes a first connection relationship between the entity detail table and the hot entity dimension table according to the primary keys of the hot entities; establishes a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the tags; and associates the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship. Since the embodiments of the present disclosure distinguish whether the entities in the entity detail table are hot entities in the form of adding tags, and do not need to divide the entity detail table into a hot entity detail table and a non-hot entity detail table, when establishing the second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner, the primary keys corresponding to the hot entity tags in the entity detail table are shuffled, so as to associate the entity dimension table with the entity detail table. Therefore, the association efficiency of large data tables can be improved, the code amount increases linearly, the maintainability of the code can be enhanced, the maintenance cost of the code can be reduced, and it can be applied to scenarios where hot entities can be enumerated or hot entities cannot be enumerated.
[0126] Exemplary Device
[0127] After introducing the media of the exemplary embodiments of the present disclosure, next, reference is made to Figure 13 to describe the data processing device of the exemplary embodiments of the present disclosure. The device of the exemplary embodiments of the present disclosure can implement each process in the foregoing data processing method embodiments and achieve the same functions and effects.
[0128] Figure 13 FIG. is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure. As Figure 13 shown, the data processing device 1300 of the embodiments of the present disclosure includes: an acquisition module 1301, an addition module 1302, a splitting module 1303, and an association module 1304. Among them:
[0129] The acquisition module 1301 is configured to obtain a hot entity table based on an entity detail table in response to an execution instruction, where the entity detail table includes entity parameters, and the hot entity table is used to store the primary keys of hot entities.
[0130] The addition module 1302 is configured to add tags to the entity detail table according to the hot entity table, where the tags are used to label whether the entities in the entity detail table are hot entities.
[0131] The splitting module 1303 is used to split the entity dimension table according to the hot entity table, obtaining a hot entity dimension table and a non-hot entity dimension table. The entity dimension table is used to store the dimension attribute information of entities.
[0132] The association module 1304 is used to associate the entity dimension table with the entity detail table according to tags, the hot entity dimension table, and the non-hot entity dimension table.
[0133] In a possible implementation manner, the association module 1304 may specifically be used to: establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity; establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to tags; and associate the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship.
[0134] In a possible implementation manner, when the association module 1304 is used to establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity, it may specifically be used to: broadcast the hot entity dimension table to the nodes where the entity detail table is located; and establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
[0135] In a possible implementation manner, when the association module 1304 is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to tags, it may specifically be used to: perform a shuffling process on the primary keys corresponding to the hot entity tags in the entity detail table to obtain the shuffled primary keys; and establish a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table.
[0136] In a possible implementation manner, the shuffling process includes concatenating a random value and / or other fields to the primary key.
[0137] In a possible implementation manner, when the association module 1304 is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table, it may specifically be used to: obtain a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; obtain a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; determine a target hash value in the second hash values that is the same as the first hash value; and allocate the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish the second connection relationship.
[0138] In a possible implementation, the adding module 1302 may be specifically configured to: broadcast the hot entity table to the node where the entity detail table is located, and according to the primary keys in the hot entity table and the primary keys in the entity detail table, add fields corresponding to the labels and the values of the fields to the information of each entity in the entity detail table, where the values include a first preset value or a second preset value. The first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
[0139] In a possible implementation, the splitting module 1303 may be specifically configured to: broadcast the hot entity table to the node where the entity dimension table is located, extract the dimension attributes of the hot entities and store them in a table to obtain the hot entity dimension table, and extract the dimension attributes of the non-hot entities and store them in a table to obtain the non-hot entity dimension table.
[0140] In a possible implementation, the obtaining module 1301 may be specifically configured to: based on the parameters of the entities in the entity detail table, sort the entities in the entity detail table according to the parameters to obtain a hot entity table corresponding to a preset number of hot entities.
[0141] In a possible implementation, the execution instruction includes identifiers of multiple entity dimension tables. The data processing device further includes a processing module 1305, configured to: parse the execution instruction to obtain the identifiers of the multiple entity dimension tables; obtain the corresponding entity dimension tables according to the identifiers; and respectively associate each obtained entity dimension table with the entity detail table.
[0142] The device according to the embodiments of the present disclosure can be used to execute the solutions of the data processing methods in any of the above method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.
[0143] Exemplary Medium
[0144] After introducing the methods of the exemplary embodiments of the present disclosure, next, reference is made to Figure 14 to describe the storage media of the exemplary embodiments of the present disclosure.
[0145] Figure 14 FIG. is a schematic diagram of a storage medium provided by an embodiment of the present disclosure. Referring to Figure 14 as shown, the storage medium 1400 stores a program product for implementing the above method according to the embodiments of the present disclosure. It may be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.
[0146] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples of the readable storage medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium.
[0148] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN).
[0149] Exemplary Computing Device
[0150] After introducing the methods, media, and apparatuses of the exemplary embodiments of the present disclosure, next, reference is made to Figure 15 illustrate the computing device of the exemplary embodiments of the present disclosure.
[0151] Figure 15 The shown computing device 1500 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0152] Figure 15 The structural schematic diagram of the computing device provided for an embodiment of the present disclosure is as shown in Figure 15As shown, the computing device 1500 is presented in the form of a general-purpose computing device. The components of the computing device 1500 may include, but are not limited to: at least one processing unit 1501, at least one storage unit 1502, and a bus 1503 that connects different system components (including the processing unit 1501 and the storage unit 1502). Among them, computer-executable instructions are stored in at least one storage unit 1502; at least one processing unit 1501 includes a processor that executes the computer-executable instructions to implement the method described above.
[0153] The bus 1503 includes a data bus, a control bus, and an address bus.
[0154] The storage unit 1502 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 15021 and / or a cache memory 15022, and may further include a readable medium in the form of non-volatile memory, such as a read-only memory (ROM) 15023.
[0155] The storage unit 1502 may also include a program / utility 15025 having a set (at least one) of program modules 15024. Such program modules 15024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0156] The computing device 1500 may also communicate with one or more external devices 1504 (such as a keyboard, a pointing device, etc.). Such communication may be carried out through an input / output (I / O) interface 1505. And, the computing device 1500 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1506. As Figure 15 shown, the network adapter 1506 communicates with other modules of the computing device 1500 through the bus 1503. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in combination with the computing device 1500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0157] It should be noted that although several units / modules or sub-units / sub-modules of the data processing device are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.
[0158] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the illustrated operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.
[0159] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. Such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A data processing method, comprising: In response to an execution instruction, obtaining a hot entity table based on an entity detail table, where the entity detail table contains parameters of entities, and the hot entity table is used to store primary keys of hot entities; Adding a label to the entity detail table according to the hot entity table, where the label is used to mark whether the entities in the entity detail table are hot entities; Splitting an entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table, where the entity dimension table is used to store dimension attribute information of entities; Establishing a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity; Establishing a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label; Associating the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship.
2. The data processing method according to claim 1, where establishing the first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity includes: Broadcasting the hot entity dimension table to the node where the entity detail table is located; Establishing the first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
3. The data processing method according to claim 1, where establishing the second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label includes: Performing a shuffling process on the primary keys corresponding to the hot entity labels in the entity detail table to obtain shuffled primary keys; Establishing the second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table.
4. The data processing method according to claim 3, where the shuffling process includes concatenating a random value and / or other fields to the primary key.
5. The data processing method according to claim 3, where establishing the second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table includes: Obtaining a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; Obtaining a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; Determining a target hash value in the second hash values that is the same as the first hash value; Assigning the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish the second connection relationship.
6. The data processing method according to claim 1, where adding a label to the entity detail table according to the hot entity table includes: Broadcast the hot entity table to the node where the entity detail table is located. According to the primary key in the hot entity table and the primary key in the entity detail table, add the field corresponding to the label and the value of the field to the information of each entity in the entity detail table. The value includes a first preset value or a second preset value. The first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
7. The data processing method according to claim 1, wherein the splitting the entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table includes: Broadcast the hot entity table to the node where the entity dimension table is located, extract the dimension attributes of the hot entities and write them to a table to obtain the hot entity dimension table, and extract the dimension attributes of the non-hot entities and write them to a table to obtain the non-hot entity dimension table.
8. The data processing method according to any one of claims 1 to 7, wherein the obtaining the hot entity table based on the entity detail table includes: Based on the parameters of the entities in the entity detail table, sort the entities in the entity detail table according to the parameters to obtain a hot entity table corresponding to a preset number of hot entities.
9. The data processing method according to any one of claims 1 to 7, wherein the execution instruction includes identifiers of multiple entity dimension tables, and the data processing method further includes: Parse the execution instruction to obtain the identifiers of the multiple entity dimension tables; Obtain the corresponding entity dimension tables according to the identifiers; Associate each obtained entity dimension table with the entity detail table respectively.
10. A data processing device, comprising: An obtaining module, configured to obtain a hot entity table based on an entity detail table in response to an execution instruction. The entity detail table includes parameters of entities, and the hot entity table is used to store the primary keys of hot entities; An adding module, configured to add a label to the entity detail table according to the hot entity table. The label is used to mark whether the entities in the entity detail table are hot entities; A splitting module, configured to split an entity dimension table according to the hot entity table to obtain a hot entity dimension table and a non-hot entity dimension table. The entity dimension table is used to store entity dimension attribute information; An associating module, configured to establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity; Establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label; Associate the entity dimension table with the entity detail table according to the first connection relationship and the second connection relationship.
11. The data processing device according to claim 10, when the associating module is configured to establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity, specifically configured to: Broadcast the hot entity dimension table to the node where the entity detail table is located; Establish a first connection relationship between the entity detail table and the hot entity dimension table according to the primary key of the hot entity.
12. The data processing device according to claim 10, when the association module is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table in a Shuffle connection manner according to the label, specifically: Perform a shuffling process on the primary keys corresponding to the hot entity labels in the entity detail table to obtain the shuffled primary keys; Establish a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table.
13. The data processing device according to claim 12, wherein the shuffling process includes concatenating a random value and / or other fields to the primary key.
14. The data processing device according to claim 12, when the association module is used to establish a second connection relationship between the non-hot entity dimension table and the entity detail table according to the primary keys in the shuffled entity detail table, specifically: Obtain a first hash value corresponding to the primary key of a first entity in the non-hot entity dimension table, where the first entity is any entity in the non-hot entity dimension table; Obtain a second hash value corresponding to the primary key of a second entity in the entity detail table, where the second entity is any entity in the entity detail table; Determine a target hash value in the second hash values that is the same as the first hash value; Assign the first entity in the non-hot entity dimension table corresponding to the first hash value and the second entity in the entity detail table corresponding to the target hash value to the same node to establish the second connection relationship.
15. The data processing device according to claim 10, the adding module specifically: Broadcast the hot entity table to the nodes where the entity detail table is located, and according to the primary keys in the hot entity table and the primary keys in the entity detail table, add a field corresponding to the label and the value of the field to the information of each entity in the entity detail table, where the value includes a first preset value or a second preset value, the first preset value is used to indicate that the entity included in the entity detail table is a hot entity, and the second preset value is used to indicate that the entity included in the entity detail table is a non-hot entity.
16. The data processing device according to claim 10, the splitting module specifically: Broadcast the hot entity table to the nodes where the entity dimension table is located, extract the dimensional attributes of the hot entities and store them in a table to obtain the hot entity dimension table, and extract the dimensional attributes of the non-hot entities and store them in a table to obtain the non-hot entity dimension table.
17. The data processing device according to any one of claims 10 to 16, the obtaining module specifically: Based on the parameters of the entities in the entity detail table, sort the entities in the entity detail table according to the parameters to obtain a hot entity table corresponding to a preset number of hot entities.
18. The data processing device according to any one of claims 10 to 16, the execution instruction includes identifiers of multiple entity dimension tables, and the data processing device further includes a processing module for: Parse the execution instruction to obtain the identifiers of the multiple entity dimension tables; Obtain the corresponding entity dimension table according to the said identifier; Associate each obtained entity dimension table with the said entity detail table respectively.
19. A computing device, comprising: A processor, and a memory communicatively connected to the said processor; The said memory stores computer-executable instructions; The said processor executes the computer-executable instructions stored in the said memory to implement the data processing method as described in any one of claims 1 to 9.
20. A storage medium, in which computer program instructions are stored, and when the said computer program instructions are executed, the data processing method as described in any one of claims 1 to 9 is implemented.
21. A computer program product, including a computer program, and when the said computer program is executed by a processor, the data processing method as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Database table processing method and device, electronic equipment and storage medium
CN111324604A
Data query method and device and electronic equipment
CN112835966A