Data Processing Method, Device, Readable Medium and Equipment Based on Data Warehouse
By selecting core data tables with a large number of associated tables and high field similarity in the data warehouse, building temporary tables and performing data processing, the low efficiency and high cost problems caused by multiple associations in the data warehouse are solved, and more efficient data processing and resource utilization are achieved.
Patent Information
- Application Number
- CN202210281609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-21
AI Technical Summary
Multiple associations of data tables in data warehouses lead to problems such as low development efficiency, huge cost of SQL execution and excessive execution time.
By obtaining the data table set from the data warehouse, selecting core data tables with the number of associated tables exceeding the threshold and the field similarity is high, building temporary tables, and responding to data processing operations based on temporary tables.
It improves the development efficiency of data warehouses, shortens the response time of SQL execution, reduces the cost of SQL execution, and saves memory resources and computing resources.
Smart Images

Figure CN114610758B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to a data processing method, device, readable medium and equipment based on a data warehouse. Background Art
[0002] A data warehouse is a strategic collection that provides support for all types of data for the decision-making processes at all levels of an enterprise. It is a single data storage created for analytical reporting and decision support purposes.
[0003] The data in a data warehouse often exists in the form of tables. In order to execute Structured Query Language (SQL) in the data warehouse, it is often necessary to perform multiple joins on the tables, which brings great research and development time to developers and extremely low development efficiency. Moreover, due to the large number of joins, SQL execution also brings problems such as huge machine consumption costs and excessive execution time.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of this application is to provide a data processing method, device, readable medium and equipment based on a data warehouse, which at least to some extent overcomes the technical problems in the related art such as low development efficiency of the data warehouse, huge consumption costs of SQL execution in the data warehouse, and excessive execution time.
[0006] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.
[0007] According to one aspect of the embodiments of this application, a data processing method based on a data warehouse is provided, including:
[0008] Obtain a data table set from the data warehouse, where the data table set includes the data tables in the data warehouse;
[0009] Select at least one core data table from the data table set, where the number of associated tables of the core data table exceeds a first threshold, and the associated tables of the core data table are data tables whose field similarity with the core data table exceeds a second threshold;
[0010] Construct a temporary table according to the fields included in the at least one selected core data table;
[0011] Respond to a data processing operation based on the temporary table.
[0012] In some embodiments of the present application, based on the above technical solutions, after constructing a temporary table according to the fields included in at least one selected core data table, the method further includes:
[0013] Obtain other data tables in the data table set except the at least one core data table;
[0014] Select a target data table with the shortest path between the other data tables and the at least one core data table;
[0015] Add the fields included in the selected target data table to the temporary table.
[0016] In some embodiments of the present application, based on the above technical solutions, the method for selecting at least one core data table from the data table set includes:
[0017] Use a clustering algorithm to cluster the data tables in the data table set to obtain multiple data table subsets;
[0018] Select data tables with the number of associated tables exceeding a first threshold from the multiple data table subsets as core data tables.
[0019] In some embodiments of the present application, based on the above technical solutions, before selecting at least one core data table from the data table set, the method further includes:
[0020] Use a locality-sensitive hashing algorithm to calculate the field similarity between each data table in the data table set;
[0021] Determine data tables with the field similarity exceeding a second threshold as associated tables;
[0022] Construct an association matrix based on the associated tables of each data table, and the association matrix is used to represent the association situation between each data table in the data table set and other data tables.
[0023] In some embodiments of the present application, based on the above technical solutions, the method for selecting at least one core data table from the data table set includes:
[0024] Convert the association matrix into association key-value pairs;
[0025] Extract node pairs from the association key-value pairs, and the node pairs include combinations of two associated tables of the data table;
[0026] Construct an undirected graph neural network with the associated tables included in the node pairs as nodes, and train the undirected graph neural network to obtain the feature vectors of the associated tables;
[0027] Obtain at least one core data table based on the feature vectors.
[0028] In some embodiments of the present application, based on the above technical solutions, a method for obtaining at least one core data table based on the feature vectors includes:
[0029] Classify the feature vectors by using a clustering algorithm, and determine the category of the association table included in the node pair according to the classification result of the feature vectors;
[0030] Calculate the field similarity between any one of the association tables in each category and other association tables, and perform a descending order sorting;
[0031] Obtain the association tables in any one of the association tables in each category that are sorted before a first preset value, and form at least one node list in each category;
[0032] Calculate the intersection of at least one node list in each category to obtain at least one intersection node;
[0033] Obtain the number of association tables of the at least one intersection node, and perform a descending order sorting on the at least one intersection node according to the number of association tables to form an intersection sorting table;
[0034] Use the data tables corresponding to the intersection nodes in the intersection sorting table that are sorted before a second preset value as the core data tables.
[0035] In some embodiments of the present application, based on the above technical solutions, before responding to a data processing operation based on the temporary table, the method further includes:
[0036] Verify the temporary table. If the number of temporary tables in the temporary table exceeds a third preset value, or the number of fields in the temporary table set exceeds a fourth preset value, or any two data tables in the data table set cannot be directly associated through the temporary table, then discard the temporary table.
[0037] According to one aspect of the embodiments of the present application, the present application provides a data processing device based on a data warehouse, including:
[0038] An acquisition module, configured to acquire a data table set from the data warehouse, where the data table set includes data tables in the data warehouse;
[0039] A selection module, configured to select at least one core data table from the data table set, where the number of association tables of the core data table exceeds a first threshold, and the association tables of the core data table are data tables whose field similarity with the core data table exceeds a second threshold;
[0040] A construction module, configured to construct a temporary table according to the fields included in the at least one selected core data table;
[0041] A processing module, configured to respond to a data processing operation based on the temporary table.
[0042] In some embodiments of the present application, based on the above technical solution, the data processing device further includes an adding module, and the adding module includes:
[0043] A clustering unit, configured to obtain other data tables in the data table set except the at least one core data table;
[0044] A selection unit, configured to select a target data table with the shortest path between the other data tables and the at least one core data table;
[0045] An adding unit, configured to add the fields included in the selected target data table to the temporary table.
[0046] In some embodiments of the present application, based on the above technical solution, the selection module includes:
[0047] A clustering unit, configured to cluster the data tables in the data table set by using a clustering algorithm to obtain multiple data table subsets;
[0048] A selection unit, configured to respectively select data tables with the number of associated tables exceeding a first threshold from the multiple data table subsets as core data tables.
[0049] In some embodiments of the present application, based on the above technical solution, the data processing device further includes a preprocessing module, and the preprocessing module includes:
[0050] A similarity calculation unit, configured to calculate the field similarity between each data table in the data table set by using a locality-sensitive hashing algorithm;
[0051] An association unit, configured to determine data tables with the field similarity exceeding a second threshold as associated tables;
[0052] A matrix construction unit, configured to construct an association matrix based on the associated tables of each data table, and the association matrix is used to represent the association situation between each data table in the data table set and other data tables.
[0053] In some embodiments of the present application, based on the above technical solution, the construction module includes:
[0054] A conversion unit, configured to convert the association matrix into association key-value pairs;
[0055] An extraction unit, configured to extract node pairs from the association key-value pairs, and the node pairs include combinations of two associated tables of the data table;
[0056] A training unit, configured to construct an undirected graph neural network with the association tables included in the node pairs as nodes, and train the undirected graph neural network to obtain feature vectors of the association tables;
[0057] A construction unit, configured to obtain at least one core data table based on the feature vectors.
[0058] In some embodiments of the present application, based on the above technical solutions, the construction unit includes:
[0059] A vector classification unit, configured to classify the feature vectors using a clustering algorithm, and determine the categories of the association tables included in the node pairs according to the classification results of the feature vectors;
[0060] A first sorting unit, configured to calculate the field similarity between any one of the association tables in each category and other association tables, and perform a descending order sorting;
[0061] A list acquisition unit, configured to acquire the association tables in each category that are sorted before a first preset value for any one of the association tables, and form at least one node list in each category;
[0062] An intersection acquisition unit, configured to calculate the intersection of at least one node list in each category to obtain at least one intersection node;
[0063] A second sorting unit, configured to obtain the number of association tables of the at least one intersection node, and perform a descending order sorting on the at least one intersection node according to the number of association tables to form an intersection sorting table;
[0064] A construction subunit, configured to use the data tables corresponding to the intersection nodes sorted before a second preset value in the intersection sorting table as core data tables.
[0065] In some embodiments of the present application, based on the above technical solutions, the data processing device further includes a verification module, and the verification module is configured to verify the temporary table. If the number of temporary tables in the temporary table exceeds a third preset value, or the number of fields in the temporary table set exceeds a fourth preset value, or any two data tables in the data table set cannot be directly associated through the temporary table, then discard the temporary table.
[0066] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method based on a data warehouse as in the above technical solutions.
[0067] According to one aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes: a processor; and a memory for storing executable instructions of the processor. Wherein, the processor is configured to execute the data processing method based on the data warehouse in the above technical solution by executing the executable instructions.
[0068] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method based on the data warehouse in the above technical solution.
[0069] In the technical solution provided by the embodiments of the present application, the present application obtains a data table set from the data warehouse, and then selects at least one core data table from the data table set. Wherein, the number of associated tables of the core data table exceeds a first threshold, and a temporary table is constructed according to the fields included in the selected at least one core data table; and a data processing operation is responded based on the temporary table. Through the method of the present application, data tables with a relatively large number of associated tables in the data table set can be selected as core data tables, and a temporary table is constructed according to the fields included in the core data tables, and the constructed temporary table is used to respond to the data processing operation. Since the temporary table contains more fields, the data can directly arrive after two associations. Therefore, the development efficiency of the data warehouse can be improved, the SQL execution response time of the data warehouse can be accelerated, and the cost consumed by SQL execution can be reduced; at the same time, the temporary table constructed by the present application has a small memory, and the present application can maximize the improvement of the model association efficiency on the premise of sacrificing the least memory resources, greatly saving the computing resources for data modeling.
[0070] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0071] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0072] Figure 1 Schematically shows a schematic diagram of a data table in a data warehouse.
[0073] Figure 2Schematically shows an exemplary system architecture block diagram to which the technical solution of the present application is applied.
[0074] Figure 3 Schematically shows a flowchart of the data processing method based on a data warehouse according to the present application.
[0075] Figure 4 Schematically shows a flowchart of the method for selecting core data tables according to the present application.
[0076] Figure 5 Schematically shows a flowchart of the method for preprocessing data tables in a data table set according to the present application.
[0077] Figure 6 Schematically shows a schematic diagram of the association matrix constructed according to the present application.
[0078] Figure 7 Schematically shows a flowchart of the method for selecting core data tables by using the preprocessed data tables according to the present application.
[0079] Figure 8 Schematically shows a schematic diagram of the graph neural network constructed according to the present application.
[0080] Figure 9 Schematically shows a flowchart of the method for obtaining at least one core data table based on feature vectors according to the present application.
[0081] Figure 10 Schematically shows a flowchart of the method for updating a temporary table according to the present application.
[0082] Figure 11 Schematically shows a schematic diagram of the structure of the association network according to the present application.
[0083] Figure 12 Schematically shows a block diagram of the structure of the data processing device based on a data warehouse provided by an embodiment of the present application.
[0084] Figure 13 Schematically shows a block diagram of the computer system structure of an electronic device for implementing an embodiment of the present application. Detailed implementation manners
[0085] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0086] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0087] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0088] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0089] A data warehouse is a strategic collection that provides all types of data support for the decision-making processes at all levels of an enterprise. It is a single data store created for analytical reporting and decision support purposes.
[0090] The data in a data warehouse often exists in the form of tables, and in order to execute Structured Query Language (SQL) in the data warehouse, it is often necessary to perform multiple associations on the tables. For example, as Figure 1 shown, Figure 1 Schematically shows a schematic diagram of data tables in a data warehouse. There are five data tables in a certain data warehouse, namely data table T1 - data table T5. Among them, each group of data tables contains different field names. For example, data table T1 includes id (identification number), dialogid (order identification number), and dialogamout (order quantity).
[0091] When developing a data model for the Figure 1 corresponding data warehouse, it is often necessary to perform associations through multiple tables in order to achieve the association of two data tables. For example, when calculating the total order volume of unrefunded orders, the implemented SQL statement is as follows:
[0092] Select sum(T1.dialog_amout)as sum_amount
[0093] from T1
[0094] Join T2 on T1.dialog_id = T2.dialog_id
[0095] Join T3 on T2.cate1_id = T3.cate1_id
[0096] Join T4 on T3.sale_id = T4.sale_id
[0097] Join T5 on T5.fas_id = T4.fas_id
[0098] where T5.status_id = 0 and T5.is_pop = 1
[0099] In the above example, in order to obtain the flag of whether to return cash, four associations are made. Among them, each association uses the same field names between different tables for association. That is, the order quantity in the data table T1 is selected using SQL, and then the data table T2 and the data table T1 are associated using the dialog id (dialog identification number), and then the data table T3 and the data table T2 are associated, and so on until the data table T5 is also associated. Finally, the total order quantity of the orders without cash return is output.
[0100] And what is disclosed above is the number of associations when the number of data tables is small. When the number of data tables in the data warehouse is large, there may even be more than ten multiple associations. This brings great research and development time to developers, the development efficiency is extremely low, and due to the large number of associations, the SQL execution will also bring problems such as huge machine consumption costs and too long execution time.
[0101] To solve the above problems, all data tables related to all data models in the data warehouse can be associated to obtain a super-wide table, so as to achieve that all data models in the data warehouse can reach 100% through two associations between two tables. However, when associating all data models in the data warehouse to obtain a super-wide table, the super-wide table will occupy extremely large memory, and when performing table association queries based on the super-wide table, the query efficiency is extremely low, even much slower than multiple associations of multiple small tables, or even unable to query at all. Therefore, there is currently no method to solve the problem that multiple associations of data tables consume a lot of time costs and even waste a lot of resources.
[0102] To solve the above technical problems, the present application discloses a data processing method based on a data warehouse, a data processing device based on a data warehouse, a computer-readable medium, and an electronic device. The content of the present application will be further described below in various aspects.
[0103] Figure 2 An exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.
[0104] As Figure 2 shown, the system architecture 200 may include a terminal device 210, a network 220, and a server 230. The terminal device 210 may include various electronic devices such as a smart phone, a tablet computer, a laptop computer, and a desktop computer. The server 230 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 220 may be a communication medium of various connection types capable of providing a communication link between the terminal device 210 and the server 230, for example, it may be a wired communication link or a wireless communication link.
[0105] According to the implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 230 may be a server group composed of multiple server devices. In addition, the technical solution provided by the embodiments of the present application may be applied to the terminal device 210, or may be applied to the server 230, or may be jointly implemented by the terminal device 210 and the server 230. The present application does not make special limitations on this.
[0106] The terminal device 210 or the server 230 may establish a connection with the data warehouse through the network 220. After establishing the connection, the terminal device 210 or the server 230 may obtain a data table set from the data warehouse, and then select at least one core data table from the data table set, where the number of associated tables of the core data table exceeds a first threshold, and the associated table of the core data table is a data table whose field similarity with the core data table exceeds a second threshold; then construct a temporary table according to the fields included in the selected at least one core data table; and finally respond to a data processing operation based on the temporary table. Through the data processing of the data warehouse by the terminal device 210 or the server 230, the development efficiency of the data warehouse can be improved. When responding to a data processing operation, for example, when SQL is executed, the SQL execution response time of the data warehouse can be accelerated, and the cost consumed by SQL execution can be reduced.
[0107] The above part introduces the content of the exemplary system architecture applying the technical solution of the present application. Next, the data processing method based on the data warehouse of the present application will be continued to be introduced.
[0108] As Figure 3 shown, Figure 3 A flowchart of the data processing method based on the data warehouse of the present application is schematically shown. According to one aspect of the embodiments of the present application, the present application provides a data processing method based on a data warehouse, including step S310-step S340.
[0109] In step S310: Obtain a data table set from the data warehouse, where the data table set includes data tables in the data warehouse.
[0110] The data warehouse contains many data tables. For example, Figure 1 as shown in multiple data tables. In step S310 of this application, obtaining a data table set from the data warehouse can be obtaining all data tables from the data warehouse and forming a data table set with all the data tables. This application can also obtain some data tables in the data warehouse according to actual needs and form a data table set with the some data tables. For example, when only data processing operations (SQL operations) need to be performed between certain types of data tables, this application can obtain some data tables related to the corresponding type in the data warehouse to form a data table set.
[0111] In an embodiment of this application, this application can directly connect the server 130 to the data warehouse through the network 120, so that the data tables in the data warehouse can be directly processed. At this time, this application can directly obtain data tables from the data warehouse and perform data processing in the data warehouse. Using this method, there is no need to obtain a data table set from the data warehouse. Therefore, the data processing process can be simplified and the data processing efficiency can be improved.
[0112] In step S320: Select at least one core data table from the data table set. The number of associated tables of the core data table exceeds a first threshold. The associated table of the core data table is a data table whose field similarity with the core data table exceeds a second threshold.
[0113] The selection strategy for selecting core data tables in this application can include a first strategy and a second strategy. Among them, the first strategy is that the selected core data table should have as large an association with other data tables as possible, corresponding to a relatively large number of associated tables of the core data table. The second strategy is that the similarity of the tables that can be associated between the selected core data tables should be small, corresponding to a relatively large difference between the associated tables of the core data tables.
[0114] Among them, this application can select core data tables only using the first strategy, or can also use both strategies for selection. When using both strategies for selection at the same time, it can be avoided that the difference between the associated tables of the selected core data tables is small. Therefore, there will be no problem that the number of duplicate data tables in the subsequently constructed temporary table is too large, resulting in memory occupation.
[0115] The following specifically explains the selection methods corresponding to the two core table selection strategies.
[0116] The first strategy of this application is that the correlation between the selected core data table and other data tables should be as large as possible. Correspondingly, the number of associated tables of the core data table exceeds the first threshold. The associated table of the core data table is a data table whose field similarity with the core data table exceeds the second threshold. Among them, for the calculation of the field similarity between two data tables in this application, it can include calculating any one of the field names, field data, and the names of the two data tables, or considering the three comprehensively to calculate the field similarity between the two data tables.
[0117] In an embodiment of this application, the method for determining that the field similarity exceeds the second threshold can be: converting the fields of two data tables into field vectors, then calculating the distance between the two field vectors, and then setting the threshold of the distance between the field vectors as the second threshold. When the vector distance corresponding to the fields of the two data tables exceeds this second threshold, it indicates that the two data tables are associated.
[0118] Take Figure 1 the first four data tables in as an example. Since there is the same field dialog_id in data table T1 and data table T2, therefore, data table T1 and data table T2 are associated. The corresponding data table T2 is the associated table of data table T1. And since there are no same fields between data table T3 - data table T4 and data table T1, therefore, data table T3 - data table T4 are not the associated tables of data table T1. Correspondingly, using the same method, the associated tables of data table T2 can be obtained as data table T1, data table T3, and data table T4; the associated tables of data table T3 are data table T2 and data table T4; the associated tables of data table T4 are data table T2 and data table T3. Therefore, it can be concluded that the number of associated tables of data table T2 is the largest. Therefore, if the first threshold set in this application is 2, then only data table T2 exceeds this first threshold. Therefore, the corresponding data table T2 is the selected core data table.
[0119] Among them, the first threshold of this application can be determined according to the number of data tables included in the data warehouse or the number of data tables included in the data table set. At the same time, the first threshold of this application can also be determined according to the memory size of the temporary table constructed later. When the number of data tables included in the data table set is larger, the first threshold can be set larger correspondingly. When the first threshold is set larger, the number of selected core data tables will be smaller, and the memory of the temporary table constructed later will be smaller.
[0120] The content of the first strategy in the core data table selection strategy of this application is disclosed above. Next, continue to disclose the corresponding selection method of the second strategy in the core data table selection strategy of this application.
[0121] The second strategy of this application is that the similarity of the related tables among the selected core data tables should be small, which corresponds to a large difference among the related tables of each core data table. Among them, the first strategy compares the field similarity of each data table, while the second strategy compares the differences between each data table itself.
[0122] Continuing with the data tables in Figure 1 as an example, the related tables of data table T2 are data tables T1, T3, and T4; the related tables of data table T3 are data tables T2 and T4; assume there is another data table T6, and its related tables are data tables T5 and T2. Taking data table T2 as an example, there is a same related table T4 in the related tables of data tables T2 and T3, while there is no same related table between data tables T2 and T6. Therefore, the difference between the related tables of data table T2 and data table T6 is greater than the difference between the related tables of data table T2 and data table T3. At this time, we can select data tables T2 and T6 and exclude data table T3.
[0123] In an embodiment of this application, the second strategy of this application can be further selected after the first strategy is selected. For example, when multiple core data tables are selected through the first strategy, for the core data tables with the same number of related tables, the core data tables can be further selected by comparing the differences between the related tables in each core data table.
[0124] In an embodiment of this application, the second strategy of this application can be carried out simultaneously with the first strategy. For example, continuing with the data tables in Figure 1 as an example, according to the first strategy of this application, data table T2 can be selected. And according to the second strategy, compared with data tables T2 and T3, the difference between the related tables of data table T2 and data table T1 is greater. Therefore, data table T1 is also selected as a core data table by using the second strategy.
[0125] When this application uses the second strategy and the first strategy simultaneously, this application can use the following method to select core data tables.
[0126] In an embodiment of this application, as shown in Figure 4 shown, Figure 4 schematically shows a flowchart of the method for this application to select core data tables. The method for this application to select at least one core data table from a data table set includes step S410 - step S420.
[0127] Step S410: Use a clustering algorithm to cluster the data tables in the data table set to obtain multiple data table subsets.
[0128] The purpose of clustering is to partition a data table set into different classes or clusters according to a specific criterion (such as a distance criterion), so that the similarity of data objects within the same cluster is as large as possible, while the difference of data objects not in the same cluster is also as large as possible. That is, after clustering, the data of the same class are gathered together as much as possible, and different data are separated as much as possible. In this application, the data tables in the data table set can be regarded as individual data points, and then the K-means algorithm or the Agglomerative algorithm can be used to cluster the data tables in the data table set to obtain multiple data table subsets.
[0129] Through the clustering algorithm, this application can cluster the data tables in the data warehouse according to the field similarity of the data tables. After clustering, similarity calculation is performed, which can effectively reduce the amount of computation and avoid the waste of computer computing power caused by comparing the similarities of data tables with large differences in the data warehouse.
[0130] In an embodiment of this application, this application can calculate the field similarity between any two data tables in the data warehouse and complete the clustering process through the Locality-Sensitive Hashing (LSH) algorithm. The basic idea of the Locality-Sensitive Hashing algorithm is that after two adjacent data points in the original data space are transformed through the same mapping or projection transformation, the probability that these two data points are still adjacent in the new data space is very high, while the probability that non-adjacent data points are mapped to the same bucket is very small. That is to say, if we perform some hash mappings on the original data, we hope that two originally adjacent data points can be hashed into the same bucket and have the same bucket number. After performing hash mappings on all the data in the original data set, we obtain a hash table. These original data sets are scattered into the buckets of the hash table, and each bucket will contain some original data. The data belonging to the same bucket are very likely to be adjacent. And this application regards each data table in the data warehouse as a data point in the Locality-Sensitive Hashing algorithm. Using the Locality-Sensitive Hashing algorithm, it is possible to find one or some data points that are approximately (with a large field similarity) the closest to the query data point (any data table) in a massive high-dimensional data set (data warehouse). It should be noted that using the Locality-Sensitive Hashing algorithm does not guarantee that the data closest to the data table can definitely be found, but it reduces the number of data points that need to be matched while ensuring a high probability of finding the nearest neighbor data points.
[0131] Step S420: Select the data tables with the number of associated tables exceeding the first threshold from multiple data table subsets as core data tables.
[0132] After clustering is completed, the data tables with the number of associated tables exceeding the first threshold can be selected from multiple data table subsets as core data tables, and the selection method is the same as the above-mentioned first strategy.
[0133] Through step S410 of the present application, the data table set can be pre-classified according to the similarity of each data table in advance, avoiding the problem that the core data tables selected subsequently are all of the same category, resulting in waste of memory resources, and realizing the simultaneous execution of the above-mentioned second strategy and the first strategy.
[0134] In an embodiment of the present application, as Figure 5 shown, Figure 5 A flowchart showing a method for preprocessing data tables in a data table set according to the present application is schematically shown. Before selecting at least one core data table from the data table set, the present application can preprocess the data tables in the data table set. The specific preprocessing method includes steps S510 - S530.
[0135] Step S510: Calculate the field similarity between each data table in the data table set by using the locality - sensitive hashing algorithm.
[0136] The present application can calculate the field similarity between any two data tables in the data warehouse by using the locality - sensitive hashing algorithm. Use the locality - sensitive hashing algorithm LSH to build an index (Hash table) for the data tables and perform approximate nearest - neighbor search through the index. Through the locality - sensitive hashing algorithm, one or some data points that are approximately the nearest neighbors to the query data point can be found. Among them, the present application can use any data table as the query data point, so that the approximate nearest - neighbor data tables of each data table can be found as some data points. The present application can use the approximate nearest - neighbor data tables found for each data table as the associated tables of the data table corresponding to the query data point. The present application can also, after finding the approximate nearest - neighbor data tables, further calculate the field similarity between the approximate nearest - neighbor data tables and the data tables corresponding to the query data points, so as to obtain the field similarity between each data table. Among them, to calculate the similarity between two fields, first convert the two fields into vectors, and then use the cosine similarity calculation formula to calculate their similarity.
[0137] The formula corresponding to the cosine similarity algorithm is as follows:
[0138]
[0139] Where similarity represents the similarity, A and B respectively represent the vectors of the two fields. After calculating the field similarity between each data table in the data table set through the above steps, continue with step S520.
[0140] Step S520: Determine the data tables with field similarity exceeding the second threshold as associated tables.
[0141] After obtaining the field similarity, the data tables with field similarity exceeding the second threshold can be determined as associated tables. Among them, the second threshold can be determined according to the number of core data tables to be selected. When the second threshold is smaller, the corresponding number of selected core data tables is larger.
[0142] For example, continuing with Figure 1 as an example, through field similarity comparison, when other data tables have exactly the same fields as the core data table, that data table is taken as the associated table of the core data table. For example, the associated table of data table T1 is data table T2.
[0143] Step S530: Construct an association matrix based on the associated tables of each data table. The association matrix is used to represent the association situation between each data table and other data tables in the data table set.
[0144] As Figure 6 shown, Figure 6 schematically shows the association matrix diagram constructed in this application.
[0145] Through the association matrix, the association situation between each data table can be clearly understood. Among them, Y represents that there is an association relationship between the two, and N represents that there is no association relationship between the two.
[0146] After preprocessing the data tables in the data set, the processed data tables can be selected.
[0147] As Figure 7 shown, Figure 7 schematically shows the method flow chart for selecting core data tables using the preprocessed data tables in this application. In an embodiment of this application, the method for selecting at least one core data table from the data table set includes steps S710 - step S740.
[0148] Step S710: Convert the association matrix into association key - value pairs.
[0149] According to the associated table corresponding to each data table, convert the association matrix into association key - value pairs. For example, Figure 6 the corresponding association matrix can be converted into the following association key - value pairs: T1: [T2], T2: [T1, T3, T4], T3: [T2, T4], T4: [T2, T3]. Among them, the association key - value pairs include the candidate data tables that may be selected as core data tables and the associated tables associated with the candidate data tables. For example, the association key - value pair T2: [T1, T3, T4] includes the candidate data table T2 and the associated tables T1, T3, T4 associated with the candidate data table.
[0150] Step S720: Extract node pairs from the association key - value pairs. The node pairs include combinations of two associated tables of the data table.
[0151] Among them, a node pair is a combination of two association tables. Therefore, for the association key-value pair T1: [T2], since there is only one association table, no node pair can be extracted. For the association key-value pair T2: [T1, T3, T4], the extractable node pairs include [T1, T3], [T1, T4], and [T3, T4]. Therefore, in this application, by extracting node pairs from the association key-value pairs, multiple groups of node pairs can be extracted, and step S730 is continued.
[0152] Step S730: Using the association tables included in the node pairs as nodes, construct an undirected graph neural network, and train the undirected graph neural network to obtain the feature vectors of the association tables.
[0153] As Figure 8 shown, Figure 8 Schematically shows the schematic diagram of the graph neural network constructed in this application. In this application, the association tables included in the node pairs are used as nodes. For example, data tables T1, T2, T3, and T4 are used as nodes to construct an undirected graph neural network, and the connections between each node are randomly undirected. Among them, this application can also set a weight value, for example, set the weight value to 1. Then define a random walk strategy, and train the undirected graph neural network based on the random walk strategy to obtain the feature vectors of each association table.
[0154] In an embodiment of this application, this application can obtain the feature vector values of each association table through the method of Graph Embedding. The idea of Graph Embedding is to find a mapping function to convert each node in the undirected graph neural network of this application into a low-dimensional dense embedding representation, requiring that similar nodes in the graph are close in the low-dimensional space.
[0155] This application can use the deepwalk algorithm for Graph Embedding. The deepwalk algorithm learns the community representations of the graph network nodes through truncated random walk. The Deepwalk algorithm includes two main steps: the first step is to sample node sequences using the Random Walk algorithm, and the second step is to use the skip-gram algorithm to learn the expression vectors. The feature vectors of each association table can be obtained through the two steps of the Deepwalk algorithm.
[0156] In one embodiment of the present application, the present application can also perform graph embedding through the node-vec algorithm. The node-vec algorithm modifies the way of random walks based on the deepwalk algorithm, adds conditions for forming sequences, and uses the ideas of depth-first sampling (DFS) and breadth-first sampling (BFS) to calculate the probability of reaching the next node to select nodes. The remaining steps are the same as those of deepwalk.
[0157] After obtaining the feature vectors of each association table through the above method, continue with step S740.
[0158] Step S740: Obtain at least one core data table based on the feature vectors.
[0159] As Figure 9 shown, Figure 9 Schematically shows a flowchart of the method for obtaining at least one core data table based on the feature vectors in the present application.
[0160] In one embodiment of the present application, the method for obtaining at least one core data table based on the feature vectors includes steps S910 - S960.
[0161] Step S910: Classify the feature vectors using a clustering algorithm and determine the category of the association table included in the node pair according to the classification result of the feature vectors.
[0162] After the present application obtains the feature vectors of each association table, it can classify the feature vectors using a clustering algorithm to determine the category of the association table included in the node pair. Among them, the clustering algorithm of the present application can use the K-means algorithm or the Agglomerative algorithm. Through step S910, the association tables corresponding to multiple feature vectors are divided into multiple categories. For example, after classifying the association tables included in the node pair corresponding to Figure 8 , the vector corresponding to the data table T1 can be classified into one category as category one. The vectors corresponding to the data tables T2, T3, and T4 can be classified into one category as category two.
[0163] Step S920: Calculate the field similarity between any association table in each category and other association tables, and perform a descending order sorting.
[0164] Continuing with Category 2 in Step S910 as an example, after classification, it is obtained that Category 2 includes data table T2, data table T3, and data table T4. Among them, in Step S920 of the present application, calculating the field similarity between any associated table and other associated tables in each category can be calculating the field similarity between any data table and other data tables in Category 2, and sorting them in descending order. The calculation method of the field similarity is the same as that in Step S510. For example, the similarity sorting obtained through calculation is as follows:
[0165] T2: [T3, T4], T3: [T2, T4], T3: [T2, T3]
[0166] In the above example, Category 2 includes three associated tables. When the number of associated tables is n, the following sorting results will be obtained for each associated table correspondingly:
[0167] Ti: {t1, ti + i...}, i = {1, 2, 3... n}
[0168] For example, when n is 10, the descending sorting results of the other nine data tables are included in T1 correspondingly, and any one of the other nine data tables may be ranked in the front. Therefore, through Step S920, the field similarity sorting results of each associated table and other associated tables in each category can be obtained, and Step S930 is continued.
[0169] Step S930: Obtain the associated tables in any associated table in each category that are ranked before the first preset value, and form at least one node list in each category.
[0170] The first preset value of the present application can be modified as needed. For example, the first preset value can be 5, then the top five associated tables in any associated table in each category are obtained correspondingly, and multiple node lists are formed.
[0171] Correspondingly, when the first preset value is N, by obtaining any associated table in Step S920, the following node list can be obtained.
[0172] Ti: {t1, ti + i...}, i = {1, 2, 3... N}
[0173] When Step S930 is performed on each associated table in each category, multiple node lists within the category can be obtained, and then the same steps are performed on each other associated table in different categories. Finally, multiple node lists can be obtained.
[0174] For example, a data warehouse has nine data tables, which are divided into three categories through the above steps. Among them, data tables T1 and T3 belong to category one, data tables T2, T6, T8, and T9 belong to category two, and data tables T4, T5, and T7 belong to category three. Take the first preset value as 2. Among them, the data table with the highest field similarity to any data table is itself. The node list corresponding to each data table includes: Category one: T1: [T1, T3], T3: [T3, T1]. Category two: T2: [T2, T6], T6: [T6, T8], T8: [T8, T6], T9: [T9, T6]. Category three: T4: [T4, T7], T5: [T5, T4], T7: [T7, T5]. Through the above method, the above nine node lists can be formed, and each node list contains the data tables ranked in the top two in terms of similarity to any data table.
[0175] Step S940: Calculate the intersection of at least one node list in each category to obtain at least one intersection node.
[0176] Calculate the associated tables commonly included in at least one node list in each of the above categories, and take the commonly included associated tables as intersection nodes. For example, through category one in step S930, the intersection nodes obtained are data tables T1 and data table T3; through category two, the intersection node obtained is data table T4; there are no intersection nodes in category three. Therefore, there are three intersection nodes obtained through the above steps, namely data table T1, data table T3, and data table T4.
[0177] Step S950: Obtain the number of associated tables of at least one intersection node, and sort at least one intersection node in descending order according to the number of associated tables to form an intersection sorting table.
[0178] After obtaining at least one intersection node, the number of associated tables of each intersection node can be calculated by referring to step S520 of this application, and at least one intersection node is sorted in descending order according to the number of associated tables.
[0179] For example, the obtained intersection nodes include data table T1, data table T3, and data table T4. Data table T1 has three associated tables, data table T3 has two associated tables, and data table T4 has five associated tables. The corresponding sorting result can be obtained as: data table T4, data table T3, data table T1.
[0180] Step S960: Use the data tables corresponding to the intersection nodes ranked before the second preset value in the intersection sorting table as core data tables.
[0181] Among them, the second preset value can be determined according to actual needs. For example, when the second preset value is 2, the data tables corresponding to the top two intersection nodes in the intersection sorting table can be taken as the core data tables. That is, the data tables T4 and T3 in step S950 are taken out as the core data tables.
[0182] After the core data tables are selected through the above steps in this application, step S330 is continued.
[0183] In step S330: A temporary table is constructed according to the fields included in the at least one selected core data table.
[0184] After at least one core data table is selected in this application, the fields included in the at least one core data table can be directly extracted, and then these fields are combined together to construct a temporary table, where a temporary table containing all the fields in the at least one core data table is obtained.
[0185] This application can also combine and save the at least one core data table to construct a temporary table, where the temporary table contains the at least one core data table. For example, after the core data tables obtained by this application are the data tables T2 and T1, the corresponding temporary table contains two data tables.
[0186] After the temporary table is constructed through the above steps in this application, since the temporary table of this application is composed of at least one core data table, there may be a situation where some data tables in the data warehouse cannot be associated with the core data table. Based on this problem, after step S330 in this application, the following steps are also included.
[0187] As Figure 10 shown, Figure 10 A flowchart of the method for updating the temporary table in this application is schematically shown.
[0188] In an embodiment of this application, after step S330, this application also includes updating the temporary table, and the specific updating method includes steps S1010 - step S1030.
[0189] Step S1010: Obtain other data tables in the data table set except the at least one core data table.
[0190] Obtain other data tables except the at least one core data table selected through step S950. For example, continue to use Figure 1 the first four data tables as an example. When the selected core data tables are the data tables T2 and T1, the corresponding other data tables are the data tables T3 and T4.
[0191] Step S1020: Select a target data table with the shortest path between at least one core data table from other data tables.
[0192] Among them, the shortest path in this application is determined based on the association relationship. This application can construct an association network according to the associations between all tables in the data warehouse to Figure 1 take the corresponding first four data tables as an example. As Figure 11 shown, Figure 11 schematically shows the structural diagram of the association network of this application. This application can construct an association network of data table T1 - data table T4 based on the association relationship. From Figure 11 it can be seen that data table T1 and data table T2 are associated with each other, so they are connected by a double arrow. Taking Figure 11 as an example, if the selected core data table is data table T3, then the corresponding other data tables at this time include data table T1, data table T2, and data table T4. Correspondingly, the shortest path from data table T1 to the core data table which is data table T3 is from data table T1 to data table T2, and then from data table T2 to data table T3. Therefore, the corresponding target data table is data table T2.
[0193] After constructing the association network and selecting the target data table through the above method, continue with step S1030.
[0194] Step S1030: Add the fields included in the selected target data table to the temporary table.
[0195] Among them, this application can add the target data table to the temporary table so that the number of data tables in the temporary table is at least the number of one core data table plus the number of the target data table. This application can also directly add the fields included in the target data table to the temporary table so that the temporary table contains the fields of the target data table.
[0196] After completing the construction of the temporary table through the above steps, this application can also verify the temporary table. The specific steps are as follows.
[0197] In an embodiment of this application, before step S340 and after constructing the temporary table, this application can also verify the temporary table. If the number of temporary tables in the temporary table exceeds the third preset value, or the number of fields in the temporary table set exceeds the fourth preset value, or any two data tables in the data table set cannot be directly associated through the temporary table, then discard the temporary table.
[0198] This step is to verify whether the constructed temporary table is qualified. Among them, there are three verification conditions. If one of the three conditions is not met, it means that the temporary table is unqualified. If the temporary table is unqualified, then discard the temporary table and return to step S320 to re - select at least one core data table.
[0199] Among them, the three conditions of this application are respectively the verification of the accessibility of the data table, the number of data tables, and the number of fields. The accessibility of the data table means that any two data tables in the data table set cannot be directly associated through a temporary table, indicating that based on the temporary table, data cannot directly reach after two associations. The verification of the number of data tables and the number of fields is to avoid the problem that the memory of the constructed temporary table is too large and occupies resources.
[0200] After passing the verification through the above steps, step S340 can be continued.
[0201] In step S340: Respond to data processing operations based on the temporary table.
[0202] This application can perform various data processing operations based on the temporary table. Among them, the most basic processing operation is to complete the association of any two data tables in the data warehouse and achieve direct access to the data after two associations. Since the temporary table contains the fields of all data tables in the data warehouse, therefore, data table A can be associated with the temporary table and then the temporary table can be associated with data table B to achieve direct access.
[0203] After obtaining the temporary table through the above steps, this application conducts tests on real data tables through the data warehouse. The test results can achieve the association between any two tables, and the accessibility reaches 100%. Moreover, the number of finally constructed temporary tables is less than 10% of the total number of data tables, and the average number of fields of all constructed temporary tables is basically the same as the average value of other non-temporary tables. Therefore, it can meet the requirements of this application and solve the problems of this application.
[0204] In addition, with the advancement of the 5G messaging platform, structured data such as users and messages will gradually become huge. In order to improve the service ability and intelligence of the platform, it is necessary to perform intelligent analysis based on this incremental data, and a large amount of data model development work is involved in the intelligent analysis process. Through the solution of this application, it can be achieved that any two tables can reach 100% through two associations. Therefore, the method of this application can be applied to intelligent analysis. At the same time, in the big data era, the corresponding method of this application can also be applicable in business scenarios involving a large number of table association operations. This application does not limit this.
[0205] Through the method of the present application, data tables with a relatively large number of associated tables in the data table set can be selected as core data tables, and a temporary table can be constructed based on the fields included in the core data tables. The constructed temporary table is used to respond to data processing operations. Since the temporary table contains a relatively large number of fields, data can directly reach after two associations. Therefore, the development efficiency of the data warehouse can be improved, the SQL execution response time of the data warehouse can be accelerated, and the cost consumed by SQL execution can be reduced. At the same time, compared with the existing large wide table, the memory of the temporary table constructed in the present application occupies less memory. Therefore, the present application can maximize the improvement of the model association efficiency on the premise of sacrificing the least memory resources, and greatly save the computing resources for data modeling.
[0206] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0207] The above part introduced the content of the data processing method based on the data warehouse of the present application. Next, the content of the data processing device based on the data warehouse of the present application will be continued to be introduced.
[0208] As Figure 12 shown, Figure 12 Schematically shows the structural block diagram of the data processing device based on the data warehouse provided by the embodiment of the present application.
[0209] According to one aspect of the embodiment of the present application, the present application provides a data processing device 1200 based on a data warehouse, including:
[0210] An acquisition module 1210, configured to acquire a data table set from the data warehouse, where the data table set includes data tables in the data warehouse;
[0211] A selection module 1220, configured to select at least one core data table from the data table set, where the number of associated tables of the core data table exceeds a first threshold, and the associated table of the core data table is a data table whose field similarity with the core data table exceeds a second threshold;
[0212] A construction module 1230, configured to construct a temporary table based on the fields included in the at least one selected core data table;
[0213] A processing module 1240, configured to respond to data processing operations based on the temporary table.
[0214] In one embodiment of the present application, the data processing device 1200 of the present application further includes an adding module, and the adding module includes:
[0215] A clustering unit configured to obtain other data tables in the data table set except for at least one core data table;
[0216] A selecting unit configured to select a target data table with the shortest path between the other data tables and at least one core data table;
[0217] An adding unit configured to add the fields included in the selected target data table to the temporary table.
[0218] In one embodiment of the present application, the selecting module 1220 of the present application includes:
[0219] A clustering unit configured to cluster the data tables in the data table set by using a clustering algorithm to obtain a plurality of data table subsets;
[0220] A selecting unit configured to respectively select the data tables with the number of associated tables exceeding a first threshold from the plurality of data table subsets as core data tables.
[0221] In one embodiment of the present application, the data processing device 1200 of the present application further includes a preprocessing module, and the preprocessing module includes:
[0222] A similarity calculation unit configured to calculate the field similarity between each data table in the data table set by using a locality-sensitive hashing algorithm;
[0223] An association unit configured to determine the data tables with the field similarity exceeding a second threshold as associated tables;
[0224] A matrix construction unit configured to construct an association matrix based on the associated tables of each data table, and the association matrix is used to represent the association situation between each data table in the data table set and other data tables.
[0225] In one embodiment of the present application, the construction module 1230 of the present application includes:
[0226] A conversion unit configured to convert the association matrix into association key-value pairs;
[0227] An extraction unit configured to extract node pairs from the association key-value pairs, and the node pairs include combinations of two associated tables of the data table;
[0228] A training unit configured to construct an undirected graph neural network with the associated tables included in the node pairs as nodes and the connections of the node pairs as edges, and train the undirected graph neural network to obtain the feature vectors of the associated tables;
[0229] A construction unit for obtaining at least one core data table based on feature vectors.
[0230] In an embodiment of the present application, the construction unit of the present application includes:
[0231] A vector classification unit configured to classify the feature vectors using a clustering algorithm and determine the category of the association table included in the node pair according to the classification result of the feature vectors;
[0232] A first sorting unit configured to calculate the field similarity between any one of the association tables in each category and other association tables and perform a descending order sorting;
[0233] A list acquisition unit configured to acquire the association tables ranked before a first preset value in any one of the association tables in each category and form at least one node list in each category;
[0234] An intersection acquisition unit configured to calculate the intersection of at least one node list in each category and obtain at least one intersection node;
[0235] A second sorting unit configured to obtain the number of association tables of the at least one intersection node and perform a descending order sorting on the at least one intersection node according to the number of association tables to form an intersection sorting table;
[0236] A construction subunit configured to use the data tables corresponding to the intersection nodes ranked before a second preset value in the intersection sorting table as core data tables.
[0237] In an embodiment of the present application, the data processing device 1200 of the present application further includes a verification module, and the verification module is configured to verify the temporary table. If the number of temporary tables in the temporary table exceeds a third preset value, or the number of fields in the temporary table set exceeds a fourth preset value, or any two data tables in the data table set cannot be directly associated through the temporary table, then discard the temporary table.
[0238] Through the data processing device 1200 of the present application, data tables with a large number of association tables in the data table set can be selected as core data tables, and temporary tables can be constructed according to the fields included in the core data tables, and the constructed temporary tables are used to respond to data processing operations. Since the temporary tables contain more fields, data can reach directly after two associations. Therefore, the development efficiency of the data warehouse can be improved, the SQL execution response time of the data warehouse can be accelerated, and the cost consumed by SQL execution can be reduced; at the same time, the temporary tables constructed by the present application have a small memory and can avoid resource waste.
[0239] Details of the data processing device based on the data warehouse provided in the embodiments of the present application have been described in detail in the corresponding method embodiments, and will not be repeated here.
[0240] The above part introduced the content of the data processing device based on the data warehouse of the present application. Next, other aspects of the present application will be continued to be introduced.
[0241] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method based on the data warehouse in the above technical solution.
[0242] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the data processing method based on the data warehouse in the above technical solution by executing the executable instructions.
[0243] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method based on the data warehouse in the above technical solution.
[0244] Figure 13 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.
[0245] It should be noted that Figure 13 The computer system 1300 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0246] Such as Figure 13As shown, the computer system 1300 includes a central processing unit 1301 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1302 (ROM) or the program loaded from the storage section 1308 into the random access memory 1303 (RAM). In the random access memory 1303, various programs and data required for system operation are also stored. The central processing unit 1301, the read-only memory 1302, and the random access memory 1303 are connected to each other via a bus 1304. An input / output interface 1305 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1304.
[0247] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a local area network card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that the computer program read from it can be installed into the storage section 1308 as needed.
[0248] Specifically, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the central processing unit 1301, various functions defined in the system of the present application are executed.
[0249] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0250] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0251] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0252] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0253] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0254] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A data processing method based on a data warehouse, characterized in that, including: obtaining a data table set from a data warehouse, where the data table set includes data tables in the data warehouse; calculating the field similarity between each data table in the data table set by using a locality-sensitive hashing algorithm; determining data tables with the field similarity exceeding a second threshold as associated tables; constructing an association matrix based on the associated tables of each data table, where the association matrix is used to represent the association situation between each data table in the data table set and other data tables; converting the association matrix into association key-value pairs; extracting node pairs from the association key-value pairs, where the node pairs include combinations of two associated tables of the data table; constructing an undirected graph neural network with the associated tables included in the node pairs as nodes, and training the undirected graph neural network to obtain feature vectors of the associated tables; obtaining at least one core data table based on the feature vectors, where the number of associated tables of the core data table exceeds a first threshold, and the associated tables of the core data table are data tables with the field similarity exceeding the second threshold to the core data table; constructing a temporary table based on the fields included in the at least one selected core data table; responding to a data processing operation based on the temporary table.
2. The data processing method based on a data warehouse according to claim 1, wherein After constructing the temporary table based on the fields included in the at least one selected core data table, the method further includes: obtaining other data tables in the data table set except the at least one core data table; selecting a target data table with the shortest path between the other data tables and the at least one core data table; adding the fields included in the selected target data table to the temporary table.
3. The data processing method based on a data warehouse according to claim 1, wherein selecting at least one core data table from the data table set, including: clustering the data tables in the data table set by using a clustering algorithm to obtain a plurality of data table subsets; selecting, from the plurality of data table subsets, data tables with the number of associated tables exceeding the first threshold as core data tables.
4. The data processing method based on a data warehouse according to claim 1, wherein obtaining at least one core data table based on the feature vectors, including: classifying the feature vectors by using a clustering algorithm, and determining the categories of the associated tables included in the node pairs according to the classification results of the feature vectors; calculating the field similarity between any one of the associated tables in each category and other associated tables, and sorting them in descending order; obtaining the associated tables ranked before a first preset value in any one of the associated tables in each category to form at least one node list in each category; calculating the intersection of at least one node list in each category to obtain at least one intersection node; obtaining the number of associated tables of the at least one intersection node, and sorting the at least one intersection node in descending order according to the number of associated tables to form an intersection sorting table; using the data tables corresponding to the intersection nodes ranked before a second preset value in the intersection sorting table as core data tables.
5. The data processing method based on a data warehouse according to claim 2, wherein Before responding to the data processing operation based on the temporary table, the method further includes: verifying the temporary table, and if the number of temporary tables in the temporary table exceeds a third preset value, or the number of fields in the temporary table set exceeds a fourth preset value, or any two data tables in the data table set cannot be directly associated through the temporary table, then discard the temporary table.
6. A data processing device based on a data warehouse, characterized in that, including: An acquisition module, configured to acquire a data table set from a data warehouse, where the data table set includes data tables in the data warehouse; A similarity calculation unit, configured to calculate the field similarity between each data table in the data table set by using a locality-sensitive hashing algorithm; An association unit, configured to determine data tables with a field similarity exceeding a second threshold as associated tables; A matrix construction unit, configured to construct an association matrix based on the associated tables of each data table, where the association matrix is used to represent the association situation between each data table in the data table set and other data tables; A conversion unit, configured to convert the association matrix into association key-value pairs; An extraction unit, configured to extract node pairs from the association key-value pairs, where the node pairs include combinations of two associated tables of the data table; A training unit, configured to construct an undirected graph neural network with the associated tables included in the node pairs as nodes, and train the undirected graph neural network to obtain feature vectors of the associated tables; A construction unit, configured to obtain at least one core data table based on the feature vectors, where the number of associated tables of the core data table exceeds a first threshold, and the associated tables of the core data table are data tables with a field similarity exceeding the second threshold with the core data table; A construction module, configured to construct a temporary table according to the fields included in at least one selected core data table; A processing module, configured to respond to a data processing operation based on the temporary table.
7. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method based on a data warehouse according to any one of claims 1 to 5.
8. An electronic device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the data processing method based on a data warehouse according to any one of claims 1 to 5 by executing the executable instructions.
Citation Information
Patent Citations
Method and device for processing data
CN113111084A