Data processing method, device and equipment and computer readable storage medium
By optimizing the data table connection order and selecting the target physical node, the problem of data summary table acquisition time is solved, and the efficiency of message delivery is improved.
Patent Information
- Application Number
- CN202510675364.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
The traditional data summary table acquisition time is too long, resulting in low packet delivery efficiency.
By generating the optimal connection order and selecting the target physical node, the data table connection process is optimized, including generating the connection order of the data table at the optimization target at the minimum connection cost, and selecting the physical node in the distributed database for the optimization target at the minimum execution cost.
It shortens the time to obtain data summary tables and improves the timely rate of message delivery.
Smart Images

Figure CN120596528A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of financial technology, and in particular to a data processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] With the growing development of internet finance, the public's direct and indirect participation in various financial activities is increasing, and financial behavior is becoming increasingly diverse. To maintain normal economic order and social stability, the state has placed greater emphasis on regulation and is continuously strengthening its oversight. Various regulatory agencies require financial institutions to promptly report daily activities that may involve transactional and customer risks, generating reports in a prescribed format.
[0003] Traditionally, multiple data tables involved in a message have been combined into a data summary table, and the message has been generated based on the data summary table. However, this traditional approach can take a long time to obtain the data summary table, resulting in low message delivery efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method, apparatus, device, and computer-readable storage medium, which can shorten the time for obtaining a data summary table and improve the timeliness of message reporting.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, comprising:
[0006] According to the current reporting requirements, obtain multiple data tables to be connected;
[0007] With minimizing the connection cost as the optimization goal, based on the first preset search rule, an optimal connection order of the multiple data tables is generated; the connection cost is related to the connection time and the size of the connection result table;
[0008] With minimizing execution cost as an optimization goal, selecting a target physical node for executing each connection operation in the optimal connection sequence from a distributed database based on a second preset search rule; the execution cost is related to the amount of data transmission, memory consumption, and load generated by executing the connection operation;
[0009] Performing a join operation on the multiple data tables on the target physical node according to the optimal join order.
[0010] In a second aspect, an embodiment of the present application provides a data processing device, including:
[0011] The acquisition module is used to obtain multiple data tables to be connected according to the current reporting requirements;
[0012] A generating module, configured to generate an optimal connection sequence for the plurality of data tables based on a first preset search rule with minimizing connection cost as an optimization goal; the connection cost is related to the connection time and the size of the connection result table;
[0013] a selection module configured to select, from a distributed database, a target physical node for executing each connection operation in the optimal connection sequence based on a second preset search rule, with minimizing execution cost as an optimization goal; wherein the execution cost is related to the amount of data transmitted, memory consumption, and load generated by executing the connection operation;
[0014] A connection module is used to perform a connection operation on the multiple data tables according to the optimal connection order on the target physical node.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the data processing method provided in the first aspect of the embodiment of the present application are implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the data processing method provided in the first aspect of the embodiment of the present application.
[0017] The technical solution provided by the embodiment of the present application obtains multiple data tables to be connected according to the current reporting requirements; takes minimizing the connection cost as the optimization goal, based on the first preset search rule, generates the optimal connection order of multiple data tables; takes minimizing the execution cost as the optimization goal, selects the target physical node for executing the optimal connection order from the distributed database based on the second preset search rule; and performs the connection operation on the multiple data tables according to the optimal connection order on the target physical node. It can be seen that before connecting multiple data tables, the embodiment of the present application generates the optimal connection order with the minimum connection cost through the first preset search rule, and generates each target physical node for executing the optimal connection order with the minimum execution cost through the second preset search rule, and connects multiple data tables in the shortest time on each target physical node, thereby shortening the acquisition time of the data summary table and reducing the size of the connection result table, thereby improving the timeliness of message reporting. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a data processing method provided in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of a flow chart of a process for determining an optimal connection sequence according to an embodiment of the present application;
[0020] Figure 3A schematic diagram of a process for determining a target physical node according to an embodiment of the present application;
[0021] Figure 4 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0022] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the technical solutions in the embodiments of this application are further described in detail through the following embodiments and in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Those skilled in the art can make adjustments to them as needed to suit specific application scenarios. It should also be noted that, for ease of description, the drawings only show parts related to this application rather than all structures.
[0024] Typically, commercial banks need to splice several database tables related to message reporting into a message data summary table, and convert the message data summary table into a text file. After that, all relevant information such as transactions, customers, institutions, warnings, etc. involving risky behaviors are extracted from the text file and a message is generated. When multiple data tables are connected to generate a message data summary table, due to the large number of data tables involved, the large amount of data, and the left outer join type, if the connection order of multiple data tables is not selected properly, an overly large multi-table connection result table will be generated, occupying too much memory space. At the same time, it will also cause the multi-table connection time to be too long, resulting in low message reporting efficiency. To this end, the embodiment of the present application provides a data processing method, which aims to solve the above-mentioned technical problems.
[0025] It should be noted that the information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0026] Figure 1 A flow chart of the data processing method provided in the embodiment of the present application. Figure 1 As shown, the method may include:
[0027] S101. According to current reporting requirements, multiple data tables to be connected are obtained.
[0028] In this embodiment of the present application, commercial banks need to screen daily transactions and corresponding customer transactions that may involve regulatory risks, generate messages, and submit them promptly. Since the data submitted daily must be spliced from relevant data tables to generate a data summary table, the data summary table is exported as a text file. The message generation program then scans this text file line by line to read the relevant information to generate a message, which is then submitted to the People's Bank of China for supervision. Therefore, based on the current reporting requirements, multiple data tables to be connected can be obtained, where the data in the multiple data tables to be connected can be one or more of relevant risk transaction data, customer information, and early warning information.
[0029] S102 : Taking minimizing the connection cost as an optimization goal, based on a first preset search rule, generate an optimal connection sequence of multiple data tables.
[0030] Among them, the connection cost is related to the connection time and the size of the connection result table. In some embodiments, the connection cost can be a weighted sum of the connection time and the size of the connection result table. The connection time refers to the time required to perform a left outer join operation on multiple data tables at the level of only considering the connection logic itself (that is, not relying on specific hardware performance). Optionally, the connection time is obtained by at least one of the following parameters: the full scan time of the left table, the full scan time of the right table, the connection processing time of the matching rows of the left table, and the placeholder filling time of the unmatched rows of the left table. For example, the placeholder can be the character "NULL".
[0031] The size of the join result table refers to the size of the join result table generated after performing a left outer join operation on multiple data tables. For a left outer join, the contents of the left table need to be completely retained, and the matching fields in the right table need to be filled into the left table. Unmatched fields need to be filled with corresponding placeholders in the left table. Accordingly, the size of the join result table can optionally be determined by the following process: determining the number of matching rows and unmatched rows in the left table that meet the join conditions; determining the storage size of the left table's matching rows based on the left table row size and the right table's matching field size; determining the storage size of the left table's unmatched rows based on the left table row size and the placeholder size of the right table's unmatched fields; and determining the size of the join result table based on the number of matching rows, the number of unmatched rows, the left table's matching row storage size, and the left table's unmatched row storage size.
[0032] Specifically, according to the join condition, the number of matching rows that meet the join condition is determined in the left table, and the number of unmatched rows is equal to the difference between the total number of rows in the left table and the number of matching rows; further, for the matching rows of the left table, the storage size of the matching rows of the left table after the left table and the right table are joined can be determined based on the sum of the left table row size and the right table matching field size; for the unmatched rows of the left table, the storage size of the unmatched rows of the left table after the left table and the right table are joined can be determined based on the sum of the left table row size and the placeholder size of the unmatched field of the right table; further, the product of the number of matching rows of the left table and the storage size of the matching rows of the left table is determined as the total storage size of the matching rows of the left table, the product of the number of unmatched rows of the left table and the storage size of the unmatched rows of the left table is determined as the total storage size of the unmatched rows of the left table, and the sum of the total storage size of the matching rows of the left table and the total storage size of the unmatched rows of the left table is determined as the size of the join result table.
[0033] In the embodiment of the present application, the optimal connection order of multiple data tables can be generated based on some optimization algorithms, such as dynamic programming algorithm, particle swarm optimization algorithm, genetic algorithm, etc., with minimizing the connection cost as the optimization goal. The optimal connection order includes the optimal order of sequentially connecting the data tables.
[0034] S103 : Taking minimizing the execution cost as the optimization goal, select a target physical node for executing each connection operation in the optimal connection sequence from the distributed database based on a second preset search rule.
[0035] Among them, the execution cost mainly considers the cost generated at the hardware execution level. Optionally, the execution cost is proportional to the amount of data transmission, memory consumption and load generated by executing the connection operation. In an embodiment of the present application, minimizing the execution cost can be used as the optimization goal, and based on some optimization algorithms, such as greedy algorithms, reinforcement learning, load balancing algorithms, etc., the target physical nodes for executing each connection operation in the optimal connection sequence can be generated, thereby forming an optimal execution path. The optimal execution path includes each target physical node that sequentially executes each connection operation in the optimal connection sequence.
[0036] S104: Perform a join operation on the multiple data tables according to the optimal join order on the target physical node.
[0037] After obtaining the optimal connection sequence and all target physical nodes that execute the optimal connection sequence, a left outer join operation is performed on multiple data tables according to the optimal connection sequence on each target physical node, which shortens the acquisition time of the data summary table and improves the timeliness of message delivery.
[0038] For example, assume that four data tables related to message reporting need to be connected into a data summary table, and the optimal connection order of these four data tables is obtained through the optimization algorithm as T1T2T3T4. The target physical nodes for executing each connection operation in the optimal connection order are node A, node B and node C, respectively. Then, a left outer join operation of T1 and T2 is performed on node A to obtain the connection result table T5. A left outer join operation of T5 and T3 is performed on node B to obtain the connection result table T6. A left outer join operation of T6 and T4 is performed on node C to obtain the final data summary table.
[0039] Furthermore, the data summary table may be sent to a downstream message generation program, which generates a message based on the data summary table and sends the message to the regulatory agency.
[0040] In one embodiment, optionally, the number of data tables to be connected is N, and N is an integer greater than 2, such as Figure 2 As shown, the above S102 may include:
[0041] S201 , taking each of the N data tables as a first-level solution space.
[0042] S202. Combine N data tables in pairs to obtain multiple table pairs. For each table pair, change the connection order of the data tables in the table pair and calculate the connection cost of each connection order. The connection order with the lowest connection cost is determined as the optimal connection order of the table pair. The optimal connection order of all table pairs is determined as the second-level solution space.
[0043] S203 : Generate a k-th layer solution space according to the i-th layer solution space and the j-th layer solution space.
[0044] Here, i and j are both integers greater than or equal to 1, and i+j=k, i≤j, and k increases layer by layer from 3 to N.
[0045] S204 : Perform a global comparison on the solutions in the solution space of the Nth layer, and determine the connection sequence with the minimum connection cost as the optimal connection sequence of the N data tables.
[0046] Specifically, each of the N data tables is used as the first-level solution space, that is, each data table is a solution in the solution space; the solutions in the first-level solution space are combined in pairs to obtain multiple table pairs; for each table pair, the connection order of the data tables in the table pair is changed; for each connection order, the connection time and connection result table generated by connecting the table pair based on the connection order are estimated, and the connection cost of the connection order is calculated based on the connection time and connection result table; the connection order with the smallest connection cost is used as the optimal connection order for the table pair, and the optimal connection order of all table pairs is determined as the second-level solution space. The solution in the second-level solution space is the optimal connection order for connecting two tables. It can be understood that there are multiple optimal connection orders for connecting two tables in the second-level solution space, which can ensure that all possible connection orders are retained, thereby finding the global optimal solution in the last level (i.e., the Nth level). Next, the solution spaces of the first and second layers are combined pairwise to obtain multiple table pairs in the third layer. For each table pair in the third layer, the join order of the data tables in the table pair is changed. For each join order, the join cost of the join order is calculated based on the join time and join result table generated by the join order. The join order with the lowest join cost is determined as the optimal join order for the table pair. The optimal join order for all table pairs in the third layer is determined as the solution space of the third layer. The solution in the third layer solution space is the optimal join order for the three-table join. Similarly, the solution space of the Nth layer is generated based on the solution space of the i-th layer and the solution space of the j-th layer. The solutions in the solution space of the first N-1 layers are all local optimal solutions. That is, for each layer in the first N-1 layers, the join costs between different table pairs are not compared. After obtaining the solution space of the last layer (i.e., the Nth layer), the solutions in the solution space of the Nth layer represent candidate join orders for the N data tables. The solutions in the solution space of the Nth layer are globally compared, and the candidate join order with the lowest join cost is determined as the optimal join order for the N data tables.
[0047] Continuing with the example of four data tables (such as TI, T2, T3, and T4), these four data tables are used as solutions in the first-level solution space. In the second level, the four data tables in the first-level solution space are combined in pairs to obtain six table pairs (such as (T1, T2), (T1, T3), (T1, T4), (T2, T3), (T2, T4), and (T3, T4)). For each table pair, the connection order of the data tables in the table pair is changed and the connection cost of each connection order is calculated. For example, for the table pair (T1, T2), there are two corresponding connection orders, such as T1T2 and T2T1. From these two connection orders, the connection order with the smallest cost is selected as the optimal connection order for the table pair. By analogy, we can get The optimal connection order of the above 6 table pairs is determined as the solution in the second-level solution space; based on the solution space of the first and second levels, the solution space of the third level is generated, and so on, until the solution space of the fourth level is obtained. The solutions in the solution space of the fourth level are obtained by combining the solutions in the solution space of the first and second levels in pairs, or by combining the solutions in the solution space of the second level in pairs. The solutions in the solution space of the first three levels are all local optimal solutions, that is, the connection costs between different table pairs are not compared. After obtaining the solution space of the fourth level, the solutions in the solution space of the fourth level are globally compared, and the connection order with the smallest connection cost is determined as the optimal connection order of the four data tables.
[0048] In an embodiment of the present application, with the goal of minimizing the connection cost, a local optimal solution for each layer is generated in a hierarchical manner, and the local optimal solution is used as a candidate subset for the next layer until the solution space of the Nth layer is obtained. The solutions in the solution space of the Nth layer are then globally compared to obtain the optimal connection order of the N data tables. The dynamic programming method can quickly and accurately find the global optimal connection order.
[0049] In one embodiment, optionally, Figure 3 As shown, the above S103 may include:
[0050] S301: For a current connection operation in an optimal connection sequence, determine a candidate physical node corresponding to the current connection operation.
[0051] For the current connection operation in the optimal connection sequence, all possible candidate physical nodes are identified from the distributed database. These candidate physical nodes can be physical nodes in the distributed database that store relevant data tables to be connected or can efficiently access these data tables.
[0052] Optionally, the candidate physical nodes may include at least one of the following: the physical node where the data table to be connected associated with the current connection operation is located, the physical node where the intermediate connection result table generated by the previous connection operation is located, the active physical node, the physical node whose current resource occupancy meets the preset conditions, and the physical node whose operating reliability meets the preset requirements.
[0053] Analyze the locations of the data tables to be connected involved in the current connection operation and the intermediate connection result tables of the previous connection operation, determine which physical nodes in the distributed database store the data tables required for the current connection operation, and identify these physical nodes as candidate physical nodes. This can reduce the amount of data transmission during the connection operation, thereby reducing network communication costs and improving data connection efficiency;
[0054] Active physical nodes refer to physical nodes that are currently running and can execute tasks. Selecting active physical nodes as candidate physical nodes can ensure that connection operations can be effectively executed.
[0055] Analyze the historical operating status of all nodes in the distributed database and select those physical nodes with stable performance and low failure rate as candidate physical nodes to improve the stability of connection operations;
[0056] Evaluating the current resource usage of all physical nodes in the distributed database and selecting physical nodes with sufficient resources as candidate physical nodes can shorten the total execution time of the connection operation, that is, shorten the time to obtain the data summary table, and thus improve the timeliness of message delivery.
[0057] S302: Calculate the data transmission volume and memory consumption generated by the candidate physical node executing the current connection operation.
[0058] If the relevant data table involved in the current connection operation is not stored on the candidate physical node, the candidate physical node needs to obtain the data table from the physical node where the data table is located, which will generate data transmission between the candidate physical node and the physical node where the data table is located. The data transmission volume is estimated, which is related to the size of the data table to be transferred; at the same time, the memory consumption refers to the memory size occupied by the candidate physical node when performing the connection operation, which is related to the size of the left table and the right table to be connected.
[0059] S303: Determine the execution cost of the candidate physical node for executing the current connection operation based on the data transmission volume, memory consumption, and the load of the candidate physical node.
[0060] Optionally, the execution cost C of the candidate physical node may be calculated according to the following formula:
[0061] C=V / (1-L)*M;
[0062] Among them, V is the data transmission volume, L is the load, and M is the memory consumption.
[0063] It can be understood that the smaller the data transmission volume, memory consumption and load, the smaller the execution cost of the corresponding candidate physical node, and the more suitable it is for executing the current connection operation; the larger the data transmission volume, memory consumption and load, the greater the execution cost of the corresponding candidate physical node, and the less suitable it is for executing the current connection operation.
[0064] S304. Determine the candidate physical node with the lowest execution cost as the target physical node for executing the current connection operation, use the next connection operation in the optimal connection sequence as the new current connection operation, and repeat the step of determining the candidate physical node corresponding to the current connection operation until the last connection operation in the optimal connection sequence is reached, and obtain all target physical nodes for executing the optimal connection sequence.
[0065] Compare the execution costs of all candidate physical nodes, select the candidate physical node with the smallest execution cost as the target physical node for executing the current connection operation, use the next connection operation in the optimal connection sequence as the new current connection operation, and repeat the steps S301-S304 above until all connection operations are assigned to corresponding target physical nodes.
[0066] On the basis of the optimal connection sequence, by estimating the execution cost of the candidate physical nodes, the optimal target physical node is selected for each connection operation in the optimal connection sequence based on the execution cost, thereby obtaining the optimal execution path with the minimum execution cost for executing the optimal connection sequence. By executing connection operations on multiple data tables through the target physical nodes on the optimal execution path, the efficiency of data table connection can be improved, thereby improving the efficiency of obtaining data summary tables, and further improving the timeliness of message delivery.
[0067] Figure 4 A structural diagram of a data processing device provided in an embodiment of the present application. Figure 4 As shown, the apparatus may include an acquisition module 401 , a generation module 402 , a selection module 403 and a connection module 404 .
[0068] Specifically, the acquisition module 401 is used to acquire multiple data tables to be connected according to the current reporting requirements;
[0069] The generation module 402 is configured to generate an optimal connection sequence for multiple data tables based on a first preset search rule with minimizing the connection cost as an optimization goal; the connection cost is related to the connection time and the size of the connection result table;
[0070] The selection module 403 is configured to select a target physical node for executing each connection operation in the optimal connection sequence from the distributed database based on a second preset search rule with minimizing the execution cost as an optimization goal; the execution cost is related to the amount of data transmission, memory consumption, and load generated by executing the connection operation;
[0071] The connection module 404 is used to perform a connection operation on multiple data tables according to the optimal connection order on the target physical node.
[0072] Based on the above embodiment, optionally, the number of data tables is N, and N is an integer greater than 2; the generation module 402 is further used to use each data table in the N data tables as the solution space of the first layer; the N data tables are combined in pairs to obtain multiple table pairs; for each table pair, the connection order of the data tables in the table pair is changed and the connection cost of each connection order is calculated, and the connection order with the minimum connection cost is determined as the optimal connection order of the table pair; the optimal connection order of all table pairs is determined as the solution space of the second layer; the solution space of the kth layer is generated according to the solution space of the i-th layer and the solution space of the j-th layer; the solutions in the solution space of the N-th layer are globally compared, and the connection order with the minimum connection cost is determined as the optimal connection order of the N data tables; wherein i and j are both integers greater than or equal to 1, and i+j=k, i≤j, and k increases layer by layer from 3 to N.
[0073] Based on the above embodiment, optionally, based on the above embodiment, optionally, the connection time is obtained by at least one of the following parameters:
[0074] Full scan time of the left table, full scan time of the right table, join processing time of matching rows of the left table, and time to fill placeholders for unmatched rows of the left table.
[0075] Based on the above embodiment, optionally, the size of the join result table is determined by the following process: determining the number of matching rows and the number of unmatched rows in the left table that meet the join condition; determining the storage size of the matching rows of the left table based on the left form row size and the matching field size of the right table; determining the storage size of the unmatched rows of the left table based on the left form row size and the placeholder size of the unmatched fields of the right table; and determining the size of the join result table based on the number of matching rows, the number of unmatched rows, the storage size of the matching rows of the left table, and the storage size of the unmatched rows of the left table.
[0076] Based on the above embodiment, optionally, the selection module 403 is also used to determine the candidate physical node corresponding to the current connection operation in the optimal connection sequence; calculate the data transmission volume and memory consumption generated by the candidate physical node performing the current connection operation; determine the execution cost of the candidate physical node performing the current connection operation based on the data transmission volume, memory consumption and the load of the candidate physical node; determine the candidate physical node with the smallest execution cost as the target physical node for performing the current connection operation, use the next connection operation in the optimal connection sequence as the new current connection operation, and repeat the step of determining the candidate physical node corresponding to the current connection operation until the last connection operation in the optimal connection sequence is reached, and all target physical nodes for executing the optimal connection sequence are obtained.
[0077] Based on the above embodiment, optionally, the candidate physical node includes at least one of the following:
[0078] The physical node where the data table to be connected associated with the current connection operation is located, the physical node where the intermediate connection result table generated by the previous connection operation is located, the active physical node, the physical node whose current resource occupancy meets the preset conditions, and the physical node whose operation reliability meets the preset requirements.
[0079] Based on the above embodiment, optionally, the execution cost is proportional to the data transmission amount, memory consumption and load.
[0080] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the device includes a processor 50, a memory 51, an input device 52 and an output device 53; the number of processors 50 in the device can be one or more. Figure 5 In the embodiment, a processor 50 is used as an example; the processor 50, the memory 51, the input device 52 and the output device 53 in the device can be connected by a bus or other means. Figure 5 The bus connection is taken as an example.
[0081] The memory 51, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method in the embodiments of the present application (for example, the acquisition module 401, the generation module 402, the selection module 403, and the connection module 404 in the data processing device). The processor 50 executes the various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 51, that is, implementing the above-mentioned data processing method.
[0082] The memory 51 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created during data processing, etc. Furthermore, the memory 51 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 51 may further include memory remotely located relative to the processor 50, and these remote memories may be connected to the device / terminal / server via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0083] The input device 52 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. The output device 53 may include a display device such as a display screen.
[0084] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0085] According to the current reporting requirements, obtain multiple data tables to be connected;
[0086] With minimizing the connection cost as the optimization goal, based on the first preset search rule, an optimal connection order of multiple data tables is generated; the connection cost is related to the connection time and the size of the connection result table;
[0087] With minimizing execution cost as an optimization goal, selecting a target physical node for executing each connection operation in the optimal connection sequence from the distributed database based on a second preset search rule; the execution cost is related to the amount of data transferred, memory consumed, and load generated by executing the connection operation;
[0088] Perform join operations on multiple data tables on the target physical node in the optimal join order.
[0089] The data processing devices, electronic devices, and computer-readable storage media provided in the above embodiments can execute the data processing methods provided in any embodiment of the present application, and have the corresponding functional modules and beneficial effects of executing the methods. For technical details not fully described in the above embodiments, please refer to the data processing methods provided in any embodiment of the present application.
[0090] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present application can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer's floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0091] It is worth noting that the various units and modules included in the above embodiments are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application.
[0092] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.
Claims
1. A data processing method, characterized in that: include: According to the current reporting requirements, obtain multiple data tables to be connected; Taking minimizing the connection cost as an optimization goal, based on a first preset search rule, generating an optimal connection order of the multiple data tables; The connection cost is related to the connection time and the size of the connection result table; With minimizing the execution cost as the optimization goal, selecting a target physical node for executing each connection operation in the optimal connection sequence from a distributed database based on a second preset search rule; The execution cost is related to the data transmission volume, memory consumption and load generated by executing the connection operation; Performing a join operation on the multiple data tables according to the optimal join order on the target physical node.
2. The method according to claim 1, characterized in that The number of the data tables is N, and N is an integer greater than 2; the optimization goal is to minimize the connection cost, and based on the first preset search rule, the optimal connection order of the multiple data tables is generated, including: Each data table in the N data tables is used as a solution space of the first layer; Combining the N data tables in pairs to obtain a plurality of table pairs; for each table pair, changing the connection order of the data tables in the table pair and calculating the connection cost of each connection order, and determining the connection order with the minimum connection cost as the optimal connection order for the table pair; and determining the optimal connection order of all table pairs as the second-level solution space; Generate the solution space of the kth layer according to the solution space of the ith layer and the solution space of the jth layer; i and j are both integers greater than or equal to 1, and i+j=k, i≤j, and k increases layer by layer from 3 to N; A global comparison is performed on the solutions in the solution space of the Nth layer, and a connection sequence with the minimum connection cost is determined as the optimal connection sequence of the N data tables.
3. The method according to claim 1, characterized in that The connection time is obtained by at least one of the following parameters: Full scan time of the left table, full scan time of the right table, join processing time of matching rows of the left table, and time to fill placeholders for unmatched rows of the left table.
4. The method according to claim 1, wherein The size of the connection result table is determined by the following process: Determine the number of matching rows and unmatched rows in the left table that meet the join condition; Determine the left table matching row storage size based on the left table row size and the right table matching field size; Determine the left table's unmatched row storage size based on the left table's row size and the right table's unmatched field's placeholder size; The size of the join result table is determined based on the number of matching rows, the number of unmatched rows, the storage size of the matching rows of the left table, and the storage size of the unmatched rows of the left table.
5. The method according to claim 1, wherein With minimizing the execution cost as the optimization goal, selecting a target physical node for executing each connection operation in the optimal connection sequence from a distributed database based on a second preset search rule includes: For a current connection operation in the optimal connection sequence, determining a candidate physical node corresponding to the current connection operation; Calculating the data transmission volume and memory consumption generated by the candidate physical node executing the current connection operation; Determining an execution cost of the candidate physical node performing the current connection operation based on the data transmission volume, the memory consumption, and the load of the candidate physical node; The candidate physical node with the lowest execution cost is determined as the target physical node for executing the current connection operation, the next connection operation in the optimal connection sequence is used as the new current connection operation, and the step of determining the candidate physical node corresponding to the current connection operation is repeated until the last connection operation in the optimal connection sequence is reached, and all target physical nodes for executing the optimal connection sequence are obtained.
6. The method according to claim 5, characterized in that The candidate physical nodes include at least one of the following: The physical node where the data table to be connected associated with the current connection operation is located, the physical node where the intermediate connection result table generated by the previous connection operation is located, the active physical node, the physical node whose current resource occupancy meets the preset conditions, and the physical node whose operation reliability meets the preset requirements.
7. The method according to claim 5, characterized in that The execution cost is proportional to the data transmission amount, memory consumption and load.
8. A data processing device, characterized in that: include: The acquisition module is used to obtain multiple data tables to be connected according to the current reporting requirements; A generating module, configured to generate an optimal connection sequence of the plurality of data tables based on a first preset search rule with minimizing connection cost as an optimization goal; The connection cost is related to the connection time and the size of the connection result table; A selection module is configured to select a target physical node for executing each connection operation in the optimal connection sequence from a distributed database based on a second preset search rule with minimizing execution cost as an optimization goal; The execution cost is related to the data transmission volume, memory consumption and load generated by executing the connection operation; A connection module is used to perform a connection operation on the multiple data tables according to the optimal connection order on the target physical node.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Connection operation execution method and device, storage medium and electronic equipment
CN121350039A