Optimization method and system for data processing

CN122838459APending Publication Date: 2026-09-29FUJIAN HUAYU EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610690020.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,在多轮次的表关联计算过程中,每轮计算后的中间结果数据量可能发生显著变化

Benefits of technology

[0006]本发明的有益效果在于:本发明通过获取待执行的表关联任务,依据初始数据规模确定当前关联方式并执行属于当前计算批次的表关联任务。该步骤在首轮计算中采用与初始数据特征相匹配的关联策略。在当前计算批次完成后,收集该批次产生的中间结果数据。基于该中间结果数据进行分析,以确定下一计算批次的目标关联方式。按照该目标关联方式执行下一计算批次的计算。由此,表关联任务在多轮次计算过程中的关联方式实现动态调整。现有技术通常在任务开始阶段根据初始表的数据规模选择一种关联方式,并在整个计算过程中持续使用该固定关联方式。本发明克服了在多轮次表关联计算过程中因每轮计算后的中间结果数据量发生显著变化而导致的效率下降问题。当中间结果数据量较初始表大幅增长时,目标关联方式可相应地从广播连接切换为哈希分区连接。这避免了大量数据在网络中传输所造成的网络带宽占用显著增加。该方案降低了分布式计算集群的网络传输开销,提升了大数据表关联计算在多轮次执行场景下的整体计算效率,同时减少了单节点内存压力,提高了分布式计算集群的资源利用率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838459A_ABST
    Figure CN122838459A_ABST
Patent Text Reader

Abstract

The application provides an optimization method and system for data processing. The method comprises: obtaining a table association task to be executed; determining a current association mode of the table association task, and executing the table association task belonging to a current computing batch according to the current association mode; collecting intermediate result data generated by the current computing batch; analyzing the intermediate result data to determine a target association mode of the table association task in a next computing batch; and executing the next computing batch according to the target association mode. The application can adapt to changes in data size and improve the efficiency of big data table association calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet data processing technology, and in particular to optimization methods and systems for data processing. Background Technology

[0002] With the rapid development of internet technology and the continuous growth of data scale, big data computing has become one of the core technologies of internet-related products. In big data computing scenarios, table join operations are the most common data processing tasks, widely used in data warehouse queries, real-time analysis, machine learning feature engineering, and other fields. Depending on the data scale of the tables involved in the join, table join operations can be categorized into several types, such as small table join with large table, large table join with small table, small table join with small table, and large table join with large table. Currently, distributed computing systems typically use fixed join methods when handling table join operations. Common join methods include broadcast join and hash partition join. A broadcast join involves copying all data from a table and distributing it to multiple computing nodes, with each node's subtasks performing join calculations and queries with that table. This method is suitable for scenarios where a small table joins a large table, avoiding data redistribution of the large table by broadcasting the small table to all nodes, thereby reducing network transmission overhead. A hash partition join involves partitioning the two tables involved in the join according to the hash value of the join field, with data with the same hash value assigned to the same computing node for join calculations. This approach is suitable for scenarios involving large table joins, achieving parallel computation across nodes by simultaneously redistributing data from two large tables. In practical applications, existing technologies typically select a join method based on the initial table's data size at the start of the task and continue using that method throughout the computation. For example, when joining a smaller table with a larger table, the system chooses a broadcast join method, broadcasting the smaller table to all nodes for the first round of computation. After the first round, if it's necessary to continue joining intermediate results with other tables, the system still uses the broadcast join method for subsequent calculations. However, during multiple rounds of table joins, the amount of intermediate result data after each round can change significantly. After the initial smaller table has undergone join computation, the amount of intermediate result data may increase dramatically. Continuing to use the broadcast join method in this case would result in a large amount of data being transmitted over the network, significantly increasing network bandwidth usage and reducing computational efficiency. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide an optimized data processing method that can adapt to changes in data scale and improve the efficiency of large data table association calculations.

[0004] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: An optimization method for data processing, applied to a distributed computing cluster, includes: obtaining table join tasks to be executed; determining the current join method of the table join tasks and executing the table join tasks belonging to the current computing batch according to the current join method; collecting intermediate result data generated by the current computing batch; analyzing the intermediate result data to determine the target join method of the table join tasks in the next computing batch; and executing the computing of the next computing batch according to the target join method.

[0005] To solve the above-mentioned technical problems, another technical solution adopted by the present invention is as follows: An optimized system for data processing includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the table join tasks to be executed; determine the current join method of the table join tasks, and execute the table join tasks belonging to the current calculation batch according to the current join method; collect the intermediate result data generated by the current calculation batch; analyze the intermediate result data to determine the target join method of the table join tasks in the next calculation batch; execute the calculation of the next calculation batch according to the target join method.

[0006] The beneficial effects of this invention are as follows: This invention obtains the table join task to be executed, determines the current join method based on the initial data size, and executes the table join task belonging to the current calculation batch. This step employs a join strategy matching the characteristics of the initial data in the first round of calculation. After the current calculation batch is completed, the intermediate result data generated by that batch is collected. Based on the analysis of this intermediate result data, the target join method for the next calculation batch is determined. The calculation of the next calculation batch is executed according to this target join method. Thus, the join method of the table join task is dynamically adjusted during multiple rounds of calculation. Existing technologies typically select a join method based on the initial table data size at the beginning of the task and continuously use this fixed join method throughout the entire calculation process. This invention overcomes the efficiency decline problem caused by significant changes in the amount of intermediate result data after each round of calculation during multi-round table join calculations. When the amount of intermediate result data increases significantly compared to the initial table, the target join method can be switched from broadcast connection to hash partition connection accordingly. This avoids a significant increase in network bandwidth consumption caused by the transmission of large amounts of data over the network. This solution reduces network transmission overhead in distributed computing clusters, improves the overall computing efficiency of big data table join calculations in multi-round execution scenarios, reduces memory pressure on single nodes, and improves resource utilization of distributed computing clusters. Attached Figure Description

[0007] Figure 1 A flowchart illustrating the steps of an optimization method for data processing provided in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an optimized data processing system provided in an embodiment of the present invention; Detailed Implementation To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0008] In existing technologies, with the rapid development of internet technology and the continuous growth of data scale, big data computing has become one of the core technologies of internet-related products. In big data computing scenarios, table join operations are the most common data processing tasks, widely used in data warehouse queries, real-time analysis, machine learning feature engineering, and other fields. During multiple rounds of table join calculations, the amount of intermediate result data after each round may change significantly. When the initial small table undergoes join calculations, the amount of intermediate result data may increase dramatically. Continuing to use broadcast connections at this point will result in a large amount of data being transmitted over the network, significantly increasing network bandwidth consumption and reducing computational efficiency.

[0009] To at least address the aforementioned issues, this invention collects and analyzes the intermediate result data scale in real time during multi-round table join calculations, and dynamically adjusts the join method based on changes in data volume. This enables the system to adaptively switch between broadcast joins and hash partition joins, thereby achieving adaptation to changes in data scale, reducing network overhead, and improving the efficiency of large data table join calculations.

[0010] The following describes in detail an optimization method for data processing according to the present invention, applied to a distributed computing cluster, as shown in the appendix. Figure 1 ,include: Step 101: Obtain the table join task to be executed. A table join task refers to a computational task executed in a distributed computing cluster that performs a join operation on at least two data tables. The join operation includes matching and combining rows from one data table with rows from another data table based on the join key. In big data computing scenarios, table join tasks include types such as small table join with large table, large table join with small table, small table join with small table, and large table join with large table.

[0011] Step 102: Determine the current join method for the table join task and execute the table join task belonging to the current calculation batch according to the current join method. The current join method refers to the distributed table join execution strategy selected when starting the current calculation batch, such as broadcast join or reduced node join. The current calculation batch refers to the calculation unit currently being executed when the entire execution process of the table join task is divided into multiple sequential execution stages. Each batch completes the join calculation of a portion of the data tables and produces intermediate result data.

[0012] Step 103: Collect intermediate result data generated by the current calculation batch; where intermediate result data refers to the partial association result dataset obtained after the table association task finishes an execution batch, and this result dataset will be used as input data for the next batch of subsequent association calculations.

[0013] Step 104: Analyze the intermediate result data to determine the target association method for the table association task in the next calculation batch; where the target association method refers to the distributed table connection execution strategy suitable for the next calculation batch, which is determined after analyzing and judging the amount of data in the intermediate result data and the amount of data in the target data table to be associated in the next batch, such as broadcast association method or node reduction association method.

[0014] Step 105: Execute the calculation for the next batch of calculations according to the target association method; As described above, this embodiment collects intermediate result data from the current computation batch and analyzes this data to determine the target association method for the next computation batch, thus achieving dynamic adjustment of the association method in the table association task. This process does not require interrupting the current computation batch and can automatically switch to a better association method in subsequent batches, thereby reducing data skew and computational redundancy caused by a fixed association method and improving the overall execution efficiency of table association tasks in a distributed computing cluster.

[0015] In one embodiment of this application, step 102, determining the current association method of the table association task and executing the table association task belonging to the current calculation batch according to the current association method, includes: Step 201: Obtain the first table in the table association task to be executed; wherein, the first table refers to the data table that is determined to be small in data volume and suitable for broadcast transmission in the network in the table association task, and this table will be associated with the second table stored on each child node of the distributed computing cluster.

[0016] Step 202: If the current association method for the table association task is determined to be broadcast connection, then the data of the first table is broadcast to multiple child nodes of the distributed computing cluster. This allows multiple child nodes to perform association calculations between the first table and the second table within their child nodes. The data size of the first table is smaller than that of the second table. Broadcast connection is a distributed table join execution strategy that transmits all data from the smaller table to all child nodes of the distributed computing cluster via the network. Each child node then performs an association calculation between this table and the larger, locally stored second table. Broadcast refers to the process of sending all data from the first table to all designated child nodes. Child nodes are the worker nodes in the distributed computing cluster that undertake data storage and computation tasks. The second table, in the broadcast connection method, refers to the table stored locally on a child node and with a larger data size than the first table.

[0017] As described above, this embodiment broadcasts the first table, which has a smaller data volume, to multiple child nodes of the distributed computing cluster, allowing each child node to complete the association calculation between the first and second tables locally. This avoids the network overhead caused by frequent cross-node transmission of the second table data in traditional solutions, improves the execution efficiency of table association tasks under the broadcast connection method, and reduces the cluster communication load.

[0018] In one embodiment of this application, step 103, collecting intermediate result data generated by the current calculation batch, includes: Step 301: Receive intermediate result data from multiple child nodes of the distributed computing cluster in the current computing batch through the intermediate computing cache layer in the distributed computing cluster; wherein, the intermediate computing cache layer is a service layer independently deployed in the distributed computing cluster, used to receive and cache statistical information of the intermediate result data generated by each subtask in each computing batch, and the cache layer has the ability to communicate with the task allocation node and child nodes.

[0019] Step 302: Write the intermediate result data into the cache record structure in the intermediate computation cache layer to collect the intermediate result data. The cache record structure is a data organization format used to record the statistical information of intermediate results for each computation batch, and includes at least the fields of total task identifier, batch identifier, subtask identifier, and the amount of computation result data generated by the subtask. The total task identifier uniquely identifies a table-related task, the batch identifier identifies which computation batch under this task, the subtask identifier identifies the subtask assigned to a certain child node within this batch, and the amount of computation result data records records the number or size of intermediate result data produced by the subtask.

[0020] As described above, this embodiment receives and stores intermediate result data fed back by each child node in the current computing batch through an intermediate computing cache layer, avoiding the direct writing of intermediate result data to external storage devices by each child node. This reduces data transmission paths and lowers network bandwidth usage. Simultaneously, the writing method of the cache record structure improves the real-time performance and consistency of data collection, thereby enhancing the data aggregation efficiency of the distributed computing cluster during batch processing.

[0021] In one embodiment of this application, step 104 involves analyzing intermediate result data to determine the target association method for the table association task in the next calculation batch, including: Step 401: Obtain the data volume of the intermediate result data and the data volume of the target data table to be associated in the next calculation batch; whereby the data volume of the intermediate result data refers to the number of records or the number of bytes of storage space occupied by the intermediate result dataset generated in the current calculation batch. The target data table to be associated refers to the data table that needs to continue the association operation with the intermediate result data in the next calculation batch.

[0022] Step 402: If the amount of intermediate result data is greater than or equal to the first preset threshold, and the amount of data in the target data table to be associated is less than the second preset threshold, then the target association method is determined to be the node reduction association method. The node reduction association method is a distributed table join execution strategy that executes the next batch of association calculations only on the specific set of child nodes that currently hold the intermediate result data, and synchronizes the small table data to be associated to these child nodes, thereby avoiding broadcasting data to all child nodes and reducing the amount of network input and output data transmission.

[0023] Step 403: If the amount of intermediate result data is less than the first preset threshold and the amount of data in the target data table to be associated is greater than or equal to the second preset threshold, then the target association method is determined to be the broadcast association method; wherein, the broadcast association method is an execution strategy in which the intermediate result data is broadcast as a small table to all child nodes of the distributed computing cluster, and each child node performs association calculation with the broadcast received data and the large table stored locally.

[0024] As described above, this embodiment obtains the amount of intermediate result data and the amount of data in the target data table to be associated, and dynamically determines the target association method according to preset threshold conditions. This embodiment can reduce network I / O overhead by reducing the number of nodes when the amount of intermediate result data is large, and use the broadcast association method to ensure computational parallelism when the amount of intermediate result data is small. This realizes flexible adjustment of the association method according to the actual amount of data, and effectively improves data processing efficiency.

[0025] In one embodiment of this application, step 105, performing the calculation of the next calculation batch according to the target association method, includes: Step 501: Tables in the target data table to be associated that have a data volume less than a second preset threshold are classified as first-type tables, and tables in the target data table to be associated that have a data volume greater than or equal to the second preset threshold are classified as second-type tables. The first-type tables refer to small data tables with a data volume less than the second preset threshold among all data tables involved in the next calculation batch. The second-type tables refer to large data tables with a data volume greater than or equal to the second preset threshold.

[0026] Step 502: Assign a preset number of child nodes to perform association calculations between the first target data tables contained in the first type of table, generating intermediate aggregation results for the first type of table; wherein, the preset number of child nodes refers to a number of child nodes selected based on the data size of the first type of table. The intermediate aggregation result refers to the aggregated dataset obtained after performing association calculations on the specified number of child nodes for multiple tables in the first type of table.

[0027] Step 503: Broadcast the intermediate aggregation result of the first type of table to all child nodes of the distributed computing cluster so that all child nodes can perform association calculations with the intermediate aggregation result and the second type of table. As described above, this embodiment can adopt the most suitable association method for tables with different data volumes, reduce the number of nodes when performing association calculations on small tables, reduce network I / O overhead, and at the same time ensure the parallelism when performing association calculations on large tables, effectively improving data processing efficiency.

[0028] In one embodiment of this application, step 403, after determining that the target association method is the node association reduction method, includes: Step 601: Determine the target child node in the distributed computing cluster that generates intermediate result data; wherein, the target child node refers to the child node that locally stores the intermediate result data after completing the calculation of the previous batch of calculations.

[0029] Step 602: Synchronize the data of the target data table to be associated to the target child nodes; where synchronization means copying and transmitting all the data of the target data table to be associated to the local storage of each target child node through the network.

[0030] Step 603: Assign the table association task to the target child node in the next calculation batch, so that the target child node can perform the association calculation between the target data table to be associated and the intermediate result data, thereby reducing the amount of network input and output data transmission between the target child nodes; where the amount of network input and output data transmission refers to the total amount of data communication generated by the child nodes through network data transmission during the distributed computing process.

[0031] As described above, this embodiment determines the target child node that generates intermediate result data and synchronizes the data of the target data table to be associated to the target child node, so that the association calculation is performed only locally on the target child node. This embodiment significantly reduces the amount of network input and output data transmission between target child nodes, reduces network bandwidth consumption, and reduces the latency caused by cross-node data transmission, thereby improving the execution efficiency of table association calculation.

[0032] In one embodiment of this application, after obtaining the table association task to be executed in step 101, the method further includes: Step 701: Determine whether the data volume of the current data tables to be associated in the table association task is less than the second preset threshold. The second preset threshold is a pre-set value used to determine whether the current data tables to be associated in the table association task are small tables. If the data volume of all data tables is less than this threshold, the task belongs to the scenario of small table association with small table, and the dynamic adjustment of association method is not started.

[0033] Step 702: If the amount of data in the current data tables to be associated by the table association task is less than the second preset threshold, then skip the steps of collecting intermediate result data generated by the current calculation batch and analyzing the intermediate result data to determine the target association method of the table association task in the next calculation batch. Step 703: If there is a data table with a data volume greater than or equal to the second preset threshold in the current data table to be associated with the table association task, then the step of determining the current association method of the table association task is triggered. As described above, after obtaining the table join task to be executed, this embodiment first determines whether the data volume of all tables to be joined is less than a second preset threshold. If all data volume is less than the threshold, the step of collecting intermediate result data and determining the next batch of target join methods based on the data are skipped. If there is a data table with a data volume greater than or equal to the threshold, the step of determining the current join method is triggered. In this way, unnecessary intermediate result collection and analysis processing can be avoided when the data volume is small, thereby reducing the waste of computing resources and improving the execution efficiency of the table join task.

[0034] In one embodiment of this application, step 103, before collecting intermediate result data generated by the current computation batch, further includes: Step 801: Deploy an intermediate computing cache layer in the distributed computing cluster; Step 802: Configure the communication connection between the intermediate computing cache layer and the task allocation node and multiple child nodes in the distributed computing cluster, so that the intermediate computing cache layer can receive the task instructions issued by the task allocation node and the intermediate result data reported by the multiple child nodes; wherein, the task allocation node is the control node in the distributed computing cluster responsible for splitting the table association task into multiple batches and allocating them to various child nodes for execution.

[0035] As described above, by deploying an intermediate computing cache layer and establishing communication connections between it and the task allocation node and multiple child nodes, the intermediate computing cache layer can directly receive task instructions and intermediate result data. This avoids frequent transmission and temporary storage of intermediate result data between multiple nodes, thereby reducing network I / O overhead and data dependencies between nodes. Therefore, it effectively reduces the data transmission latency of distributed computing tasks, improves the efficiency of intermediate result data retrieval, and enhances the overall computing performance of the system.

[0036] In one embodiment of this application, step 102, executing the table join task belonging to the current calculation batch according to the current join method, further includes: Step 901: Determine whether all table association tasks have been completed. If so, stop obtaining the next calculation batch. As described above, the ability to stop acquiring the next batch of data when the task is completed ahead of schedule avoids unnecessary batch acquisition and data processing operations, saves computing resources and system overhead, and improves the overall efficiency of data processing.

[0037] The data processing optimization method and system described above are applicable to massive data processing, especially massive data processing in distributed computing systems. The following is a description of specific implementation methods.

[0038] This distributed computing system includes task allocation nodes, multiple computing sub-nodes, and an intermediate computing cache layer. The task allocation nodes are responsible for receiving table join calculation requests, decomposing the calculation task into multiple sub-tasks, and allocating them to the respective computing sub-nodes; the computing sub-nodes are responsible for executing the specific table join calculation operations; the intermediate computing cache layer is responsible for recording the intermediate result data of each round of calculation and analyzing and processing the data to adjust the subsequent join method.

[0039] In this embodiment, the intermediate computation cache layer is implemented using a distributed in-memory database, such as a Redis cluster, to store the computation result statistics for each subtask. The data structure of this cache layer includes four fields: total task identifier, batch identifier, subtask identifier, and computation result data volume. The total task identifier uniquely identifies a complete table-association computation task, the batch identifier identifies the round of task execution, the subtask identifier identifies the specific subtask assigned to each computation sub-node, and the computation result data volume records the number of intermediate result data entries generated after the subtask is executed.

[0040] Step A: In scenarios involving massive data computation, a new intermediate computation cache layer is added. This intermediate computation cache layer is used to record the results of each round of computation and perform subsequent analysis. This corresponds to step 101 above.

[0041] Step B: When a data table join operation involves a large table, a special processing procedure is initiated; when only a small table is joined with another small table, this special processing procedure is not initiated. This corresponds to step 102 above.

[0042] Step C: During the first round of calculation, the intermediate computation cache layer determines the initial connection method based on the amount of data in the tables to be joined. For scenarios involving a small table joining a large table, if the small table is in the millions or tens of millions and the large table is in the hundreds of millions, the intermediate computation cache layer determines to use a broadcast connection method, and the task allocation node broadcasts the small table data to multiple child nodes. This corresponds to step 102 above. Optionally, the table to the left of the join position indicates that the data has already been retrieved, so if it is a small table, it can be broadcast to allow multiple nodes to process it concurrently. If the table to the left of the join position is a large table, the data from the corresponding small table is synchronously transmitted back to the individual nodes where these large tables reside for processing, thereby reducing the amount of data transmitted.

[0043] Step D: Each child node in the current round performs table join calculations according to the determined connection method, and after the subtasks in this round are completed, it feeds back its respective calculation result data volume to the intermediate calculation cache layer. This corresponds to step 102 above.

[0044] Step E: The intermediate computation cache layer caches the computation results of the subtasks in this round. The cache structure includes the total task identifier, batch identifier, subtask identifier, and the amount of computation result data. This corresponds to step 103 above.

[0045] Step F: After the intermediate computation cache layer collects the computation results of one round, it compares the current batch result data volume with the data volume of the table to be associated in the next batch to determine the connection method for the next round. If the current batch result data volume reaches the hundreds of millions and the next batch is associated with a small table, then it is determined that the next round will not use a broadcast connection, and the data of the small table will be synchronized to the partial nodes where the result data is located for computation; if the current batch result data volume and the data volume of the table to be associated in the next batch are both in the millions, then it is determined that the next round will first perform computation on partial nodes, and then broadcast the computation results to the hundreds of millions data table for full node computation. This corresponds to step 104 above.

[0046] Step G: The intermediate computation cache layer adjusts the subsequent computation connection method based on the analysis results, generates a new task plan, and controls the child nodes to execute the next round of table join calculations according to the adjusted connection method. This corresponds to step 105 above.

[0047] Step H: Repeat steps D to G, adjusting the join method of table join calculations round by round, until all table join calculations required for this query task are completed.

[0048] Please refer to Figure 2 The present invention also provides a data processing optimization system 210, including a memory 211, a processor 212, and a computer program stored on the memory 211 and executable on the processor 212, wherein the processor 212 executes the computer program to implement the various steps of a data processing optimization method as described above.

[0049] The beneficial effects of the terminal of the present invention are the same as those of the method described above, and will not be repeated here.

[0050] In summary, this invention acquires the table join task to be executed, determines the current join method based on the initial data size, and executes the table join task belonging to the current calculation batch. This step employs a join strategy matching the characteristics of the initial data in the first round of calculation. After the current calculation batch is completed, the intermediate result data generated by that batch is collected. This intermediate result data is analyzed to determine the target join method for the next calculation batch. The next calculation batch is then executed according to this target join method. Thus, the join method of the table join task is dynamically adjusted during multiple rounds of calculation. Existing technologies typically select a join method based on the initial table data size at the beginning of the task and continuously use this fixed join method throughout the entire calculation process. This invention overcomes the efficiency decline problem caused by significant changes in the amount of intermediate result data after each round of calculation during multi-round table join calculations. When the amount of intermediate result data increases significantly compared to the initial table, the target join method can be switched from broadcast connection to hash partition connection accordingly. This avoids a significant increase in network bandwidth consumption caused by the transmission of large amounts of data over the network. This solution reduces network transmission overhead in distributed computing clusters, improves the overall computing efficiency of big data table join calculations in multi-round execution scenarios, reduces memory pressure on single nodes, and improves resource utilization of distributed computing clusters.

[0051] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention's specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. An optimization method for data processing, characterized in that, Applied to distributed computing clusters, the method includes: Retrieve table-related tasks to be executed; Determine the current association method of the table association task, and execute the table association task belonging to the current calculation batch according to the current association method; Collect intermediate result data generated by the current calculation batch; Based on the analysis of the intermediate result data, the target association method of the table association task in the next calculation batch is determined; The calculations for the next batch of calculations are performed according to the target association method described above.

2. The method according to claim 1, characterized in that, The step of determining the current association method of the table association task and executing the table association task belonging to the current calculation batch according to the current association method includes: Obtain the first table in the table association task to be executed; If it is determined that the current association method of the table association task is broadcast connection, then the data of the first table is broadcast to multiple child nodes of the distributed computing cluster, so that the multiple child nodes can perform association calculations between the first table and the second table in the child nodes, and the data volume of the first table is less than the data volume of the second table.

3. The method according to claim 1, characterized in that, The collection of intermediate result data generated by the current calculation batch includes: The intermediate computing cache layer in the distributed computing cluster receives intermediate result data fed back by multiple child nodes of the distributed computing cluster in the current computing batch. The intermediate result data is written into the cache record structure in the intermediate computing cache layer to collect the intermediate result data.

4. The method according to claim 1, characterized in that, The analysis based on the intermediate result data to determine the target association method for the table association task in the next calculation batch includes: The amount of data obtained from the intermediate result data, and the amount of data from the target data table to be associated in the next calculation batch of the table association task; If the amount of intermediate result data is greater than or equal to the first preset threshold, and the amount of data in the target data table to be associated is less than the second preset threshold, then the target association method is determined to be the node reduction association method. If the amount of intermediate result data is less than the first preset threshold, and the amount of data in the target data table to be associated is greater than or equal to the second preset threshold, then the target association method is determined to be broadcast association method.

5. The method according to claim 4, characterized in that, The step of performing the calculation of the next batch of calculations according to the target association method includes: The table with a data volume less than the second preset threshold is classified as a first type of table, and the table with a data volume greater than or equal to the second preset threshold is classified as a second type of table. A preset number of child nodes are assigned to perform association calculations between the first target data tables contained in the first type of table, generating intermediate aggregation results for the first type of table; The intermediate aggregation results of the first type of table are broadcast to all child nodes of the distributed computing cluster, so that all child nodes can perform association calculations between the intermediate aggregation results and the second type of table.

6. The method according to claim 4, characterized in that, The determination of the target association method after reducing node association methods includes: Determine the target child node in the distributed computing cluster that generates the intermediate result data; Synchronize the data of the target data table to be associated to the target child node; The table association task will be assigned to the target child node in the next computation batch.

7. The method according to claim 1, characterized in that, After obtaining the table join task to be executed, the method further includes: Determine whether the data volume of the currently to-be-associated data table associated with the table association task is less than the second preset threshold; If the amount of data in the current data table to be associated with the table association task is less than the second preset threshold, then skip the steps of collecting intermediate result data generated by the current calculation batch and the steps of analyzing the intermediate result data to determine the target association method of the table association task in the next calculation batch. If there is a data table with a data volume greater than or equal to the second preset threshold in the current data table to be associated with the table association task, then the step of determining the current association method of the table association task is triggered.

8. The method according to claim 1, characterized in that, Before collecting the intermediate result data generated by the current computation batch, the method further includes: Deploy an intermediate computing cache layer in a distributed computing cluster; Configure the communication connection between the intermediate computing cache layer and the task allocation node and multiple child nodes in the distributed computing cluster, so that the intermediate computing cache layer can receive the task instructions issued by the task allocation node and the intermediate result data reported by the multiple child nodes.

9. The method according to claim 1, characterized in that, The step of executing the table join task belonging to the current calculation batch according to the current join method further includes: Determine whether all the tasks associated with the tables have been completed. If so, stop obtaining the next batch of calculations.

10. An optimization system for data processing, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the optimized method for data processing as described in any one of claims 1 to 9.