A response method and device for distributed computing tasks
By determining the target fields in the distributed computing task and merging the lookup tables using broadcast merging, the hardware resource consumption problem when the data is large is solved and the computing efficiency is improved.
Patent Information
- Application Number
- CN202010054782.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-01-17
AI Technical Summary
The existing distributed computing technology requires a large amount of hardware resources to sort out data when integrating lookup tables, resulting in increased processing time and inefficient efficiency.
By determining the target fields of the distributed computing task, the target data is extracted from the benchmark query table to generate a field lookup table, and when the data amount is less than the preset threshold, it is merged with the target configuration table of the distributed node to avoid correlation of the field numbers.
It reduces the number of data combing operations, reduces the data read and write pressure of distributed nodes, and improves the response rate of distributed computing.
Smart Images

Figure CN111241163B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and device for responding to distributed computing tasks. Background Art
[0002] With the continuous advancement of the electronic process, most files can be digital files and stored in a cloud database. To ensure the access efficiency of the database, most data storage modes adopt distributed storage, dividing electronic files belonging to the same into multiple different databases and handing them over to each distributed node for storage. Therefore, when responding to data calculation tasks, a distributed computing framework, such as the Spark engine, needs to be used to extract data from each distributed node based on a query table.
[0003] In the existing distributed computing technology, when using a distributed computing engine, due to data changes, it is often necessary to integrate multiple query tables in the database and perform calculation responses based on the integrated data table. However, in the process of integrating the query table, it is necessary to re-comb each data in the table. When the data volume of the data table is large, it is necessary to consume more hardware resources to perform the combing operation of the data table, increasing the processing time and reducing the efficiency of distributed computing. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and device for responding to distributed computing tasks to solve the problem that in the existing distributed computing technology, in the process of integrating query tables, it is necessary to re-comb each data in the table. When the data volume of the data table is large, it is necessary to consume more hardware resources to perform the combing operation of the data table, increasing the processing time and having low efficiency of distributed computing.
[0005] The first aspect of the embodiments of the present invention provides a method for responding to a distributed computing task, including:
[0006] If a distributed computing task is received, determine the target field of the distributed computing task;
[0007] Extract the target data of the target field from a reference query table to generate a field query table;
[0008] Count the field data volume of the field query table;
[0009] If the field data volume is less than a preset broadcast trigger threshold, broadcast and send the field query table to each distributed node, so that the distributed node merges the field query table with a local target configuration table based on a broadcast merging method;
[0010] Execute the distributed computing task based on the target configuration table merged by each distributed node.
[0011] The second aspect of the embodiments of the present invention provides a response device for distributed computing tasks, including:
[0012] A target field recognition unit, configured to determine the target field of the distributed computing task if a distributed computing task is received;
[0013] A field query table generation unit, configured to extract the target data of the target field from a reference query table and generate a field query table;
[0014] A field data volume statistics unit, configured to count the field data volume of the field query table;
[0015] A broadcast merge trigger unit, configured to broadcast and send the field query table to each distributed node if the field data volume is less than a preset broadcast trigger threshold, so that the distributed node merges the field query table with a local target configuration table based on a broadcast merge method;
[0016] A distributed computing task response unit, configured to execute the distributed computing task based on the target configuration table merged by each distributed node.
[0017] The third aspect of the embodiments of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the first aspect are implemented.
[0018] The fourth aspect of the embodiments of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the first aspect are implemented.
[0019] Implementing a response method and device for distributed computing tasks provided by the embodiments of the present invention has the following beneficial effects:
[0020] In the embodiments of the present invention, by parsing the distributed computing tasks, the target fields required for the computing operations are determined, and the target data related to the target fields is extracted from the reference query table to generate a field query table. Thus, a total data table is split into sub-tables related to the tasks, a large amount of invalid data unrelated to the current computation is removed, the data volume of the sub-tables is reduced, and when the data volume of the fields in the field query table is less than the broadcast trigger threshold, the field query table is sent to each distributed storage node by means of broadcast transmission, so that the distributed storage nodes can merge the field query table with the target configuration table based on the broadcast merge (Broadcast Join) method. Since the merge method of Broadcast Join does not require associating the field number (Key) values and sorts out the entire data table, it belongs to a fast merge method of the data table, which improves the efficiency of distributed computing. Compared with the existing response technologies for distributed computing, the present embodiment can filter out the valid data, that is, the target data of the target fields related to the current computation, when generating the data table, reduce the data volume of the data table to be sent. In the case of a small data volume of the data table, the field query table and the target configuration table are merged by the simple merge method of Broadcast Join, greatly reducing the number of data sorting operations and the data read / write pressure on the distributed nodes, thereby improving the response rate of distributed computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a flowchart of the implementation of a method for responding to distributed computing tasks provided by the first embodiment of the present invention;
[0023] Figure 2 It is a specific implementation flowchart of a method for responding to distributed computing tasks provided by the second embodiment of the present invention;
[0024] Figure 3 It is a specific implementation flowchart of S202 of a method for responding to distributed computing tasks provided by the third embodiment of the present invention;
[0025] Figure 4 It is a specific implementation flowchart of S101 of a method for responding to distributed computing tasks provided by the fourth embodiment of the present invention;
[0026] Figure 5It is a specific implementation flowchart of a method for responding to distributed computing tasks provided by the fifth embodiment of the present invention;
[0027] Figure 6 It is a specific implementation flowchart of a method for responding to distributed computing tasks provided by the sixth embodiment of the present invention;
[0028] Figure 7 It is a specific implementation flowchart of a method for responding to distributed computing tasks S102 provided by the seventh embodiment of the present invention;
[0029] Figure 8 It is a structural block diagram of a device for responding to distributed computing tasks provided by an embodiment of the present invention;
[0030] Figure 9 It is a schematic diagram of a terminal device provided by another embodiment of the present invention. Detailed implementation manners
[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0032] In the embodiments of the present invention, by analyzing the distributed computing tasks, the target fields required for the computing operations are determined, and the target data related to the target fields is extracted from the reference query table to generate a field query table, so as to split a general data table into sub-tables related to the tasks, eliminating a large amount of invalid data irrelevant to the current calculation, reducing the data volume of the sub-tables, and when the data volume of the fields in the field query table is less than the broadcast trigger threshold, the field query table is sent to each distributed storage node by means of broadcast transmission, so that the distributed storage node can merge the field query table with the target configuration table based on the broadcast merge (Broadcast Join) method. Since the merge method of Broadcast Join does not require associating the field number (Key) values and sorts out the entire data table, it belongs to a fast merge method of the data table, improving the efficiency of distributed computing and solving the problems of the existing distributed computing technology that, during the process of integrating the query table, it is necessary to re-sort out each data in the table, and when the data volume of the data table is large, it is necessary to consume more hardware resources to execute the sorting operation of the data table, increasing the processing time and resulting in low efficiency of distributed computing.
[0033] In an embodiment of the present invention, the execution subject of the process is a terminal device. The terminal device includes, but is not limited to: devices such as servers, computers, smart phones, and tablet computers that can respond to distributed computing tasks. Specifically, the terminal device may be a server in a distributed computing system deployed based on the Spark engine. The server and a plurality of different distributed storage nodes together form a distributed computing system, which is used to store data uploaded by each user terminal and respond to distributed computing tasks. Figure 1 The flowchart of the implementation of the method for responding to a distributed computing task provided in the first embodiment of the present invention is shown and described in detail as follows:
[0034] In S101, if a distributed computing task is received, the target field of the distributed computing task is determined.
[0035] In this embodiment, the user can generate a distributed computing task on the local terminal and send the distributed computing task to the terminal device through the client corresponding to the distributed computing system. In this case, the distributed computing task carries the program identifier of the client. After receiving the distributed computing task, the terminal device can identify the program identifier to determine whether the user terminal is a legal terminal; if so, the operation of S101 is executed; otherwise, it is identified as an invalid task. The terminal device can also be set with a timing task. When the preset calculation trigger condition is met, a distributed computing task is automatically created and the operation of S101 is executed. For example, a calculation task triggered periodically such as setting the sales record of the current month on the last day of each month can be configured with a trigger script for this type of distributed computing task. When it is detected that the current moment meets the preset trigger period, the corresponding trigger script is executed to generate the corresponding distributed computing task.
[0036] Specifically, in this embodiment, the distributed computing task is specifically a computing system built with the Spark engine. Among them, the framework for processing Stream data on the Spark distributed computing system has a basic principle of dividing the Stream data into multiple small data segments and processing the data segments in a manner similar to batch processing. Since Spark Streaming is built on the Spark distributed computing system, on the one hand, due to the low-latency execution of Spark, it can be used for real-time computing. On the other hand, compared with other processing frameworks based on records (such as Storm), the resilient distributed dataset RDD with narrow dependencies in the Spark distributed computing system can recompute from the source data to achieve the purpose of fault tolerance. In addition, because the data is divided into multiple small data segments and the small-batch processing method is adopted, the Spark distributed computing system can be compatible with the logic and algorithms of both batch and real-time data processing at the same time, facilitating some specific application scenarios that require the joint analysis of historical data and real-time data. The Spark distributed computing system can be divided into at least one computing driver device, that is, the terminal device in this embodiment, and several schedulers, that is, executors. The scheduler is on each node where the RDD is distributed, that is, the distributed nodes in this embodiment. Connect to the Spark cluster, create RDDs, accumulators, and broadcast variables through SparkContext. The computing driver device will divide the computing task into a series of small shards, that is, tasks, and then send them to the distributed nodes for execution. The distributed nodes can communicate. After each distributed node completes its shard task, all the information will be sent to the computing driver device, and the computing driver device will send the response result to the user terminal. Optionally, if the distributed computing system is a distributed computing system built with the Spark engine, the above computing task can be a task based on the Spark-SQL language.
[0037] In this embodiment, the distributed computing task may include computing content, and the terminal device parses the computing content. Optionally, by determining the computing type corresponding to the computing content and the target object requesting the computation, the target field corresponding to the distributed computing task is determined.
[0038] In S102, the target data of the target field is extracted from the reference query table to generate a field query table.
[0039] In this embodiment, the terminal device stores a reference query table, which records the existing fields of all objects, that is, it belongs to the total data table. And a distributed computing task may only involve some fields in the reference query table. Therefore, in order to reduce the amount of data when sending the data table, it can be divided into sub-data tables based on the reference query table, that is, only extract the target fields required for this calculation and the data of each record associated with the target fields, without sending the entire reference query table to the distributed nodes.
[0040] Exemplarily, the existing fields of the reference query data packet include "user ID", "user age", "user address", "associated user list", and "contact information", and the calculation content of the distributed computing task received by the terminal device is to calculate the average age of users. Then the target fields are "user ID" and "user age". At this time, the terminal device only needs to calculate the average age of users based on the data of "user ID" and "user age" in the reference query table, and generate a field query table according to the target data corresponding to the two fields of "user ID" and "user age", and send the field query table to the distributed nodes, without sending the reference query table composed of all existing fields to the distributed nodes, reducing the data transmission volume by 60%.
[0041] In this embodiment, the reference query table may contain object information of multiple different objects, and each object information contains parameter values of all existing fields in the reference query table. Therefore, when extracting the target data of the target fields, in fact, it is to extract the parameter values of each object regarding the target fields to form the above-mentioned target data.
[0042] In S103, count the field data volume of the field query table.
[0043] In this embodiment, after the terminal device obtains the field query table extracted from the reference query table, it is necessary to determine the field data volume of the field query table. Since the distributed computing system uses different data table merging methods according to the different data volumes of the field query table. Therefore, before sending the field query table to each distributed node, identify the data volumes of the target data corresponding to each target field, and accumulate all the target data volumes, that is, the field data volume of the entire data table can be calculated.
[0044] In this embodiment, the data table merging methods of the distributed computing system can be at least divided into Broadcast Join, Shuffle Hash Join, and Sort Merge Join. Among them, since both Shuffle Hash Join and Sort Merge Join require restructuring the data tables, that is, partitioning the two tables according to the keys corresponding to the data, and then joining the data with the same key value in each partition, so as to merge the data with the same key value in the two data tables, thus realizing the merging of the two data tables. However, the above methods involve a large amount of data transmission between different distributed nodes and occupy the network resources of the network input / output (IO) interface of the distributed nodes. Therefore, merging the data tables through Shuffle Hash Join and Sort Merge Join will reduce the response efficiency of the distributed computing and increase the computing duration. In order to reduce network resource consumption and improve the response efficiency of computing tasks, the Broadcast Join method should be increased to merge the data tables. The trigger condition for the Broadcast Join method is that the data volume of the data table is less than the broadcast trigger threshold. When the data volume of the data table is greater than or equal to the broadcast trigger threshold, the distributed computing system will merge the data table through the two methods of Shuffle Hash Join and Sort Merge Join. By extracting the field query table from the benchmark query table, a large amount of invalid data is filtered, thereby reducing the data volume of the data table and increasing the probability of using the Broadcast Join method.
[0045] In S104, if the field data volume is less than the preset broadcast trigger threshold, broadcast and send the field query table to each distributed node, so that the distributed node merges the field query table with the local target configuration table based on the Broadcast Join method.
[0046] In this embodiment, if it is detected that the field data volume of the field query table is less than the broadcast trigger threshold, the Broadcast Join method can be used to merge the data tables. Therefore, the field query table can be sent to each distributed node. After the distributed node receives the field query table sent by broadcast, it is determined that the current merging method uses the Broadcast Join method to merge the field query table and the local target configuration table.
[0047] Preferably, in this embodiment, the broadcast trigger threshold can be dynamically adjusted according to the current number of tasks. Specifically, if the number of tasks in the current distributed computing task is large, the network resources allocated to each computing task are small. At this time, the broadcast trigger threshold can be increased, thereby increasing the probability of using the broadcast merging method; and in idle time, that is, when the current number of tasks is small, the broadcast trigger threshold can be lowered, and merged through hash switching merge and sort connection merge. In this case, before executing the operation of S104, the terminal device can obtain the number of tasks of the distributed computing task currently being processed, and calculate the broadcast trigger threshold corresponding to the number of tasks through the preset broadcast trigger threshold conversion algorithm, and compare the broadcast trigger threshold with the field data volume.
[0048] In S105, the distributed computing task is executed based on the target configuration table merged by each distributed node.
[0049] In this embodiment, the terminal device can send each field query table to multiple distributed nodes connected to it, and then the distributed nodes can merge the field query table with the local target configuration table. After each distributed node performs the data table merge operation, it returns the merged target configuration table to the terminal device. The terminal device, as a management device for managing multiple distributed nodes, is used to parse the computing tasks and determine the multiple computing operations contained in the multiple distributed computing tasks. The computing operations include extraction of target fields, merging of extracted data, and calculation after data merging. Different computing types are handled by different distributed nodes. Therefore, after obtaining the merged target configuration table, the terminal device can determine the distributed nodes associated with the target fields required by the distributed computing tasks, and send data query tasks to each distributed node. The distributed nodes can feed back the query data to the terminal device, and then the terminal device can send the received data to the distributed nodes that perform data merging and data calculation for subsequent computing tasks, and feed back the calculation results to the terminal device, and respond to the distributed computing tasks through the above process.
[0050] As can be seen from the above, a method for responding to a distributed computing task provided by an embodiment of the present invention parses the distributed computing task, determines the target fields required for the computing operation, extracts the target data related to the target fields from the reference query table, and generates a field query table, thereby splitting a total data table into sub-tables related to the task, eliminating a large amount of invalid data irrelevant to the current calculation, reducing the data volume of the sub-tables, and when the data volume of the fields in the field query table is less than the broadcast trigger threshold, sending the field query table to each distributed storage node by means of broadcast, so that the distributed storage node merges the field query table with the target configuration table based on the broadcast merge (Broadcast Join). Since the merge method of Broadcast Join does not need to associate the field number (Key) value and sort the entire data table, it belongs to a fast merge method of the data table, which improves the efficiency of distributed computing. Compared with the existing distributed computing response technology, this embodiment can filter out the valid data, that is, the target data of the target fields related to the current calculation, when generating the data table, reduce the data volume of the required data table, and when the data volume of the data table is small, merge the field query table and the target configuration table by the simple merge method of Broadcast Join, greatly reducing the number of data sorting operations, reducing the data reading and writing pressure of the distributed nodes, and thus improving the response rate of distributed computing.
[0051] Figure 2 FIG. shows a specific implementation flowchart of a method for responding to a distributed computing task provided by the second embodiment of the present invention. Refer to Figure 2 , relative to Figure 1 the embodiment described above, before the step of if the data volume of the fields is less than a preset broadcast trigger threshold, broadcasting and sending the field query table to each distributed node in a method for responding to a distributed computing task provided by this embodiment, the following steps S201 to S205 are further included, which are specifically described in detail as follows:
[0052] Further, before the step of if the data volume of the fields is less than a preset broadcast trigger threshold, broadcasting and sending the field query table to each distributed node, the following steps are further included:
[0053] In S201, obtain the network resource parameters at the current moment and the historical operation records associated with the preset time period in which the current moment is located.
[0054] In this embodiment, the terminal device can dynamically adjust the broadcast trigger threshold according to the current network conditions, and specifically, it can be determined based on the historical broadcast threshold configured in the past and the current network conditions. Therefore, the terminal device can obtain the network resource parameters and historical operation records at the current moment, and calculate the broadcast trigger threshold corresponding to the current moment based on the two parameters obtained above. Among them, the current moment is specifically the moment when the distributed computing task is received. Multiple network resource parameters can be obtained, and the network resource parameters include but are not limited to: network packet loss rate, network transmission rate, bit error rate, network delay, etc.
[0055] In this embodiment, the terminal device can be pre-divided into different characteristic time periods. The terminal device can determine the characteristic time period in which the current moment falls, and obtain the historical operation records created within this characteristic time period as the historical operation records associated with the current moment.
[0056] In S202, the network resource parameters are imported into a preset threshold factor conversion model to calculate the first threshold factor.
[0057] In this embodiment, the terminal device can be provided with a threshold factor conversion model. The terminal device imports the network resource parameters into this threshold factor conversion model and outputs the first threshold factor corresponding to the network resource parameters at the current moment. Specifically, the larger the value of the network resource parameter, the more network resources are available currently, and at this time, the value of the corresponding first threshold factor is larger, increasing the probability of using the broadcast merging method; conversely, if the value of the network resource parameter is smaller, it means that the available network resources currently are smaller, and the value of the corresponding first threshold factor is smaller, reducing the probability of using the broadcast merging method. This threshold factor conversion model can be a hash function.
[0058] In S203, based on the creation time of each of the historical operation records, the weight value of the historical operation record is configured.
[0059] In this embodiment, the historical operation record includes the creation time of this record. The terminal device can configure the weight value of this historical operation record according to the difference between each creation time and the current moment. Among them, the smaller the difference between the current moment and the creation time, the higher the weight value of the corresponding historical operation record; conversely, the larger the difference between the current moment and the creation time, the lower the weight value of the corresponding historical operation record. Since the smaller the specific difference between the creation time and the current moment, the smaller the difference in the system structure of the distributed computing system and the total amount of data in the database at the moment corresponding to the historical operation record and the system structure and data volume at the current moment, the higher the reference value of the corresponding historical broadcast threshold, and thus the larger the corresponding weight value, which can improve the accuracy of the current broadcast trigger threshold.
[0060] In S204, according to the historical broadcast thresholds of each of the historical operation records and the weight value, a second threshold factor is calculated.
[0061] In this embodiment, the historical operation record includes the historical broadcast threshold compared when responding to a historical calculation task. The terminal device can perform weighted accumulation on the historical broadcast thresholds and weight values of each historical operation record, so as to calculate the second threshold factor.
[0062] In S205, according to the first threshold factor and the second threshold factor, the broadcast trigger threshold is calculated.
[0063] In this embodiment, after the terminal device calculates the first threshold factor related to the network resource parameter tube and the second threshold factor related to the historical broadcast threshold, it can calculate the broadcast trigger threshold corresponding to the current moment based on the above two parameters, so as to achieve the purpose of dynamically adjusting the broadcast trigger threshold.
[0064] In the embodiment of the present invention, by obtaining the current network resource parameter and the historical operation record, the broadcast trigger threshold at the current moment is calculated, so that the broadcast trigger threshold compared with the field data volume at the current moment matches the current network load condition, improving the accuracy of the broadcast trigger threshold.
[0065] Figure 3 FIG. shows a specific implementation flowchart of a response method S202 for a distributed computing task provided by the third embodiment of the present invention. Refer to Figure 3 relative to Figure 2 the embodiment described above, a response method S202 for a distributed computing task provided by this embodiment includes: S2021 to S2022, which are specifically described in detail as follows:
[0066] Further, the importing the network resource parameter into a preset threshold factor conversion model and calculating the first threshold factor includes:
[0067] In S2021, the maximum available resource parameter of the network where it is located is obtained.
[0068] In this embodiment, the terminal device can obtain the current accessed network, that is, the above-mentioned network where it is located, and the maximum available resource parameter, that is, the upper limit value of each network resource parameter. Exemplarily, the network resource parameter includes an uplink rate and a downlink rate, then the maximum available resource parameter includes the highest uplink rate and the highest downlink rate; for network resource parameters such as bit error rate and packet loss rate, they are converted into positive parameters, such as the maximum correct rate of data transmission, that is, the corresponding minimum bit error rate, and the packet sending success rate, that is, the corresponding minimum packet loss rate.
[0069] In S2022, import the maximum available resource parameter and the network resource parameter into a preset threshold factor conversion model to calculate the first threshold factor. The threshold factor conversion model is specifically as follows:
[0070]
[0071] where FirstBrdcst is the first threshold factor; CurrentResource i is the i-th network resource parameter; MaxWebResource i is the i-th maximum available resource parameter; BaseLv is a preset reference coefficient; and n is the total number of network resource parameters.
[0072] In this embodiment, the terminal device can calculate the ratio between the current network resource parameter and the maximum available resource parameter. If the network resource parameter is closer to the maximum available resource parameter, it indicates that the current network environment is better and can be used to transmit the data table of big data. Therefore, the value of the corresponding first threshold factor is also larger. On the contrary, if the difference between the network resource parameter and the maximum available resource parameter is larger, it indicates that the current network environment is worse. At this time, the value of the corresponding first threshold factor is also smaller.
[0073] In the embodiment of the present invention, the terminal device obtains the maximum available resource parameter of the current network where it is located, and compares the network resource parameter with the maximum available resource parameter to calculate the first threshold factor, so as to be able to normalize each network resource parameter and improve the accuracy of the first threshold factor.
[0074] Figure 4 The specific implementation flowchart of a response method S101 for a distributed computing task provided in the fourth embodiment of the present invention is shown. Refer to Figure 4 , relative to Figure 1 the above embodiment, a response method S101 for a distributed computing task provided in this embodiment includes: S1011 to S1014, which are specifically described in detail as follows:
[0075] Further, the determining the target field of the distributed computing task if the distributed computing task is received includes:
[0076] In S1011, parse the distributed computing task to obtain the task type of the distributed computing task.
[0077] In this embodiment, the terminal device can determine the target field associated with the distributed computing task through self-identification without manual setting by the user. Different computing tasks correspond to different task types, and different task types require different data to be called during response. Therefore, the terminal device can parse the distributed computing task and identify the task type corresponding to the computing task. Specifically, the terminal device can extract the computing content of the distributed computing task, extract the computing keywords in the computing content, and determine the task type associated with the computing keywords.
[0078] In S1012, extract the historical response results matching the task type from the computing response database, and identify the historical fields included in each of the historical response results.
[0079] In this embodiment, the computing response database stores all the historical response results of the terminal device. The historical response results include the target fields associated with the historical computing tasks, that is, the above-mentioned historical fields. The terminal device can extract the historical response results matching the task type from the computing response database according to the task type of the distributed computing task, that is, the task type of the extracted historical response results is the same as the task type of the currently required distributed computing task, so it can be determined that the target fields associated in the historical response results may also be the target fields associated with the currently required computing task.
[0080] In S1013, calculate the association degree of each of the historical fields based on the occurrence times and occurrence time of each of the historical fields in all the historical response results.
[0081] In this embodiment, the terminal device can count the occurrence times of each historical field in all the historical response results. If the occurrence times are larger, it means that the historical field has a higher association degree with the task type; conversely, if the occurrence times of the historical field in all the historical response results are smaller, the association degree with the task type is lower. And the terminal device can calculate the occurrence frequency according to the occurrence time of the historical field in each historical response result. Based on the occurrence times and the occurrence frequency, the association degree between each historical field and the task type can be identified.
[0082] In S1014, select the historical fields with the association degree greater than the preset association threshold as the target fields.
[0083] In this embodiment, after calculating the association degree between each historical field and the task type, the terminal device can select the historical fields with the association degree greater than the association threshold as the target fields, so as to achieve the purpose of automatically identifying the target fields of the distributed computing task.
[0084] In an embodiment of the present invention, by obtaining historical response results that match the task type of the current computing task, the target fields corresponding to the computing task are automatically extracted through the historical fields included in the historical response results, reducing user operations and improving the response efficiency of distributed computing tasks.
[0085] Figure 5 FIG. shows a specific implementation flowchart of a response method S101 for a distributed computing task provided in the fifth embodiment of the present invention. Refer to Figure 5 , relative to Figure 1 In the embodiment described above, a response method S101 for a distributed computing task provided in this embodiment includes: S1015 to S1017, which are specifically described in detail as follows:
[0086] Further, the step of determining the target fields of the distributed computing task if the distributed computing task is received includes:
[0087] In S1015, extract the Structured Query Language (SQL) statement included in the distributed computing task.
[0088] In this embodiment, the distributed computing task is a computing task generated based on the Spark-SQL framework. The terminal device can parse the distributed computing task and extract the SQL statement carried by the computing task. Since the SQL statement is used to query the target data in the database, that is, the SQL statement carries the target field information corresponding to the computing task, the SQL statement can be parsed to automatically determine the SQL language.
[0089] In S1016, perform semantic analysis on the SQL statement to obtain the query keywords corresponding to the SQL statement.
[0090] In this embodiment, the terminal device can extract a SQL statement library, which records multiple standard paragraphs. The terminal device can extract the feature statements for defining the query data association from the SQL statement based on the standard paragraphs, and based on the query keywords included in the feature statements.
[0091] In S1017, query the existing fields in the reference query table that match each of the query keywords, and identify the existing fields that match the query keywords as the target fields.
[0092] In this embodiment, the terminal device can identify whether each query keyword exists in the existing fields in the reference query table. If so, identify the existing field as the target field.
[0093] In an embodiment of the present invention, by performing semantic parsing on the SQL statements within a computing task, query keywords are extracted, and target fields that match are extracted from a benchmark query table based on the query keywords, achieving the purpose of automatically identifying target fields, reducing user operations, and improving the response efficiency of distributed computing tasks.
[0094] Figure 6 FIG. shows a specific implementation flowchart of a method for responding to a distributed computing task provided in the sixth embodiment of the present invention. Refer to Figure 6 , relative to Figures 1 to 5 any one of the above embodiments, after statistically calculating the field data volume of the field query table, a method for responding to a distributed computing task provided in this embodiment further includes: S601 - S602, which are specifically described in detail as follows:
[0095] Furthermore, after statistically calculating the field data volume of the field query table, it further includes:
[0096] In S601, if the field data volume is greater than or equal to the broadcast trigger threshold, the data numbers of each target data are added to the field query table.
[0097] In this embodiment, the terminal device can obtain that if it detects that the field data volume of the field query table is greater than or equal to the broadcast trigger threshold, it recognizes that it is necessary to merge two data tables by means of hash-switching merge Shuffle Hash Join, and Shuffle Hash Join needs to determine the data numbers of each target data, that is, the Key values, so that the distributed nodes can determine the associated data based on the Key values. Therefore, the terminal device needs to add the data numbers of each data to the field query table.
[0098] In S602, the field query table added with the data numbers is sent to each distributed node, so that each distributed node queries the associated data corresponding to each field data in the local data of the local target configuration table based on the data numbers, and reconstructs the target configuration table.
[0099] In this embodiment, the terminal device sends the field query table added with data labels to each distributed node. The distributed nodes can compare the key values of each target data with the key values of the local data in the local target configuration table, identify the local data and target data with the same key value as mutually associated data, and reconstruct the target configuration table according to the association relationship between the data, so as to merge the target data of the field query table into the target configuration table.
[0100] In an embodiment of the present invention, in the case of a large amount of data, by using the Shuffle Hash Join method to merge two data tables, the data tables can be sorted out, which is convenient for the management of distributed data.
[0101] Figure 7 FIG. shows a specific implementation flowchart of a response method S102 for a distributed computing task provided in the seventh embodiment of the present invention. Refer to Figure 7 , relative to Figures 1 to 5 any one of the above embodiments, a response method S102 for a distributed computing task provided in this embodiment includes: S1021 to S1022, which are specifically described in detail as follows:
[0102] In S1021, if any reference field in the reference query table matches the target field of the distributed computing task, then the reference field is identified as a valid field.
[0103] In this embodiment, after the terminal device determines the target field of the distributed computing task, it can separate the field query table required for this calculation from the reference query table, that is, separate the small table from the large table. Specifically, the terminal device can determine whether each existing field in the reference query table, that is, the above reference field, matches the target field. If so, the reference field is identified as a valid field.
[0104] In S1022, the data associated with all the valid fields is identified as target data, and a field query table is generated based on all the target data and the valid fields.
[0105] In this embodiment, after the terminal device identifies all the valid fields in the reference query table, the data associated with the valid fields is the target data required to be extracted by the target field, and based on the target data and the valid fields, a field query table is separated from the reference query table.
[0106] In an embodiment of the present invention, by identifying the valid fields in the reference query table and generating a field query table, the amount of data in the data table to be sent can be reduced, and the consumption of network resources can be reduced.
[0107] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0108] Figure 8 FIG. shows a structural block diagram of a response device for a distributed computing task provided in an embodiment of the present invention. Each unit included in the response device for the distributed computing task is used to execute Figure 1 the corresponding steps in the embodiment. For details, please refer to Figure 1 andFigure 1 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown.
[0109] See Figure 8 , the response device for the distributed computing task includes:
[0110] A target field recognition unit 81, configured to determine the target field of the distributed computing task if a distributed computing task is received;
[0111] A field query table generation unit 82, configured to extract the target data of the target field from a reference query table and generate a field query table;
[0112] A field data volume statistics unit 83, configured to count the field data volume of the field query table;
[0113] A broadcast merge trigger unit 84, configured to broadcast and send the field query table to each distributed node if the field data volume is less than a preset broadcast trigger threshold, so that the distributed nodes merge the field query table with a local target configuration table based on a broadcast merge method;
[0114] A distributed computing task response unit 85, configured to execute the distributed computing task based on the target configuration tables merged by each distributed node.
[0115] Optionally, the response device for the distributed computing task further includes:
[0116] A network resource determination unit, configured to obtain network resource parameters at the current moment and historical operation records associated with a preset time period in which the current moment is located;
[0117] A first threshold factor calculation unit, configured to import the network resource parameters into a preset threshold factor conversion model and calculate a first threshold factor;
[0118] A weight value determination unit, configured to configure weight values of the historical operation records based on the creation times of the historical operation records;
[0119] A second threshold factor calculation unit, configured to calculate a second threshold factor according to historical broadcast thresholds of the historical operation records and the weight values;
[0120] A broadcast trigger threshold calculation unit, configured to calculate the broadcast trigger threshold according to the first threshold factor and the second threshold factor.
[0121] Optionally, the first threshold factor calculation unit includes:
[0122] A maximum available resource parameter acquisition unit, configured to acquire the maximum available resource parameters of the network where it is located;
[0123] A first threshold factor conversion unit, configured to import the maximum available resource parameters and the network resource parameters into a preset threshold factor conversion model, and calculate the first threshold factor; the threshold factor conversion model is specifically:
[0124]
[0125] wherein, FirstBrdcst is the first threshold factor; CurrentResource i is the i-th network resource parameter; MaxWebResource i is the i-th maximum available resource parameter; BaseLv is a preset reference coefficient; n is the total number of network resource parameters.
[0126] Optionally, the target field recognition unit 81 includes:
[0127] A task type recognition unit, configured to analyze the distributed computing task to obtain the task type of the distributed computing task;
[0128] A historical field acquisition unit, configured to extract historical response results matching the task type from a calculation response database, and identify historical fields included in each of the historical response results;
[0129] An association degree calculation unit, configured to calculate the association degree of each historical field respectively based on the number of occurrences and occurrence time of each historical field in all the historical response results;
[0130] A target field selection unit, configured to select the historical fields with the association degree greater than a preset association threshold as the target fields.
[0131] Optionally, the target field recognition unit 81 includes:
[0132] An SQL statement extraction unit, configured to extract a structured query language SQL statement included in the distributed computing task;
[0133] A query keyword acquisition unit, configured to perform semantic analysis on the SQL statement to obtain query keywords corresponding to the SQL statement;
[0134] A query keyword screening unit, configured to query existing fields matching each of the query keywords in a reference query table, and identify the existing fields matching the query keywords as the target fields.
[0135] Optionally, the response device for the distributed computing task further includes:
[0136] A field query table adjustment unit, configured to add the data numbers of each target data to the field query table if the amount of field data is greater than or equal to the broadcast trigger threshold;
[0137] A hash sorting and merging unit, configured to send the field query table added with the data numbers to each distributed node, so that each distributed node queries associated data corresponding to each field data in the local data of the local target configuration table based on the data numbers, and reconstructs the target configuration table.
[0138] Optionally, the field query table generation unit 81 includes:
[0139] A valid field identification unit, configured to identify a reference field as a valid field if any reference field in the reference query table matches the target field of the distributed computing task;
[0140] A target data selection unit, configured to identify the data associated with all the valid fields as target data, and generate a field query table according to all the target data and the valid fields.
[0141] Therefore, the response device for the distributed computing task provided by the embodiment of the present invention can also filter out valid data, that is, the target data of the target field related to the current calculation, when generating a data table, reduce the amount of data in the data table to be sent. In the case where the amount of data in the data table is small, the field query table and the target configuration table are merged by a simple merging Broadcast Join method, greatly reducing the number of data sorting operations, reducing the data reading and writing pressure of the distributed nodes, and thus improving the response rate of the distributed computing.
[0142] Figure 9 is a schematic diagram of a terminal device provided by another embodiment of the present invention. As Figure 9 shown, the terminal device 9 of this embodiment includes: a processor 90, a memory 91, and a computer program 92 stored in the memory 91 and executable on the processor 90, such as a response program for a distributed computing task. When the processor 90 executes the computer program 92, the steps in the above-mentioned embodiments of the response method for each distributed computing task are implemented, such as Figure 1 shown S101 to S105. Alternatively, when the processor 90 executes the computer program 92, the functions of each unit in the above-mentioned device embodiments are implemented, such as Figure 8 shown module 81 to 85 functions.
[0143] Exemplarily, the computer program 92 may be divided into one or more units, which are stored in the memory 91 and executed by the processor 90 to implement the present invention. The one or more units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 92 in the terminal device 9. For example, the computer program 92 may be divided into a target field recognition unit, a field query table generation unit, a field data volume statistics unit, a broadcast merge trigger unit, and a distributed computing task response unit, and the specific functions of each unit are as described above.
[0144] The terminal device 9 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 90 and a memory 91. Those skilled in the art can understand that Figure 9 merely examples of the terminal device 9, which do not constitute a limitation on the terminal device 9, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.
[0145] The so-called processor 90 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0146] The memory 91 may be an internal storage unit of the terminal device 9, such as the hard disk or memory of the terminal device 9. The memory 91 may also be an external storage device of the terminal device 9, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 9. Further, the memory 91 may also include both the internal storage unit and the external storage device of the terminal device 9. The memory 91 is used to store the computer program and other programs and data required by the terminal device. The memory 91 may also be used to temporarily store data that has been output or will be output.
[0147] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0148] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A response method for distributed computing tasks, characterized in that Including: If a distributed computing task is received, determine the target field of the distributed computing task; Extract the target data of the target field from the reference query table to generate a field query table; Count the amount of field data in the field query table; If the amount of field data is less than a preset broadcast trigger threshold, broadcast and send the field query table to each distributed node, so that the distributed node merges the field query table with the local target configuration table based on the broadcast merging method; Execute the distributed computing task based on the target configuration table merged by each distributed node; Before the step of if the amount of field data is less than a preset broadcast trigger threshold, broadcast and send the field query table to each distributed node, it further includes: Obtain the network resource parameters at the current moment and the historical operation records associated with the preset time period in which the current moment is located; Import the network resource parameters into a preset threshold factor conversion model to calculate a first threshold factor; Configure the weight value of each historical operation record based on the creation time of each historical operation record; Calculate a second threshold factor according to the historical broadcast threshold of each historical operation record and the weight value; Calculate the broadcast trigger threshold according to the first threshold factor and the second threshold factor.
2. The response method according to claim 1, wherein The step of importing the network resource parameters into a preset threshold factor conversion model to calculate a first threshold factor includes: Obtain the maximum available resource parameters of the network where it is located; Import the maximum available resource parameters and the network resource parameters into a preset threshold factor conversion model to calculate the first threshold factor; the threshold factor conversion model is specifically: wherein, FirstBrdcst is the first threshold factor; CurrentResource i is the i-th network resource parameter; MaxWebResource i is the i-th maximum available resource parameter; BaseLv is a preset reference coefficient; n is the total number of network resource parameters.
3. The response method according to claim 1, characterized in that, The step of if a distributed computing task is received, determine the target field of the distributed computing task includes: Analyze the distributed computing task to obtain the task type of the distributed computing task; Extract the historical response results matching the task type from the computing response database and identify the historical fields included in each historical response result; Calculate the association degree of each historical field respectively based on the number of occurrences and occurrence time of each historical field in all historical response results; Select the historical fields with the association degree greater than a preset association threshold as the target fields.
4. The response method according to claim 1, wherein The step of if a distributed computing task is received, determine the target field of the distributed computing task includes: Extract the Structured Query Language (SQL) statement included in the distributed computing task; Perform semantic analysis on the SQL statement to obtain the query keywords corresponding to the SQL statement; Query the existing fields matching each query keyword in the reference query table, and identify the existing fields matching the query keywords as the target fields.
5. The response method according to any one of claims 1-4, characterized in that, After the step of counting the amount of field data in the field query table, it further includes: If the amount of field data is greater than or equal to the broadcast trigger threshold, add the data numbers of each target data to the field query table; Send the field query table with the added data numbers to each distributed node, so that each distributed node queries the associated data corresponding to each target data in the local data of the local target configuration table based on the data numbers, and reconstructs the target configuration table.
6. The response method according to any one of claims 1-4, characterized in that Extracting the target data of the target field from the reference query table to generate a field query table includes: If any reference field in the reference query table matches the target field of the distributed computing task, identify the reference field as a valid field; Identify the data associated with all the valid fields as target data, and generate a field query table according to all the target data and the valid fields.
7. A response device for distributed computing tasks, characterized in that, Includes: A target field identification unit, configured to determine the target field of the distributed computing task if a distributed computing task is received; A field query table generation unit, configured to extract the target data of the target field from the reference query table to generate a field query table; A field data volume statistics unit, configured to count the field data volume of the field query table; A broadcast merge trigger unit, configured to broadcast and send the field query table to each distributed node if the field data volume is less than a preset broadcast trigger threshold, so that the distributed node merges the field query table with the local target configuration table based on the broadcast merge method; A distributed computing task response unit, configured to execute the distributed computing task based on the target configuration tables merged by each distributed node; The response device of the distributed computing task further includes: A network resource determination unit, configured to obtain the network resource parameters at the current moment and the historical operation records associated with the preset time period in which the current moment is located; A first threshold factor calculation unit, configured to import the network resource parameters into a preset threshold factor conversion model to calculate a first threshold factor; A weight value determination unit, configured to configure the weight value of the historical operation record based on the creation time of each historical operation record; A second threshold factor calculation unit, configured to calculate a second threshold factor according to the historical broadcast threshold of each historical operation record and the weight value; A broadcast trigger threshold calculation unit, configured to calculate the broadcast trigger threshold according to the first threshold factor and the second threshold factor.
8. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.