Method, apparatus, and electronic device for determining optimal path for data reading

By obtaining computing resources and storing data information in the K8S platform, and combining multiple algorithms to determine the optimal path for reading Presto data from different dimensions, the problem of low path accuracy in the existing technology is solved, and more efficient data reading and writing and SQL query efficiency is achieved.

CN116028217BActive Publication Date: 2025-06-13中国邮政储蓄银行股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211625496.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-06-13
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

In the prior art, the efficiency of calculating the optimal path for reading data is low, resulting in a low accuracy of the optimal path, which in turn increases the read and write costs and reduces the query efficiency of SQL statements.

Method used

By obtaining the computing resource information and storage data information in the K8S platform, using the decision tree algorithm, weighted statistical algorithm, k-means algorithm and edge algorithm, the optimal path is determined from the engine dimension and storage dimension, and the worker is assigned to the optimal path to perform tasks.

Benefits of technology

It improves the accuracy of data reading paths, reduces the reading and writing costs of data, and thus improves the query efficiency of SQL statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028217B_ABST
    Figure CN116028217B_ABST
Patent Text Reader

Abstract

The present application provides a method, an apparatus, and an electronic device for determining an optimal path for data reading. The method includes: obtaining computing resource information and stored data information; determining a first path based on the computing resource information and the stored data information, and determining a second path based on the computing resource information; determining an optimal path for reading data according to the first path and the second path, and allocating a worker to the optimal path, where the worker is used to execute tasks using the optimal path. In this solution, by deploying Presto in the K8S platform, computing resource information and stored data information can be obtained, and an optimal path for reading data can be determined according to the computing resource information and the stored data information. The accuracy of the optimal path determined by this solution through various information is relatively high, thereby reducing the data reading and writing costs and improving the query efficiency of SQL statements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of K8S data processing. Specifically, it relates to a method, an apparatus, and an electronic device for determining an optimal path for data reading. Background Art

[0002] In large-scale Presto cluster deployments, since the optimal path for Presto to read data is determined only by the rack location (multiple physical machines are installed in the computer room, and each physical machine has a different installation location, and the rack location is the location of the physical machine), and at the same time, there will be an impact on data reading when a data warehouse tool (such as hive) executes tasks on the resource manager (Yet Another Resource Negotiator, abbreviated as yarn). All of the above will prevent the container (worker) from being assigned the most suitable shard data for reading and calculation. Therefore, in the current solution, due to the low efficiency of calculating the optimal path for Presto to read data, the accuracy rate of the optimal path will be low, which will lead to too high read and write costs and ultimately result in low query efficiency of SQL statements. Summary of the Invention

[0003] The main objective of this application is to provide a method, an apparatus, and an electronic device for determining an optimal path for data reading, so as to solve the problem in the prior art that due to the low efficiency of calculating the optimal path for Presto to read data, the accuracy rate of the optimal path is low, which will lead to too high read and write costs and ultimately result in low query efficiency of SQL statements.

[0004] According to one aspect of an embodiment of the present invention, a method for determining an optimal path for data reading is provided, including: obtaining computing resource information and stored data information, where the computing resource information is information related to the running status of a component for reading data and the data reading volume of a data node storing the data, and the stored data information is information related to the type, backup, and performance of the data, and the performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. Both the computing resource information and the stored data information are information in the K8S platform; determining a first path according to the computing resource information and the stored data information, and determining a second path according to the computing resource information, where the starting point of the first path is the location of the engine, the ending point of the first path is the target location, the starting point of the second path is the location of the data node, and the ending point of the second path is the target location, and the engine is a component for the worker to execute a task application; determining an optimal path for reading data according to the first path and the second path, and allocating the worker to the optimal path, where the worker is used to execute a task using the optimal path.

[0005] Optionally, determining a first path according to the computing resource information and the stored data information, and determining a second path according to the computing resource information includes: using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to process the computing resource information and the stored data information to obtain the first path; using an edge algorithm to process the computing resource information to obtain the second path.

[0006] Optionally, using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to process the computing resource information and the stored data information to obtain the first path includes: inputting the computing resource information and the stored data information into a decision tree for classification to obtain a classification result, determining a target classification result that meets a predetermined condition in the classification result, and determining a plurality of the workers according to the target classification result, where the positions of any two of the workers are different; obtaining a plurality of weight information according to the weighted statistical algorithm, and one piece of weight information corresponds to the importance of one of the workers in data reading; according to the worker and the weight information corresponding to the worker, allocating the engine to the worker closest to the engine, and determining the shortest path among a plurality of paths from the positions of the plurality of engines to the target position as the first path.

[0007] Optionally, using an edge algorithm to process the computing resource information to obtain the second path includes: determining the positions of a plurality of the data nodes according to the computing resource information; determining a plurality of paths from the positions of the plurality of data nodes to the target position; and determining the shortest path among the plurality of paths as the second path.

[0008] Optionally, determining an optimal path for reading data according to the first path and the second path includes: comparing the lengths of the first path and the second path; and determining the path with the shortest length among the first path and the second path as the optimal path.

[0009] Optionally, allocating the worker to the optimal path includes: splitting target data to obtain a plurality of shard data, where the target data is the data at the target location; generating a plurality of shard tasks according to the plurality of shard data, and allocating the plurality of shard tasks to the corresponding worker; controlling each worker to execute the corresponding shard task, where the path adopted by each worker to execute the shard task is an optimal sub-path, and the optimal sub-path refers to the optimal path for the worker to execute the shard task; aggregating the plurality of shard tasks to obtain a target task, and determining a target worker corresponding to the target task, where the path adopted by the target worker to execute the target task is the optimal path.

[0010] Optionally, the method further includes: determining whether the number of tasks executed by the worker is greater than or equal to a quantity threshold; in the case where the number of tasks executed by the worker is greater than or equal to the quantity threshold, adding the worker to a queue; in the case where the number of tasks executed by the worker is less than the quantity threshold, adding a new task to the worker.

[0011] Optionally, before obtaining the computing resource information and the storage data information, the method further includes: determining whether a forced scheduling function is effective; in the case where the forced scheduling function is effective, traversing the locations of all the workers, and in the case where the location of a worker is the same as the target location, adding the worker to the queue; in the case where the forced scheduling function fails, obtaining the computing resource information and the storage data information.

[0012] According to another aspect of the embodiments of the present invention, there is also provided an apparatus for determining an optimal path for data reading, including: a first acquisition unit, configured to acquire computing resource information and stored data information, where the computing resource information is information related to the running state of a component for reading data and the data reading volume of a data node storing the data, and the stored data information is information related to the type, backup, and performance of the data, and the performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data, and both the computing resource information and the stored data information are information in the K8S platform; a first determination unit, configured to determine a first path according to the computing resource information and the stored data information, and determine a second path according to the computing resource information, where the starting point of the first path is the position of the engine, the ending point of the first path is the target position, the starting point of the second path is the position of the data node, the ending point of the second path is the target position, and the engine is a component for the worker to execute a task application; a second determination unit, configured to determine an optimal path for reading data according to the first path and the second path, and allocate the worker to the optimal path, where the worker is used to execute a task by using the optimal path.

[0013] According to yet another aspect of the embodiments of the present invention, there is also provided an electronic device, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods.

[0014] In the embodiments of the present invention, first, computing resource information and stored data information are acquired, then a first path is determined according to the computing resource information and the stored data information, and finally, an optimal path for reading data is determined according to the first path and the second path, and the worker is allocated to the optimal path. In this solution, by deploying Presto in the K8S platform, computing resource information and stored data information can be acquired, and an optimal path for reading data can be determined according to the computing resource information and the stored data information. The accuracy of the optimal path determined by this solution through multiple pieces of information is relatively high, thereby reducing the data reading and writing costs and improving the query efficiency of SQL statements. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0016] Figure 1 A flowchart showing a method for determining an optimal path for data reading according to an embodiment of this application is shown;

[0017] Figure 2 Shows a schematic flow chart for determining the optimal path;

[0018] Figure 3 Shows a schematic flow chart for allocating workers to the optimal path;

[0019] Figure 4 Shows a schematic structural diagram of a device for determining the optimal path of a data reading according to an embodiment of the present application;

[0020] Figure 5 Shows a schematic flow chart of a solution for determining the optimal path of another data reading. Detailed implementation manners

[0021] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0024] It should be understood that when an element (such as a layer, film, region, or substrate) is described as being "on" another element, the element can be directly on the other element, or there may also be an intermediate element. Moreover, in the specification and the claims, when it is described that an element is "connected" to another element, the element can be "directly connected" to the other element, or "connected" to the other element through a third element.

[0025] For ease of description, some nouns or terms related to the embodiments of the present application are described below:

[0026] K8S platform: Kubernetes (abbreviated as K8S) is an open-source version of the large-scale container management technology Borg created by Google. The K8S platform is a container cluster management system, an open-source platform that can achieve functions such as automatic deployment, automatic scaling, and maintenance of container clusters.

[0027] Presto: Presto is a distributed SQL query engine developed by Facebook, a product specifically designed for real-time query computing of big data. Presto is designed to solve problems such as the slow running speed of the MapReduce model of Hive and the inability to directly display HDFS data through BI or Dashboards.

[0028] In some solutions, the high availability of Presto can be achieved, and the data sharding function can also be achieved, but the current solutions cannot accurately calculate the optimal path for reading data; in Presto, the Coordinator will split the stage (tasks in multiple stages) into multiple tasks and submit them to each worker for parallel execution; the input data for each task is one or more splits (sharded data), and a split is a part of the data of a table. For example, a hive table is a file on hdfs (Hadoop Distributed File System, file system). Since the worker needs to read the hdfs file to read the split data, if the split can be exactly allocated to the worker node where the data is located for reading and calculation, a lot of network transmission consumption can be saved, which is beneficial to accelerating the query performance.

[0029] There are two split allocation scheduling methods provided in Presto for selection. One is SimpleNodeSelector, and the other is TopologyAwareNodeSelector based on network topology. The default scheduling method is SimpleNodeSelector. In addition, Presto also provides two scheduling functions, namely node-scheduler.optimized-local-scheduling and hive.force-local-scheduling. When node-scheduler.optimized-local-scheduling is effective, Presto tries to select a worker with a light task on the same node as the split data as much as possible. When hive.force-local-scheduling is effective, Presto will force the task to be executed on the worker on the same node as the split data, otherwise an error will be reported.

[0030] As mentioned in the background art, in the prior art, due to the low efficiency of calculating the optimal path for Presto to read data, the accuracy of the optimal path is low, which will lead to too high read and write costs and ultimately result in low query efficiency of SQL statements. To solve the above problems, in a typical embodiment of the present application, a method, apparatus, and electronic device for determining the optimal path of data reading are provided.

[0031] According to an embodiment of the present application, a method for determining the optimal path of data reading is provided.

[0032] Figure 1 It is a flowchart of the method for determining the optimal path of data reading according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0033] Step S101, obtain computing resource information and stored data information. The above computing resource information is information related to the running state of the component for reading data and the data reading volume of the data node storing the data. The above stored data information is information related to the type, backup, and performance of the data. The performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. The above computing resource information and the above stored data information are both information in the K8S platform;

[0034] Specifically, the computing resource information may include: the running information of the yarn queue, the global information of the K8S platform monitored by Prometheus (a monitor in the K8S platform), the relevant information of the running data nodes obtained through the container api-server for the data nodes storing the labels in etcd (used to store the running status in the K8S platform), the data reading volume of the data nodes where the worker may run, and the hardware configuration information of the data nodes.

[0035] Specifically, the stored data information may include: the metadata information (data type) stored in hive, the location of hdfs, the data backup information (the number of data backups and the backup locations), and the performance of the data obtained through SNMP (a protocol in the K8S platform) (including at least one of the following: the size of the data volume, the concurrency of the data, the throughput of the data, the sharing function of the data).

[0036] In the above step S101, various information in the K8S platform can be obtained. Compared with the prior art method of only obtaining the rack location to calculate the optimal path, this solution can calculate the optimal path based on various information in the K8S platform, which can ensure a relatively high accuracy rate of the optimal path determined by this solution.

[0037] Step S102: Determine a first path according to the above computing resource information and the above stored data information, and determine a second path according to the above computing resource information. Among them, the starting point of the first path is the location of the engine, the end point of the first path is the target location, the starting point of the second path is the location of the data node, the end point of the second path is the target location, and the engine is a component for the worker to execute the task application;

[0038] Specifically, the first path is the optimal path determined from the engine dimension, and the second path is the optimal path determined from the storage dimension. Of course, it is not limited to the above two paths. In actual situations, it can also be the optimal path determined from the dimension of the worker executing the task, and the optimal path of other dimensions can be set according to the actual situation. In the first path, from the location of the engine to the target location, it may pass through multiple data nodes and may also pass through multiple locations for storing data. Of course, the location of the engine and the target location may also be on the same path. In the second path, from the location of the data node to the target location, it may pass through multiple data nodes, and the target location may also be on the same path as the data node.

[0039] In the above step S102, the corresponding optimal paths can be determined respectively from the engine dimension and the storage dimension. Compared with the prior art method of only calculating the optimal path through the rack location, this solution can determine the optimal path according to different dimensions, which can ensure a relatively high accuracy rate of the optimal path determined by this solution.

[0040] Step S103: Determine the optimal path for reading data based on the above first path and the above second path, and allocate the above worker to the above optimal path, where the above worker is used to execute tasks using the above optimal path.

[0041] In the above step S103, when the optimal path for reading data at the target location is determined, the worker is allocated to the optimal path. The worker uses the engine to read the data at the target location using the optimal path. Since the optimal path determined by this solution is relatively accurate, the read / write cost when the worker reads the data at the target location is relatively low.

[0042] In the above method, first obtain the computing resource information and the stored data information, then determine the first path based on the computing resource information and the stored data information, and finally determine the optimal path for reading data based on the first path and the second path, and allocate the worker to the optimal path. In this solution, by deploying Presto in the K8S platform, the computing resource information and the stored data information can be obtained. Based on the computing resource information and the stored data information, the optimal path for reading data can be determined. The accuracy of the optimal path determined by this solution through multiple pieces of information is relatively high, which can further reduce the read / write cost of the data, thereby improving the query efficiency of SQL statements.

[0043] Specifically, the solution of this application determines the optimal path by deploying Presto in the K8S platform and obtaining the cluster global information (including computing resource information and stored data information), and can also control data splitting and task running, realizing the ability to customize resources such as network, storage, and computing on demand for Presto. At the same time, it realizes the optimal computing path of Presto to read remote data, reducing the read / write cost.

[0044] In some solutions, first each component needs to initialize the operating system, download media, configure ports, set parameters, and also handle the pre-dependencies, which is time-consuming and laborious and cannot guarantee consistency in different environments. However, by adopting the solution of this application, only the Charts package needs to be maintained, avoiding manual errors and cumbersome operations.

[0045] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0046] In order to further accurately determine the first path and the second path to ensure that the optimal path can be more accurately determined based on the first path and the second path subsequently, in an embodiment of the present application, the first path is determined according to the above-mentioned computing resource information and the above-mentioned stored data information, and the second path is determined according to the above-mentioned computing resource information. The specific steps are as follows:

[0047] Step S1021, using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm, process the above-mentioned computing resource information and the above-mentioned stored data information to obtain the above-mentioned first path;

[0048] In a specific embodiment of the present application, using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to process the above-mentioned computing resource information and the above-mentioned stored data information to obtain the above-mentioned first path includes: inputting the above-mentioned computing resource information and the above-mentioned stored data information into a decision tree for classification to obtain a classification result, determining a target classification result that meets a predetermined condition in the above-mentioned classification result, and determining a plurality of the above-mentioned workers according to the above-mentioned target classification result, where the positions of any two of the above-mentioned workers are different; obtaining a plurality of weight information according to the above-mentioned weighted statistical algorithm, and one of the above-mentioned weight information corresponds to the importance of one of the above-mentioned workers in data reading; according to the above-mentioned worker and the above-mentioned weight information corresponding to the above-mentioned worker, allocate the above-mentioned engine to the above-mentioned worker closest to the above-mentioned engine, and determine that the shortest path among the plurality of paths from the positions of the plurality of the above-mentioned engines to the above-mentioned target position is the above-mentioned first path. In this embodiment, the decision tree algorithm can be used to first determine the indicators that need to be concerned, and then determine the weight information of the selected indicators. For example, for the importance of the rack position in data reading and the CPU usage in data reading, the engines are allocated according to the weight information of different indicators, so that a more accurate first path in the engine dimension can be obtained.

[0049] Optionally, the decision tree algorithm can be used to select the corresponding classification result. For example, the data is divided into 20% on the same rack, 10% on the same pod, 15% on the same node, 25% under the same CPU usage threshold, 15% under the same memory, and 15% under the same HDFS path, and the important indicators that need to be concerned are obtained, such as the rack position, pod, node, storage position, and HDFS path; then, through the weighted statistical algorithm, weight information is given to the indicators that need to be focused on, and then through the k-means algorithm, the engines can be allocated in combination with historical data. For example, 40% of the engines are on different pods of the same rack, 15% of the engines are on different nodes of the same rack, 5% of the engines are under different CPU usage thresholds of the same rack, 20% of the engines are on different storages of the same rack, and 20% of the engines are on different HDFS paths of the same rack.

[0050] Step S1022: Process the above computing resource information using an edge algorithm to obtain the above second path.

[0051] In a specific embodiment of the present application, processing the above computing resource information using an edge algorithm to obtain the above second path includes: determining the positions of multiple above data nodes according to the above computing resource information; determining multiple paths from the positions of multiple above data nodes to the above target position; and determining the shortest path among multiple above paths as the above second path. In this embodiment, the second path can be determined from the storage dimension through an edge algorithm to obtain the optimal path in the storage dimension, thereby ensuring a relatively high accuracy of the optimal path determined by the subsequent second path.

[0052] Optionally, an edge algorithm is used to determine the optimal path for reading data from the storage dimension. For example, the data is in the same rack and the same pod, the data is in the same rack and the same node, the data is in the same rack and the same storage, or the data is in the same rack and the same HDFS path.

[0053] In the above steps S1021 to S1022, a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm can be used to determine the first path, and an edge algorithm can be used to determine the second path. By using different algorithms to determine the optimal path from different dimensions, the accuracy of the optimal path calculated by this solution is further ensured to be relatively high.

[0054] In the case where the first path in the engine dimension and the second path in the storage dimension have been determined, a selection can be made again according to the first path and the second path, and the shortest path among the first path and the second path is used as the optimal path. In this way, the accuracy of the optimal path determined from multiple dimensions is further ensured to be relatively high. In another embodiment of the present application, according to the above first path and the above second path, the optimal path for reading data is determined, which specifically includes the following steps:

[0055] Step S1031: Compare the length of the above first path and the length of the above second path;

[0056] Step S1032: Determine the path with the shortest length among the above first path and the above second path as the above optimal path.

[0057] Specifically, a shortest path algorithm can be used to select the optimal path from the first path and the second path. Subsequently, the optimal path can be used to execute tasks, and each task can extract data content.

[0058] In an alternative solution, the process of determining the optimal path is as Figure 2 shown and includes the following steps:

[0059] S1: Obtain the optimal path conversion data, where the data sources include:

[0060] 1. Computing resource information:

[0061] Yarn: Includes field information such as memory total (total cluster memory), MemAvail (available memory per single node), and active nodes (number of active slave nodes in the cluster).

[0062] Prometheus: Includes field information such as cpu_used (CPU usage rate) and desk_io (disk I / O).

[0063] ETCD: Includes field information such as lables (label information in the cluster), pods (pod information in the cluster), and nodes (node information in the cluster).

[0064] Presto: Includes field information such as running queries (number of queries currently running in the cluster) and queued queries (number of queries currently queued and waiting in the cluster).

[0065] 2. Storage data information:

[0066] Hive: Includes field information such as cd_id (field information ID), location (HDFS path), and table_size (table size).

[0067] HDFS: Includes field information such as dir (physical directory of the file), location (physical path), filenu (number of files), dir_bak (backup directory), and size (file size).

[0068] SMNP: Includes field information such as SysLocation (shared storage location), SysName (shared storage name), CPU usage rate, and diskio (read / write speed).

[0069] S2: Use decision tree algorithm, weighted statistical algorithm, k-means algorithm, and edge algorithm to obtain the optimal path, and slice the data reading task to obtain multiple split tasks;

[0070] S3: Customize the task according to the information such as the optimal path obtained in S2 after the split task, and determine the priority order of the tasks;

[0071] S4: Add the customized task to the Presto execution candidate queue.

[0072] By expanding the task scheduling method for data reading in Presto, the performance metrics of the entire K8S platform can be adopted, and tasks can be specified in combination with a performance evaluation algorithm, so that the worker can execute tasks using the optimal path. In another embodiment of the present application, the above worker is assigned to the above optimal path, which specifically includes the following steps:

[0073] Step S1033: Split the target data to obtain multiple shard data, where the target data is the data at the above target position;

[0074] Step S1034: Generate multiple shard tasks based on the multiple above shard data, and assign the multiple above shard tasks to the corresponding above worker;

[0075] Step S1035: Control each of the above workers to execute the corresponding above shard task. Among them, the path used by each of the above workers to execute the above shard task is the optimal sub-path, and the optimal sub-path refers to the optimal path for the above worker to execute the above shard task;

[0076] Step S1036: Aggregate the multiple above shard tasks to obtain a target task, and determine the target worker corresponding to the above target task. Among them, the path used by the above target worker to execute the above target task is the optimal path.

[0077] In the above steps S1033 to S1036, data can be sharded, shard tasks can be generated according to the shard data, the path for each worker to execute the shard task is the optimal sub-path, and then the shard tasks can be aggregated and executed. The target task can be executed by the finally determined target worker, thereby ensuring that the path used by the worker in this solution to execute the task and read data is the optimal path, ensuring that the execution efficiency of the worker in this solution is relatively high, and further improving the query efficiency of the SQL statement.

[0078] In an optional solution, the process of assigning the worker to the optimal path is as Figure 3 shown, including the following steps:

[0079] The client jdbc (client) calls the Coordinator service of Presto to perform an SQL query;

[0080] S1: The Coordinator service of Presto analyzes the SQL statement at the front end, generates a task plan and performs task scheduling. The specific process is to analyze the SOL statement using Parser&Analyzer, generate a task plan using Planner, and schedule tasks using Scheduler to control the worker;

[0081] S2: The worker controlled by the Coordinator obtains data information from HDFS, PG, Kafka, ES, etc., and after algorithms such as weighted statistics, edge computing, shortest path, k-means, decision tree, etc. according to calculation metrics, the optimal path is obtained. The specific process is to obtain data sources such as HDFS, PG, Kafka, ES, etc., and get sharded data through connecting to the data source plugin Connector Plugin. The calculation metrics include: hive mate, Prometheus, Yarn, api-server, snmp, etc., and the optimal path is obtained through algorithms such as weighted statistics, edge computing, shortest path, k-means, decision tree, etc.;

[0082] S3: Customize the worker after splitting according to the information obtained above and the calculated optimal path;

[0083] S4: The worker submits split tasks according to the number of task tasks;

[0084] S5: The workers after multiple split tasks are aggregated and executed;

[0085] S6: After the worker executes, it returns the execution result to the Coordinator, and the Coordinator returns the final data (SQL query return value) to the client;

[0086] Figure 2 The process of Figure 3 is a detailed description of step S2 in

[0087] In the process of the worker actually executing tasks, in fact, the number of tasks that some workers can execute can be dynamically adjusted. In the case where the number of tasks that a worker can execute is less than the number threshold, new tasks can be added to the worker, so as to ensure that the tasks executed by the worker are relatively saturated. In the case where the number of tasks that a worker can execute is greater than or equal to the number threshold, the worker can be added to the queue, so that the worker can wait for tasks to be executed, so as to ensure that the tasks assigned to each worker are relatively balanced and avoid the situation that some workers are assigned too many tasks and cause the worker to fail. In another embodiment of the present application, the above method further includes the following steps:

[0088] Step S104, determine whether the number of the above tasks executed by the above worker is greater than or equal to the number threshold;

[0089] Specifically, the quantity threshold can be 50, 100, 200, etc. Of course, it is not limited to these cases, and those skilled in the art can also select an appropriate quantity threshold according to the actual situation.

[0090] Step S105, when the quantity of the tasks executed by the above worker is greater than or equal to the above quantity threshold, add the above worker to the queue;

[0091] Step S106, when the quantity of the tasks executed by the above worker is less than the above quantity threshold, add a new task to the above worker.

[0092] In an alternative embodiment of the present application, before obtaining the computing resource information and the stored data information, the above method further includes:

[0093] Step S107, determine whether the forced scheduling function is effective;

[0094] Step S108, when the above forced scheduling function is effective, traverse the positions of all the above workers, and when there is a position of the above worker that is the same as the above target position, add the above worker to the queue;

[0095] Step S109, when the above forced scheduling function fails, obtain the above computing resource information and the above stored data information.

[0096] In the above steps S107 to S109, if the forced scheduling function is effective, find the worker with the same position as the target position by traversing. If the forced scheduling function fails, the algorithm for determining the optimal path in this solution can be used to calculate the optimal path, and let the worker run and execute tasks on the optimal path. In this way, different methods can be used to let the worker execute tasks, which further ensures that the execution efficiency of the worker in this solution is relatively high.

[0097] The embodiment of the present application further provides a device for determining the optimal path of data reading. It should be noted that the device for determining the optimal path of data reading in the embodiment of the present application can be used to execute the method for determining the optimal path of data reading provided by the embodiment of the present application. The following introduces the device for determining the optimal path of data reading provided by the embodiment of the present application.

[0098] Figure 4 is a schematic diagram of the device for determining the optimal path of data reading according to the embodiment of the present application. As Figure 4 shown, the device includes:

[0099] The first acquisition unit 10 is configured to acquire computing resource information and stored data information. The computing resource information is information related to the running status of the component for reading data and the data read volume of the data node storing the data. The stored data information is information related to the type, backup, and performance of the data. The performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. Both the computing resource information and the stored data information are information in the K8S platform;

[0100] The above-mentioned first acquisition unit can acquire various information in the K8S platform. Compared with the prior art method of only acquiring the rack position to calculate the optimal path, this solution can calculate the optimal path based on various information in the K8S platform, which can ensure a relatively high accuracy of the optimal path determined by this solution.

[0101] The first determination unit 20 is configured to determine a first path according to the computing resource information and the stored data information, and determine a second path according to the computing resource information. Wherein, the starting point of the first path is the position of the engine, the end point of the first path is the target position, the starting point of the second path is the position of the data node, the end point of the second path is the target position, and the engine is a component for the worker to execute the task application;

[0102] The above-mentioned first determination unit can determine the corresponding optimal paths from the engine dimension and the storage dimension respectively. Compared with the prior art method of only calculating the optimal path through the rack position, this solution can determine the optimal path according to different dimensions, which can ensure a relatively high accuracy of the optimal path determined by this solution.

[0103] The second determination unit 30 is configured to determine the optimal path for reading data according to the first path and the second path, and allocate the worker to the optimal path, where the worker is used to execute the task using the optimal path.

[0104] The above-mentioned second determination unit, when determining the optimal path for reading the data at the target position, allocates the worker to the optimal path. The worker uses the engine to read the data at the target position using the optimal path. Since the optimal path determined by this solution is relatively accurate, the read and write cost when the worker reads the data at the target position is relatively low.

[0105] In the above-mentioned device, the first acquisition unit acquires computing resource information and stored data information. The first determination unit determines a first path according to the computing resource information and the stored data information. The second determination unit determines an optimal path for reading data according to the first path and the second path, and allocates workers to the optimal path. In this solution, by deploying Presto in the K8S platform, computing resource information and stored data information can be acquired. According to the computing resource information and the stored data information, the optimal path for reading data can be determined. The accuracy of the optimal path determined by this solution through multiple pieces of information is relatively high, thereby reducing the read and write costs of data and improving the query efficiency of SQL statements.

[0106] In order to further accurately determine the first path and the second path to ensure that the optimal path can be more accurately determined according to the first path and the second path subsequently, in an embodiment of the present application, the second determination unit includes a first processing module and a second processing module. The functions of the first processing module and the second processing module are as follows:

[0107] The first processing module is used to process the above-mentioned computing resource information and the above-mentioned stored data information by using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to obtain the above-mentioned first path;

[0108] In a specific embodiment of the present application, the first processing module includes a first processing sub-module, an acquisition sub-module, and a second processing sub-module. The first processing sub-module is used to input the above-mentioned computing resource information and the above-mentioned stored data information into a decision tree for classification, obtain a classification result, determine a target classification result that meets a predetermined condition in the classification result, and determine multiple above-mentioned workers according to the above-mentioned target classification result, where the positions of any two of the above-mentioned workers are different; the acquisition sub-module is used to obtain multiple weight information according to the above-mentioned weighted statistical algorithm, and one above-mentioned weight information corresponds to the importance of one above-mentioned worker in data reading; the second processing sub-module is used to allocate the above-mentioned engine to the above-mentioned worker closest to the above-mentioned engine according to the above-mentioned worker and the above-mentioned weight information corresponding to the above-mentioned worker, and determine that the shortest path among multiple paths from the positions of multiple above-mentioned engines to the above-mentioned target position is the above-mentioned first path. In this embodiment, the decision tree algorithm can be used to first determine the indicators that need to be concerned, and then determine the weight information of the selected indicators. For example, for the importance of the rack position in data reading and the CPU usage in data reading, the engine is allocated according to the weight information of different indicators, so that a relatively accurate first path in terms of the engine dimension can be obtained.

[0109] The second processing module is used to process the above-mentioned computing resource information by using an edge algorithm to obtain the above-mentioned second path.

[0110] In a specific embodiment of the present application, the second processing module includes a first determination sub-module, a second determination sub-module, and a third determination sub-module. The first determination sub-module is used to determine the positions of multiple said data nodes according to the above-mentioned computing resource information; the second determination sub-module is used to determine multiple paths from the positions of multiple said data nodes to the above-mentioned target position; the third determination sub-module is used to determine that the shortest path among multiple said paths is the above-mentioned second path. In this embodiment, the second path can be determined from the storage dimension through an edge algorithm to obtain the optimal path in the storage dimension, thereby ensuring a relatively high accuracy rate of the optimal path determined by the subsequent second path.

[0111] For the above-mentioned first processing module and second processing module, three algorithms, namely, a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm, can be used to determine the first path, and an edge algorithm can be used to determine the second path. By using different algorithms to determine the optimal path from different dimensions, it further ensures a relatively high accuracy rate of the optimal path calculated by this solution.

[0112] In the case where the first path in the engine dimension and the second path in the storage dimension have been determined, a selection can be made again according to the first path and the second path, and the shortest path among the first path and the second path is used as the optimal path. In this way, it further ensures a relatively high accuracy rate of the optimal path determined from multiple dimensions. In another embodiment of the present application, the second determination unit includes a comparison module and a determination module, and the functions of the comparison module and the determination module are as follows:

[0113] The comparison module is used to compare the length of the above-mentioned first path and the length of the above-mentioned second path;

[0114] The determination module is used to determine that the path with the shortest length among the above-mentioned first path and the above-mentioned second path is the above-mentioned optimal path.

[0115] By expanding the task scheduling method for data reading in Presto, the performance metrics of the entire K8S platform can be adopted, and tasks can be specified in combination with a performance evaluation algorithm, so that the worker can execute tasks using the optimal path. In another embodiment of the present application, the second determination unit includes a splitting module, a generating module, a control module, and a third processing module, and the functions of the splitting module, the generating module, the control module, and the third processing module are as follows:

[0116] The splitting module is used to split the target data to obtain multiple shard data, and the above-mentioned target data is the data at the above-mentioned target position;

[0117] The generating module is used to generate multiple shard tasks according to multiple said shard data and allocate multiple said shard tasks to the corresponding said worker;

[0118] A control module, configured to control each of the above-mentioned workers to execute the corresponding above-mentioned sharding task. Among them, the path adopted by each of the above-mentioned workers to execute the above-mentioned sharding task is the optimal sub-path, and the above-mentioned optimal sub-path refers to the optimal path for the above-mentioned worker to execute the above-mentioned sharding task.

[0119] A third processing module, configured to aggregate a plurality of the above-mentioned sharding tasks to obtain a target task, and determine a target worker corresponding to the above-mentioned target task. Among them, the path adopted by the above-mentioned target worker to execute the above-mentioned target task is the optimal path.

[0120] The above-mentioned splitting module, generating module, control module and third processing module can shard data, generate sharding tasks according to the sharded data, the path for each worker to execute the sharding task is the optimal sub-path, and then aggregate and execute the sharding tasks. The target task can be executed by the finally determined target worker, thereby ensuring that the path adopted by the worker in this solution to execute the task and read data is the optimal path, ensuring that the execution efficiency of the worker in this solution is relatively high, and further improving the query efficiency of the SQL statement.

[0121] During the actual execution of tasks by workers, in fact, the number of tasks that some workers can execute can be dynamically adjusted. When the number of tasks that a worker can execute is less than the quantity threshold, new tasks can be added to the worker, so as to ensure that the tasks executed by the worker are relatively saturated. When the number of tasks that a worker can execute is greater than or equal to the quantity threshold, the worker can be added to the queue, so that the worker can wait to execute tasks, which can ensure that the tasks assigned to each worker are relatively balanced and avoid the situation that some workers are assigned too many tasks and cause the worker to fail. In another embodiment of the present application, the above-mentioned device further includes a third determination unit, a first addition unit and a second addition unit, and the functions of the third determination unit, the first addition unit and the second addition unit are as follows:

[0122] A third determination unit, configured to determine whether the number of the above-mentioned tasks executed by the above-mentioned worker is greater than or equal to the quantity threshold;

[0123] A first addition unit, configured to add the above-mentioned worker to the queue when the number of the above-mentioned tasks executed by the above-mentioned worker is greater than or equal to the above-mentioned quantity threshold;

[0124] A second addition unit, configured to add new tasks to the above-mentioned worker when the number of the above-mentioned tasks executed by the above-mentioned worker is less than the above-mentioned quantity threshold.

[0125] In an alternative embodiment of the present application, the above-mentioned device further includes a fourth determination unit, a third addition unit, and a second acquisition unit. The functions of the fourth determination unit, the third addition unit, and the second acquisition unit are as follows:

[0126] The fourth determination unit is configured to determine whether the forced scheduling function is effective before acquiring the computing resource information and the stored data information;

[0127] The third addition unit is configured to, when the above-mentioned forced scheduling function is effective, traverse the positions of all the above-mentioned workers, and add the above-mentioned worker to the queue when the position of the above-mentioned worker is the same as the above-mentioned target position;

[0128] The second acquisition unit is configured to, when the above-mentioned forced scheduling function fails, acquire the above-mentioned computing resource information and the above-mentioned stored data information.

[0129] For the above-mentioned fourth determination unit, third addition unit, and second acquisition unit, if the forced scheduling function is effective, the worker with the same position as the target position is found through traversal. If the forced scheduling function fails, the algorithm for determining the optimal path in this solution can be used to calculate the optimal path, so that the worker runs and executes tasks on the optimal path. In this way, different methods can be used to make the worker execute tasks, further ensuring that the execution efficiency of the worker in this solution is relatively high.

[0130] The above-mentioned device for determining the optimal path of data reading includes a processor and a memory. The above-mentioned first acquisition unit, first determination unit, second determination unit, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above-mentioned program units stored in the memory.

[0131] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem in the prior art that due to the low efficiency of calculating the optimal path for Presto to read data, the accuracy rate of the optimal path is relatively low, which will lead to too high reading and writing costs and ultimately low query efficiency of SQL statements can be solved.

[0132] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0133] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the method for determining the optimal path of data reading as described above is implemented.

[0134] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, it executes a method for determining an optimal path for data reading.

[0135] This application also provides an electronic device, including one or more processors, a memory, and one or more programs. Among them, the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include those for executing any one of the above methods.

[0136] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:

[0137] Step S101: Obtain computing resource information and stored data information. The computing resource information is information related to the running status of the component for reading data and the data reading volume of the data node storing the data. The stored data information is information related to the type, backup, and performance of the data. The performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. Both the computing resource information and the stored data information are information in the K8S platform;

[0138] Step S102: Determine a first path based on the computing resource information and the stored data information, and determine a second path based on the computing resource information. Among them, the starting point of the first path is the location of the engine, the ending point of the first path is the target location, the starting point of the second path is the location of the data node, the ending point of the second path is the target location, and the engine is a component for the worker to execute a task application;

[0139] Step S103: Determine an optimal path for reading data based on the first path and the second path, and allocate the worker to the optimal path, where the worker is used to execute a task using the optimal path.

[0140] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0141] This application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:

[0142] Step S101: Obtain computing resource information and stored data information. The computing resource information is information related to the running status of the component for reading data and the data reading volume of the data node storing the data. The stored data information is information related to the type, backup, and performance of the data. The performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. Both the computing resource information and the stored data information are information in the K8S platform;

[0143] Step S102: Determine the first path according to the computing resource information and the stored data information, and determine the second path according to the computing resource information. Among them, the starting point of the first path is the location of the engine, the end point of the first path is the target location, the starting point of the second path is the location of the data node, the end point of the second path is the target location, and the engine is the component for the worker to execute the task application;

[0144] Step S103: Determine the optimal path for reading data according to the first path and the second path, and allocate the worker to the optimal path, where the worker is used to execute the task using the optimal path.

[0145] To enable those skilled in the art to more clearly understand the technical solution of the present application, the technical solution and technical effects of the present application will be described below in conjunction with specific embodiments.

[0146] Embodiment

[0147] This embodiment relates to a method for determining the optimal path for data reading. Through cluster deployment by K8S, two implemented paths are as follows: Path 1: If forced scheduling is enabled, traverse all hdfs addresses of splits, convert the task priorities, and finally generate the task to be executed and send it to the Worker service; Path 2: If non-forced scheduling is enabled, create a computing cluster at the data location to implement the principle of obtaining data nearby. As Figure 5 shown, the method includes:

[0148] S1: Determine whether the forced scheduling function is effective.

[0149] If the forced scheduling function is effective, traverse all hdfs addresses of splits. If the address of a certain worker is the same as the address where the split is located, add the worker to the candidate queue.

[0150] S2: If the forced scheduling function fails, start the best location conversion process, where the information obtained includes two aspects: obtaining computing resource information and obtaining stored data information.

[0151] S3: The computing resource information is respectively sourced from the running information of the yarn queue, the global information monitored by Prometheus, the label and host information stored in etcd obtained through the container api-server, and the host resource usage information where the worker may run.

[0152] S4: The stored data information is respectively sourced from the metadata information stored in hive, the hdfs location and backup count information, and the shared storage performance metric information obtained through SNMP.

[0153] S5: After obtaining the above-mentioned information data, the optimal path is calculated.

[0154] S6: Allocate the worker to the optimal path, judge the task priority, and add the task to the candidate worker queue for waiting to be executed.

[0155] S7: After completing step S2 or S6, judge whether the number of split tasks running by the candidate worker split has not exceeded the upper limit threshold.

[0156] S8: If it does not exceed the upper limit threshold, submit the split task to the worker.

[0157] S9: If it exceeds the upper limit threshold, add it to the split queue and wait for subsequent scheduling.

[0158] In the above embodiments of the present invention, the descriptions of the various embodiments each have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0159] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0160] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0162] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0163] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0164] 1) The method for determining the optimal path for data reading in the present application first obtains computing resource information and stored data information, then determines the first path according to the computing resource information and the stored data information, and finally determines the optimal path for reading data according to the first path and the second path, and allocates workers to the optimal path. In this solution, by deploying Presto in the K8S platform, computing resource information and stored data information can be obtained, and the optimal path for reading data can be determined according to the computing resource information and the stored data information. The accuracy rate of the optimal path determined by this solution through various information is relatively high, and thus the read and write costs of data can be reduced, thereby improving the query efficiency of SQL statements.

[0165] 2) The apparatus for determining the optimal path for data reading in this application. The first acquisition unit acquires computing resource information and stored data information. The first determination unit determines the first path based on the computing resource information and the stored data information. The second determination unit determines the optimal path for reading data according to the first path and the second path, and allocates workers to the optimal path. In this solution, by deploying Presto in the K8S platform, computing resource information and stored data information can be acquired. According to the computing resource information and the stored data information, the optimal path for reading data can be determined. The accuracy of the optimal path determined by this solution through multiple types of information is relatively high, thereby reducing the data reading and writing costs and improving the query efficiency of SQL statements.

[0166] The foregoing are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A method for determining an optimal path for data reading, characterized in that, it includes: Obtain computing resource information and stored data information. The computing resource information is information related to the running status of the component for reading data and the data reading volume of the data node storing the data. The stored data information is information related to the type, backup, and performance of the data. The performance includes at least one of the following: the size of the data volume, the concurrency of the data, and the throughput of the data. Both the computing resource information and the stored data information are information in the K8S platform; Determine a first path according to the computing resource information and the stored data information, and determine a second path according to the computing resource information. Wherein, the starting point of the first path is the position of the engine, the ending point of the first path is the target position, the starting point of the second path is the position of the data node, the ending point of the second path is the target position, and the engine is a component for the worker to execute a task application; Determine the optimal path for reading data according to the first path and the second path, and allocate the worker to the optimal path, where the worker is used to execute the task using the optimal path; Determine a first path according to the computing resource information and the stored data information, and determine a second path according to the computing resource information, including: using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to process the computing resource information and the stored data information to obtain the first path; using an edge algorithm to process the computing resource information to obtain the second path.

2. The method according to claim 1, characterized in that, using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to process the computing resource information and the stored data information to obtain the first path, including: Input the computing resource information and the stored data information into a decision tree for classification to obtain a classification result, determine the target classification result that meets a predetermined condition in the classification result, and determine multiple workers according to the target classification result, where the positions of any two workers are different; Obtain multiple weight information according to the weighted statistical algorithm, and one weight information corresponds to the importance of one worker in data reading; According to the worker and the weight information corresponding to the worker, allocate the engine to the worker closest to the engine, and determine the shortest path among the multiple paths from the positions of the multiple engines to the target position as the first path.

3. The method according to claim 1, characterized in that, using an edge algorithm to process the computing resource information to obtain the second path, including: Determine the positions of multiple data nodes according to the computing resource information; Determine multiple paths from the positions of multiple data nodes to the target position; Determine the shortest path among the multiple paths as the second path.

4. The method according to claim 1, characterized in that, Determine an optimal path for reading data according to the first path and the second path, including: Compare the length of the first path and the length of the second path; Determine the path with the shortest length among the first path and the second path as the optimal path.

5. The method according to claim 1, wherein, Assign the worker to the optimal path, including: Split the target data to obtain a plurality of shard data, where the target data is the data at the target position; Generate a plurality of shard tasks according to the plurality of shard data, and assign the plurality of shard tasks to the corresponding workers; Control each worker to execute the corresponding shard task, where the path adopted by each worker to execute the shard task is an optimal sub-path, and the optimal sub-path refers to the optimal path for the worker to execute the shard task; Aggregate the plurality of shard tasks to obtain a target task, determine the target worker corresponding to the target task, where the path adopted by the target worker to execute the target task is the optimal path.

6. The method according to claim 1, wherein, The method further includes: Determine whether the number of tasks executed by the worker is greater than or equal to a quantity threshold; In the case where the number of tasks executed by the worker is greater than or equal to the quantity threshold, add the worker to the queue; In the case where the number of tasks executed by the worker is less than the quantity threshold, add a new task to the worker.

7. The method according to claim 1, wherein, Before obtaining the computing resource information and the stored data information, the method further includes: Determine whether the forced scheduling function is effective; In the case where the forced scheduling function is effective, traverse the positions of all the workers, and in the case where the position of a worker is the same as the target position, add the worker to the queue; In the case where the forced scheduling function fails, obtain the computing resource information and the stored data information.

8. An apparatus for determining an optimal path for data reading, wherein, Comprising: A first acquisition unit, configured to acquire computing resource information and stored data information, where the computing resource information is information related to the running state of a component for reading data and the data reading volume of a data node storing data, and the stored data information is information related to the type, backup, and performance of the data, and the performance includes at least one of the following: the size of the data volume, the concurrency of the data, the throughput of the data, and both the computing resource information and the stored data information are information in the K8S platform; A first determination unit, configured to determine a first path according to the computing resource information and the stored data information, and determine a second path according to the computing resource information, where a starting point of the first path is a position of an engine, an ending point of the first path is a target position, a starting point of the second path is a position of the data node, an ending point of the second path is the target position, and the engine is a component for a worker to execute a task application; A second determination unit, configured to determine an optimal path for reading data according to the first path and the second path, and allocate the worker to the optimal path, where the worker is configured to execute a task by using the optimal path; The second determination unit includes a first processing module and a second processing module. The first processing module is configured to process the computing resource information and the stored data information by using a decision tree algorithm, a weighted statistical algorithm, and a k-means algorithm to obtain the first path; the second processing module is configured to process the computing resource information by using an edge algorithm to obtain the second path.

9. An electronic device, characterized in that, it includes: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.

Citation Information

Patent Citations

  • Software and hardware hybrid deployment MySQL cluster scheduling method and device

    CN112905337A

  • Data capture engine development method and execution method, equipment and storage medium

    CN115408595A