Spark offline task resource scheduling optimization method based on input data volume
Through preset resource rule list and large language model to assist in identifying computing resource rules, Spark offline task resources are dynamically scheduled, which solves the problem of low resource scheduling efficiency in the existing technology, and realizes efficient computing resource allocation and task operation.
Patent Information
- Application Number
- CN202411666798.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing Spark offline task resource scheduling lacks dynamic configuration solutions, resulting in low resource scheduling efficiency and low task operation completion rate, which cannot effectively meet customers' needs for computing resources.
By setting the list of resource rules, the calculation resource rules are matched according to the number of rows of the Spark offline task's data table, and the task is dynamically scheduled to be executed to the corresponding execution node. The method includes collecting and parsing task data, matching resource rules, and activating the execution node through the node manager.
It realizes dynamic configuration of resource parameters based on the data volume input to the data table of Spark task, optimizes the allocation of computing resources, improves computing efficiency, promotes the efficient operation of Spark offline tasks, and meets customers' needs for computing resources.
Smart Images

Figure CN119201474B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of task management, and in particular to a Spark offline task resource scheduling optimization method, system and electronic device in combination with input data volume. Background Art
[0002] During the running of Spark offline tasks, if the computing resources (driver-memory, driver-cores, num-executors, executor-memory, executor-cores) of a single task are not specially specified, the driver resources will be set according to the default configuration, and the executor resources will be dynamically specified according to the actual input data volume and number of files.
[0003] However, in actual application scenarios, there are the following situations:
[0004] Tasks with large computational loads will inevitably occupy too many computing resources, but the execution time of these tasks may not be urgent, so smaller computing resources can be set for such tasks;
[0005] Data tasks for business or senior management are high-priority tasks that are expected to be completed faster, so more computing resources are needed to ensure that the computing time is not too long.
[0006] For individual customers, they expect the characteristics of their business data to be stable. Therefore, for Spark tasks of different data sizes, a relative standard can be set according to the data volume. In the actual calculation process, resource parameters can be configured dynamically according to preset rules based on the data volume of the Spark task input data table.
[0007] However, the existing Spark offline task resource scheduling link does not have the above-mentioned dynamic resource scheduling configuration solution for offline tasks, and therefore cannot effectively meet customer needs, resulting in inefficient Spark offline task resource scheduling and low task running and execution completion rates. Summary of the invention
[0008] In order to solve the above problems, the present application proposes a Spark offline task resource scheduling optimization method, system and electronic device combined with input data volume.
[0009] On the one hand, the present application proposes a Spark offline task resource scheduling optimization method combined with input data volume, comprising the following steps:
[0010] S1. Preset and build a resource rule list consisting of different numbers of rows and corresponding computing resource rules;
[0011] S2. Collect and parse the Spark offline task, obtain the number of rows in the data table of the Spark offline task, and calculate the number of rows in the data table;
[0012] S3. Based on the resource rule list, match the number of rows in the data table to obtain the corresponding computing resource rule;
[0013] S4. According to the computing resource rule, the Spark offline task is scheduled and sent to the corresponding execution node for execution.
[0014] As an optional implementation scheme of the present application, optionally, S1, pre-constructing a resource rule list consisting of different numbers of rows and corresponding computing resource rules, including:
[0015] Build a large model prompt word for extracting line numbers and identifying computing resource rules, and configure it in the preset LLM large language model;
[0016] Collecting historical execution logs of several Spark offline tasks from a backend database;
[0017] Traversing the historical execution log, the LLM large language model identifies and extracts the number of rows of the data table of different Spark offline tasks and the computing resource rules for executing the Spark offline tasks from the historical execution log based on the large model prompt words;
[0018] Counting the number of rows of different Spark offline tasks and the corresponding computing resource rules, and automatically filling them into a preset rule table by the LLM large language model to obtain the resource rule list;
[0019] The resource rule list is configured in a resource scheduler.
[0020] As an optional implementation scheme of the present application, optionally, the computing resource rule includes the following rule elements:
[0021] driver-memory: driver memory
[0022] driver-cores: driver cores;
[0023] num-executors: number of executors;
[0024] executor-memory: executor memory;
[0025] executor-cores: executor cores.
[0026] As an optional implementation scheme of the present application, optionally, S2, collecting and parsing the Spark offline task, obtaining the number of rows of the data table of the Spark offline task, and calculating the number of rows of the data table, includes:
[0027] Collecting the Spark offline tasks reported by the client to the resource manager;
[0028] Parsing the Spark offline task through a resource manager to obtain the number of rows of the data table in the Spark offline task;
[0029] The enumerate function is used in combination with a file iteration method to calculate the number of rows in the data table and input it into the resource scheduler.
[0030] As an optional implementation scheme of the present application, optionally, S3, based on the resource rule list, matching the number of rows of the data table to obtain the corresponding computing resource rule includes:
[0031] Read the number of rows of the data table in the current Spark offline task through the resource manager;
[0032] Call the resource rule list, perform matching retrieval on the number of rows of the data table in the current Spark offline task, and find the computing resource rule corresponding to the number of rows of the data table in the current Spark offline task;
[0033] Bind the computing resource rule to the current Spark offline task.
[0034] As an optional implementation scheme of the present application, optionally, when calling the resource rule list to match and retrieve the number of rows of the data table in the current Spark offline task, it includes:
[0035] Generate a corresponding search task through the resource manager;
[0036] Sending the search task to the LLM large language model, the LLM large language model executes the search task, performs matching search on the number of rows of the data table in the current Spark offline task and feeds back the corresponding computing resource rule to the resource manager;
[0037] The resource manager forwards the current Spark offline task and the corresponding computing resource rule to the corresponding node manager through a router according to the task attribute of the current Spark offline task.
[0038] As an optional implementation scheme of the present application, optionally, S4, according to the computing resource rule, scheduling the Spark offline task to the corresponding execution node for execution, includes:
[0039] Receiving, through the node manager, the computing resource rule forwarded by the resource manager and the current Spark offline task;
[0040] The node manager reads the computing resource rule, identifies the computing resource attribute in the computing resource rule, and activates the corresponding execution node according to the computing resource attribute;
[0041] The Spark offline task is forwarded to the activated execution node to perform task execution of the Spark offline task.
[0042] On the other hand, the present application proposes a system for implementing the Spark offline task resource scheduling optimization method combined with the input data volume, including:
[0043] Client, used to report Spark offline tasks;
[0044] A resource manager, configured to collect and parse the Spark offline task, obtain the number of rows in the data table of the Spark offline task, and calculate the number of rows in the data table; and, based on the resource rule list, match the number of rows in the data table to obtain the corresponding computing resource rule;
[0045] A node manager is used to manage each execution node; and, according to the computing resource rules, schedule the Spark offline task to be sent to the corresponding execution node for execution;
[0046] A router, configured for the resource manager to forward the current Spark offline task and the corresponding computing resource rule to the corresponding node manager according to the task attribute of the current Spark offline task;
[0047] LLM large language model API, used by the resource manager to call the LLM large language model;
[0048] Backend database, used for backend data storage;
[0049] The resource manager, router, node manager, LLM large language model API and backend database are all deployed on the backend server;
[0050] The client is in communication connection with the backend server.
[0051] In another aspect, the present application further provides an electronic device, comprising:
[0052] processor;
[0053] a memory for storing processor-executable instructions;
[0054] Wherein, the processor is configured to implement the Spark offline task resource scheduling optimization method combined with the input data volume when executing the executable instructions.
[0055] Technical effects of the present invention:
[0056] This application collects and parses Spark offline tasks to obtain the number of rows in the data table of the Spark offline task, and calculates the number of rows in the data table; based on the preset resource rule list, the number of rows in the data table is matched to obtain the corresponding computing resource rules; according to the computing resource rules, the Spark offline task is scheduled and sent to the corresponding execution node for execution. It can optimize the scheduling of computing resources in combination with the attribute of the number of rows in the Spark offline task data table, so that in the actual calculation process, according to the amount of data in the Spark task input data table, resource parameters can be dynamically configured according to preset rules, so as to optimize the allocation of computing resources, improve computing efficiency, promote the efficient operation of Spark offline tasks, and effectively meet customer needs.
[0057] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0059] Figure 1 It is shown as a schematic diagram of the implementation process of the present invention;
[0060] Figure 2 It is a schematic diagram showing the scheduling and execution of computing resources of the present invention;
[0061] Figure 3 Shown is a schematic diagram of the system composition structure of the present invention;
[0062] Figure 4 It is a schematic diagram showing the application of the electronic device of the present invention. DETAILED DESCRIPTION
[0063] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0064] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0065] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure. Example 1
[0066] like Figure 1 As shown, on the one hand, the present application proposes a Spark offline task resource scheduling optimization method combined with the input data volume, comprising the following steps:
[0067] S1. Preset and build a resource rule list consisting of different numbers of rows and corresponding computing resource rules;
[0068] S2. Collect and parse the Spark offline task, obtain the number of rows in the data table of the Spark offline task, and calculate the number of rows in the data table;
[0069] S3. Based on the resource rule list, match the number of rows in the data table to obtain the corresponding computing resource rule;
[0070] S4. According to the computing resource rule, the Spark offline task is scheduled and sent to the corresponding execution node for execution.
[0071] The present invention mainly combines the number of rows (that is, the amount of input data) of the Spark offline task data table to optimize the scheduling of computing resources, so that in the actual calculation process, resource parameters can be dynamically configured according to preset rules based on the amount of data in the Spark task input data table, thereby optimizing the allocation of computing resources and improving computing efficiency.
[0072] First, design resource rules and set detailed configuration of computing resources under different data row number ranges according to actual business scenarios and optimization experience (see Figure 2 list of resource rules shown).
[0073] Secondly, when the Spark offline task is started, the total number of rows in the input data table is calculated as the basis for resource selection. Specifically, according to the number of rows in the calculated data table, a match is made in the resource rule list to find a resource rule that meets the row number range, which is used as the parameter setting for the final task run submission.
[0074] The Spark task execution platform schedules and executes the Spark offline tasks reported by each client according to the above method, so as to optimize the computing resources of the allocation platform.
[0075] Spark task execution platform, or the system's application architecture such as Figure 2 As shown, the system background is configured with a resource manager, a node manager, and a background database (which can be understood in conjunction with the existing Spark task execution platform or the architecture of the Spark running mode).
[0076] The present invention also configures a router for task scheduling and routing forwarding of computing resources. The router is deployed in the background to realize routing forwarding of tasks and data. The router has the functions of computing resource task scheduling and routing forwarding. Specifically:
[0077] Deploy routers in the background to implement routing forwarding. Routing is the process of forwarding data packets, and the path is selected based on the routing table. The router is a device that performs routing functions and maintains the routing table.
[0078] 1. Router deployment:
[0079] Configure routing table: add static routing or dynamic routing according to network topology (in the present invention, the Spark operating system or Spark operating platform deploys several node managers that can handle different task attributes / types, so the routing path and address between the resource manager and the corresponding node manager can be configured in the router, which can be specifically configured by the administrator).
[0080] Set priority: Use the management distance value (AD value) to set the routing priority (the resource manager can identify whether the Spark offline task has a priority flag. If so, the Spark offline task will be executed first).
[0081] Interface configuration: Make sure the router interface is turned on and configured with the correct IP.
[0082] 2. Routing and forwarding process:
[0083] Matching routing table: Select the best path based on the longest mask match principle.
[0084] Next-hop forwarding: forwards data packets according to the next-hop address pointed to by the routing table.
[0085] Failure to match processing: If no routing entry is matched, the packet is discarded.
[0086] 3. Static routing configuration:
[0087] Applicable to small-scale networks and manual addition of routing entries.
[0088] Configuration command example: ip route-static target network subnet mask next hop.
[0089] Through the above steps, the router can be successfully deployed in the background, and the routing forwarding of tasks and data can be realized.
[0090] In addition, an LLM large language model API is configured to call the LLM large language model from a third-party platform to assist the resource manager in performing corresponding tasks.
[0091] The present invention can send a retrieval or matching instruction (including specific task requirements, logic or keywords) to the large language model, requiring the large language model to perform the corresponding task according to the large model prompt words in the instruction, and return the corresponding retrieval or matching results to the resource manager. Therefore, an LLM large language model API port is configured on the background, and a third-party application platform can be logged in through the API, so that the resource manager can call the LLM large language model to perform the corresponding task, improve the task execution efficiency of the resource manager, and reduce the operating pressure of the resource manager (traditional task management is concentrated on the background, which will cause greater management pressure on the resource manager).
[0092] The implementation principle of the present invention will be further described below.
[0093] As an optional implementation scheme of the present application, optionally, S1, pre-constructing a resource rule list consisting of different numbers of rows and corresponding computing resource rules, including:
[0094] Build a large model prompt word for extracting line numbers and identifying computing resource rules, and configure it in the preset LLM large language model;
[0095] Collecting historical execution logs of several Spark offline tasks from a backend database;
[0096] Traversing the historical execution logs, the LLM large language model identifies and extracts the number of rows of the data table of different Spark offline tasks and the computing resource rules for executing the Spark offline tasks from the historical execution logs based on the large model prompt words;
[0097] Counting the number of rows of different Spark offline tasks and the corresponding computing resource rules, and automatically filling them into a preset rule table by the LLM large language model to obtain the resource rule list;
[0098] The resource rule list is configured in a resource scheduler.
[0099] like Figure 2As shown, it is a resource rule list corresponding to different numbers of rows. The resource rule list contains computing resource rules corresponding to different ranges of numbers of rows, and the computing resource rules include the following rule elements:
[0100] driver-memory: driver memory
[0101] driver-cores: driver cores;
[0102] num-executors: number of executors;
[0103] executor-memory: executor memory;
[0104] executor-cores: executor cores.
[0105] For example, when facing a data table with a row number ranging from 0 to 50 million rows, the system can match resource rule 1 (driver-memory: 4g, driver-cores: 2, num-executors: 10, executor-memory: 4g, executor-cores: 2), notify the node manager of this configuration, and assign and execute Spark offline tasks for the data table with a row number ranging from 0 to 50 million rows.
[0106] Administrators can summarize the reasonable resource quotas for different tasks based on historical task operation information; determine the resource quota requirements for different tasks based on business timeliness requirements; determine the reasonable resource quotas for different tasks based on task performance test observations; and configure the above conclusions into the preset rule table, which is the resource rule list.
[0107] There are many types of tasks, and the execution data of the tasks is also much. In order to obtain the computing resource information matched by different tasks (executor data of the tasks) as much as possible, the present invention uses the LLM large language model to assist in the construction of the rule table. In this way, the efficiency of the rule table construction is improved, and the execution information of a large number of Spark offline tasks is analyzed and identified as much as possible, and the computing resource information occupied by tasks with different data volumes is obtained when they are executed.
[0108] Here, the LLM large language model (or LLM or language model) is used to construct a source rule list, which is operated by an administrator to construct prompt words (a type of model prompt words in the large language model) and input them into the LLM large language model.
[0109] Large language models include: GPT-3 / 4, LaMDA, Megatron-Turing NLG, PAI-NG, WPSAI, or Baidu's Wenxin Yiyan and other various open source and commercial large language models. Users can connect to the corresponding model themselves.
[0110] As for the range of rows for executing different tasks and the corresponding computing resource rules, the present invention uses the LLM large language model to identify the execution logs of different Spark offline tasks from the historical execution logs in the background database, and identifies and extracts the number of rows of different Spark offline tasks and the computing resource rules for executing the tasks from the execution logs. Because the platform can generate corresponding execution logs when executing tasks, the present invention can use the LLM large language model to identify and extract the number of rows in the data table of the corresponding task and the computing resource rules used to execute the task from the historical Spark execution logs. In the present invention, each resource rule in the resource rule list is Figure 2 The above is just an example. The number of specific resource rules, the range of lines in the resource rules, and the corresponding computing resource conditions can be revised by the administrator. Alternatively, the rules can be configured directly according to the number of lines in the execution log and the content of the corresponding computing resource rules.
[0111] The LLM large language model can perform corresponding data retrieval based on the prompt words entered by the user administrator, retrieve the historical execution log associated with the prompt words from the background database, and identify and extract the number of rows corresponding to the execution of each historical task and the corresponding computing resource rules from the log. After the LLM large language model feedbacks the corresponding retrieval results, the administrator (or the resource manager) counts the number of rows of different tasks and the corresponding computing resource rules, and fills them into the preset rule table (which can be simplified by using an Excel table and then constructed using a rule template) to generate a resource rule list. After the resource rule list is constructed, it can be configured in the resource manager. For example, the administrator logs in to the background through the client to configure the resource rule list.
[0112] The steps to let the LLM large language model identify and extract corresponding data from the log based on the prompt words can refer to the following steps:
[0113] 1. Data preparation:
[0114] Collect and prepare log files containing the required information.
[0115] 2. Word segmentation and embedding:
[0116] Use a tokenizer to split the log text into small text blocks (tokens).
[0117] These tokens are mapped to specific integer codes and converted into numerical representations (embeddings) of high-dimensional vectors.
[0118] 3. Model prediction:
[0119] The embedded vectors are processed using LLM’s multi-layer neural network and attention mechanism.
[0120] Generate prediction results related to the log content based on the prompt words.
[0121] 4. Data Extraction:
[0122] Parse and extract information related to the prompt word from the output of the model.
[0123] The extracted data will be organized for subsequent use or analysis.
[0124] This process ensures that LLM can effectively identify and extract relevant information based on the prompt words from the logs.
[0125] Specifically, it can be understood in conjunction with the application of the LLM large language model.
[0126] As an optional implementation scheme of the present application, optionally, S2, collecting and parsing the Spark offline task, obtaining the number of rows of the data table of the Spark offline task, and calculating the number of rows of the data table, includes:
[0127] Collecting the Spark offline tasks reported by the client to the resource manager;
[0128] Parsing the Spark offline task through a resource manager to obtain the number of rows of the data table in the Spark offline task;
[0129] The enumerate function is used in combination with a file iteration method to calculate the number of rows in the data table and input it into the resource scheduler.
[0130] First, perform task analysis to obtain the number of rows in the data table in the task (analysis of the task package file).
[0131] The format and structure of the task package file can be determined first, and the resource manager can read the task package file content, extract key data fields in the task package, calculate the data volume, such as the number of rows and records, and generate data volume statistics.
[0132] Here, the enumerate function is used in combination with the file iteration method to calculate the number of rows in the data table and input it into the resource scheduler. The following enumerate function combined with the file iteration python code can be executed:
[0133] with open('path / to / your / data.csv', 'r') as file:
[0134] row_count = sum(1 for line in file),
[0135] print(f"Row count: {row_count}").
[0136] principle:
[0137] Use the with statement to open the CSV file to ensure that the file will be properly closed after reading;
[0138] Use the sum(1 for line in file) expression to count the number of lines in a file; here, 1 for line in file is a generator expression that generates a number 1 for each line in the file, and then the sum() function adds up these numbers to get the total number of lines in the file;
[0139] Use the print statement to output the calculated number of rows;
[0140] Note that this method counts every line in the file, including empty lines and possible header lines; if you do not want to count empty or header lines, you will need to add appropriate logic to filter these lines when reading the file. For example, you could check whether each line is empty or starts with a specific header line and adjust the line count accordingly.
[0141] By the above method, the number of rows of the corresponding data volume can be quickly obtained. If the number of rows in the data table in the task package is too large, the resource manager can sub-package the task package, calculate the number of rows of each sub-package according to the above method, and then sum the number of rows.
[0142] As an optional implementation scheme of the present application, optionally, S3, based on the resource rule list, matching the number of rows of the data table to obtain the corresponding computing resource rule includes:
[0143] Read the number of rows of the data table in the current Spark offline task through the resource manager;
[0144] Call the resource rule list, perform matching retrieval on the number of rows of the data table in the current Spark offline task, and find the computing resource rule corresponding to the number of rows of the data table in the current Spark offline task;
[0145] Bind the computing resource rule to the current Spark offline task.
[0146] After calculating the number of rows in the current task data table, the resource manager can match the range of rows based on the content in the resource rule list to match the corresponding resource rule for the current task. The matched computing resource rule is bound to the current task, and then the resource manager sends the bound task and rule to the node manager, so that the node manager can select the execution node (executor) configured with the computing resource information according to the computing resource information in the computing resource rule to execute the current task.
[0147] The computing resource rules specify the types and core information of the driver and executor that execute the current offline task. Therefore, the subsequent node manager can refer to this rule and select the corresponding execution node to execute the current offline task, realizing the service function of scheduling the corresponding computing resources for task execution based on the number of rows in the data table, and realizing the optimized scheduling and execution of tasks.
[0148] As an optional implementation scheme of the present application, optionally, when calling the resource rule list to match and retrieve the number of rows of the data table in the current Spark offline task, it includes:
[0149] Generate a corresponding search task through the resource manager;
[0150] Sending the search task to the LLM large language model, the LLM large language model executes the search task, performs matching search on the number of rows of the data table in the current Spark offline task and feeds back the corresponding computing resource rule to the resource manager;
[0151] The resource manager forwards the current Spark offline task and the corresponding computing resource rule to the corresponding node manager through a router according to the task attribute of the current Spark offline task.
[0152] Because the resource manager can call the LLM model from a third party through the large language model API port to perform the corresponding resource management task, the present invention can generate a row number retrieval instruction for the corresponding task when the resource manager performs retrieval and matching of Spark offline tasks. The instruction contains the number of rows of the current task and the corresponding retrieval requirements and logic (constructed by the administrator), and sends it to the large language model. The large language model executes the retrieval task, and retrieves the computing resource rules that match the number of rows from the resource rule list according to the task row number requirements required in the task and feeds them back to the resource manager. Therefore, it can save the computing running memory of the resource manager, reduce the running pressure of the resource manager, allow the resource manager to perform task management on massive tasks, and improve the efficiency of task execution by combining the retrieval and matching of the corresponding computing resource rules by LLM.
[0153] As an optional implementation scheme of the present application, optionally, S4, according to the computing resource rule, scheduling the Spark offline task to be sent to the corresponding execution node for execution, includes:
[0154] Receiving, through the node manager, the computing resource rule forwarded by the resource manager and the current Spark offline task;
[0155] The node manager reads the computing resource rule, identifies the computing resource attribute in the computing resource rule, and activates the corresponding execution node according to the computing resource attribute;
[0156] The Spark offline task is forwarded to the activated execution node to perform task execution of the Spark offline task.
[0157] Different node managers can execute tasks of different departments or different attributes. Therefore, the resource manager can forward the tasks of different departments to the corresponding node manager through the router. In the present invention, the node manager on the task platform can be deployed in a distributed manner, and each node manager includes a number of execution nodes, each of which is configured with an executor of different computing resources (corresponding to each resource rule). Therefore. Therefore, the node manager needs to manage each execution node, and needs to record and save the computing resource attributes of each execution node, so as to facilitate the subsequent allocation of the corresponding execution node for the current task according to the current task forwarded by the resource manager and the corresponding computing resource rules. Specifically, the node manager needs to read the computing resource information in the computing resource rules, and match the information with the computing resource attributes of each execution node, so as to match the corresponding execution node for the current task, and send the current task to the matched execution node, and activate the execution node to perform task execution on the current task. It can be understood specifically in conjunction with the operation process of the Spark task platform, and will not be repeated here.
[0158] Obviously, those skilled in the art should understand that the implementation of all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. Those skilled in the art can understand that the implementation of all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (Hard Disk Drive, abbreviated as: HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory. Example 2
[0159] like Figure 3 As shown, based on the implementation principle of Example 1, on the other hand, the present application proposes a system for implementing the Spark offline task resource scheduling optimization method combined with the input data volume, including:
[0160] Client, used to report Spark offline tasks;
[0161] A resource manager, configured to collect and parse the Spark offline task, obtain the number of rows in the data table of the Spark offline task, and calculate the number of rows in the data table; and, based on the resource rule list, match the number of rows in the data table to obtain the corresponding computing resource rule;
[0162] A node manager is used to manage each execution node; and, according to the computing resource rules, schedule the Spark offline task to be sent to the corresponding execution node for execution;
[0163] A router, configured for the resource manager to forward the current Spark offline task and the corresponding computing resource rule to the corresponding node manager according to the task attribute of the current Spark offline task;
[0164] LLM large language model API, used by the resource manager to call the LLM large language model;
[0165] Backend database, used for backend data storage;
[0166] The resource manager, router, node manager, LLM large language model API and backend database are all deployed on the backend server;
[0167] The client is in communication connection with the backend server.
[0168] The functional architecture and interaction process of this system can be understood in conjunction with the method steps of Example 1, and will not be described in detail in this example.
[0169] The modules or steps of the present invention described above can be implemented by a general-purpose computing system, they can be concentrated on a single computing system, or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by a program code executable by a computing system, so that they can be stored in a storage system and executed by the computing system, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software. Example 3
[0170] like Figure 4 As shown, further, in another aspect, the present application also proposes an electronic device, including:
[0171] processor;
[0172] a memory for storing processor-executable instructions;
[0173] The processor is configured to implement a Spark offline task resource scheduling optimization method combined with input data volume as described in Example 1 when executing the executable instructions.
[0174] The electronic device of the embodiment of the present disclosure includes a processor and a memory for storing processor executable instructions. The processor is configured to implement a Spark offline task resource scheduling optimization method combined with input data volume as described in the above embodiment 1 when executing the executable instructions.
[0175] Here, it should be noted that the number of processors can be one or more. At the same time, the electronic device of the embodiment of the present disclosure may also include an input system and an output system. Among them, the processor, memory, input system and output system may be connected through a bus or in other ways, which are not specifically limited here.
[0176] As a computer-readable storage medium, the memory can be used to store software programs, computer executable programs and various modules, such as: the program or module corresponding to the Spark offline task resource scheduling optimization method combined with the input data volume in the embodiment of the present disclosure. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.
[0177] The input system can be used to receive input numbers or signals. The signal can be a key signal related to user settings and function control of the device / terminal / server. The output system can include display devices such as display screens.
[0178] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A Spark offline task resource scheduling optimization method combined with input data volume, characterized in that: The steps include: S1. Preset and build a resource rule list consisting of different numbers of rows and corresponding computing resource rules, including: Build a large model prompt word for extracting line numbers and identifying computing resource rules, and configure it in the preset LLM large language model; Collecting historical execution logs of several Spark offline tasks from a backend database; Traversing the historical execution logs, the LLM large language model identifies and extracts the number of rows of the input data volume of different Spark offline tasks and the computing resource rules for executing the Spark offline tasks from the historical execution logs based on the large model prompt words; Counting the number of rows of different Spark offline tasks and the corresponding computing resource rules, and automatically filling them into a preset rule table by the LLM large language model to obtain the resource rule list; Configuring the resource rule list in a resource scheduler; The LLM large language model identifies the execution logs of different Spark offline tasks from the historical execution logs in the background database, identifies and extracts the number of rows of different Spark offline tasks and the computing resource rules for executing the tasks from the execution logs, including the following steps: 1). Data preparation: Collect and prepare log files containing the required information 2). Word segmentation and embedding: Use a tokenizer to split the log text into small text blocks called tokens. Map these tokens to specific integer codes and convert them into numerical representations of high-dimensional vectors called embeddings. 3). Model prediction: Use LLM's multi-layer neural network and attention mechanism to process the embedded vector; Generate prediction results related to the log content based on the prompt words; 4). Data Extraction: Parse and extract information related to the prompt word from the model output; S2. Collect and parse the Spark offline task, obtain the input data volume of the Spark offline task, and calculate the number of rows of the input data volume; S3. Based on the resource rule list, match the number of rows of the input data volume to obtain the corresponding computing resource rule; S4. According to the computing resource rule, the Spark offline task is scheduled and sent to the corresponding execution node for execution.
2. The Spark offline task resource scheduling optimization method combined with input data volume according to claim 1 is characterized in that: S2. Collecting and parsing the Spark offline task, obtaining the input data volume of the Spark offline task, and calculating the number of rows of the input data volume, including: Collecting the Spark offline tasks reported by the client to the resource manager; Parsing the Spark offline task through a resource manager to obtain the input data volume in the Spark offline task; The enumerate function is used in combination with a file iteration method to calculate the number of rows of the input data volume and input it into the resource scheduler.
3. The Spark offline task resource scheduling optimization method combined with input data volume according to claim 1 is characterized in that: S3. Based on the resource rule list, matching the number of rows of the input data volume to obtain the corresponding computing resource rule includes: Read the number of rows of the input data in the current Spark offline task through the resource manager; Calling the resource rule list, performing a matching search on the number of rows of the input data volume in the current Spark offline task, and searching for the computing resource rule corresponding to the number of rows of the input data volume in the current Spark offline task; Bind the computing resource rule to the current Spark offline task.
4. The Spark offline task resource scheduling optimization method according to claim 3 is characterized in that: When calling the resource rule list to match and retrieve the number of rows of the input data volume in the current Spark offline task, it includes: Generate a corresponding search task through the resource manager; Send the search task to the LLM large language model, and the LLM large language model executes the search task, performs matching search on the number of rows of the input data volume in the current Spark offline task, and feeds back the corresponding computing resource rules to the resource manager; The resource manager forwards the current Spark offline task and the corresponding computing resource rule to the corresponding node manager through a router according to the task attribute of the current Spark offline task.
5. The Spark offline task resource scheduling optimization method combined with input data volume according to claim 3 is characterized in that: S4. According to the computing resource rule, the Spark offline task is scheduled and sent to the corresponding execution node for execution, including: Receiving, through a node manager, the computing resource rule forwarded by the resource manager and the current Spark offline task; The node manager reads the computing resource rule, identifies the computing resource attribute in the computing resource rule, and activates the corresponding execution node according to the computing resource attribute; The Spark offline task is forwarded to the activated execution node to execute the Spark offline task.
6. The Spark offline task resource scheduling optimization method according to claim 1, characterized in that: The computing resource rules include the following rule elements: driver-memory: driver memory driver-cores: driver cores; num-executors: number of executors; executor-memory: executor memory; executor-cores: executor cores.
7. A system for implementing the Spark offline task resource scheduling optimization method in combination with input data volume as described in any one of claims 1 to 6, characterized in that: include: Client, used to report Spark offline tasks; A resource manager is used to collect and parse the Spark offline task, obtain the input data volume of the Spark offline task, and calculate the number of rows of the input data volume; And, based on the resource rule list, matching the number of rows of the input data volume to obtain the corresponding computing resource rule; Node manager, used to manage each execution node; And, according to the computing resource rule, the Spark offline task is scheduled and sent to the corresponding execution node for execution; A router, configured for the resource manager to forward the current Spark offline task and the corresponding computing resource rule to the corresponding node manager according to the task attribute of the current Spark offline task; LLM large language model API, used by the resource manager to call the LLM large language model; Backend database, used for backend data storage; The resource manager, router, node manager, LLM large language model API and backend database are all deployed on the backend server; The client is in communication connection with the backend server.
8. An electronic device, characterized in that include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the Spark offline task resource scheduling optimization method combined with input data volume as described in any one of claims 1-6 when executing the executable instructions.
Citation Information
Patent Citations
Big data computing engine task parameter determination method and device
CN114003388A
Data consistency detection and repair method and device and electronic equipment
CN118689887A