Task processing method and device, electronic equipment, storage medium and program product
By generating an initial DAG using a large model agent and dynamically adjusting it using a reinforcement learning model, the problem of low execution efficiency of big data tasks in dynamic environments is solved, and efficient task processing is achieved.
Patent Information
- Application Number
- CN202511710069.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies cannot respond promptly to dynamic environmental changes such as cluster resource fluctuations, sudden changes in data volume, and adjustments to task priorities when performing big data tasks, resulting in low task execution efficiency.
By inputting the user's processing requirements into the large model agent, an initial DAG is generated. Then, a reinforcement learning model is used to extract real-time state information, dynamically adjust the nodes and edges of the DAG, and generate a second DAG that adapts to the current environment, thereby optimizing task execution.
It enables timely response to resource fluctuations and sudden data events, significantly improving the execution efficiency of big data tasks in dynamic environments and avoiding efficiency losses caused by static scheduling.
Smart Images

Figure CN121542002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a task processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Big data tasks refer to specific processing of massive amounts of data, such as extraction, cleaning, transformation, analysis, modeling, and report generation, to meet business needs. These tasks rely on various data processing tools and cluster resources for execution. Related technologies generally employ static rules or predefined algorithms when orchestrating big data tasks, relying on fixed, unchanging built-in rules and models. Taking mainstream data flow orchestration tools like Apache NiFi and Airflow as examples, their scheduling logic is primarily based on pre-defined, fixed dependencies between tasks. When encountering dynamic environmental changes such as cluster resource fluctuations, sudden changes in data volume, or adjustments to task priorities, they often cannot respond promptly, ultimately leading to low task execution efficiency. Summary of the Invention
[0003] This application provides a task processing method, apparatus, electronic device, storage medium, and program product that can solve the problem of low task execution efficiency.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a task processing method, the method comprising: The user's processing requirements for the target task are input into the large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools. The edges of the first DAG are used to represent the execution order of the N sub-tasks, where N is an integer greater than 1. Real-time state information is extracted based on a reinforcement learning model, and the real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. Based on the real-time state information, the nodes and edges of the first DAG are adjusted using the reinforcement learning model to obtain the second DAG; The target task is processed based on the second DAG.
[0005] Secondly, embodiments of this application provide a task processing apparatus, the apparatus comprising: The first processing module is used to input the user's processing requirements for the target task into the large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools, and the edges of the first DAG are used to represent the execution order of the N sub-tasks, where N is an integer greater than 1. The first extraction module is used to extract real-time state information based on the reinforcement learning model. The real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. The first adjustment module is used to adjust the nodes and edges of the first DAG based on the real-time state information using the reinforcement learning model to obtain the second DAG; The first processing module is used to process the target task based on the second DAG.
[0006] Thirdly, embodiments of this application provide an electronic device, including a processor, the processor being used for: The user's processing requirements for the target task are input into the large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools. The edges of the first DAG are used to represent the execution order of the N sub-tasks, where N is an integer greater than 1. Real-time state information is extracted based on a reinforcement learning model, and the real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. Based on the real-time state information, the nodes and edges of the first DAG are adjusted using the reinforcement learning model to obtain the second DAG; The target task is processed based on the second DAG.
[0007] Fourthly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the task processing method as described in the first aspect.
[0008] Fifthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the task processing method as described in the first aspect.
[0009] Sixthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the task processing method described above.
[0010] In this embodiment, by inputting the user's processing requirements for the target task into a large model agent, a first directed acyclic graph (DAG) corresponding to the target task is obtained. The nodes of the first DAG represent the mapping relationship between the N sub-tasks of the target task and their corresponding execution tools, and the edges of the first DAG represent the execution order of the N sub-tasks, where N is an integer greater than 1. Real-time state information is extracted based on a reinforcement learning model. The real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. Based on the real-time state information, the nodes and edges of the first DAG are adjusted using the reinforcement learning model to obtain a second DAG. The target task is then processed based on the second DAG. In this way, the large model agent generates a first DAG based on the user's information on the processing requirements of the target task. The reinforcement learning model extracts the real-time state information of the target task when it is processed based on the first DAG. Then, based on the real-time state information, it dynamically adjusts the nodes and edges of the first DAG to generate a second DAG that adapts to the current environment. This enables timely response to resource fluctuations, data bursts, and other situations. Finally, the task is executed with the optimized second DAG, avoiding the efficiency loss caused by static scheduling and significantly improving the execution efficiency of big data tasks in dynamic environments. Attached Figure Description
[0011] Figure 1 Flowchart of the task processing method provided in the embodiments of this application Figure 1 ; Figure 2 Flowchart of the task processing method provided in the embodiments of this application Figure 2 ; Figure 3 This is a schematic diagram of the structure of the task processing device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "a" or "one," and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked," and similar terms, are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.
[0014] The following description, in conjunction with the accompanying drawings, further illustrates the task processing method, apparatus, device, storage medium, and program product proposed in the embodiments of the application.
[0015] Please see Figure 1 , Figure 1 A flowchart illustrating a task processing method provided in this application embodiment is shown in the figure. The method includes: Step 110: Input the user's processing requirements for the target task into the large model agent to obtain the first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools. The edges of the first DAG are used to represent the execution order of the N sub-tasks. Here, N is an integer greater than 1. In this step, the target task can be understood as a big data task. The processing requirement information can be understood as information proposed by the user in natural language, including elements such as the target task type, data source, processing conditions, and output requirements, used to clarify the processing goals and constraints of the big data task. For example, the user's processing requirement information for the target task could be "Analyze last week's user login logs, exclude test accounts, and generate a retention rate report." The large-scale intelligent agent can be understood as an intelligent system based on a large language model, capable of autonomously understanding tasks, planning steps, calling tools, and continuously iterating to independently complete complex goals.
[0016] After the processing requirements are input into the large model agent, the large model agent will identify the task type, decompose the sub-task elements, select suitable execution tools for the sub-tasks, and output a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG represent the mapping relationship between the N sub-tasks of the target task and their corresponding execution tools. For example, the data extraction sub-task corresponds to the HiveQuery tool, and the filtering sub-task corresponds to the SparkSQL tool. The edges specify the execution order of these N sub-tasks, such as data extraction → filtering → calculating retention rate, to ensure that the target task proceeds logically and smoothly.
[0017] Step 120: Extract real-time state information based on the reinforcement learning model. The real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. In this step, while the system processes the target task based on the first DAG, it collects real-time state information in real time through a reinforcement learning model. The real-time state information may include: Chinese processor (Central Processing Unit, CPU) utilization: reflects the actual CPU usage in the current cluster and helps to determine the resource load level; Priority distribution of the queue to be scheduled: Displays the priority distribution of tasks to be scheduled, which is used to analyze the urgency and importance of the current task scheduling.
[0018] Remaining time for the longest path in the current DAG: This indicates the remaining execution time of the longest path in the current DAG, helping to assess the expected time for the overall task to be completed.
[0019] Historical execution data: Average time for similar tasks in historical executions: Provides the average time for similar tasks in historical executions, serving as a reference for estimating the time of the current task.
[0020] These real-time status information need to be normalized: Min-Max normalization is used for heterogeneous indicators to ensure that the value range of each dimension is in [0,1].
[0021] Step 130: Based on the real-time state information, adjust the nodes and edges of the first DAG using the reinforcement learning model to obtain the second DAG; In this step, after obtaining the real-time state information of the first DAG processing target task, the reinforcement learning model will make targeted adjustments to the nodes and edges of the first DAG based on this real-time state information and through the decision-making ability of the policy network. Node adjustments include replacing execution tools and adjusting parallelism, while edge adjustments include adding or deleting data dependencies and inserting cache nodes. Ultimately, a second DAG that is adapted to the dynamic environment, has better resource utilization, and higher execution efficiency is generated.
[0022] Step 140: Process the target task based on the second DAG.
[0023] In this step, resources are allocated, subtask execution order is scheduled, and exceptions are handled in real time according to the optimized process defined by the second DAG, so as to complete the target task efficiently and stably.
[0024] In one implementation, user requirements for processing a target task are input into a large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG represent the mapping relationship between N subtasks of the target task and their corresponding execution tools, and the edges of the first DAG represent the execution order of the N subtasks, where N is an integer greater than 1. Real-time state information is extracted based on a reinforcement learning model; this real-time state information corresponds to the real-time state information when processing the target task based on the first DAG. Based on the real-time state information, the nodes and edges of the first DAG are adjusted using the reinforcement learning model to obtain a second DAG. The target task is then processed based on the second DAG.
[0025] In this implementation, the large model agent generates a first DAG based on the user's processing requirements for the target task. The reinforcement learning model extracts the real-time state information of the target task when it is processed based on the first DAG. Then, based on the real-time state information, the nodes and edges of the first DAG are dynamically adjusted to generate a second DAG adapted to the current environment. This enables timely response to resource fluctuations, data bursts, and other situations. Finally, the task is executed with the optimized second DAG, avoiding the efficiency loss caused by static scheduling and significantly improving the execution efficiency of big data tasks in dynamic environments.
[0026] Optionally, the step of inputting the user's processing requirements for the target task into the large model agent to obtain the first directed acyclic graph (DAG) corresponding to the target task includes: The user's processing requirements for the target task are input into the semantic parsing module of the large model agent to obtain an initial DAG corresponding to the target task. The nodes of the initial DAG are used to represent the mapping relationship between the N sub-tasks of the target task and their corresponding execution tools, and the edges of the initial DAG are used to represent the execution order of the N sub-tasks. The dynamic programming module of the large model agent constructs an extensible structure for the initial DAG, and injects check nodes and rollback nodes into the initial DAG to obtain the first DAG corresponding to the target task. The check nodes are used to check and save the data of subtasks whose processing time exceeds a preset time, and the rollback nodes are used to delete the data that failed to be written during the processing of the subtask.
[0027] In one implementation, the user first inputs the processing requirements of the target task into the semantic parsing module of the large model agent. The module parses the requirements, breaks them down into N specific sub-tasks, matches an appropriate execution tool for each sub-task, and finally outputs an initial DAG. The nodes of the initial DAG correspond to the mapping relationship between sub-tasks and execution tools, and the edges of the initial DAG specify the execution order of the sub-tasks. The dynamic programming module of the large model agent optimizes the initial DAG: on the one hand, it builds an extensible structure to support the dynamic adjustment of the subsequent reinforcement learning model. This extensible structure includes task_id (task ID), task_type (task type), pre_conditions (precondition expressions, which can be understood as the logical conditions that must be met before a task node starts execution), and post_actions (trigger actions, which can be understood as the subsequent operations automatically triggered by the system after a task node succeeds or fails). On the other hand, two types of fault-tolerant nodes are injected: checkpoint nodes, which are used to save the current data state of subtasks whose processing time exceeds the preset time, avoiding re-execution after failure. For example, when the compute node executes to 50% (at the 5th minute), the checkpoint node automatically records the storage path of the intermediate calculation results and the current task progress (50%). If the subsequent compute node fails at the 8th minute, the execution engine layer will read the state recorded by the checkpoint node and directly resume execution from the intermediate data at 50% progress, without having to run from the beginning. The failure recovery time is shortened from 10 minutes to less than 5 minutes. Rollback nodes are used to automatically delete the written erroneous data in the case of subtask data writing failure, ensuring data consistency. After these two optimizations, the first DAG corresponding to the target task is finally obtained.
[0028] In this implementation, the semantic parsing module of the large model agent transforms the user's natural language requirements into a structured initial DAG, clarifying the mapping and execution order between subtasks and execution tools, solving the problems of ambiguous requirements and difficulty in direct execution, and making the task process feasible. The scalable structure built by the dynamic programming module of the large model agent provides a foundation for the dynamic adjustment of the DAG in subsequent reinforcement learning layers, breaking through the adaptation limitations of traditional static DAGs. The injection of check nodes and rollback nodes enables automated fault tolerance, avoiding the need to start from scratch after the failure of high-time-consuming subtasks, reducing repeated time consumption, and deleting erroneous data that failed to be written, ensuring data consistency, significantly reducing the cost of manual intervention, and improving the efficiency and stability of task execution.
[0029] Optionally, the step of inputting the user's processing requirements for the target task into the semantic parsing module of the large model agent to obtain the initial DAG corresponding to the target task includes... The user's processing requirements for the target task are input into the semantic parsing module of the large model agent to determine the task details of the target task. The task details include at least one of the following: the task type of the target task, the data source of the target task, the execution tool of the target task, and the output format of the processing result of the target task. Based on the task details, determine the execution tools corresponding to the N sub-tasks of the target task and the execution order of the N sub-tasks; Based on the execution tools corresponding to the N sub-tasks of the target task and the execution order of the N sub-tasks, construct the initial DAG corresponding to the target task.
[0030] In one implementation, the user's processing requirements for the target task are input into the semantic parsing module of the large model agent. The semantic parsing module first processes the input natural language text, focusing on identifying verb phrases. The goal is to understand the type of task the user wishes to complete. Task types include the following: ETL (Extract, Transform, Load) tasks include data extraction, data cleaning, and data transformation. Report generation tasks include: data aggregation, indicator calculation, and data visualization. Data analysis tasks: exploratory data analysis, statistical analysis; Machine learning tasks: feature engineering, model training, and saving; Big data processing tasks: batch processing and real-time data stream processing; Data export task: Export the processed data to a report or persistent storage.
[0031] The identified verb phrases are matched with predefined task type templates, and the task type of the target task is initially determined based on the matching results. Currently supported templates include, but are not limited to: ETL templates: covering data extraction, cleaning, and transformation operations; report generation templates: including data aggregation, indicator calculation, and visualization steps.
[0032] After confirming the task type of the target task, further analyze the detailed elements of the target task, including: Input data source identifier: Clearly define the unique identifier of the data source for the target task, such as the data storage path (e.g., Hadoop Distributed File System (HDFS path)) or message queue topic (e.g., Kafka Topic). Conditional branch triggering rules: Define the tool selection or operation execution logic in specific scenarios, such as automatically enabling the Spark distributed computing engine when the amount of data to be processed exceeds 1TB; Output target constraints: Clearly define the format specifications of the task output, such as report data format, file storage format, and core business requirements, such as Service Level Agreement (SLA) service level time limits, data accuracy standards, etc.
[0033] Next, based on the parsed task details and the registered tools in the system library, a suitable execution tool is selected. Different tasks may correspond to different toolsets. For data cleaning, Pandas or Spark SQL can be chosen. For feature computation, tools such as NumPy or TensorFlow can be used. The tool selection strategy prioritizes tools whose memory consumption is less than 30% of the available resources to ensure system stability and efficiency. Considering reasonable resource allocation and tool compatibility, the optimal tool combination in terms of execution efficiency and resource utilization is selected.
[0034] The output of the semantic parsing module is an initial DAG with semantic annotations. This DAG contains the mapping relationship between the target task and the toolchain, specifically explaining the execution tools and related execution order for each subtask. Combining the parsed task details and format requirements, the most suitable tool is selected to execute each subtask according to business needs. These decisions constitute the nodes in the initial DAG.
[0035] In this implementation, the format requirements and business needs identified by the semantic parsing module provide clear guidance for toolchain selection, ensuring that the selected toolchain can accurately match specific output requirements. On the other hand, it further clarifies the technical specifications of data output, ultimately ensuring that the generated DAG not only fully meets the business scenario requirements but also strictly complies with the technical specifications of data output.
[0036] Optionally, the real-time state information includes numerical feature information and topological feature information, and the reinforcement learning model includes a policy network module; The step of adjusting the nodes and edges of the first DAG based on the real-time state information using the reinforcement learning model to obtain the second DAG includes: The numerical feature information is input into the first channel of the policy network module for mapping processing to obtain the first feature information; The topological feature information is input into the second channel of the policy network for mapping processing to obtain the second feature information; The first feature information and the second feature information are input into the joint decision layer of the policy network to obtain the adjustment action corresponding to the first DAG; Based on the adjustment action, the nodes and edges of the first DAG are adjusted to obtain the second DAG.
[0037] In one implementation, the real-time state information includes numerical feature information, such as CPU utilization and remaining time, as well as topological feature information, such as the node dependencies of the first DAG. The reinforcement learning model includes a policy network module, which comprises a dual-channel structure and a joint decision layer.
[0038] The first channel of the dual-channel structure is the numerical feature channel. A fully connected layer processes numerical metrics such as CPU utilization and remaining time, capturing the quantized state information of the system (i.e., the first feature information). The fully connected layer accepts input features containing four dimensions and maps them to 16 dimensions. This layer is followed by a ReLU activation function to introduce nonlinearity.
[0039] The second channel of the dual-channel structure is the topological feature channel. The node dependencies of the DAG are processed by the Graph Attention Network (GAT) convolutional layer to capture the structural feature information of the task flow (i.e., the second feature information). The node feature dimension processed by this channel is 8, which is mapped to 16 dimensions.
[0040] Then, the first feature information and the second feature information are input into the joint decision layer of the policy network to generate adjustment actions that can be directly applied to the first DAG. The input of the joint decision layer is 32-dimensional (from two 16-dimensional outputs), and the output is 8-dimensional.
[0041] The adjustment actions directly correspond to the optimization direction of task orchestration, ensuring that the strategy is implemented and executable. Specifically: node-level actions include adjusting task parallelism and replacing execution tools; edge-level actions include adding / removing data dependency edges (such as removing redundant dependencies) and inserting cache nodes (such as inserting cache between filter and compute to avoid redundant computation). Then, based on the adjustment actions, the nodes and edges of the first DAG are adjusted to obtain the second DAG.
[0042] In this implementation, real-time state information is categorized into numerical and topological features, and processed separately through a dual-channel mapping mechanism of the policy network. This avoids interference between different types of features, accurately preserves the core information of each feature, and provides high-quality data support for decision-making. The joint decision layer integrates the two types of feature information, ensuring that adjustment actions take into account both real-time numerical states and the DAG topology, resulting in more comprehensive and scientific decision-making and avoiding adjustment biases caused by single features. Based on the adjustment actions generated by accurate feature mapping and comprehensive decision-making, the nodes and edges of the first DAG can be specifically optimized, making the second DAG more adaptable to the dynamic execution environment, improving task execution efficiency, resource utilization, and process stability.
[0043] Optionally, after processing the target task based on the second DAG, the method further includes: Obtain the processing indicator information of the target task, wherein the processing indicator information includes at least one of the following: processing time of the target task, resource consumption during processing of the target task, and data quality during processing of the target task; Based on the processing index information of the target task, the parameters of the policy network module are adjusted according to the multi-objective reward function of the pre-built reinforcement learning model.
[0044] In one implementation, the reinforcement learning model further includes a reward calculator for evaluating and optimizing task scheduling and node operations. This part is designed with a multi-objective reward function, which is as follows: ; Where reward represents the reward function; base represents the baseline processing time; actual represents the processing time of the target task; lim represents the resource overrun ratio, which is obtained based on the resource consumption during the processing of the target task; and data represents the data quality score, which is obtained based on the data quality during the processing of the target task.
[0045] Then, based on the pre-built reinforcement learning model multi-objective reward function, the collected processing index information is used as input, and the task performance is comprehensively evaluated through the reward function. The parameters of the policy network module are adjusted in reverse according to the evaluation results to achieve adaptive iteration of the model.
[0046] In addition, for long-term tasks, a reward delay allocation mechanism with qualification traceback is adopted: during the operation of DAG, the execution status of each node is evaluated in real time by the reward calculator, and the correctness of tool selection and execution efficiency directly determine the reward value; the evaluation information fed back by the reward calculator will indirectly affect the semantic parsing module, providing optimization reference for its subsequent task processing plan.
[0047] In this implementation, the multi-objective reward function balances core business requirements, such as efficiency versus resource conservation and quality versus time consumption, ensuring that the parameter adjustments of the policy network meet the actual business scenario needs, rather than favoring a single objective. Based on metric feedback, the policy network parameters are adjusted in reverse, enabling continuous iterative optimization of the model. This makes subsequent adjustments more precise, further improving the stability of task execution, resource utilization, and business adaptability.
[0048] Optionally, processing the target task based on the second DAG includes: The application interface API call module of the large model agent is used to call the execution tool corresponding to the first subtask based on the second DAG to process the first subtask, where the first subtask is any one of the N subtasks.
[0049] In one implementation, the large model agent also includes an Application Programming Interface (API) module. This module serves as the core hub connecting the large model agent with the underlying execution tools, responsible for achieving seamless integration with the open tool ecosystem. First, tools are registered on the system. Then, metadata such as the tool's input / output modes, resource requirements, and error handling strategies are defined via YAML files. Finally, a multi-objective scoring function (historical success rate × resource matching degree × performance benchmark) is used for real-time optimization. During the execution of the target task, the execution tools for each subtask can be adjusted based on monitored real-time status data.
[0050] In this implementation, API calls are compatible with various execution tools, breaking down barriers to tool usage and enhancing task processing flexibility.
[0051] Optionally, the method further includes: Obtain the priority of the second subtask, where the second subtask is any one of the N subtasks; If the second subtask times out and is not executed, the priority of the second subtask is increased to obtain the first update priority; If the resource requirement of the second subtask is greater than the available resource quantity of the cluster where the target task is located, the priority of the second subtask is reduced to obtain the second update priority. The execution order and resource allocation of the second subtask are adjusted according to the first update priority or the second update priority.
[0052] In one implementation, the initial priority of any subtask (taking the second subtask as an example) among the N subtasks after the target task is broken down is first determined. The initial priority is usually set based on rules such as business importance, task dependency, and preset SLA.
[0053] If the second subtask has reached the preset execution time window, but has not been executed due to resource consumption, scheduling queuing or other reasons, its priority will be automatically increased to generate the first update priority, ensuring that the task is not delayed. If the assessment finds that the total amount of CPU, memory, storage, and other resources required by the second subtask exceeds the remaining available resources in the cluster where the current target task is located, its priority will be reduced to avoid resource contention that could block the entire task, and a second update priority will be generated.
[0054] Based on the first or second update priority generated above, the two core scheduling configurations of the second subtask are further adjusted synchronously: first, the execution order, with high-priority subtasks entering the scheduling queue in advance and being executed first; second, the resource allocation, with high-priority tasks receiving more available resource quotas, and low-priority tasks being reduced as needed or waiting for resource release, thus achieving dynamic adaptation of subtask scheduling.
[0055] In addition, this section can also adopt a resource reservation mechanism to reserve resource slots in advance for critical path tasks.
[0056] In this implementation, priority adjustment is based on real-time execution status and resource status, rather than static preset, which allows the scheduling strategy to be more flexible in adapting to dynamic changes in the cluster.
[0057] Optionally, the method further includes: Calculate the partition data skew of the third subtask, where the partition data skew is the ratio of the maximum partition data size to the average partition data size of the third subtask, and the third subtask is any one of the N subtasks. If the partition data skew of the third subtask is greater than the preset skew, adjust the data volume of each partition of the third subtask until the partition data skew of the third subtask is less than or equal to the preset skew.
[0058] In one implementation, taking the third subtask as an example, the partition data skew of the third subtask is calculated. The partition data skew is calculated as "the maximum partition data size of the subtask ÷ the average partition data size". The larger the ratio, the more uneven the data distribution across partitions. A reasonable skew threshold is preset, such as 3. When the calculated partition data skew of the third subtask exceeds this threshold, it indicates a serious data skew problem, triggering an automatic adjustment mechanism. Through data redistribution, partition splitting / merging, etc., the data size of each partition of the subtask is adjusted, and the adjustment is continuously iterated until the partition data skew drops to the preset threshold or below, ensuring that the data size of each partition tends to be balanced.
[0059] In this implementation, data skew can easily lead to problems such as partition processing timeouts and task failures. By balancing the amount of data in each partition, the risk of such anomalies can be reduced.
[0060] Optionally, the method further includes: Monitor the number of retries for the fourth subtask, which is any one of the N subtasks, and the number of retries is used to indicate the number of times the fourth subtask has been executed; If the number of retries exceeds the preset number, the execution tool corresponding to the fourth subtask will be replaced.
[0061] In one implementation, taking the fourth subtask as an example, the system monitors the number of retries during its execution in real time. This number directly reflects the failure status of the subtask. Each failure and re-execution accumulates one retry count. When the number of retries for the subtask exceeds a preset threshold (e.g., 3 times), it indicates that the currently adapted execution tool may have compatibility or performance issues. In this case, the system will automatically switch the execution tool corresponding to the subtask to retry the subtask execution.
[0062] In this implementation, there is no need for manual troubleshooting of failures and tool replacement, thus automating exception handling and improving the intelligence level of task orchestration.
[0063] See Figure 2 , Figure 2 This is a flowchart of a task processing method provided in an embodiment of this application. The following will be presented through... Figure 2 This application provides a detailed description of the complete workflow of the task processing method. The task processing method is primarily based on a large-scale intelligent agent, a reinforcement learning model, a task execution engine model, and a feedback data channel. Specifically: I. The large-scale intelligent agent comprises a language parsing module, a dynamic programming module, and an API calling module. The user's processing requirements for the target task are input into the semantic parsing module of the large-scale intelligent agent, resulting in an initial Directed Acyclic Graph (DAG) corresponding to the target task. The nodes of the initial DAG represent the mapping relationship between the N subtasks of the target task and their corresponding execution tools, and the edges of the first DAG represent the execution order of the N subtasks. Then, based on the dynamic programming module of the large-scale intelligent agent, an extensible structure is constructed from the initial DAG, and check nodes and rollback nodes are injected into the initial DAG to obtain the first DAG corresponding to the target task. Using the API calling module of the large-scale intelligent agent, based on the second DAG, the execution tool corresponding to the first subtask is called to process the first subtask, which can be any one of the N subtasks.
[0064] II. The reinforcement learning model includes a state feature extraction module, a policy network module, and a reward calculator. The state feature extraction module is used to extract real-time state information corresponding to the processing of the target task based on the first DAG; the policy network module is used to adjust the nodes and edges of the first DAG based on the real-time state information to obtain the second DAG; the reward calculator adjusts the parameters of the policy network module through a multi-objective reward function.
[0065] III. The task execution engine layer includes a resource manager, an adaptive scheduler, and a runtime monitor: 1. In the task execution engine layer, the resource manager plays a core role in managing and scheduling computing resources, enabling efficient task execution in a cluster environment. The following are the functions of the resource manager and its specific role in the system: Resource allocation: Dynamically manage cluster resources, allocating CPU, memory, storage, and network bandwidth according to task scheduling needs to ensure that each task can be executed with the required resources; Load balancing: The resource manager ensures the effectiveness of resource utilization on different nodes and optimizes the overall performance of the cluster by balancing the load; Node monitoring and management: Responsible for real-time monitoring of the status of each node in the cluster (such as health, availability, and resource utilization); Resource priority adjustment: Based on task priority and real-time demand, dynamically adjust the resource priority of tasks in the cluster to ensure that critical tasks have priority access to resources and ensure service quality.
[0066] This component works in conjunction with the adaptive scheduler to dynamically adjust the execution order of tasks and emergency resource allocation strategies. Simultaneously, anomaly detection data from the monitor can trigger the resource manager to reallocate resources, ensuring the smooth implementation of the plan.
[0067] 2. The adaptive scheduler relies on real-time resource data provided by the resource manager (such as the available CPU, memory, and network bandwidth of each node) to formulate scheduling strategies. Only by understanding the resource status can the scheduler make effective task allocation and scheduling decisions. The adaptive scheduler uses a dynamic priority calculation execution process: it defines a `calc_priority` function to calculate task priorities, with the following logic: Base Priority: The function first obtains the base priority of the task, which is derived from the task's SLA urgency level.
[0068] Prioritize Timed-out Tasks: If the current time exceeds the task's estimated start time plus a tolerance threshold, indicating that the task has timed out, its priority is doubled to increase the urgency of processing the task.
[0069] Resource Insufficiency Task Degradation: If a task requires more resources than the current cluster has available, it indicates insufficient resources. The function will reduce the priority of the task, specifically to half of the base priority, to reflect the actual difficulty of its execution.
[0070] This section also employs a resource reservation mechanism, reserving resource slots in advance for critical path tasks.
[0071] 3. The monitor uses anomaly detection rules, and its execution flow is as follows: The detector is used to detect abnormal situations and trigger corresponding processing measures based on specific conditions. The specific logic is as follows: Data skew detection: The function calculates the skewness of the partitioned data, which is the ratio of the largest partition's data size to the average partition's data size. If this ratio exceeds 3.0, it indicates a severely uneven data distribution, and the system will trigger a dynamic repartition operation to rebalance the data allocation.
[0072] Task retry monitoring: The function also monitors the number of retries for the same task. If the number of retries exceeds 3, it indicates that the task is continuously failing, and an alternative tool switching process will be initiated to try to execute the task using other tools, thereby increasing the probability of success.
[0073] Through the above logic, this function can monitor the system status in real time and quickly respond to potential abnormal situations, thereby ensuring the stable execution of tasks.
[0074] IV. Data Feedback Channel The data feedback channel is primarily responsible for collecting and transmitting important data during system operation to support continuous model learning and decision optimization. Specifically, this includes: Model Training Channel: This channel focuses on feeding real-time and historical data into the model to continuously optimize algorithm performance through training. This data includes task execution results, resource usage, and response times, enabling the model to adapt to changing environments and improve decision-making quality.
[0075] Execution Log Database: This section stores log information during task execution, recording the input, output, and execution status of each task. These logs not only provide a basis for subsequent troubleshooting and performance analysis but also provide a large amount of effective data for model training.
[0076] Real-time monitoring and evaluation: By analyzing execution logs and real-time metrics, the system can monitor the effectiveness of the current strategy. If monitoring results show performance degradation or anomalies, it will prompt adjustments to the strategy, which will in turn affect the direction and content of model updates.
[0077] Feedback influences decision-making: Data feedback channels affect upper-level task execution and management. By obtaining feedback in real time, the system can adjust resource allocation strategies based on the latest information, achieving more efficient task scheduling.
[0078] The data feedback channel acts as a bridge in the entire system architecture, supporting not only model learning and optimization but also providing a basis for dynamic adjustments to task decisions. Specifically, the collected data helps identify system bottlenecks, optimize resource allocation, and improve the efficiency and effectiveness of task execution.
[0079] Next, the embodiments of this application will be further described in detail through two specific examples: Example 1: Step 1: Requirement Semantic Parsing (Large Model Agent): 1. Input requirements: Analyze user login logs from the previous week, exclude test accounts, calculate retention rates for each channel, and generate an Excel report before 10:00 AM every day.
[0080] 2. Processing procedure: (1) Identifying key task entities based on Few-shot learning: Time range: last week (automatically converted to a specific date range); Filtering condition: account_type != 'test' (exclude test accounts); Calculation metric: Day N retention rate (automatically matched to the standard calculation formula).
[0081] (2) Generate an initial DAG with semantic labels.
[0082] Step 2: Dynamic Optimization (Reinforcement Learning Model): 1. State feature input: (1) Cluster CPU utilization is 65%; (2) Number of tasks currently pending: 12; (3) Historical time taken for similar tasks: average 38 minutes.
[0083] 2. Policy network decision output: (1) Tool replacement: Replace the execution tool of the compute node from Pandas to Dask (goal: improve parallel computing efficiency). (2) Adding new nodes: Insert a cache node between the filter and compute nodes (goal: avoid redundant calculations and shorten the time consumption).
[0084] Step 3: Task Execution and Model Feedback 1. Monitoring and dynamic adjustment of the execution process: (1) Anomaly detection: Data skew occurs during the filter stage (the maximum partition data volume is 5 times that of the average partition). (2) Automatic processing: Trigger dynamic Repartition operation (by increasing the number of shuffle partitions, the amount of data in each partition is balanced).
[0085] 2. Feedback data collection: (1) Execution efficiency: The actual total time was 42 minutes; (2) Resource consumption: CPU peak 85%; (3) Data quality: retention rate calculation error < 0.1%.
[0086] Model iterative update: 3. Large Model: The optimization decision will be incorporated into the Few-shot example library to supplement small sample data and improve the accuracy of subsequent requirement analysis and DAG generation; 4. Reinforcement learning model: Update policy network weights based on TD-error (temporal difference error) to optimize the rationality of subsequent DAG adjustment decisions.
[0087] Example 2: 1. User input requirements: Process user behavior events every hour to generate feature input recommendation models; if an anomaly occurs, automatically fall back to the Redis cached results.
[0088] 2. Large-scale intelligent agent parsing process: (1) Extraction of key elements: Execution frequency: Executed periodically every hour; Data source: Kafka (user behavior event data); Core tasks: generating features required for the recommendation model and model inference; Exception handling rules: In case of an exception, the result will be rolled back to the Redis cache.
[0089] (2) Automatic tool matching: Data consumption: Flink (real-time reading from Kafka data source); Feature processing: Spark (generates input features for the model); Model inference: PyTorch (load the recommendation model and perform inference); Cache storage: Redis (stores backup cached results).
[0090] 3. Generation of task flow diagram (DAG).
[0091] 4. Key technology optimization solutions: (1) Dynamic parameter adjustment: Batch auto-optimization: Flexible adjustment of batch size based on real-time data volume (e.g., automatically increasing batch size during peak periods to improve processing efficiency). Model elastic scaling: Based on traffic prediction, model service instances are expanded in advance to cope with peak loads.
[0092] (2) Intelligent handling of anomalies: Anomaly detection trigger conditions: Model inference delay > 5 seconds, or service execution failure; Seamless degradation mechanism: Automatically switches to preheated Redis cached data to ensure uninterrupted service.
[0093] (3) Resource optimization strategy: Reinforcement learning dynamic scheduling: Real-time adjustment of computing resource allocation, reducing redundant resource usage by 30%; Cache preheating mechanism: Based on user behavior prediction, hot data is preloaded to Redis.
[0094] Through the above embodiments, the system achieves: 1. Dynamic adaptability: Automatically adjusts computing resources and task parameters based on real-time load to adapt to traffic fluctuations.
[0095] 2. Intelligent degradation capability: Switches to the cache backup path in milliseconds to ensure service availability.
[0096] 3. Resource utilization efficiency: Through reinforcement learning optimization, redundant computing consumption is reduced by more than 20%.
[0097] 4. Fully automated operation and maintenance: From requirement analysis and task orchestration to anomaly recovery, no manual intervention is required throughout the entire process.
[0098] Please see Figure 3 , Figure 3 This is a schematic diagram of a task processing device provided in an embodiment of this application. As shown in the figure, the device 300 includes: The first processing module 310 is used to input the user's processing requirements for the target task into the large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools, and the edges of the first DAG are used to represent the execution order of the N sub-tasks, where N is an integer greater than 1. The first extraction module 320 is used to extract real-time state information based on a reinforcement learning model. The real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. The first adjustment module 330 is used to adjust the nodes and edges of the first DAG based on the real-time state information and using the reinforcement learning model to obtain the second DAG; The second processing module 340 is used to process the target task based on the second DAG.
[0099] Optionally, the first processing module 310 includes: The first processing unit is used to input the user's processing requirements for the target task into the semantic parsing module of the large model agent to obtain an initial DAG corresponding to the target task. The nodes of the initial DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools, and the edges of the initial DAG are used to represent the execution order of the N sub-tasks. The second processing unit is used to construct an extensible structure for the initial DAG based on the dynamic programming module of the large model agent, and inject check nodes and rollback nodes into the initial DAG to obtain a first DAG corresponding to the target task. The check nodes are used to check and save the data of subtasks whose processing time exceeds a preset time, and the rollback nodes are used to delete the data that failed to be written during the processing of the subtask.
[0100] Optionally, the first processing unit includes: The first determining subunit is used to input the user's processing requirements for the target task into the semantic parsing module of the large model agent to determine the task details of the target task. The task details include at least one of the following: the task type of the target task, the data source of the target task, the execution tool of the target task, and the output format of the processing result of the target task. The second determining subunit is used to determine the execution tools corresponding to the N subtasks of the target task and the execution order of the N subtasks based on the task details information. The first construction subunit is used to construct the initial DAG corresponding to the target task based on the execution tools corresponding to the N subtasks of the target task and the execution order of the N subtasks.
[0101] Optionally, the real-time state information includes numerical feature information and topological feature information, and the reinforcement learning model includes a policy network module; The first adjustment module 330 includes: The third processing unit is used to input the numerical feature information into the first channel of the policy network module for mapping processing to obtain the first feature information; The fourth processing unit is used to input the topology feature information into the second channel of the policy network for mapping processing to obtain the second feature information; The fifth processing unit is used to input the first feature information and the second feature information into the joint decision layer of the policy network to obtain the adjustment action corresponding to the first DAG; The first adjustment unit is used to adjust the nodes and edges of the first DAG based on the adjustment action to obtain the second DAG.
[0102] Optionally, the device further includes: The first acquisition module is used to acquire the processing indicator information of the target task, wherein the processing indicator information includes at least one of the processing time of the target task, the resource consumption during the processing of the target task, and the data quality during the processing of the target task. The second adjustment module is used to adjust the parameters of the policy network module based on the processing index information of the target task and the multi-objective reward function of the pre-built reinforcement learning model.
[0103] Optionally, the second processing module 340 includes: The sixth processing unit is used to call the execution tool corresponding to the first subtask based on the second DAG using the application interface API call module of the large model agent. The first subtask is any one of the N subtasks.
[0104] Optionally, the device further includes: The second acquisition module is used to acquire the priority of the second subtask, wherein the second subtask is any one of the N subtasks; The first priority module is used to increase the priority of the second subtask if it times out and is not executed, so as to obtain the first update priority. The first reduction module is used to reduce the priority of the second subtask to obtain a second update priority when the required resources of the second subtask are greater than the available resources of the cluster where the target task is located. The third adjustment module is used to adjust the execution order of the second subtask and the resource allocation of the second subtask according to the first update priority or the second update priority.
[0105] Optionally, the device further includes: The first calculation module is used to calculate the partition data skew of the third subtask, wherein the partition data skew is the ratio of the maximum partition data volume to the average partition data volume of the third subtask, and the third subtask is any one of the N subtasks. The fourth adjustment module is used to adjust the data volume of each partition of the third subtask when the partition data skewness of the third subtask is greater than the preset skewness, until the partition data skewness of the third subtask is less than or equal to the preset skewness.
[0106] Optionally, the device further includes: The first monitoring module is used to monitor the number of retries for the fourth subtask, which is any one of the N subtasks, and the number of retries is used to indicate the number of times the fourth subtask has been executed. The first replacement module is used to replace the execution tool corresponding to the fourth subtask if the number of retries exceeds a preset number.
[0107] The task processing device provided in this application embodiment can achieve... Figure 1The various processes implemented in the method embodiments shown achieve the same technical effects, and will not be described again here to avoid repetition.
[0108] Specifically, see Figure 4 As shown, this application embodiment also provides an electronic device, including a bus 401, an antenna 403, a bus interface 404, a processor 405, and a memory 406. The processor 405 is used for: The user's processing requirements for the target task are input into the large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task. The nodes of the first DAG are used to represent the mapping relationship between the N sub-tasks of the target task and the corresponding execution tools. The edges of the first DAG are used to represent the execution order of the N sub-tasks, where N is an integer greater than 1. Real-time state information is extracted based on a reinforcement learning model, and the real-time state information is the real-time state information corresponding to the processing of the target task based on the first DAG. Based on the real-time state information, the nodes and edges of the first DAG are adjusted using the reinforcement learning model to obtain the second DAG; The target task is processed based on the second DAG.
[0109] exist Figure 4 In this context, a bus architecture (represented by bus 401) is used. Bus 401 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 405 and memory represented by memory 406. Bus 401 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Data processed by processor 405 is transmitted over a wireless medium via antenna 403. Furthermore, antenna 403 also receives data and transmits data to processor 405.
[0110] Processor 405 is responsible for managing bus 401 and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 406 can be used to store data used by processor 405 during operation.
[0111] Alternatively, the processor 405 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0112] Optionally, the processor 405 is further configured to: The user's processing requirements for the target task are input into the semantic parsing module of the large model agent to obtain an initial DAG corresponding to the target task. The nodes of the initial DAG are used to represent the mapping relationship between the N sub-tasks of the target task and their corresponding execution tools, and the edges of the initial DAG are used to represent the execution order of the N sub-tasks. The dynamic programming module of the large model agent constructs an extensible structure for the initial DAG, and injects check nodes and rollback nodes into the initial DAG to obtain the first DAG corresponding to the target task. The check nodes are used to check and save the data of subtasks whose processing time exceeds a preset time, and the rollback nodes are used to delete the data that failed to be written during the processing of the subtask.
[0113] Optionally, the processor 405 is further configured to: The user's processing requirements for the target task are input into the semantic parsing module of the large model agent to determine the task details of the target task. The task details include at least one of the following: the task type of the target task, the data source of the target task, the execution tool of the target task, and the output format of the processing result of the target task. Based on the task details, determine the execution tools corresponding to the N sub-tasks of the target task and the execution order of the N sub-tasks; Based on the execution tools corresponding to the N sub-tasks of the target task and the execution order of the N sub-tasks, construct the initial DAG corresponding to the target task.
[0114] Optionally, the real-time state information includes numerical feature information and topological feature information, and the reinforcement learning model includes a policy network module; The processor 405 is also used for: The numerical feature information is input into the first channel of the policy network module for mapping processing to obtain the first feature information; The topological feature information is input into the second channel of the policy network for mapping processing to obtain the second feature information; The first feature information and the second feature information are input into the joint decision layer of the policy network to obtain the adjustment action corresponding to the first DAG; Based on the adjustment action, the nodes and edges of the first DAG are adjusted to obtain the second DAG.
[0115] Optionally, the processor 405 is further configured to: Obtain the processing indicator information of the target task, wherein the processing indicator information includes at least one of the following: processing time of the target task, resource consumption during processing of the target task, and data quality during processing of the target task; Based on the processing index information of the target task, the parameters of the policy network module are adjusted according to the multi-objective reward function of the pre-built reinforcement learning model.
[0116] Optionally, the processor 405 is further configured to: The application interface API call module of the large model agent is used to call the execution tool corresponding to the first subtask based on the second DAG to process the first subtask, where the first subtask is any one of the N subtasks.
[0117] Optionally, the processor 405 is further configured to: Obtain the priority of the second subtask, where the second subtask is any one of the N subtasks; If the second subtask times out and is not executed, the priority of the second subtask is increased to obtain the first update priority; If the resource requirement of the second subtask is greater than the available resource quantity of the cluster where the target task is located, the priority of the second subtask is reduced to obtain the second update priority. The execution order and resource allocation of the second subtask are adjusted according to the first update priority or the second update priority.
[0118] Optionally, the processor 405 is further configured to: Calculate the partition data skew of the third subtask, where the partition data skew is the ratio of the maximum partition data size to the average partition data size of the third subtask, and the third subtask is any one of the N subtasks. If the partition data skew of the third subtask is greater than the preset skew, adjust the data volume of each partition of the third subtask until the partition data skew of the third subtask is less than or equal to the preset skew.
[0119] Optionally, the processor 405 is further configured to: Monitor the number of retries for the fourth subtask, which is any one of the N subtasks, and the number of retries is used to indicate the number of times the fourth subtask has been executed; If the number of retries exceeds the preset number, the execution tool corresponding to the fourth subtask will be replaced.
[0120] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described task processing method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0121] This application provides a readable storage medium that stores a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the task processing method embodiments described above and achieve the same technical effects. To avoid repetition, further details are omitted here.
[0122] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0123] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0124] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A task processing method characterized by, The method comprises: inputting user processing demand information of a target task to a large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task, wherein a node of the first DAG is used to represent a mapping relationship between N sub-tasks of the target task and corresponding execution tools, and an edge of the first DAG is used to represent an execution order of the N sub-tasks, wherein N is an integer greater than 1; extracting real-time state information based on a reinforcement learning model, wherein the real-time state information is corresponding real-time state information when the target task is processed based on the first DAG; adjusting the nodes and edges of the first DAG based on the real-time state information by using the reinforcement learning model to obtain a second DAG; processing the target task based on the second DAG.
2. The task processing method according to claim 1, characterized by, The method comprises: inputting user processing demand information of a target task to a semantic analysis module of the large model agent to obtain an initial DAG corresponding to the target task, wherein a node of the initial DAG is used to represent a mapping relationship between N sub-tasks of the target task and corresponding execution tools, and an edge of the initial DAG is used to represent an execution order of the N sub-tasks; constructing an expandable structure for the initial DAG based on a dynamic programming module of the large model agent, and obtaining a first DAG corresponding to the target task after injecting a check node and a rollback node into the initial DAG, wherein the check node is used to check and save data of a sub-task whose processing time exceeds a preset time, and the rollback node is used to delete data written unsuccessfully during processing of a sub-task.
3. The task processing method according to claim 2, characterized by, The method comprises: inputting user processing demand information of a target task to a semantic analysis module of the large model agent to determine task detail information of the target task, wherein the task detail information comprises at least one of a task type of the target task, a data source of the target task, an execution tool of the target task, and a processing result output format of the target task; determining an execution tool corresponding to N sub-tasks of the target task and an execution order of the N sub-tasks based on the task detail information; constructing an initial DAG corresponding to the target task according to the execution tool corresponding to the N sub-tasks of the target task and the execution order of the N sub-tasks.
4. The task processing method of claim 1, wherein, The real-time state information comprises numerical feature information and topological feature information, and the reinforcement learning model comprises a policy network module. The method comprises: inputting the numerical feature information into a first channel of the policy network module for mapping processing to obtain first feature information; The topological feature information is input to a second channel of the policy network for mapping processing, to obtain second feature information; The first feature information and the second feature information are input to a joint decision layer of the policy network, to obtain an adjustment action corresponding to the first DAG; The nodes and edges of the first DAG are adjusted based on the adjustment action, to obtain a second DAG.
5. The task processing method according to claim 4, characterized by, After the target task is processed based on the second DAG, the method further comprises: Obtaining processing index information of the target task, the processing index information comprising at least one of processing time consumption of the target task, resource consumption amount during processing of the target task, and data quality during processing of the target task; According to the processing index information of the target task, adjusting parameters of the policy network module based on a multi-objective reward function of a pre-constructed reinforcement learning model.
6. The task processing method of claim 1, wherein, The processing of the target task based on the second DAG comprises: An application program interface (API) calling module of the large model intelligent agent calls an execution tool corresponding to a first subtask to process the first subtask based on the second DAG, the first subtask being any one of the N subtasks.
7. The task processing method according to any one of claims 1 to 6, characterized by, The method further comprises: Obtaining a priority of a second subtask, the second subtask being any one of the N subtasks; In a case where the second subtask is not executed within a timeout period, the priority of the second subtask is raised to obtain a first updated priority; In a case where a required resource amount of the second subtask is greater than an available resource amount of a cluster where the target task is located, the priority of the second subtask is lowered to obtain a second updated priority; The execution order of the second subtask and the resource allocation amount of the second subtask are adjusted according to the first updated priority or the second updated priority.
8. The task processing method according to any one of claims 1 to 6, characterized by, The method further comprises: Calculating a partition data skew of a third subtask, the partition data skew being a ratio of a maximum partition data amount to an average partition data amount of the third subtask, the third subtask being any one of the N subtasks; In a case where the partition data skew of the third subtask is greater than a preset skew, the data amount of each partition of the third subtask is adjusted until the partition data skew of the third subtask is less than or equal to the preset skew.
9. The task processing method of claim 1, wherein, The method further comprises: Monitoring a retry number of a fourth subtask, the fourth subtask being any one of the N subtasks, the retry number indicating the number of executions of the fourth subtask; In a case where the retry number is greater than a preset number, an execution tool corresponding to the fourth subtask is replaced.
10. A task processing apparatus characterized by comprising: The device comprises: The first processing module is configured to input user processing requirement information for a target task into a large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task, wherein a node of the first DAG is configured to represent a mapping relationship between N sub-tasks of the target task and corresponding execution tools, and an edge of the first DAG is configured to represent an execution order of the N sub-tasks, where N is an integer greater than 1. The first extraction module is configured to extract real-time state information based on a reinforcement learning model, wherein the real-time state information is corresponding real-time state information when the target task is processed based on the first DAG. The first adjustment module is configured to adjust nodes and edges of the first DAG based on the real-time state information by using the reinforcement learning model to obtain a second DAG. The first processing module is configured to process the target task based on the second DAG.
11. An electronic device, comprising: The processor is configured to: input user processing requirement information for a target task into a large model agent to obtain a first directed acyclic graph (DAG) corresponding to the target task, wherein a node of the first DAG is configured to represent a mapping relationship between N sub-tasks of the target task and corresponding execution tools, and an edge of the first DAG is configured to represent an execution order of the N sub-tasks, where N is an integer greater than 1; extract real-time state information based on a reinforcement learning model, wherein the real-time state information is corresponding real-time state information when the target task is processed based on the first DAG; adjust nodes and edges of the first DAG based on the real-time state information by using the reinforcement learning model to obtain a second DAG; process the target task based on the second DAG.
12. An electronic device, comprising: The computer program is stored on the computer readable storage medium and is executed by the processor to implement the steps of the task processing method according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium and is executed by the processor to implement the steps of the task processing method according to any one of claims 1 to 9.
14. A computer program product, characterised in that, The computer program is stored on the computer readable storage medium and is executed by the processor to implement the steps of the task processing method according to any one of claims 1 to 9.