A data processing method based on SQL script

By building a task dependency graph and applying topological sorting, dynamically adjusting parallelism and resource allocation, combined with game theory and queuing theory models, the problems of unreasonable task execution order and low resource utilization in complex data processing systems are solved, and efficient and fast data processing is achieved.

CN120123062BActive Publication Date: 2025-08-08COASTAL RONGXIN (BEIJING) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510217976.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-08-08
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing technology fails to fully consider the dependencies and resource allocation between tasks in complex and large-scale data processing systems, resulting in unreasonable task execution order and low resource utilization, which makes it impossible to effectively deal with high concurrency and high load conditions, and the system response time is long.

Method used

By building a task dependency graph, the loop is adjusted using topological sorting and depth-first search algorithms, the task parallelism and resource allocation are dynamically adjusted, resource allocation is optimized in combination with game theory models, and the queuing theory model is used to optimize the processing rate of real-time data flow tasks.

Benefits of technology

It realizes precise control of task execution order, reduces waiting time and resource conflicts, improves system processing efficiency and response speed, and ensures that the system operates efficiently under different loads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123062B_ABST
    Figure CN120123062B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology and discloses a data processing method based on SQL scripts, comprising the following steps: Step 1: constructing a task dependency graph, wherein each SQL task is regarded as a node and the data dependencies between tasks are directed edges; Step 2: topologically sorting the task dependency graph. During the topological sorting process, if there are loops in the task dependencies, a depth-first search algorithm is used to detect and adjust the loops, or the loops are eliminated by task decomposition; Step 3: based on the execution order of the topological sorting, the execution time of each task is calculated and the execution order of each task is adjusted. The present invention adopts a topological sorting optimization technology based on the task dependency graph, achieving the technical effect of accurately controlling the execution order of tasks and reducing waiting time. Through in-depth analysis of task dependencies, redundant calculations and resource conflicts in task execution are avoided, effectively shortening the overall processing time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a data processing method based on SQL scripts. Background Art

[0002] With the rapid development of information technology, especially in the fields of big data, cloud computing, and real-time data processing, the requirements for data processing systems are becoming increasingly stringent. In these application scenarios, data volumes are enormous, processing tasks are complex, and real-time performance is critical. Therefore, improving task scheduling efficiency and optimizing resource allocation are key to enhancing system performance and responsiveness. To this end, various optimization techniques have been proposed and applied to data processing, including task scheduling, resource allocation, and real-time data stream processing.

[0003] Currently, many existing technologies focus primarily on single-task scheduling and static resource allocation, often relying on simple sequential execution or resource scheduling methods based on preset rules. While these technologies can meet requirements in certain simple application scenarios, they struggle to achieve optimal performance in complex, large-scale systems due to inter-task dependencies and the dynamic changes in system resources. Existing scheduling algorithms often ignore inter-task dependencies or fail to flexibly adjust parallelism and resource allocation based on real-time load. This results in low system efficiency, long response times, and low resource utilization under high concurrency or high load conditions.

[0004] The existing technology has the following problems: First, many task scheduling schemes fail to fully consider the dependencies between tasks, resulting in an unreasonable order of task execution, increased waiting time between tasks, and redundant calculations. Second, resource allocation is usually static and cannot be dynamically adjusted according to the real-time requirements of the task and the load conditions of the system. Therefore, system resources cannot be effectively utilized, resulting in low task execution efficiency. Furthermore, there is a lack of dynamic scheduling mechanisms for real-time data stream tasks. In high-concurrency situations, tasks may wait in line for a long time, resulting in system response delays. Therefore, there is still much room for improvement in the existing technology in optimizing task execution order, parallelism, and resource allocation. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a data processing method based on SQL scripts, which solves the problem of improving system processing efficiency and response speed by optimizing task execution sequence, parallelism and resource allocation in complex data processing tasks.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data processing method based on SQL scripts, comprising the following steps:

[0007] Step 1: Build a task dependency graph, where each SQL task is considered a node and the data dependencies between tasks are directed edges.

[0008] Step 2: Topologically sort the task dependency graph. If there are loops in the task dependency during the topological sorting process, the depth-first search algorithm will detect and adjust the loops, or eliminate the loops by task decomposition.

[0009] Step 3: Based on the execution order of topological sorting, calculate the execution time of each task and adjust the execution order of each task;

[0010] Step 4: Dynamically adjust the parallelism and computing resource allocation of each task and adjust the total execution time of all tasks;

[0011] Step 5: Based on the game theory model, calculate the resource allocation strategy for each task. Model the task scheduling problem as a non-cooperative game. Each task selects the optimal resource allocation strategy based on the resource allocation of other tasks. The resource allocation strategy is adjusted using the Nash equilibrium theory of game theory.

[0012] Step 6: Apply queuing theory models to real-time data stream tasks and adjust the data processing rate.

[0013] Preferably, in step 1, constructing a task dependency graph includes:

[0014] Treat each SQL task as a node, and the data dependencies between tasks as directed edges to construct a directed acyclic graph.

[0015] In the directed acyclic graph, edges represent dependencies between tasks, and nodes represent SQL query tasks.

[0016] Preferably, each SQL task is considered as a node, and the data dependencies between tasks are considered as directed edges to construct a directed acyclic graph;

[0017] In the directed acyclic graph, edges represent dependencies between tasks, and nodes represent SQL query tasks.

[0018] Preferably, in step 3, the execution time optimization includes:

[0019] Topologically sort the task dependency graph to ensure that each task is executed in sequence;

[0020] The execution order of each task is determined through the results of topological sorting, and the execution time between tasks is optimized.

[0021] Preferably, the dynamic adjustment of task parallelism and computing resource allocation includes:

[0022] The dynamic adjustment of task parallelism and computing resource allocation includes:

[0023] A nonlinear optimization model is used to calculate the optimal parallelism and computing resources for each task to minimize the total execution time of all tasks. There is a nonlinear relationship between the execution time of a task and its parallelism and computing resources, and there is a nonlinear relationship between the task execution time and parallelism and resource allocation.

[0024] Preferably, the step of minimizing the total task execution time includes:

[0025] Adjust the parallelism and computing resource allocation of each task based on the execution time model of each task;

[0026] The optimization process adjusts the execution time of each task.

[0027] Preferably, in step 5, the calculation step of the game theory model includes:

[0028] The task scheduling problem is modeled as a non-cooperative game. Each task selects a resource allocation strategy based on the resource allocation of other tasks. The payoff function of each task is the ratio of its execution time to the allocated resources. The resource allocation strategy is adjusted through the Nash equilibrium theory of game theory.

[0029] Preferably, in step six, the step of applying the queuing theory model includes:

[0030] Model the real-time data flow task as an M / M / 1 queue system;

[0031] The task processing rate is dynamically adjusted according to the arrival rate and service rate of the real-time data stream.

[0032] Preferably, in step 2, the shortest path is calculated by topological sorting of the task dependency graph or a topological sorting algorithm is used to optimize the task execution order.

[0033] In step 4, the optimization step of computing task parallelism and resource allocation is solved by nonlinear optimization, using the Lagrange multiplier method, considering the constraints of computing resources, to obtain the optimal parallelism and resource allocation for each task.

[0034] Preferably, the queuing theory model optimization step in the real-time data stream processing is performed by adjusting the ratio of the service rate to the arrival rate.

[0035] The present invention provides a data processing method based on SQL scripts. It has the following beneficial effects:

[0036] 1. This invention utilizes a topological sorting optimization technique based on a task dependency graph, achieving the technical effect of precisely controlling the order of task execution and reducing waiting time. Compared to existing approaches that simply schedule tasks sequentially, this invention avoids redundant computation and resource conflicts during task execution through in-depth analysis of task dependencies, effectively shortening overall processing time.

[0037] 2. By introducing a nonlinear optimization model for parallelism and resource allocation, this invention maximizes system processing power by rationally allocating computing resources and parallelism. Compared to existing approaches that use static resource allocation or a single optimization strategy, this invention dynamically adjusts based on task complexity and real-time resource availability, ensuring efficient system operation under varying loads.

[0038] 3. This invention utilizes queuing theory models to optimize the processing rate of real-time data stream tasks, achieving the technical effect of reducing task wait times and improving system response speed under high concurrency conditions. Compared to existing solutions that lack dynamic adjustment mechanisms, this invention can adjust the service rate in real time based on system load and task priority, avoiding resource bottlenecks and processing delays in high-load systems.

[0039] 4. This invention combines game theory models to optimize resource allocation strategies, achieving the technical effect of intelligently scheduling system resources and improving task execution efficiency. Compared to the simple resource allocation methods used in existing technologies, this invention uses non-cooperative games and Nash equilibrium theory to select the optimal resource allocation strategy for each task, ensuring the rational sharing of resources between tasks and maximizing overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] Please see the attached Figure 1 , an embodiment of the present invention provides a data processing method based on SQL scripts, including:

[0043] Construction of task dependency graph

[0044] The construction of a task dependency graph clearly defines the dependencies between tasks and provides a foundation for subsequent task scheduling, execution sequence optimization, and resource allocation. The task dependency graph helps the system accurately identify the execution order between tasks, avoiding conflicts and redundant operations during execution. By constructing a reasonable task dependency graph, we can ensure that every task in the data processing flow is executed in the correct order, thereby improving overall processing efficiency.

[0045] In this embodiment, the task dependency graph is constructed using a directed acyclic graph (DAG) structure. A DAG is a graph structure in which each node represents a SQL task and each directed edge represents the dependency between tasks. Specifically, each task has a corresponding node in the graph, and the execution order and data flow dependencies between tasks are connected by directed edges. This structure ensures that tasks are executed in the order of their dependencies while also avoiding the occurrence of circular dependencies.

[0046] SQL task dependency graph construction

[0047] When building the task dependency graph, each SQL task must first be converted into a node. Each node contains task-related information, such as the task ID, required computing resources, and execution time.

[0048] The data flow relationship between tasks is represented by directed edges. A directed edge from task A to task B means that the execution of task B depends on the output of task A.

[0049] Typically, the dependency between tasks is unidirectional, meaning that a task can only depend on the execution results of the previous task. This unidirectional dependency is usually explicitly expressed in the form of directed edges.

[0050] Dependencies between tasks are not just sequential. Some tasks may need to rely on the results of multiple tasks before they can begin their own execution. In this case, a node in the graph will have multiple incoming edges, indicating that the task depends on multiple predecessor tasks.

[0051] Alternatively, a task dependency graph can be constructed by analyzing the relationships between SQL script operations such as SELECT, INSERT, and UPDATE. For example, the input data for a query operation may come from multiple tables. The connections between these tables are represented in the graph through edges. Each edge represents the flow of tasks from one task to the next.

[0052] Specifically, the process of building a task dependency graph includes the following steps:

[0053] 1. Identify the execution relationships of SQL tasks: The system parses SQL scripts and identifies dependencies between tasks by analyzing the input and output relationships of each SQL task. Input data comes from the output of the predecessor task, and output data serves as input for subsequent tasks.

[0054] Build a directed acyclic graph (DAG): In a task dependency graph, tasks are connected by edges, where each edge represents the output data of one task serving as the input for the next. The graph is free of loops, ensuring unidirectional data flow and avoiding circular dependencies when executing tasks.

[0055] 3. Marking task information: Each task node will contain necessary execution information, such as the computing resources required for task execution, execution time, priority, etc. This will provide support for subsequent execution order optimization.

[0056] 4. Determine inter-task dependencies: The key to a task dependency graph is correctly expressing inter-task dependencies. The system analyzes the execution order and data flow of SQL scripts to determine inter-task dependencies and represents them in the graph through directed edges.

[0057] In one possible implementation, if a circular dependency exists between nodes in the task dependency graph, the system will detect the loop using a depth-first search (DFS) algorithm. Once a loop is detected, the system prompts the user to make manual adjustments or automatically splits the task using task decomposition to eliminate the circular dependency. Task decomposition can break down complex SQL tasks into multiple subtasks, each with a clear input-output relationship, eliminating the original circular dependency problem.

[0058] During implementation, a task dependency graph can also be automatically generated by reasoning about the execution order of SQL scripts. The system first parses the SQL script, identifies the input-output relationships of each operation, and constructs a task dependency graph based on these relationships. This allows the system to automatically identify the execution order of tasks within the SQL script.

[0059] Specifically, the direction of each edge in the task dependency graph represents the direction of data flow. During execution, the system ensures that each task begins only after all predecessor tasks have completed, based on the topological order in the dependency graph. For nodes in the graph with multiple incoming edges, this indicates that the task must wait for all predecessor tasks represented by these incoming edges to complete before executing.

[0060] In one possible implementation, the task dependency graph is constructed automatically. The system parses SQL scripts or other task definition files, automatically identifies inter-task dependencies, and generates a corresponding DAG. Each node in the graph represents a SQL task, and edges represent data dependencies between tasks. The graph generation process also considers factors such as task execution order, resource consumption, and data flow between tasks.

[0061] For each task T i , its execution time is T i (p i ,r i ) is expressed by the following formula:

[0062]

[0063] in:

[0064] p i Represents task T i The degree of parallelism;

[0065] r i Represents task T i Required computing resources;

[0066] c i It is task T i An intrinsic constant representing the computational effort of the task;

[0067] f(r i ) is a resource scheduling function, which can usually be a linear or logarithmic function, representing the relationship between the execution time of a task and resources.

[0068] In the diagram, the execution time and dependencies of task nodes jointly influence the execution order of subsequent tasks. When optimizing the execution order of tasks, we must not only consider the dependencies between tasks but also comprehensively consider the execution time of tasks, and rationally arrange parallelism and resource allocation to reduce the overall execution time.

[0069] In some embodiments, the task dependency graph can be dynamically adjusted based on system load and execution efficiency. When certain tasks are complex, the system can prioritize those with lower resource requirements and shorter execution times, thereby improving overall execution efficiency.

[0070] By constructing a reasonable DAG and combining task execution time and resource requirements, the system can efficiently schedule the execution order of tasks. By optimizing the task nodes and dependencies in the graph, tasks can be executed in the optimal order, reducing computational redundancy and improving the overall processing efficiency of the system.

[0071] Topological sorting optimizes execution order

[0072] Now that the task dependency graph (DAG) has been constructed and the execution dependencies between tasks have been clearly defined, the tasks need to be sorted optimally to ensure they are executed in the optimal order and minimize overall execution time. Topological sorting is key to this process, ensuring that each task is executed only after its dependent tasks have completed, thus avoiding execution conflicts and resource contention.

[0073] In this embodiment, topological sorting is used to determine the order in which tasks are executed, optimizing resource consumption and time delays during execution. Topological sorting ensures that dependencies between tasks are correctly handled by traversing the task dependency graph. In a graph structure, each task depends on the execution results of the previous task, and topological sorting is an effective method for resolving these dependencies.

[0074] In its implementation, the system first traverses the task dependency graph using a depth-first search (DFS) or breadth-first search (BFS) algorithm to determine the execution order. When performing topological sorting, the execution order of each task must consider not only the dependencies between tasks but also the execution time of each task to further optimize the use of system resources.

[0075] In general, the topological sorting step in the task dependency graph is optimized in the following ways:

[0076] Dependency graph construction: By analyzing the input-output relationships between each SQL task, a task dependency graph is constructed. In the graph, the dependencies between tasks are represented by directed edges, with the direction of the edges indicating the order in which the tasks are executed. The nodes in the graph represent SQL tasks, while the edges represent the data flow between tasks.

[0077] Topological sorting process: Topological sorting traverses the task dependency graph to ensure that the predecessor tasks of each task are executed first. This step ensures that tasks are executed in the order of their dependencies, thus avoiding data conflicts and execution errors.

[0078] Optimizing task execution time: Topological sorting not only focuses on the order of tasks, but also needs to optimize task execution time. In one possible implementation, the execution time of a task is related to the execution time of its dependent predecessor tasks. Therefore, the topological sorting process may need to adjust based on execution time to ensure that the execution order of tasks minimizes the total execution time.

[0079] Specifically, tasks typically execute only after all predecessor tasks have completed. Task execution time is closely related to system resource allocation. During task scheduling, a task may not execute immediately due to insufficient resources or incomplete task dependencies. Therefore, optimizing task execution order must also consider the proper allocation of task execution time and resources.

[0080] In some embodiments, the optimization of task execution order also involves dynamic adjustment of task parallelism. The system can dynamically adjust the parallelism p according to the task dependencies and execution time. i and computing resources i , in order to minimize the execution time. The optimization formula here can be expressed as:

[0081]

[0082] in:

[0083] T i (p i ,r i ) is task T i The execution time depends on the parallelism p i and resource allocation i ;

[0084] c i is the computational constant of the task;

[0085] f(r i ) is the resource scheduling function, representing the relationship between task execution time and resources. During task scheduling, the system optimizes this formula, adjusting each task's parallelism and resource allocation to minimize overall execution time. This process is closely related to topological sorting, which provides the foundation for optimization.

[0086] Alternatively, topological sorting can be combined with other sorting algorithms. For example, topological sorting based on a shortest path algorithm can further optimize the order of task execution. By selecting a shortest path algorithm, the system can optimize task execution time and improve task efficiency.

[0087] In one possible implementation, the task dependency graph is dynamically adjusted based on the task's computational resource requirements. In this approach, topological sorting not only prioritizes the order in which tasks are executed, but also dynamically adjusts the order based on information such as the system's real-time load and resource allocation. This approach maximizes resource utilization during task execution, effectively reducing overall execution time.

[0088] For example, when a task requires fewer computing resources, the system can prioritize its execution. For tasks with higher computing resource requirements, the system adjusts their execution order based on resource availability, thus avoiding resource competition and task execution blockage. Through this dynamic adjustment, the system can achieve optimal resource allocation and further optimize task execution time.

[0089] During implementation, topological sorting optimization not only considers the order of tasks but also changes in the execution environment. For example, the execution time of tasks may vary depending on changes in system load, so the execution order of topological sorting may need to be adjusted in real time based on the system state. Specifically, when the system load is low, some tasks can be executed in parallel. When the system load is high, the system can adjust the execution order of tasks to avoid resource competition and conflicts between tasks.

[0090] Through these optimizations, topological sorting not only ensures the correctness of inter-task dependencies but also further reduces total execution time by dynamically adjusting parallelism and resource allocation. In practice, the system can optimize the execution order of tasks in real time based on their characteristics and system status, achieving optimal data processing performance.

[0091] Topological sorting optimizes execution order by clarifying the dependencies between tasks and optimizing their execution order. This effectively reduces redundant computations, shortens execution time, and improves system resource utilization. By dynamically adjusting task parallelism and resource allocation, this method enables efficient data processing in large-scale data processing and high-concurrency environments.

[0092] Execution time optimization

[0093] In the present invention, execution time optimization is used. As the number of tasks increases, relying solely on sequential execution can lead to wasted resources, excessive waiting times, and reduced overall execution efficiency. Therefore, the present invention further improves the system's processing capabilities through execution time optimization, especially in environments with large-scale data and high concurrency. The goal of execution time optimization is to reasonably adjust the execution order of tasks based on the dependencies between tasks, ensuring that each task is executed in parallel as much as possible while avoiding resource conflicts and excessive waiting times.

[0094] In this embodiment, execution time optimization is achieved by analyzing the task dependency graph and the execution time of the tasks. In the aforementioned steps, the task dependency graph has been constructed, and the execution order of the tasks has been determined by topological sorting. During the execution time optimization process, the system first calculates the execution time of each task and, based on this, calculates the optimal execution order of the tasks. The execution time of a task is not only related to the computational complexity of the task, but also to the system's resource allocation and the parallelism of the task. Therefore, execution time optimization not only takes into account the execution order of the tasks, but also involves the adjustment of the parallelism of the tasks and resource allocation.

[0095] In general, the execution time of a task can be expressed as the following formula:

[0096]

[0097] in:

[0098] T i (p i ,r i ) represents task T i The execution time depends on the parallelism of the task p i and resources i ;

[0099] c i Is related to task T i Inherent computationally expensive constants;

[0100] f(r i ) is a resource scheduling function that represents the relationship between task execution time and allocated resources. Specifically, f(r i ) can be designed according to the resource allocation of the system, usually a linear or logarithmic function. In the process of execution time optimization, the system calculates the execution time T of each task. i (p i ,r i ), and optimize the execution order of tasks based on these calculation results. Specifically, the system will try to adjust the parallelism p of the tasks i and computing resources i ,to minimize the total execution time.,The execution time of each task has a nonlinear relationship with,resource allocation and parallelism.,Thus, the optimization process needs to consider multiple factors,such as the dependency between tasks, parallelism and resource,limitations.

[0101] Execution time optimization can be achieved through a nonlinear optimization model. The optimization goal is to minimize the total execution time of all tasks, which can be expressed by the following formula:

[0102]

[0103] Where n is the total number of tasks, T i (p i ,r i ) represents each task T i The execution time, p i and r i Represents tasks T i parallelism and resource allocation.

[0104] By solving this optimization problem, the system can calculate the optimal parallelism and resource allocation for each task, thereby minimizing the overall execution time.

[0105] Specifically, during the optimization process, the system first analyzes the execution time and resource requirements of each task and evaluates the parallelism relationships between tasks. For each task, the system determines its parallelism and resource allocation based on its dependencies and execution time. The system also considers resource contention between tasks to ensure that resource overallocation and task execution conflicts are avoided.

[0106] In one possible implementation, the system can also dynamically adjust parallelism and resource allocation based on task priority and resource load. For example, some compute-intensive tasks may require more resources and lower parallelism, while some I / O-intensive tasks may be more suitable for higher parallelism and fewer resources. In this way, the system can dynamically adjust the execution order of tasks based on their characteristics and system load, further improving overall execution efficiency.

[0107] During implementation, execution time optimization not only considers the order of task dependencies but also requires the rational planning of computing resources during task execution. Generally, there is a positive correlation between task execution time and the required computing resources. To prevent tasks from waiting for resources for extended periods, the system schedules resources and prioritizes them based on their computing resource requirements, thereby reducing wait times and improving execution efficiency.

[0108] Execution time optimization can be further optimized by combining it with task scheduling algorithms. During task scheduling, the system dynamically adjusts the execution order of tasks based on their computing resource requirements, execution time, and task dependencies, and adjusts their execution priority based on the availability of system resources. This allows the system to achieve more efficient resource utilization and reduce overall task execution time.

[0109] Execution time optimization also takes into account inter-task execution latency and system load. When the system load is low, tasks execute faster, and the system can increase task parallelism to shorten overall execution time. Conversely, when the system load is high, the system automatically reduces task parallelism and reschedules tasks based on resource availability. This ensures that the system maintains efficient operation under varying load conditions.

[0110] Execution time optimization also involves intelligently scheduling system resources. When optimizing task execution, the system also adjusts the execution order and resource allocation based on the current state of the system's resources. By monitoring and adjusting system resources in real time, the system maintains high processing efficiency even in complex environments such as high concurrency and large data processing.

[0111] Formula: To further illustrate the execution time optimization process, the following formula is used to describe the relationship between task execution time, parallelism, and resources:

[0112]

[0113] in:

[0114] p i It is task T i The degree of parallelism;

[0115] r i It is task T i resource allocation;

[0116] c i It is task T i Inherent computational effort;

[0117] f(r i ) is a resource scheduling function that represents the relationship between resources and task execution time, which can usually be a linear or logarithmic function.

[0118] Execution time optimization: By rationally scheduling task execution, adjusting task parallelism, and optimizing resource allocation, the system enables more efficient data processing in scenarios with high concurrency and complex task scheduling. During implementation, the system dynamically optimizes the execution order of tasks based on their execution time, dependencies, and resource requirements, minimizing total execution time and ensuring efficient and accurate task completion.

[0119] Parallelism and resource allocation optimization

[0120] In the previous steps, the task dependency graph was constructed, and the execution order of tasks was optimized through topological sorting. Further optimization will focus on task parallelism and the proper allocation of computing resources, which will directly impact task execution efficiency. By optimizing parallelism and resource allocation, the system can reduce total execution time while avoiding excessive resource consumption and resource contention among tasks.

[0121] In this embodiment, the system uses a nonlinear optimization model to optimize parallelism and resource allocation. i and computing resources i The adjustment is based on the execution time model of each task. The execution time of a task is not only related to its parallelism, but also closely related to the allocated computing resources. In order to minimize the total execution time, the system needs to dynamically adjust the parallelism and resource allocation of each task based on the characteristics of the task, its dependencies, and the availability of computing resources. Specifically, the execution time of each task T i (p i ,r i ) can be expressed by the following formula:

[0122]

[0123] in:

[0124] T i (p i ,r i ) represents task T i The execution time depends on the parallelism of the task p i and resources i ;

[0125] c i It is task T i The inherent computational constant usually reflects the basic computation time required to execute the task;

[0126] p i It is task T i The degree of parallelism, that is, the number of computing units on which tasks can be processed simultaneously;

[0127] f(r i ) is a resource scheduling function that describes the relationship between resource allocation and task execution time.

[0128] Usually f(r i ) can be a linear function or a logarithmic function, and the function form will vary depending on how resources are allocated.

[0129] Generally speaking, the relationship between task execution time, parallelism, and resources is nonlinear. Higher task parallelism can reduce execution time inversely. However, if resources are insufficient, task execution time will be constrained by resource allocation. Therefore, the system needs to consider inter-task dependencies, resource constraints, and parallelism to calculate the optimal parallelism and resource allocation for each task.

[0130] As an option, the system can use a nonlinear optimization algorithm to solve the optimal parallelism p for each task. i and resources i The optimization goal is to minimize the total execution time of the task, that is,

[0131]

[0132] Where n is the total number of tasks, T i (p i ,r i ) represents task T i The execution time, p i and r i Represents tasks T i By solving this optimization problem, the system can obtain the optimal parallelism and resource allocation for each task, thereby minimizing the overall execution time.

[0133] The system will also dynamically adjust the parallelism and resource allocation of tasks based on the current system load and resource conditions. For example, if a task encounters resource contention during execution, the system will reduce the parallelism of the task. i , to avoid system overload and ensure that the task can be completed smoothly. When resources are idle, the system can increase the parallelism p i , improve processing efficiency.

[0134] Specifically, the system will dynamically calculate the resources required for each task based on the computing requirements of the task. i And adjust the parallelism p i . For example, computationally complex tasks may require more computing resources, and the degree of parallelism may also need to be higher to reduce execution time. For simpler tasks, the system will allocate fewer resources and reduce the waste of computing resources by reducing the degree of parallelism. During the implementation process, tasks with higher dependencies will be given priority to ensure that the execution of these tasks is not affected by parallelism or resource allocation. For tasks without dependencies, the system can flexibly adjust their parallelism and resource allocation so that these tasks can be completed as soon as possible. The system will intelligently adjust the parallelism and resource allocation of tasks based on the priority, resource requirements and execution time of the tasks to minimize the overall execution time.

[0135] While dynamically adjusting parallelism and resource allocation, the system also uses adaptive mechanisms to monitor task execution progress in real time. If certain tasks are taking too long to execute, the system automatically identifies and adjusts their resource allocation to improve task processing speed. Furthermore, if resource bottlenecks occur, the system prioritizes computing resources for high-priority tasks to ensure the timely completion of critical tasks.

[0136] In some embodiments, to further improve resource utilization, the system can also analyze historical task execution data. In this way, the system can predict the task execution time and resource requirements based on the task's historical execution time, resource requirements, and other information, thereby making more accurate resource allocation and parallelism adjustment decisions.

[0137] This invention significantly improves the efficiency of task execution. While ensuring the correctness of inter-task dependencies, it rationally allocates resources and optimizes task parallelism, making data processing more efficient. Furthermore, the system can adapt to varying load conditions, dynamically adjusting parallelism and resource allocation to ensure optimal performance in all environments.

[0138] Game Theory Models and Resource Allocation

[0139] In the previous steps, we optimized the execution order of tasks by constructing a task dependency graph and performing topological sorting. However, further down the execution process, task parallelism and resource allocation are crucial for maximizing processing efficiency. By properly adjusting parallelism and resource allocation, we can effectively reduce execution time, avoid resource conflicts, and ensure optimal use of system resources when executing multiple tasks simultaneously.

[0140] The optimization of parallelism and resource allocation is achieved through a nonlinear optimization model. The system dynamically calculates the optimal parallelism and resource allocation for each task based on the execution time of each task, the computing resource requirements, and the dependencies between tasks. The execution time of each task is not only related to its computational complexity, but also closely related to the resources available to the task (such as CPU, memory, etc.) and the parallelism of the task. The optimization goal is to minimize the overall execution time of the system by reasonably arranging the task parallelism and computing resources, while avoiding excessive resource competition and waste. The execution time of the task T i (p i ,r i ) can be calculated using the following formula:

[0141]

[0142] in:

[0143] (p i ,r i ) represents task Ti The execution time depends on the parallelism p of the task i and resources i ;

[0144] c i It is task T i The inherent computational constant, representing the computational requirements of the task, p i It is task T i The degree of parallelism, which represents the number of processing units on which tasks can be performed simultaneously;

[0145] f(r i ) is a resource scheduling function, which represents the relationship between the execution time of a task and the resources allocated to the task. i The function may be in the form of a linear function, a logarithmic function, or other forms, depending on the system's resource allocation strategy.

[0146] In the actual task scheduling process, the parallelism of the task p i and resources i It will directly affect the execution time of the task. Therefore, a reasonable parallelism and resource allocation strategy can effectively reduce the execution time.

[0147] As an option, the system may use nonlinear optimization methods to calculate the optimal value when optimizing parallelism and resource allocation. The optimization goal of the system is to minimize the total execution time of all tasks, which can be expressed by the following formula:

[0148]

[0149] in:

[0150] n is the total number of tasks;

[0151] T i (p i ,r i ) is task T i The execution time depends on the parallelism of the task p i and resources i .

[0152] By solving this optimization problem, the system can calculate the optimal parallelism and resource allocation for each task. The system will prioritize resources for tasks that require greater computing resources and dynamically adjust the parallelism and resources of tasks based on their execution time, resource requirements, and task dependencies.

[0153] Specifically, during the optimization process, the system needs to consider the dependencies between tasks. i Depends on task T j The output of task T iMust wait for task T j The system must optimize the parallelism and resources of each task before it can be executed. Therefore, the system must not only optimize the parallelism and resources of each task, but also ensure that the execution order of tasks is not disrupted. By topologically sorting task dependencies, the system can clearly understand the execution order of tasks and further adjust the parallelism and resource allocation of tasks based on this.

[0154] In one possible implementation, to further improve execution efficiency, the system can dynamically adjust the degree of parallelism and resource allocation based on the characteristics of the task and the execution environment. For example, for compute-intensive tasks, the system might choose a lower degree of parallelism and allocate more computing resources. For I / O-intensive tasks, the degree of parallelism can be appropriately increased to improve overall data processing speed.

[0155] Task execution time is affected not only by parallelism and resource allocation, but also by system load and task scheduling strategies. To address this, the system monitors system load in real time and adjusts parallelism and resource allocation accordingly. When system load is low, the system increases task parallelism to improve processing speed. When system load is high, the system reduces task parallelism and adjusts resource allocation to avoid excessive resource competition.

[0156] When the system load is high, the system can also schedule tasks based on their priority. For high-priority tasks, the system will prioritize allocating more resources and higher parallelism; for low-priority tasks, the system can reduce their parallelism and resource allocation to ensure overall system efficiency.

[0157] Task parallelism and resource allocation depend not only on the task's inherent characteristics but also require dynamic adjustments based on its execution priority and system resource allocation. In real-time data processing, task priorities often need to be dynamically assigned based on their importance and real-time requirements. Therefore, the system can intelligently adjust parallelism and resource allocation based on these factors, ensuring that important tasks are prioritized and minimizing latency.

[0158] The system may predict the execution time and resource requirements of a task based on the task's execution history data. This prediction mechanism based on historical data can help the system estimate the execution time of a task more accurately and provide a basis for resource allocation. In this way, the system can predict bottlenecks in the task execution process in advance and dynamically adjust the parallelism and resource allocation of the task during the execution process to ensure efficient and stable operation of the system. The present invention can effectively optimize the execution time of tasks, avoid resource conflicts and excessive consumption, and improve the parallelism and efficiency of task processing. In practical applications, the system can adapt to different data processing environments and task types, dynamically adjust the parallelism and resource allocation, and thus maximize the performance of the overall system.

[0159] Through nonlinear optimization methods and dynamic resource scheduling, the present invention can intelligently adjust task parallelism and resource allocation based on task execution characteristics, dependencies, and system load, reducing execution time and improving resource utilization. This optimization strategy is of great significance for handling complex tasks, improving big data processing capabilities, and enhancing system stability.

[0160] Application of Queuing Theory in Real-Time Data Stream Processing: Because real-time data stream processing often faces a large number of concurrent tasks and uncertain data arrival rates, efficiently scheduling and processing these tasks becomes crucial to system performance. In the previous steps, a task dependency graph was constructed, and the execution order of tasks was optimized through topological sorting. Next, the application of queuing theory models can help further optimize the processing speed of real-time data stream tasks, ensuring that tasks are processed efficiently and with low latency.

[0161] Real-time data stream processing tasks are modeled as a queuing system, often analyzed using the M / M / 1 queuing model. The core concepts in this queuing model are the task arrival rate and service rate, two factors that directly impact task waiting time and processing time. In the M / M / 1 model, tasks arrive according to a Poisson distribution, and the service time of each task follows an exponential distribution. This model allows the system to calculate key parameters such as task waiting time, queue length, and system load, providing a theoretical basis for real-time data stream processing.

[0162] The arrival rate λ and service rate μ of real-time data stream tasks are two key performance indicators. The arrival rate λ indicates how often tasks arrive at the queue, while the service rate μ indicates the rate at which the system can process tasks. Based on the queuing theory model, the system's average waiting time W and the average number of tasks in the queue L can be calculated using the following formula:

[0163]

[0164] in:

[0165] W represents the average waiting time of the task, in time;

[0166] L represents the average number of tasks in the queue;

[0167] λ represents the arrival rate of tasks;

[0168] μ represents the service rate of the task.

[0169] To improve the efficiency of real-time data stream processing, the system can dynamically adjust the ratio between the service rate μ and the arrival rate λ based on the current task load and resource availability. For example, when the system load is low, the task arrival rate λ may increase, and the system can accordingly increase the service rate μ to accelerate task processing. Conversely, when the load is high, the system may reduce μ to avoid excessive computing resource usage and thus ensure system stability.

[0170] The queuing theory model in the present invention is mainly used to dynamically adjust the processing rate of real-time data flow tasks. When the system recognizes that the arrival rate λ of the data flow task does not match the current system resources, the system load can be balanced by adjusting the service rate μ. For example, when the data flow load increases, the service rate μ can be appropriately increased to enhance the task processing capability; and when the data flow load decreases, the service rate μ can be appropriately reduced to avoid wasting system resources. The queuing theory model in real-time data flow task processing can also be used to optimize task scheduling strategies. The system can adjust the priority and execution order of tasks by real-time monitoring of task arrival time, service time and system load. When the system detects that some tasks take a long time to execute, the priority of these tasks can be increased to ensure that they are completed as soon as possible, thereby reducing system delays.

[0171] The system may dynamically adjust the service rate based on the priority and urgency of the task. For example, the system may assign a higher service rate to urgent real-time tasks to ensure that these tasks are processed as quickly as possible. For non-real-time tasks, the system may reduce their service rate to free up more computing resources for other, more urgent tasks.

[0172] When the system load is high, queuing theory models can effectively control the system load by adjusting the ratio between the service rate and the arrival rate. In this case, the system may dynamically adjust the priority of each task and adjust the order in which tasks are processed based on the real-time system load to ensure stable operation under high load.

[0173] To prevent tasks from waiting in queues for extended periods, the system employs an intelligent scheduling mechanism that allocates high- and low-priority tasks to different processing queues based on real-time changes in arrival and service rates. This allows the system to flexibly schedule tasks based on their characteristics, ensuring that real-time tasks are completed first and minimizing delays for other tasks.

[0174] The system can also combine queuing theory models with data from real-time monitoring systems to predict task execution times and service rates, enabling more accurate scheduling decisions. This approach prevents tasks from waiting for long periods due to insufficient resources, while maximizing the use of system resources and ensuring efficient processing of real-time data stream tasks.

[0175] The application of queuing theory models to real-time data stream processing can effectively optimize task processing speed. By dynamically adjusting the arrival and service rates of tasks, the system ensures efficient and low-latency task processing, avoiding excessive resource usage and long task wait times. Furthermore, through intelligent scheduling and priority management, the system maintains stable operation under high loads and complex environments, maximizing its real-time data stream processing capabilities.

[0176] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A data processing method based on SQL script, characterized in that: The following steps are involved: Step 1: Build a task dependency graph, where each SQL task is considered a node and the data dependencies between tasks are directed edges. Step 2: Topologically sort the task dependency graph. If there are loops in the task dependency graph during the topological sorting process, the depth-first search algorithm will detect and adjust the loops, or eliminate the loops by task decomposition. Step 3: Based on the execution order of topological sorting, calculate the execution time of each task and adjust the execution order of each task; Step 4: Dynamically adjust the parallelism and computing resource allocation of each task and adjust the total execution time of all tasks; The system uses a nonlinear optimization model to optimize parallelism and resource allocation. The parallelism p i and computing resources i The adjustment is based on the execution time model of each task. The execution time of each task is T i (p i ,r i ) is expressed by the following formula: Where: T i (p i ,r i ) represents task T i The execution time depends on the parallelism of the task p i and resources i ;c i It is task T i The inherent computational constant usually reflects the basic computational time required for task execution; i It is task T i The degree of parallelism is the number of computing units that can process tasks simultaneously; f(r i ) is the resource scheduling function, which describes the relationship between resource allocation and task execution time; f(r i ) is a linear function or a logarithmic function. The function form varies depending on the resource allocation method. A nonlinear optimization algorithm is used to solve the optimal parallelism p for each task. i and resources i ;The optimization goal is to minimize the total execution time of the task, that is, Where n is the total number of tasks, T i (p i ,r i ) represents task T i The execution time, p i and r i Represents tasks T i parallelism and resource allocation; Step 5: Based on the game theory model, calculate the resource allocation strategy for each task. Model the task scheduling problem as a non-cooperative game. Each task selects the optimal resource allocation strategy based on the resource allocation of other tasks. The resource allocation strategy is adjusted using the Nash equilibrium theory of game theory. Step 6: Apply queuing theory models to real-time data stream tasks and adjust the data processing rate; In step 6, the application of the queuing theory model includes: modeling the real-time data stream task as an M / M / 1 queue system; The task processing rate is dynamically adjusted according to the arrival rate and service rate of the real-time data stream; the queuing theory model optimization step in the real-time data stream processing is performed by adjusting the ratio of the service rate to the arrival rate.

2. A data processing method based on SQL script according to claim 1, characterized in that: In the step 1, constructing the task dependency graph includes: Treat each SQL task as a node, and the data dependencies between tasks as directed edges to construct a directed acyclic graph. In the directed acyclic graph, edges represent dependencies between tasks, and nodes represent SQL query tasks.

3. The data processing method based on SQL script according to claim 1, characterized in that: In step 3, the execution time optimization includes: Topologically sort the task dependency graph to ensure that each task is executed in sequence; The execution order of each task is determined based on the results of topological sorting, and the execution time between tasks is adjusted.

4. The data processing method based on SQL script according to claim 1, characterized in that: Steps to minimize the total task execution time include: Adjust the parallelism and computing resource allocation of each task based on the execution time model of each task; The optimization process adjusts the execution time of each task.

5. The data processing method based on SQL script according to claim 1, characterized in that: In step 5, the calculation steps of the game theory model include: The task scheduling problem is modeled as a non-cooperative game. Each task selects a resource allocation strategy based on the resource allocation of other tasks. The payoff function of each task is the ratio of its execution time to the allocated resources. The resource allocation strategy is adjusted through the Nash equilibrium theory of game theory.

6. The data processing method based on SQL script according to claim 1, characterized in that: In the step 2, the shortest path is calculated by topological sorting of the task dependency graph or a topological sorting algorithm is used to optimize the task execution order.

7. The data processing method based on SQL script according to claim 1, characterized in that: In the fourth step, the optimization step of calculating task parallelism and resource allocation is solved by nonlinear optimization, using the Lagrange multiplier method to calculate resource constraints and obtain the parallelism and resource allocation of each task.

Citation Information

Patent Citations

  • Heuristic Storm node task scheduling optimization method

    CN116974724A

  • Efficient high-throughput calculation task scheduling method

    CN118656181A