Industrial algorithm model scheduling method and system for industrial big data
By building an algorithm dependency diagram, identifying bottleneck nodes, performing operation fusion and data partition optimization, and dynamically adjusting resource allocation and priorities, the problems of complex dependencies, multiple memory copies and parallel bottlenecks in industrial big data processing are solved, efficient algorithm scheduling and resource utilization are achieved, and system performance and reliability are improved.
Patent Information
- Application Number
- CN202510418712.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When facing industrial big data processing, the existing industrial algorithm scheduling methods have complex data dependencies, resulting in the inability to achieve end-to-end performance improvement of single algorithm optimization. The system overhead and high latency are high during data transmission and format conversion, and the lack of intelligent matching and scheduling of heterogeneous computing resources, resulting in low resource utilization, low parallel efficiency, and the parallel bottleneck of algorithm pipelines limits the overall throughput.
Build an algorithm dependency diagram, identify bottleneck nodes, perform operation fusion through data dependency characteristic analysis, determine the best data partitioning strategy, dynamically adjust parallelism, allocate the most matching heterogeneous computing resources, and adjust the algorithm operation priority in real time, realize a fine-grained synchronization system, reduce data transmission and copy overhead, and optimize execution paths and parallel efficiency.
Significantly reduce data transmission overhead between algorithms, reduce end-to-end processing delays, improve overall system throughput, improve resource utilization, enhance system adaptability and scalability, ensure the accuracy and reliability of calculation results, and meet the requirements of industrial control and decision-making systems.
Smart Images

Figure CN120336011A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial big data processing and algorithm scheduling, and more specifically, to an industrial algorithm model scheduling method and system for industrial big data. Background Art
[0002] With the development of industry and the popularization of intelligent manufacturing, the amount of data generated in the industrial production process has increased explosively, posing higher requirements for data processing and algorithm scheduling systems. Industrial big data has the characteristics of large volume, multiple types, and strong timeliness, and requires multiple professional algorithms to cooperate in processing to extract valuable information, assist decision-making, and optimize production.
[0003] However, the existing industrial algorithm scheduling methods mainly face the following problems: The complex data dependency relationships in the industrial algorithm pipeline lead to the inability to achieve end-to-end performance improvement with single algorithm optimization, lacking a global optimization perspective; Multiple memory copies during data transmission and format conversion result in high system overhead and latency; Different algorithms have different data access patterns, and a unified partitioning strategy leads to low parallel efficiency; There is a lack of intelligent matching and scheduling capabilities for heterogeneous computing resources, and the resource utilization rate is not high; The parallel bottleneck of the algorithm pipeline in the big data processing scenario limits the overall throughput.
[0004] Therefore, there is a need for an industrial algorithm model scheduling method specifically for the industrial big data environment, which can optimize the collaborative execution process of multiple algorithms from a global perspective, solve complex dependency relationships, reduce data transmission overhead, and improve parallel efficiency, so as to meet the high-performance requirements of industrial big data processing. Summary of the Invention
[0005] The present invention provides an industrial algorithm model scheduling method and system for industrial big data, which solves the technical problems of complex data dependencies, multiple memory copies, lack of global optimization, data access pattern differences, and parallel bottlenecks existing in the collaborative execution process of multiple algorithms in the related art.
[0006] The present invention provides an industrial algorithm model scheduling method for industrial big data, including:
[0007] Constructing an algorithm dependency graph of an execution pipeline composed of multiple industrial algorithms, where graph nodes represent algorithm operations and edges represent data dependency relationships, and identifying bottleneck nodes affecting the overall performance through a critical path analysis algorithm;
[0008] For the identified bottleneck nodes, analyzing the data exchange patterns between adjacent algorithm operations based on data dependency characteristics, and performing operation fusion on algorithm pairs with close dependencies to reduce data transmission and copy overhead;
[0009] Determine the optimal data partitioning strategy for different algorithms according to the data access pattern and parallel speedup ratio of the algorithms, and implement an elastic execution pipeline to dynamically adjust the parallelism during the data stream processing;
[0010] Based on the algorithm to calculate the eigenvector and the computing resource characteristic vector, allocate the most matching heterogeneous computing resources for the algorithm operations on the critical path, and adjust the execution priority of the algorithm operations in real time according to the system load and running scenario;
[0011] Implement a fine-grained synchronization system, decompose the data dependencies between algorithms into finer-grained partial dependencies, enable the downstream algorithm to start processing when the upstream algorithm generates partial results, reduce the waiting time and increase the concurrency.
[0012] In a preferred embodiment, the critical path analysis algorithm includes the following steps:
[0013] Allocate an execution time weight for each algorithm node, and the time weight is estimated based on historical execution data or a performance model;
[0014] Calculate all possible paths from the start node to the end node:
[0015] Path = {Path1, Path2,..., Path n};
[0016] where Path represents all possible paths from the start node to the end node, Path1, Path2, Path n represent the 1st, 2nd, nth complete execution paths respectively, and n represents the total number of paths;
[0017] For each path Path i , calculate its total execution time:
[0018] T_path(Path i ) = ∑v∈Path i w(v);
[0019] where T_path(Path i ) represents the total execution time of path Path i , v represents the algorithm node on path Path i , w(v) represents the execution time weight of this node, and ∑ represents the sum of the weights of all nodes on the path;
[0020] Determine the critical path:
[0021] CirtPath = argmax(T_path(Path i ));
[0022] i.e., the path with the longest total execution time, where CritPath represents the critical path, and T_path(Path i ) represents the total execution time of the i-th path, and argmax is used to return the variable value that makes the function reach the maximum value;
[0023] Identify the nodes on the critical path, which are the bottleneck points affecting the overall performance.
[0024] In a preferred embodiment, the data dependence characteristic is quantified by a data dependence strength matrix DS, and its calculation formula is:
[0025]
[0026] where DS(i, j) represents the data dependence strength between algorithms i and j, D size (i, J) is the data exchange size, F access (i, j) is the access frequency, and T interval (i, j) is the execution time interval between the two algorithms.
[0027] In a preferred embodiment, the data partitioning strategy is based on the speedup ratio of the algorithm for different data characteristics, and the calculation formula of the speedup ratio is:
[0028]
[0029] where S(v, D, PS i ) represents the speedup ratio of algorithm v when processing data set D using partitioning strategy PS i , T serial (v, D) is the serial execution time of algorithm v for processing data D, and T parallel (V, D, PSi, n) is the execution time when using partitioning strategy PS i with a parallelism of n.
[0030] In a preferred embodiment, the allocation of the most matching heterogeneous computing resources is determined by an algorithm resource matching degree matrix ARM, and its calculation formula is:
[0031]
[0032] where ARM(v, r) represents the matching degree score between algorithm node v and computing resource r, k represents the number of dimensions of the feature vector, w i is the weight of the i-th feature dimension, CF i (v) represents the i-th computing feature component of algorithm node v, RF i (r) represents the i-th resource feature component of computing resource r, and match(CF i (v), RFi (r) is the matching degree function of the algorithm feature vector and the resource feature vector.
[0033] In a preferred embodiment, the reduction of data transmission and copy overhead includes the following steps:
[0034] Create a shared memory area pool for data exchange between algorithms;
[0035] For each pair of algorithms that need to exchange data, allocate a shared memory area;
[0036] Modify the output interface of the algorithm so that its data is directly written into the shared memory area instead of the local memory;
[0037] Modify the input interface of the receiving algorithm so that it directly reads data from the shared memory area, avoiding data copy.
[0038] In a preferred embodiment, the implementation of the elastic execution pipeline dynamically adjusts the parallelism during the data stream processing according to the following formula:
[0039] n t+1 =n t +Δn;
[0040] Δn=f(L sys , P cur , P target );
[0041] Wherein, n t+1 is the adjusted parallelism at the next time point t + 1, Δn is the change amount of the parallelism, n t is the current parallelism, f is the adjustment function, L sys is the system load, P cur is the current performance, P target is the target performance.
[0042] In a preferred embodiment, the real-time adjustment of the execution priority of the algorithm operations is calculated by the following formula:
[0043]
[0044] Wherein, Priority(v, t) represents the real-time priority value of the algorithm node v at the time point t, BP(v) is the basic priority of the algorithm v, k represents the number of adjustment factors affecting the priority, represents the consecutive multiplication operation from i = 1 to k, α i is the adjustment factor weight, af i (v, t) is the value of the i-th adjustment factor of the algorithm v at the time t.
[0045] In a preferred embodiment, the fine-grained synchronization system determines the execution timing based on a data readiness condition, and its conditional expression is:
[0046] Ready(v) = ∧ u∈pred(v) Complete(u, output(u, v));
[0047] Wherein, Ready(v) represents the readiness state of algorithm node v, ∧ represents the logical "AND" operation, pred(v) is the set of predecessor nodes of v, output(u, v) is the output data transmitted from node u to v, and Complete(u, d) represents that node u has completed the generation of data d.
[0048] In a preferred embodiment, an industrial algorithm model scheduling system for industrial big data is used to implement an industrial algorithm model scheduling method for industrial big data, including:
[0049] An algorithm dependency graph construction module, which is used to construct a dependency graph of the algorithm execution pipeline and identify bottleneck nodes;
[0050] A zero-copy data channel module, which is used to achieve efficient data transmission between algorithms;
[0051] An adaptive partitioning and scheduling module, which is used to determine the best data partitioning strategy according to algorithm characteristics and dynamically adjust the parallelism;
[0052] A heterogeneous resource matching and priority allocation module, which is used to allocate the most suitable computing resources for algorithms and adjust the execution priority;
[0053] A fine-grained synchronization module, which is used to implement the partial data dependency and asynchronous execution mechanism between algorithms.
[0054] The beneficial effects of the present invention are as follows:
[0055] Performance improvement: By reducing the data transmission overhead between algorithms, optimizing the execution path, and improving the parallel efficiency, the data transmission overhead between algorithms is reduced by 65%, the end-to-end processing delay is reduced by 48%, the overall system throughput is increased by 52%, and the ability to process large-scale industrial data is significantly enhanced.
[0056] Optimized resource utilization: By identifying global performance bottlenecks and accurately allocating resources, the situation of some resources being overloaded while other resources are idle is avoided, the computing resource utilization rate is increased by 60%, and the system runs more balanced and efficiently.
[0057] Enhanced adaptability: By adopting an adaptive data partitioning strategy and a dynamic priority adjustment mechanism, the system can automatically select the best execution strategy according to algorithm characteristics and runtime environment changes, adapt to various industrial big data processing scenarios, and achieve the optimal execution plan without manual intervention.
[0058] Improved scalability: Based on modular design and standardized interfaces, the system supports dynamically adding new algorithm operations to existing pipelines and can automatically perform global re-optimization, enabling new algorithms to work efficiently with existing algorithms and facilitating the continuous evolution and functional expansion of industrial algorithm models.
[0059] Guaranteed stability: Through fine-grained synchronization mechanisms and data consistency management, even under high-concurrency execution conditions, the accuracy and reliability of calculation results can be ensured, meeting the strict requirements of industrial control and decision-making systems. Description of the Drawings
[0060] Figure 1 is a flowchart of the industrial algorithm model scheduling method for industrial big data of the present invention;
[0061] Figure 2 is a detailed flowchart of the algorithm dependency graph for constructing an execution pipeline composed of multiple industrial algorithms of the present invention;
[0062] Figure 3 is a detailed flowchart of the operation fusion for algorithm pairs with close dependencies of the present invention;
[0063] Figure 4 is a detailed flowchart of determining the best data partitioning strategy for different algorithms of the present invention;
[0064] Figure 5 is a detailed flowchart of allocating the most suitable heterogeneous computing resources to algorithm operations on the critical path of the present invention;
[0065] Figure 6 is a detailed flowchart of decomposing the data dependencies between algorithms into finer-grained partial dependencies of the present invention. Detailed Embodiments
[0066] Now, the industrial algorithm model scheduling method for industrial big data described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0067] In at least one embodiment of the present invention, an industrial algorithm model scheduling method for industrial big data is disclosed, as Figures 1 to 6 shown, and specifically includes the following steps:
[0068] Step 1: Construct an algorithm dependency graph for the execution pipeline composed of multiple industrial algorithms, where graph nodes represent algorithm operations and edges represent data dependencies, and identify bottleneck nodes affecting overall performance through the critical path analysis algorithm;
[0069] Specifically, it includes the following sub-steps:
[0070] Sub-step 1.1: Construct an algorithm operation dependency graph;
[0071] Use a directed acyclic graph to represent the execution pipeline composed of multiple industrial algorithms, where graph nodes represent algorithm operations and edges represent data dependencies. The specific implementation process is as follows:
[0072] Collect information on all algorithm operation nodes in the industrial environment, including algorithm identifiers, algorithm types, input data requirements, output data types, etc.;
[0073] Analyze the data flow relationship between algorithms to determine which algorithm outputs are used as inputs for other algorithms;
[0074] Create directed edges to connect related algorithm nodes to form a complete algorithm dependency graph G(V, E), where V is the set of algorithm nodes and E is the set of data dependency edges.
[0075] Sub-step 1.2: Perform critical path analysis;
[0076] Apply the critical path analysis algorithm to process the constructed dependency graph to identify bottleneck nodes affecting overall performance. The calculation process is as follows:
[0077] Assign an execution time weight w(v) to each algorithm node, and the time weight is estimated based on historical execution data or a performance model;
[0078] Calculate all possible paths from the start node to the end node:
[0079] Path = {Path1, Path2,..., Path n};
[0080] where Path represents all possible paths from the start node to the end node, and Path1, Path2, Path n represent the 1st, 2nd, and nth complete execution paths respectively, and n represents the total number of paths;
[0081] For each path Path i , calculate its total execution time:
[0082] T_path(Path i ) = ∑v∈Path i w(v);
[0083] Among them, T_path(Path i ) represents the total execution time of path Path i , v represents an algorithm node on path Path i , w(v) represents the execution time weight of this node, and ∑ represents the summation of the weights of all nodes on the path;
[0084] Determine the critical path:
[0085] CritPath = argmax(T_path(Path i )));
[0086] That is, the path with the longest total execution time, where CritPath represents the critical path, T_path(Path i ) represents the total execution time of the i-th path, and argmax is used to return the variable value that makes the function obtain the maximum value;
[0087] Identify the nodes on the critical path, and these nodes are the bottleneck points affecting the overall performance.
[0088] The specific implementation of the critical path analysis algorithm adopts an improved depth-first search method, as follows:
[0089] Algorithm: Improved critical path analysis;
[0090] Input: Algorithm dependency graph G(V, E), node execution time weight w(v);
[0091] Output: Critical path CritPath, critical node set CriticalNodes;
[0092] Initialize the earliest start time ES(v) of all nodes in the graph to 0;
[0093] Traverse all nodes in the graph in topological order:
[0094] For each node v:
[0095] ES(v) = max{ES(u) + w(u) | u ∈ predecessor node(v)};
[0096] Among them, ES(v) represents the earliest start time of node v, ES(u) represents the earliest start time of predecessor node u, w(u) represents the execution time of predecessor node u, and max represents the maximum value;
[0097] Initialize the latest start time LS(v) of all nodes in the graph to ES(terminating node);
[0098] Traverse all nodes in the graph in reverse topological order:
[0099] For each node v:
[0100] LS(v) = max{LS(u) - w(v) | u ∈ successor nodes of (v)};
[0101] Where LS(v) represents the latest start time of node v, LS(u) represents the latest start time of successor node u, w(v) represents the execution time of node v itself, and max represents the maximum value;
[0102] Calculate the time slack of each node:
[0103] Slack(v) = LS(v) - ES(v);
[0104] Where Slack(v) represents the time slack of node v, LS(v) represents the latest start time of node v, and ES(v) represents the earliest start time of node v;
[0105] Set of critical nodes:
[0106] CriticalNodes = {v | Slack(v) = 0};
[0107] Where CriticalNodes represents the set of critical nodes, v represents a node, and Slack(v) represents the time slack of node v;
[0108] The critical path CritPath is the path formed by connecting the set of critical nodes according to the dependency relationship.
[0109] Application example in the steelmaking process algorithm pipeline: A steel plant production control system uses multiple algorithms to work together, including raw material ratio optimization algorithm, temperature control algorithm, composition analysis algorithm, quality prediction algorithm, and energy consumption optimization algorithm. Through critical path analysis, it is found that the composition analysis algorithm is a bottleneck node on the critical path, and its execution time accounts for 42% of the total execution time. After optimizing this algorithm specifically, the end-to-end processing delay of the entire pipeline is reduced by 31%, significantly improving the production efficiency.
[0110] Sub-step 1.3, calculate the node resource consumption index;
[0111] Conduct resource consumption analysis on each algorithm node to determine its impact on system resources:
[0112] Collect resource indicators such as CPU usage CPU(v), memory occupancy MEM(v), and I / O operation volume IO(v) of each algorithm node v;
[0113] Calculate the comprehensive resource consumption index:
[0114] R(v) = α·CPU(v) + β·MEM(v) + γ·IO(v);
[0115] Among them, R(v) represents the comprehensive resource consumption index of algorithm node v, α represents the weight coefficient of CPU resources, β represents the weight coefficient of memory MEM resources, and γ represents the weight coefficient of IO resources;
[0116] Combine the resource consumption index with the position of the node on the critical path to calculate the performance impact factor of the node:
[0117] PIF(v) = R(v)·IsCP(v);
[0118] Among them, PIF(v) represents the performance impact factor of algorithm node v, R(v) represents the comprehensive resource consumption index of algorithm node v, and IsCP(v) represents whether node v is on the critical path.
[0119] Through the above analysis, obtain the ranking of the performance bottleneck degrees of each node in the algorithm pipeline, providing an accurate target for subsequent optimization.
[0120] Step 2, for the identified bottleneck nodes, analyze the data exchange mode between adjacent algorithm operations based on the data dependence characteristics, and perform operation fusion on algorithm pairs with strong dependence to reduce data transmission and copy overhead;
[0121] It includes the following sub-steps:
[0122] Sub-step 2.1, analyze the data dependence characteristics between algorithm operations;
[0123] Perform a fine-grained analysis of the data dependence relationship between algorithm operations to identify optimizable data transmission modes:
[0124] Analyze the data exchange characteristics between each pair of adjacent algorithm nodes (v1, v2), including the data volume size, data structure type, access frequency, etc.;
[0125] Construct a data dependence strength matrix DS, where the element DS(i,j) represents the data dependence strength between algorithm i and algorithm j, and the calculation formula is:
[0126]
[0127] Among them, DS(i, j) represents the data dependence strength between algorithm i and algorithm j, D size (i,j) is the data exchange size, F access (i,j) is the access frequency, T interval (i, j) is the execution time interval between the two algorithms;
[0128] Set the dependency strength threshold θ. When DS(i, j) > θ, mark the corresponding algorithm pair as a "strong dependency pair" and use it as a candidate for operation fusion.
[0129] Sub-step 2.2, perform algorithm operation fusion;
[0130] For the identified strong dependency algorithm pairs, perform algorithm operation fusion to reduce data transmission overhead:
[0131] For each strong dependency algorithm pair (v1, v2), analyze the compatibility of its operation characteristics and resource requirements;
[0132] Evaluate the performance improvement metric FP(v1, v2) after fusion. The calculation formula is:
[0133] FP(v1, v2) = T e xec(v1) + T e xec(v2) + T comm (v1, v2) - T fusion (v1, v2);
[0134] Among them, PF(v1, v2) represents the performance improvement after fusing algorithms v1 and v2, T e xec(v1), T e xec(v2) respectively represent the time overhead of separately executing algorithms v1 and v2, T comm (v1, v2) is the data transmission time, T fusion (v1, v2) is the execution time after fusion;
[0135] When FP(v1, v2) > 0, create a fusion operation unit F(v1, v2), merge the execution logics of the two algorithms, and eliminate the intermediate data transmission steps.
[0136] Sub-step 2.3, optimize the data transmission path;
[0137] To support more efficient data transmission between industrial algorithms, optimize the transmission path:
[0138] Construct a communication delay matrix L = [T c omm(i, j)], where T c omm(i, j) represents the data transmission time from computing node i to node j;
[0139] For transmission tasks with data volume exceeding the threshold θ, apply the following optimization measures:
[0140] When T c omm(i, j) > T threshold omm(i, j), create a dedicated zero-copy data channel DCij between nodes i and j;
[0141] When T c omm(i, j) ≤ T threshold , use the standard data transmission mechanism;
[0142] For frequently interacting node pairs, establish a persistent communication connection to reduce connection establishment overhead.
[0143] Application example in the industrial image processing algorithm pipeline: The product quality inspection system of a certain factory includes multiple algorithms such as image acquisition, preprocessing, feature extraction, defect recognition, and result analysis. Among them, a large amount of image data (about 500MB per second) needs to be transmitted between the image preprocessing and feature extraction algorithms. After implementing the zero-copy data channel, the data transmission time between the two algorithms is reduced from the original 75ms to 8ms, the CPU usage rate is reduced by 28%, the overall processing delay is reduced by 36%, and the system can process higher-resolution images without increasing hardware costs.
[0144] Step 3: According to the data access pattern and parallel speedup ratio of the algorithm, determine the best data partitioning strategy for different algorithms, and implement an elastic execution pipeline to dynamically adjust the parallelism during the data stream processing;
[0145] Including the following sub-steps:
[0146] Sub-step 3.1: Analyze the data access pattern of the algorithm;
[0147] Analyze the data access patterns of each algorithm to identify its data parallel characteristics:
[0148] Collect the access pattern information of each algorithm operation v to data D, including sequential access, random access, block access, etc.;
[0149] Analyze the data dependency types of the algorithm, including element-level dependency, block-level dependency, global dependency, etc.;
[0150] According to the access pattern and dependency type, evaluate the data parallel potential DP(v, D) of the algorithm. The higher the value, the greater the parallel potential.
[0151] Sub-step 3.2: Calculate the parallel speedup ratio and determine the best partitioning strategy;
[0152] Based on the algorithm characteristics and data characteristics, calculate the parallel speedup ratio under different partitioning strategies:
[0153] For algorithm v and data D, consider n possible partitioning strategies {PS1, PS2,..., PS n}, where PS1, PS2, PS n represent the 1st, 2nd, and nth partitioning strategies respectively, and n represents the total number of partitioning strategies;
[0154] For each partitioning strategy PS i , estimate its parallel execution time T parallel (V, D, PS i , n), where T parallel (V, D, PS i , n) represents the estimated time for algorithm v to process data D in parallel with a parallelism of n using partitioning strategy PS i ;
[0155] Calculate the speedup ratio:
[0156]
[0157] where S(v, D, PS i ) represents the speedup ratio of algorithm v when using partitioning strategy PS to process dataset D i , T serial (v, D) is the serial execution time of algorithm v to process data D, and T parallel (V, D, PS i , n) is the execution time when using partitioning strategy PS i with a parallelism of n.
[0158] Select the partitioning strategy with the highest speedup ratio:
[0159]
[0160] where PS opt represents the optimal partitioning strategy that can achieve the maximum speedup ratio among all alternative partitioning strategies, represents the independent variable when taking the maximum value, represents selecting, among all alternative partitioning strategies PS i , the partitioning strategy that can make the speedup ratio S(v, D, PS i ) reach the maximum value as the optimal strategy PS opt .
[0161] Sub-step 3.3, construct an elastic execution pipeline;
[0162] Implement an elastic execution pipeline that can dynamically adjust the parallelism:
[0163] Create a pipeline execution engine PE to support dynamic parallelism configuration;
[0164] Implement a partition data manager PDM responsible for partitioning the input data D according to the selected partitioning strategy PS opt ;
[0165] Establish a parallel task scheduler PTS responsible for allocating the partitioned data to computing resources;
[0166] Implement the Runtime Monitoring Module (RTM) to collect execution performance metrics and system load information;
[0167] Dynamically adjust the parallelism degree n and the partitioning strategy PS according to the data collected by RTM opt , the formula is:
[0168] n t+1 = n t + Δn;
[0169] Δn = f(L sys , P cur , P target );
[0170] Among them, n t+1 is the adjusted parallelism degree at the next time point t + 1, Δn is the change amount of the parallelism degree, n t is the current parallelism degree, f is the adjustment function, L sys is the system load, P cur is the current performance, P target is the target performance.
[0171] The specific composition and working process of the elastic execution pipeline system are as follows:
[0172] System composition:
[0173] Pipeline Execution Engine (PE): The core execution unit, including a task queue and a working thread pool;
[0174] Partition Data Manager (PDM): Implement various data partitioning strategies, such as hash partitioning, range partitioning, block partitioning, etc.;
[0175] Parallel Task Scheduler (PTS): A task assignment system based on the work stealing algorithm;
[0176] Runtime Monitoring Module (RTM): Collect metrics such as CPU usage, memory usage, I / O throughput, etc.;
[0177] Load Balancer (LB): Adjust the parallelism degree and partitioning strategy according to the monitoring data;
[0178] Working process:
[0179] Initialization stage: Set the initial parallelism degree n0 and the partitioning strategy PS0 according to the algorithm characteristics and data scale;
[0180] Execution stage:
[0181] PDM partitions the input data according to the current partitioning strategy PS_t;
[0182] PTS assigns the partitioned tasks to the working threads in PE;
[0183] RTM monitors the execution performance and system load in real time;
[0184] Adjustment phase:
[0185] Trigger an adjustment evaluation every fixed interval T;
[0186] Calculate the performance deviation: δ = (P_target - P_cur) / P_target; where δ represents the deviation degree between the current performance and the target performance, P_target represents the expected target performance level, and P_cur represents the actual performance of the current system;
[0187] If |δ| > threshold ε, adjust the parallelism and partitioning strategy;
[0188] If δ > 0 (performance shortage): increase the parallelism or select a more efficient partitioning strategy;
[0189] If δ < 0 (resource waste): reduce the parallelism or adjust the partitioning granularity.
[0190] Application example in industrial Internet of Things data processing: The equipment monitoring system of a large chemical enterprise needs to process data from tens of thousands of sensors per minute for fault prediction and efficiency optimization. The original system with fixed parallelism has increased response latency during peak load periods and serious resource waste during low load periods. After applying the elastic execution pipeline, the system can automatically adjust the parallel processing ability according to the data flow. The processing latency during peak periods is reduced by 52%, the resource occupancy during non-peak periods is reduced by 41%, and it can also handle sudden data flow changes. It still maintains a stable processing latency when the peak-to-valley ratio is 8:1.
[0191] Step 4, calculate the eigenvector of features and the eigenvector of computing resource characteristics based on the algorithm, allocate the most matching heterogeneous computing resources for the algorithm operations on the critical path, and adjust the execution priority of the algorithm operations in real time according to the system load and running scenario;
[0192] Including the following sub-steps:
[0193] Sub-step 4.1, analyze the adaptability between the computing characteristics of the algorithm and the resources;
[0194] Analyze the computing characteristics of the industrial algorithm and evaluate its execution efficiency on different computing resources:
[0195] Extract the computing eigenvector CF(v) of each algorithm v, including dimensions such as computing density, memory access pattern, branch prediction difficulty, parallelism, etc.;
[0196] Collect the set of available heterogeneous computing resources R = {R1, R2,..., R mThe feature vector RF(r) of {}, including processor type, number of cores, memory bandwidth, cache size, etc., where R represents the set of available heterogeneous computing resources, and R1, R2, R m represent the first, second, and m-th available heterogeneous computing resources respectively, and m represents the total number of resources;
[0197] Construct the algorithm-resource matching degree matrix ARM, where the element ARM(v, r) represents the execution efficiency of algorithm v on resource r, and the calculation formula is:
[0198]
[0199] where ARM(v, r) represents the matching degree score between algorithm node v and computing resource r, k represents the number of dimensions of the feature vector, and w i is the weight of the i-th feature dimension, CF i (v) represents the i-th computing feature component of algorithm node v, and RF i (r) represents the i-th resource feature component of computing resource r, and match(CF i (v), RF i (r)) is the matching degree function of the algorithm feature vector and the resource feature vector.
[0200] Sub-step 4.2, calculate the resource allocation score of the critical path operations;
[0201] Allocate the most suitable computing resources to the operations on the critical path:
[0202] For each operation node v on the critical path, calculate its execution performance improvement index on each resource r:
[0203]
[0204] where S(v, r) represents the performance improvement score of allocating operation v to computing resource r, T serial (v) is the serial execution time of operation v, and T parallel (v, r) is the parallel execution time on resource r, and Priority(v) is the priority weight of operation v;
[0205] Select the resource-operation pairing with the highest score: (v * , r * ) = argmax v,r S(v, r), and allocate resource r to operation v, where (v * , r * ) represents the pair with the highest S(v, r) score among all possible resource-operation pairings (v, r), and argmax v,rIt means to find the pair of combinations \((v, r)\) that can maximize the objective function \(S(v, r)\) among all possible combinations of algorithm operations \(v\) and computing resources \(r\).
[0206] Update the remaining resource set and the unallocated operation set, and repeat step 2 until all critical path operations obtain resource allocations.
[0207] Sub-step 4.3, implement the dynamic priority adjustment mechanism;
[0208] Establish a dynamic priority adjustment system to adjust the algorithm priority in real time according to the system load and execution situation:
[0209] Define the priority adjustment factor set \(AF = \{af_1, af_2,..., af\) k \(\}\), including waiting time, urgency, resource utilization rate, etc. Among them, \(AF\) represents the priority adjustment factor set, and \(af_1, af_2, af\) k represent the 1st, 2nd, and \(k\)th priority adjustment factors respectively, and \(k\) represents the total number of priority adjustment factors;
[0210] For each algorithm operation \(v\), initialize its basic priority \(BP(v)\);
[0211] Implement the priority dynamic calculation function:
[0212]
[0213] Among them, \(Priority(v, t)\) represents the dynamic priority of algorithm operation \(v\) at time point \(t\), and \(BP(v)\) represents the basic priority of operation \(v\), represents the consecutive multiplication operation from \(i = 1\) to \(k\), and \(\alpha\) i is the adjustment factor weight, and \(af\) i \((v, t)\) is the value of the \(i\)th adjustment factor of algorithm \(v\) at time \(t\);
[0214] During the resource scheduling period, execute algorithm operations according to the dynamic priority sorting.
[0215] Through heterogeneous resource matching and dynamic priority allocation, the system can make full use of diverse computing resources, respond to changes in the workload, and improve the overall execution efficiency.
[0216] Step 5, implement a fine-grained synchronization system, decompose the data dependencies between algorithms into finer-grained partial dependencies, so that downstream algorithms can start processing when upstream algorithms produce partial results, reducing waiting time and increasing concurrency;
[0217] It includes the following sub-steps:
[0218] Sub-step 5.1, construct an asynchronous data preprocessing pipeline;
[0219] During the execution of the upstream algorithm, asynchronously preprocess the data required by the downstream algorithm:
[0220] Analyze the data dependency graph G(V, E) in the algorithm execution pipeline to identify data preprocessing opportunities;
[0221] For each pair of data-dependent algorithms (v1, v2), create an asynchronous preprocessing task P(v1, v2) during the execution of v1;
[0222] Implement a preprocessing task manager PTM responsible for scheduling and executing preprocessing tasks;
[0223] Design preprocessing strategies, including data format conversion, dimension rearrangement, cache warm-up, etc., to prepare the optimal input format for the downstream algorithm v2.
[0224] Sub-step 5.2, implement a fine-grained synchronization system;
[0225] Build a fine-grained synchronization system to minimize the waiting time in parallel execution:
[0226] Decompose data dependencies into finer-grained partial dependencies;
[0227] Implement a data ready notification system to immediately notify the downstream algorithm v2 when the upstream algorithm v1 produces partial results;
[0228] Build a data version management system DVM to maintain multiple version states {D1, D2,..., D n} of the data D, supporting concurrent read and write, where D1, D2, D n represent the 1st, 2nd, and nth historical versions of the data object D respectively, and n represents the total number of historical versions;
[0229] Implement a conditional execution model based on data dependencies:
[0230] Ready(v) = ∧ u∈pred(v) Complete(u, output(u, v));
[0231] where Ready(v) represents the ready condition of the algorithm node v, pred(v) is the set of predecessor nodes of v, output(u, v) is the output data passed from node u to v, Complete(u, d) represents that node u has completed the generation of data v, and ∧ represents the logical AND operation.
[0232] The detailed architecture and implementation of the fine-grained synchronization system are as follows:
[0233] System composition:
[0234] Dependency Decomposition Unit (DDU): Responsible for splitting coarse-grained data dependencies into fine-grained partial dependencies;
[0235] Data Ready Notifier (DRN): A data ready event notification system based on the publish-subscribe model;
[0236] Version Manager (VM): Maintains the version history and status of each data object;
[0237] Condition Execution Controller (CEC): Manages the execution conditions and triggering logic of the algorithm;
[0238] Main data structures:
[0239] Data Block Descriptor (DBD):
[0240] {block_id: Unique identifier of the data block,
[0241] data_ptr: Memory pointer of the data block,
[0242] size: Size of the data block,
[0243] version: Version number,
[0244] status: Ready status,
[0245] producer: Producer algorithm identifier,
[0246] consumers: List of consumer algorithms
[0247] }
[0248] Dependency Relationship Table (DRT): A dependency relationship graph stored using an adjacency list
[0249] Key algorithms:
[0250] Partial Dependency Recognition Algorithm: Based on data access pattern analysis, identifies data blocks that can be processed in parallel;
[0251] Incremental Notification Algorithm: Triggers notifications only when the data status changes, reducing communication overhead;
[0252] Version Conflict Resolution Algorithm: Timestamp-based multi-version concurrency control;
[0253] Application example in intelligent manufacturing production line scheduling: A certain automotive parts production line uses multiple optimization algorithms to collaborate on production planning and scheduling. The traditional method requires that the previous algorithms (such as order classification and material requirements calculation) be fully executed before the subsequent algorithms (such as resource allocation and operation scheduling) can be started. After applying the fine-grained synchronization system, when the order classification algorithm has processed the first 20% of the order data, the resource allocation algorithm can start processing this part of the data without waiting for all orders to be processed. In actual application, the plan generation time is shortened from the original 15 minutes to 4 minutes, the system responsiveness is significantly improved, it can adapt to production plan changes faster, and the waiting time of the production line is reduced.
[0254] Sub-step 5.3, ensure data consistency and concurrency control;
[0255] Implement a data consistency guarantee system to ensure the correctness of high-concurrency execution:
[0256] Define the data access mode type DAT = {read shared, write exclusive, read-write mixed};
[0257] For each algorithm operation v and data object d, mark its access mode DAT(v, d);
[0258] Implement a lock allocation algorithm based on the access mode:
[0259] For the read shared mode, allocate a shared lock;
[0260] For the write exclusive mode, allocate an exclusive lock;
[0261] For the read-write mixed mode, implement multi-version concurrency control;
[0262] Build a conflict detection and resolution system. When a data access conflict is detected, use a priority or timestamp strategy to resolve the conflict.
[0263] Through asynchronous data preprocessing and fine-grained synchronization mechanisms, the system maximizes the opportunity for parallel execution while ensuring data consistency, reducing the waiting time between algorithms.
[0264] Technical effects of this embodiment
[0265] Performance improvement: By reducing the data transfer overhead between algorithms, optimizing the execution path, and improving parallel efficiency, the data transfer overhead between algorithms is reduced by 65%, the end-to-end processing delay is reduced by 48%, the overall system throughput is increased by 52%, and the ability to process large-scale industrial data is significantly enhanced.
[0266] Optimization of resource utilization: Through global performance bottleneck identification and precise resource allocation, the situation where some resources are overloaded while others are idle is avoided, the computing resource utilization rate is increased by 60%, and the system runs more balanced and efficiently.
[0267] Enhanced adaptability: By adopting an adaptive data partitioning strategy and a dynamic priority adjustment mechanism, the system can automatically select the best execution strategy according to the algorithm characteristics and runtime environment changes, adapt to various industrial big data processing scenarios, and achieve the optimal execution plan without manual intervention.
[0268] Improved scalability: Based on modular design and standardized interfaces, the system supports dynamically adding new algorithm operations to existing pipelines and can automatically perform global re-optimization, enabling new algorithms to work efficiently with existing algorithms and facilitating the continuous evolution and function expansion of industrial algorithm models.
[0269] Guaranteed stability: Through a fine-grained synchronization mechanism and data consistency management, even under high-concurrency execution conditions, the accuracy and reliability of calculation results can be ensured, meeting the strict requirements of industrial control and decision-making systems.
[0270] Real application examples of this embodiment
[0271] This embodiment has been applied in the intelligent manufacturing platform of a large steel enterprise, solving the performance bottleneck problem of multi-algorithm collaborative operation in the steel production process. This platform needs to run a variety of industrial algorithms in real time, including core algorithms such as raw material ratio optimization, furnace temperature control, quality prediction, energy consumption optimization, and defect detection. There are complex data dependencies between these algorithms, and their respective requirements for computing resources vary significantly.
[0272] Project background: In the production intelligent upgrade project of this steel enterprise, it is necessary to improve production efficiency and product quality through various industrial algorithms. Especially on high-end steel production lines, extremely high requirements are placed on the real-time performance, accuracy, and reliability of algorithms. Before the upgrade, each algorithm model ran independently, and data was exchanged through intermediate files or databases, resulting in obvious performance bottlenecks:
[0273] The data transmission delay between algorithms is high, affecting the overall response time;
[0274] Lack of global resource scheduling leads to overloading of some hardware resources while other resources are idle;
[0275] Unable to dynamically adjust the execution strategy according to different workloads, resulting in limited peak processing capacity;
[0276] It is difficult to identify and optimize the performance bottlenecks in complex algorithm pipelines;
[0277] It is difficult to ensure data consistency in high-concurrency scenarios;
[0278] System Scale: This platform needs to process approximately 20TB of industrial data per day, support the collaborative operation of more than 30 industrial algorithm models of different types, and serve the real-time control and decision-making support of 5 production lines throughout the factory. The system hardware environment includes 32 servers, configured with heterogeneous computing resources such as CPUs, GPUs, and FPGAs.
[0279] Implementation Process Example:
[0280] The industrial algorithm model scheduling method of the present invention was implemented in this project, and the main implementation process is as follows:
[0281] Construction of Algorithm Dependency Graph and Bottleneck Identification:
[0282] Firstly, for the core production line of the steel enterprise, a complete algorithm dependency graph was constructed. Table 1 shows some key algorithm node information:
[0283] Table 1: Characteristics and Bottleneck Analysis of Core Algorithm Nodes;
[0284]
[0285] Through critical path analysis, A3 (composition analysis), A5 (defect detection), and A4 (quality prediction) were identified as the main performance bottlenecks and were optimized first.
[0286] Implementation of Zero-Copy Data Channel:
[0287] Based on the bottleneck analysis results, we first implemented a zero-copy data channel for the algorithm pairs on the critical path. Table 2 shows the data dependency strength and zero-copy channel implementation between some algorithm pairs:
[0288] Table 2: Data Dependency Strength and Zero-Copy Channel Implementation;
[0289]
[0290] The implementation results show that the zero-copy data channel reduces the data transmission time by an average of 97.1%, significantly improving the data transfer efficiency between algorithms.
[0291] Implementation of Adaptive Data Partitioning and Parallel Scheduling:
[0292] An adaptive data partitioning strategy and an elastic execution pipeline were implemented for the key algorithms. Table 3 shows the specific configurations and effects:
[0293] Table 3: Adaptive Data Partitioning and Parallel Scheduling Configuration;
[0294]
[0295] The elastic execution pipeline can automatically adjust the degree of parallelism according to the system load, avoiding resource waste while maintaining high performance.
[0296] Implementation of heterogeneous resource matching and priority allocation:
[0297] The industrial intelligent manufacturing platform has a variety of heterogeneous computing resources, and targeted resource matching and dynamic priority adjustment are carried out based on algorithm characteristics:
[0298] Table 4: Heterogeneous resource matching and priority configuration;
[0299]
[0300] Through heterogeneous resource matching and dynamic priority adjustment, the system can automatically adjust the execution priority of algorithms and resource allocation strategies according to the requirements of different production scenarios.
[0301] Implementation of the fine-grained synchronization system:
[0302] A fine-grained synchronization system has been implemented for the data dependencies between key algorithms, reducing the waiting time between algorithms:
[0303] Table 5: Fine-grained synchronization configuration and effects;
[0304]
[0305] After implementing the fine-grained synchronization system, the average waiting time between key algorithms has been reduced by 69.1%, significantly improving the parallel processing ability of the system.
[0306] Verification of technical effects:
[0307] The application of the industrial algorithm model scheduling method of the present invention in the intelligent manufacturing platform of steel enterprises has achieved remarkable technical effects, which are mainly reflected in the following two aspects:
[0308] System performance improvement effect:
[0309] By implementing the scheduling method of the present invention, the overall performance of the system has been greatly improved, as shown in Table 6:
[0310] Table 6: Comparison of system performance improvement effects;
[0311]
[0312] Actual production benefits:
[0313] The implementation of the present invention has brought significant production benefits to steel enterprises, and Table 7 shows the improvement effects of the main production indicators:
[0314] Table 7: Improvement effects of production benefits;
[0315]
[0316] The implementation effect shows that the present invention not only improves the technical performance of the system, but also brings substantial economic benefits, fully verifying the practical value and innovation of the present invention in the industrial big data environment.
[0317] The embodiments of the present invention are described above, but the embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present embodiments, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of the present embodiments.
Claims
1. Industrial algorithm model scheduling method for industrial big data, characterized in that, It includes the following steps: Construct an algorithm dependency graph of an execution pipeline composed of multiple industrial algorithms, where graph nodes represent algorithm operations, edges represent data dependencies, and identify bottleneck nodes affecting overall performance through a critical path analysis algorithm; For the identified bottleneck nodes, analyze the data exchange patterns between adjacent algorithm operations based on data dependency characteristics, perform operation fusion on algorithm pairs with tight dependencies, and reduce data transmission and copy overhead; Determine the optimal data partitioning strategy for different algorithms according to the data access patterns and parallel speedup ratios of the algorithms, and implement an elastic execution pipeline to dynamically adjust the parallelism during the data stream processing; Based on the algorithm's calculated eigenvector and the computing resource eigenvector, allocate the most matching heterogeneous computing resources for the algorithm operations on the critical path, and adjust the execution priorities of the algorithm operations in real time according to the system load and running scenarios; Implement a fine-grained synchronization system, decompose the data dependencies between algorithms into finer-grained partial dependencies, enable downstream algorithms to start processing when the upstream algorithms produce partial results, reduce waiting time and increase concurrency.
2. The industrial algorithm model scheduling method for industrial big data according to claim 1, wherein, The critical path analysis algorithm includes the following steps: Assign an execution time weight to each algorithm node, and the time weight is estimated based on historical execution data or a performance model; Calculate all possible paths from the start node to the end node: Path = {Path1, Path2,..., Pathn}; Among them, Path represents all possible paths from the start node to the end node, and Path1, Path2, Path n respectively represent the 1st, 2nd, and nth complete execution paths, where n represents the total number of paths; For each path Path i , calculate its total execution time: T_path(Path i ) = ∑v∈Path i w(v); Among them, T_path(Path i ) represents the total execution time of path Path i . v represents the algorithm node on path Path i , w(v) represents the execution time weight of this node, and ∑ represents the summation of the weights of all nodes on the path; Determine the critical path: CritPath = argmax((T_path(Path i )); i.e., the path with the longest total execution time, where CritPath represents the critical path, and T_path(Path i ) represents the total execution time of the i-th path, and argmax is used to return the variable value that makes the function reach the maximum value; Identify the nodes on the critical path, and these nodes are the bottleneck points affecting overall performance.
3. The industrial algorithm model scheduling method for industrial big data according to claim 1, characterized in that The data dependency characteristics are quantified by a data dependency strength matrix DS, and its calculation formula is: Among them, DS(i, j) represents the data dependence strength between algorithm i and algorithm j, D size (i, j) is the data exchange size, F access (i, j) is the access frequency, T interval (i, j) is the execution time interval between the two algorithms.
4. The industrial algorithm model scheduling method for industrial big data according to claim 1, characterized in that The data partitioning strategy is based on the speedup ratio of the algorithm for different data characteristics, and the calculation formula of the speedup ratio is: Among them, S(v, D, PS i ) represents the speedup ratio of algorithm v when processing dataset D using partitioning strategy PS i , where T serial (v, D) is the serial execution time of algorithm v for processing data D, and T parallel (V, D, PS i , n) is the execution time when using partitioning strategy PS i with a parallelism of n.
5. The industrial algorithm model scheduling method for industrial big data according to claim 1, characterized in that The allocation of the most matching heterogeneous computing resources is determined by an algorithm resource matching degree matrix ARM, and its calculation formula is: Among them, ARM(v, r) represents the matching degree score between algorithm node v and computing resource r, k represents the number of dimensions of the feature vector, and w i is the weight of the i-th feature dimension, CF i (v) represents the i-th computing feature component of algorithm node v, RF i (r) represents the i-th resource feature component of computing resource r, match(CF i (v), RF i (r)) is the matching degree function of the algorithm feature vector and the resource feature vector.
6. The industrial algorithm model scheduling method for industrial big data according to claim 1, characterized in that The reduction of data transmission and copy overhead includes the following steps: Create a shared memory area pool for data exchange between algorithms; For each pair of algorithms that need to exchange data, allocate a shared memory area; Modify the output interface of the algorithm so that its data is directly written into the shared memory area instead of the local memory; Modify the input interface of the receiving algorithm so that it directly reads data from the shared memory area to avoid data copying.
7. The industrial algorithm model scheduling method for industrial big data according to claim 1, wherein The implementation of the elastic execution pipeline to dynamically adjust the parallelism during the data stream processing is calculated according to the following formula: n t+1 = n t + Δn; Δn = f(L sys , P cur , P target ) where n t+1 is the adjusted parallelism at the next time point t+1, Δn is the change in parallelism, and n t is the current parallelism, f is the adjustment function, L sys is the system load, and P cur is the current performance, and P target is the target performance.
8. The industrial algorithm model scheduling method for industrial big data according to claim 1, wherein The real-time adjustment of the execution priority of the algorithm operation is calculated by the following formula: Among them, Priority(v, t) represents the real-time priority value of algorithm node v at time point t, BP(v) is the basic priority of algorithm v, k represents the number of adjustment factors affecting the priority, represents the consecutive multiplication operation from i = 1 to k, α i is the adjustment factor weight, af i (v, t) is the value of the i-th adjustment factor of algorithm V at time t.
9. The industrial algorithm model scheduling method for industrial big data according to claim 1, wherein The fine-grained synchronization system determines the algorithm execution timing based on the data ready condition, and its conditional expression is: Ready(v) = ∧ u∈pred(v) Complete(u, output(u, v)); Where, Ready(v) represents the ready state of algorithm node v, ^ represents the logical "AND" operation, pred(v) is the set of predecessor nodes of v, output(u, v) is the output data passed from node u to v, and Gomplete(u, d) represents that node u has completed the generation of data d.
10. An industrial algorithm model scheduling system for industrial big data, used to implement the industrial algorithm model scheduling method for industrial big data described in any one of claims 1 to 9, includes: The algorithm dependency graph construction module is used to construct the dependency graph of the algorithm execution pipeline and identify bottleneck nodes; The zero-copy data channel module is used to achieve efficient data transmission between algorithms; The adaptive partitioning and scheduling module is used to determine the best data partitioning strategy according to algorithm characteristics and dynamically adjust the parallelism; The heterogeneous resource matching and priority assignment module is used to allocate the most suitable computing resources for the algorithm and adjust the execution priority; The fine-grained synchronization module is used to implement the partial data dependency and asynchronous execution mechanism between algorithms.
Citation Information
Patent Citations
Distributed collaborative debugging method and system for industrial robot
CN119292082A
Method for accelerating a CDVS extraction process based on a gpgpu platform
US20190139186A1
Cited By
Multi-account concurrent processing management system and method
CN121255407A
Government affair big data algorithm model scheduling method and system
CN122044839A