A video coding task dynamic allocation method for many-core heterogeneous computing environment

CN122824897APending Publication Date: 2026-09-25ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610826093.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]现有视频编码任务调度方法,在面向异构计算环境时,普遍存在资源分配与视频内容实时复杂度脱节、无法根据系统负载动态调整、能效低下等问题

Benefits of technology

现有视频编码任务调度方案多采用静态规则或仅依赖系统负载指标,无法在视频内容复杂度与异构计算资源状态之间建立实时、动态的联动,导致资源错配、能效低下与服务质量不稳定。本发明通过一种面向众核异构计算环境的视频编码任务动态分配方法,从根本上解决了上述问题,其创新性与技术优势具体体现在以下几个方面:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824897A_ABST
    Figure CN122824897A_ABST
Patent Text Reader

Abstract

The application discloses a kind of video coding task dynamic distribution methods for many-core heterogeneous computing environment.The method is dynamically elected with the dual role of coordination and calculation global coordination node, and the video stream is predicted adaptive GOP partition and calculates task complexity priority, while maintaining integrated prediction value resource state table.Based on task priority and predictive resource state, intelligent task allocation is carried out using speed ratio estimation and dynamic threshold strategy, task stealing is supported to cover the cost, and the optimal recovery scheme is selected according to criticality and multi-strategy cost model when node fails.The application realizes full-link intelligent scheduling, significantly improves the resource utilization, throughput and robustness of heterogeneous computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video coding technology, and specifically to a method for dynamically allocating video coding tasks in a many-core heterogeneous computing environment. Background Technology

[0002] With the comprehensive rollout of "Digital China" and the strategic deployment of new infrastructure for national computing power, emerging industries such as ultra-high-definition video, industrial internet, and metaverse are becoming core engines of economic growth. As a primary carrier of information age, video data processing capabilities directly impact the quality of digital economic development. However, video encoding, a core component consuming massive amounts of computing power in data processing, faces a sharp contradiction between the "data deluge" and "green and low-carbon" practices. Therefore, developing a highly efficient video encoding system capable of intelligently sensing content and dynamically allocating heterogeneous computing resources has become an urgent technical challenge to overcome industry bottlenecks and serve national strategies.

[0003] Currently, national and industry-level video processing public platforms and large-scale cloud computing service providers generally adopt static or coarse-grained resource allocation strategies. The typical approach is to fixedly allocate encoding tasks to CPU clusters or GPU accelerator cards, or simply schedule tasks based on queue length. This approach has several fundamental flaws. First, it leads to rigid resource utilization. It cannot dynamically match different computing power architectures based on the real-time complexity of the video content, resulting in "computing power mismatch," with expensive heterogeneous resources remaining inefficient or idle most of the time. Second, static allocation cannot proactively scale down under low load, nor can it preventatively migrate tasks based on device health status, creating a bottleneck for data center PUE optimization. This contradicts the green computing requirements under the national "dual-carbon" goal, resulting in persistently high energy consumption.

[0004] In academia and industry, optimization research on video coding has largely focused on a single aspect, with some concentrating on optimizing the coding algorithm itself or designing hardware encoders. While some research has attempted to schedule coding tasks, the scheduling criteria are often singular and lagging, failing to deeply and dynamically link "real-time perception of video content" with "the global state of heterogeneous resources." This disconnect between "perception" and "scheduling" prevents existing solutions from achieving true global efficiency optimization and quality of service assurance in complex and ever-changing real-world production environments.

[0005] To address this technical bottleneck, this invention provides a method for dynamically allocating video coding tasks in many-core heterogeneous computing environments. The core technology lies in the systematic integration of video content feature analysis, real-time monitoring of heterogeneous computing resources, and dynamic task scheduling control, constructing a closed-loop intelligent decision-making and execution framework. This solution is applicable to various computing platforms with stringent requirements for processing efficiency, latency, and energy consumption, including ultra-high-definition video production and broadcasting platforms, cloud video services, and real-time visual analysis systems. It provides key technical support for the efficient and green operation of large-scale video computing infrastructure. Summary of the Invention

[0006] Existing video coding task scheduling methods generally suffer from problems such as a disconnect between resource allocation and the real-time complexity of video content, an inability to dynamically adjust according to system load, and low energy efficiency when facing heterogeneous computing environments. To address these issues, this invention proposes a dynamic allocation method for video coding tasks in many-core heterogeneous computing environments, aiming to achieve fine-grained and adaptive scheduling of coding tasks among heterogeneous resources such as CPUs and GPUs.

[0007] To address the aforementioned technical problems, this invention provides a method for dynamically allocating video coding tasks in a many-core heterogeneous computing environment. The method includes the following steps: A method for dynamically allocating video coding tasks in a many-core heterogeneous computing environment, characterized by the following steps: Step 1: A many-core heterogeneous computing environment is a heterogeneous cluster consisting of multiple physical computing nodes interconnected by a network. At least some of the nodes use many-core architecture processors. Calculate the priority of all nodes in the many-core heterogeneous computing environment and elect a global coordinating node from multiple physical computing nodes based on the node priority. Step 2: Perform predictive adaptive GOP segmentation on the input video stream based on inter-frame motion information and content change information to generate multiple parallel task units; calculate the comprehensive complexity score based on the structure, temporal and spatial characteristics of each parallel task unit, and directly map the score to the initial scheduling priority of the task. Step 3: The global coordination node continuously monitors and maintains the dynamic status table of all physical computing nodes, and the dynamic status table records at least the resource status of each physical computing node; Step 4: The global coordination node, based on the initial scheduling priority obtained in Step 2 and the resource status in the dynamic status table obtained in Step 3, dynamically allocates the multiple parallel task units in Step 2 to physical computing nodes for execution through a load balancing strategy. Step 5: When any physical computing node is detected to have a performance abnormality or failure, the latest priority is obtained by dynamically weighting the task scheduling priority determined in Step 2 and combining it with the system status of the current task with the performance abnormality or failure. The global coordination node immediately marks this physical computing node as unavailable from the dynamic status table, obtains the new resource status, and triggers the rescheduling of all tasks allocated to this physical computing node. The rescheduling re-executes the dynamic task scheduling steps according to the latest priority of the task and the status of the remaining available units in the system to realize the dynamic allocation of video encoding tasks.

[0008] In step 1, calculating node priority specifically includes: on each physical computing node, using real-time telemetry data as input and parameters trained with historical telemetry data as a benchmark, generating a short-term availability prediction value for the node, and combining the short-term availability prediction value with the network topology centrality index and hardware management capability index of the physical computing node according to preset weights to calculate the priority of the node.

[0009] In step 2, predictive adaptive GOP segmentation is performed on the input video stream based on inter-frame motion information and content change information to generate multiple parallel task units, specifically including: The inter-frame motion information and content change information specifically include inter-frame motion energy and its short-term trend, motion energy gradient and scene switching information; the inter-frame motion energy is estimated by first reusing the motion vectors already calculated in the encoder, and when the motion vectors are unavailable or of insufficient quality, lightweight optical flow is called for supplementary estimation to achieve complexity analysis. The computational complexity of the lightweight optical flow is configured to not exceed 1 / 10 of that of the standard dense optical flow. Based on the estimation, the GOP segmentation boundary is dynamically determined to generate multiple parallel task units, and the segmentation is constrained by the minimum and maximum number of frames and the boundary to avoid recoding, and online re-evaluation is supported for task units that have not yet started encoding.

[0010] In step 3, the global coordination node continuously monitors and maintains the dynamic state table of all physical computing nodes, specifically including: The dynamic status table of all physical computing nodes includes short-term availability prediction, prediction confidence, real-time resource utilization, idle status, and computing role for each record item. The prediction window for the short-term prediction is 1 second to 60 minutes or 1 to 300 frames. When any prediction confidence is lower than 0.7 or the prediction variance exceeds a preset threshold, priority resampling is triggered. The physical computing node immediately re-collects real-time telemetry data, regenerates the short-term availability prediction and confidence based on the latest data, and reports it to the global coordination node to update the dynamic status table.

[0011] In step 4, the load balancing strategy specifically includes: Based on short-term smooth CPU / GPU utilization metrics, estimate the GPU / CPU speedup ratio S_t for each task; determine whether a task should use CPU or GPU based on S_t and a dynamic threshold; support idle computing units to actively pull tasks from the busy queue, and evaluate whether the benefits cover the costs before migration. If the costs are covered, execute the task; otherwise, do not execute it.

[0012] The rescheduling described in step 5 specifically includes: Based on the initial priority of the task, the current waiting time of the task and the number of dependent tasks are dynamically weighted to generate the latest priority; the cost of various recovery strategies is estimated for the affected task, including but not limited to differential recovery, recalculation and shadow switching; the target node with the lowest cost weight and that meets the node capacity and prediction confidence constraints is selected to perform the migration, and the status table is updated after completion.

[0013] Compared with the prior art, the present invention has the following advantages: Existing video coding task scheduling schemes mostly employ static rules or rely solely on system load indicators, failing to establish real-time, dynamic linkage between video content complexity and heterogeneous computing resource status. This leads to resource mismatch, low energy efficiency, and unstable service quality. This invention fundamentally solves these problems through a dynamic allocation method for video coding tasks in many-core heterogeneous computing environments. Its innovation and technical advantages are specifically reflected in the following aspects: 1) This invention innovatively proposes a dynamic election of a global coordinating node. Unlike statically designated master nodes, this invention dynamically elects a coordinating node based on network topology, hardware capabilities, and real-time load, and this node can intelligently switch between system coordinator and computing participant. This solves the pain point of idle master node resources in traditional master-slave architectures, transforming system management overhead into usable computing power and significantly improving the overall resource utilization of the cluster; 2) This invention innovatively proposes predictive adaptive GOP segmentation. Abandoning the fixed GOP length, this invention predicts and determines the optimal segmentation point in real time based on the motion energy and scene changes of the video content, generating task units with adaptive sizes. Furthermore, by directly calculating the task structure, temporal domain, and spatial domain features, its computational complexity is objectively quantified before scheduling, providing precise input for intelligent scheduling and fundamentally avoiding uneven load caused by inappropriate task granularity. 3) This invention innovatively proposes a predictive enhanced dynamic state table. It not only records the current utilization of GPU and CPU cores, but more importantly, maintains short-term availability predictions and prediction confidence levels. This allows the scheduler to predict the future workload of resources and proactively avoid assigning tasks to resources that are currently idle but about to become overloaded, thereby significantly reducing performance overhead such as task queuing and later migration, and improving system throughput and stability.

[0014] 4) This invention innovatively proposes intelligent task allocation based on speedup ratio and dynamic threshold. It estimates the speedup ratio of each video task on the GPU relative to the CPU and dynamically adjusts the allocation threshold based on real-time load. This makes task allocation decisions data-driven and adaptive, ensuring that GPU resources are used for tasks that yield the greatest benefit, achieving precise extraction of performance from heterogeneous hardware.

[0015] 5) This invention innovatively proposes a multi-strategy cost-optimal fault recovery method. When a node fails, this invention does not simply restart the task. Instead, it prioritizes criticality and quantifies the costs of various strategies such as migration, differential recovery, and recalculation. Under the constraints of target node capacity and stability, it selects the scheme with the lowest global recovery cost. This preserves the completed computational work to the greatest extent possible, minimizing the impact of the fault on the overall system performance. Attached Figure Description

[0016] Figure 1 System architecture diagram for a dynamic allocation method of video coding tasks for many-core heterogeneous computing environments.

[0017] Figure 2 A flowchart illustrating the overall process of dynamically allocating video coding tasks for many-core heterogeneous computing environments.

[0018] Figure 3 This is a diagram illustrating the election of a global coordinating node.

[0019] Figure 4 This is a predictive adaptive GOP segmentation process and decision graph.

[0020] Figure 5 This is a schematic diagram of the status table data structure and prediction fields.

[0021] Figure 6 This is a flowchart of the load balancing strategy and task stealing decision-making process.

[0022] Figure 7 This diagram illustrates anomaly recovery and multi-strategy cost optimization.

[0023] Figure 8 This is a performance comparison chart between the solution described in this invention and a traditional solution. Detailed Implementation

[0024] The following section, with reference to the accompanying drawings, further explains a method for dynamically allocating video coding tasks in a many-core heterogeneous computing environment.

[0025] like Figure 1 The diagram shows a predictive GOP segmentation and load balancing system architecture for many-core heterogeneous computing, illustrating the end-to-end processing flow of video encoding tasks on a many-core heterogeneous computing platform: 1) Input Video Stream: The input video stream can come from various sources, including but not limited to real-time video streams, stored video files, or live streams. The video stream is input into the system in the form of a frame sequence and then processed by the subsequent GOP segmentation module.

[0026] 2) The GOP segmentation module logically contains several sub-modules: This module contains the following processing logic. First, it reuses the motion vectors already calculated within the encoder to obtain inter-frame motion information. When the motion vectors are unavailable or of insufficient quality, it calls lightweight optical flow for supplementary estimation to obtain inter-frame motion and content change information. Based on this information, it performs dynamic GOP segmentation. Then, it extracts the temporal features (such as motion intensity) and spatial features (such as texture complexity and detail richness) of the segmented video segments, calculates the comprehensive complexity score for each candidate GOP, and this score is mapped to the initial scheduling priority of the task. (Comprehensive Complexity Score) The calculation formula is as follows, and this score will be mapped to the initial scheduling priority of the task: in The overall complexity score of the current GOP task unit. The higher the value, the greater the computational complexity of the task and the more computing resources are required. The weighting coefficients for structural features, temporal features, and spatial features are respectively determined by offline benchmark testing based on the performance differences of the target heterogeneous hardware platform when handling different feature loads. To extract structural features based on GOP length and B-frame ratio, Temporal features of the extracted average motion energy, motion energy variance, and scene transition markers; Spatial features of image entropy and texture density for extracted keyframes.

[0027] 3) Within the global coordination node: The global coordination node maintains a dynamic status table that records the real-time resource status of each physical computing node. For example... Figure 1 As shown in the example, the dynamic status table records CPU1 utilization at 45%, CPU2 utilization at 12%, GPU1 utilization at 72%, and GPU2 utilization at 69%. The global coordination node performs dynamic task allocation based on the task units (including initial priorities) generated by the GOP partitioning module and the node states in the dynamic status table. Simultaneously, the global coordination node continuously receives status information reported by each compute node and updates the dynamic status table in real time.

[0028] 4) Computing Node Cluster: The computing node cluster consists of multiple heterogeneous computing nodes, including multiple CPU nodes and multiple GPU nodes. At least some CPU nodes are equipped with many-core processors (e.g., 64 cores), suitable for handling low-complexity or logic control tasks. At least some GPU nodes are equipped with GPUs with thousands of cores (e.g., 3584 CUDA cores), suitable for handling computationally intensive, highly parallel video encoding tasks. Each node receives task units assigned by the global coordinating node, performs the actual video encoding work, and continuously feeds back its running status to the coordinating node, forming a closed-loop dynamic scheduling system.

[0029] like Figure 2 As shown, the overall flowchart of the video coding task dynamic allocation method for many-core heterogeneous computing environments includes: 1) Election of a global coordinating node. When the system starts up, a global coordinating node that has both coordination and computing roles is dynamically elected from the heterogeneous cluster of many cores based on multi-dimensional priority.

[0030] 2) Receive telemetry and features. The global coordination node continuously receives system telemetry data (including CPU and GPU utilization, memory usage, temperature, task queue length, etc.) from each physical computing node, as well as encoded features from the video input.

[0031] 3) Predictive GOP segmentation. Predictive adaptive GOP segmentation is performed on the input video stream based on inter-frame motion information and content change information, generating parallel task units that adapt to the complexity of the content.

[0032] 4) Prediction and Dynamic Status Table Maintenance. The global coordination node maintains a dynamic status table that integrates short-term prediction values, enabling proactive monitoring of the resource status of each physical computing node and providing predictive data support for intelligent scheduling.

[0033] 5) Node Anomaly Detection and Judgment. The process sets anomaly detection points at key stages, forming decision branches. If no anomaly is detected (the "No" branch in the diagram), the system returns to normal monitoring via the dotted path, ensuring low-overhead operation; if node performance anomalies or faults are detected (the "Yes" branch in the diagram), the subsequent cost-aware recovery process is triggered.

[0034] 6) Identify affected tasks. When an anomaly is triggered, the system identifies all affected task sets on the faulty node and analyzes their dependencies and criticality to ensure that critical path tasks are restored first.

[0035] 7) Cost estimation. The system quantifies and evaluates the combined cost of multiple recovery strategies for each affected task, including strategies such as migration, differential recovery, recalculation, and shadow switching, transforming recovery decisions into a cost optimization problem.

[0036] 8) Cost-aware scheduling. Based on the cost model, and under the constraints of target node capacity and prediction confidence, the most cost-effective recovery strategy and target node are selected for each task.

[0037] 9) Perform recovery and update the status. Execute the recovery action and update the system status. The system records the difference between the cost estimate and the actual cost for online learning and model optimization, forming a closed loop of continuous self-improvement.

[0038] like Figure 3 The diagram illustrates the election of the global coordinating node. The election of the global coordinating node is based on a multi-dimensional comprehensive evaluation of candidate nodes, specifically including the following four dimensions: Dimension 1: Network Topology Centrality Assessment. This measures the importance of a node's position in the cluster network, quantified by calculating the reciprocal of the average communication delay between that node and all other nodes in the cluster. Dimension Two: Static Hardware Computing Power Assessment. This assesses the hardware conditions of the node, including CPU and GPU configurations, memory capacity, and other factors. Dimension 3: Future Availability Time-Series Rolling Prediction: Based on the node's historical and real-time telemetry data, a lightweight prediction model running locally on the node generates a short-term availability prediction value for that node. Dimension 4: Historical Performance Confidence Score. This assesses the stability of a node's historical operation and quantifies the reliability of predictions.

[0039] When the system initializes or the original coordinating node fails, an election mechanism is triggered. A weighted average priority is calculated for each node, and the node with the highest election score becomes the global coordinating node. The election results are then broadcast. (Node overall priority) The calculation formula is as follows: in The overall priority score represents the i-th node, and the node with the highest score is selected as the global coordinating node. The network topology centrality metric is quantified by calculating the reciprocal of the average communication delay between the node and all other nodes in the cluster. The higher the value, the more "central" the node is in the network, and the higher the communication efficiency. Hardware management capability metrics quantify whether a node possesses dedicated management hardware and the stability of its historical operation; Short-term availability forecast, generated by a lightweight prediction model running locally on the node. The model takes the node's historical telemetry data and real-time operational telemetry data as input and outputs a predicted "health" or "idleness" value for the node's availability to perform coordination tasks in the next short-term period (within the next scheduling cycle). The value typically ranges from [value missing in original text]. Historical telemetry data refers to the historical records formed by persistently storing real-time telemetry data in a time series. It is stored in the time series database of the global coordination node, with a time span ranging from minutes to weeks. Real-time telemetry data refers to the current operating status data collected in real time by the monitoring agent deployed locally on the physical computing node, including but not limited to indicators such as CPU and GPU utilization, memory usage, temperature, and power consumption. These are preset weighting coefficients for network topology centrality, hardware management capability, and short-term availability prediction, respectively, and satisfy the following conditions: .

[0040] Historical telemetry data refers to the historical records formed by persistently storing real-time telemetry data in a time series. These records are stored in the time-series database of the global coordination node, with a time span ranging from minutes to weeks. Real-time telemetry data refers to the current operating status data collected in real time by the monitoring agent deployed locally on the physical computing node, including but not limited to indicators such as CPU and GPU utilization, memory usage, temperature, and power consumption.

[0041] The elected global coordinating node is software-configured to have two dynamically switchable operating modes: in coordination mode, the node primarily performs system-level coordination functions such as task distribution, status monitoring, and load balancing; in computing mode, the node can release all or part of its physical computing cores (CPU cores and GPU stream processors), allowing it to participate in specific computing tasks such as video encoding as a regular computing unit. The switching of operating modes is automatically decided by the global coordinating node based on the system's global load status. When the coordinating node itself fails, the election module can trigger a re-election.

[0042] like Figure 4 As shown, the predictive adaptive GOP segmentation process and decision graph illustrate the entire process from video input to the generation of schedulable task units in this invention. The process executes the following steps sequentially: 1) Input video frames: The starting point of the process, receiving a continuous sequence of video frames as the processing object.

[0043] 2) Motion Energy Calculation and Scene Switch Detection: The input frame sequence is analyzed. The core method prioritizes the use of motion vectors within the encoder and supplements calibration with a lightweight optical flow algorithm as needed. This efficiently and accurately quantifies inter-frame motion complexity while simultaneously detecting scene switch information. Motion energy... The calculation can be simplified to: in, Indicates the first The motion energy value of a frame represents the degree of motion intensity of that frame relative to the previous frame; the larger the value, the more intense the motion between frames. For the first The sum of the frame motion vector magnitudes, which reuses the result already calculated within the encoder to avoid redundant calculation overhead; This is a correction term for optical flow. , These are weighting coefficients used to balance the contributions of the motion vector and optical flow correction term to the final motion energy, satisfying... .

[0044] 3) Short-term trend and slope analysis: Based on the calculated motion energy sequence, its short-term trend and instantaneous slope are analyzed to dynamically perceive the stability and complexity fluctuations of video scenes. Motion energy slope. The calculation formula is as follows: It is the first The motion energy gradient of a frame represents the instantaneous rate of change of motion energy; a positive value indicates increased motion, while a negative value indicates decreased motion. No. The motion energy value of the frame; No. The motion energy value of the frame; It is the time interval between two adjacent frames.

[0045] 4) GOP Boundary Candidate Generation: Combining motion trend and slope analysis results with scene transition detection information, the system intelligently predicts and generates potential segmentation boundary points for image groups, forming a preliminary task unit partitioning scheme. Specifically, GOP boundary candidate generation is triggered when any of the following conditions are met: motion energy slope. The absolute value exceeds the preset threshold, a scene switch is detected, or the current GOP cumulative frame count reaches the preset upper limit.

[0046] 5) Constraint handling (parallel branching): Two key constraints are imposed on the candidate boundary to ensure feasibility: Minimum / maximum frame count constraint: Forces the length of each GOP to fall within a preset reasonable range (usually the minimum frame count is set to 8 frames and the maximum frame count is set to 120 frames), avoiding the generation of tasks that are too long or too short.

[0047] Avoid recoding boundaries: Ensure that the segmentation boundary is located in a frame that can be independently encoded (such as an I-frame) to prevent additional recoding computation overhead caused by improper segmentation.

[0048] 6) Online Re-evaluation: This is a feedback optimization step. The system continuously monitors the segmented task units that have not yet started encoding. If the complexity of the subsequent video content changes unexpectedly and drastically, it can trigger re-evaluation and secondary segmentation, demonstrating the adaptive capability of the process. The trigger condition for online re-evaluation is: the ratio of the motion energy of the newly arriving frame to the average motion energy of the current GOP exceeds a preset threshold (usually set to 2.0).

[0049] 7) Parallel Task Unit Output: The endpoint of the process. Outputs a series of GOPs of varying lengths and complexities, which are delivered as independent parallel task units to the subsequent scheduling system for resource allocation and execution.

[0050] like Figure 5 The diagram illustrates the composition and key fields of the core dynamic status table used for resource monitoring and scheduling in this invention. The diagram clearly presents the two-dimensional table structure of the status table, which includes the following core elements: 1) Table Body: The first column, Computation Unit, lists the specific physical computing unit identifiers, such as Node_0_CPU_Core_0 and Node_1_GPU_SM_1, representing the monitored CPU cores and GPU stream processor clusters. This serves as the finest-grained index for resource monitoring and task binding in the distributed cluster. The second column, Computation Role, clearly defines the heterogeneous attributes of each unit's logical control or computationally intensive nature. The third column, Current Value, records the latest real-time resource utilization of the corresponding computing unit (e.g., 42.3% and 89.4%), reflecting the instantaneous load of the hardware at the current time step. The fourth column, Predicted Value, and the fifth column, Confidence, together constitute the Prediction Enhancement Field. This field not only provides a short-term load window prediction (e.g., 46.0% and 93.0%) but also gives the prediction confidence score (e.g., 0.95 and 0.78), providing a quantitative basis for forward-looking scheduling decisions. When the confidence score is lower than a preset safety threshold (e.g., 0.58 as shown in row 4), it serves as a global interception anchor point for triggering state resampling and abnormal rescheduling, ensuring the execution security of dynamic task allocation in an uncertain heterogeneous environment.

[0051] 2) Functional Description: This is the core data structure of the global coordination node. For each physical computing unit in the table, its key record items (such as core utilization, memory bandwidth usage, and cache miss rate) contain three sub-fields: Real-time value: The latest value collected directly through operating system interfaces (such as the / proc file system in Linux, the nvidia-smi command, etc.); Predicted values: Based on the historical time series of this indicator, the predicted values ​​for the near future are generated by a lightweight time series prediction model (such as exponential smoothing, ARIMA, or a small LSTM network), with a prediction window of 1 second to 60 minutes or 1 to 300 frames. Confidence level: a level between The confidence level quantifies the reliability of the prediction. It can be calculated based on the error of the prediction model on recent historical data, the variance of the prediction sequence, or the uncertainty output of the model itself. The closer the confidence level is to 1, the more reliable the prediction result. Prediction Confidence Level The calculation formula is as follows: in, This refers to the current confidence score of the prediction, with a range of values. The closer the value is to 1, the more reliable the prediction result; The mean absolute error of the prediction model over a recent historical window (such as the previous 100 time steps); The baseline mean absolute error is the average absolute error over the same historical window using a simple forecasting method (such as zero-order hold or moving average) as the benchmark. A very small positive number, used to prevent division by zero errors.

[0052] When the prediction confidence of any node is lower than a preset threshold (e.g.) Figure 5 middle When the node is in a state of emergency, priority resampling is triggered, which means that the real-time telemetry data of the node is immediately re-collected, the short-term availability prediction value and confidence level are regenerated based on the latest data, and the dynamic status table is updated and reported to the global coordination node.

[0053] By integrating real-time monitoring data and predictive information, the dynamic status table provides key decision-making basis for subsequent intelligent scheduling algorithms, enabling the scheduler to predict the future busyness of resources and proactively avoid assigning tasks to resources that are "idle at the moment but about to be overloaded".

[0054] like Figure 6 The diagram illustrates the core decision-making process for task scheduling and resource optimization in this invention. It employs a clear hierarchy, presenting a two-tiered strategy: basic task allocation and proactive performance optimization. 1) Core allocation decision First, features of the GOP task to be scheduled are extracted to estimate the speedup of the task on heterogeneous nodes. Simultaneously, based on the real-time load status of the CPU and GPU in the system, a dynamically adjusted decision threshold is calculated. Then, the decision is made using a diamond-shaped decision box. ”: If the condition is met (the "Yes" branch in the diagram), the task will be assigned to the GPU queue for execution. If the condition is not met (the "No" branch in the diagram), the task will be assigned to the CPU queue for execution.

[0055] This mechanism ensures that computing tasks are intelligently directed to hardware units that can achieve higher performance gains.

[0056] acceleration ratio The calculation formula is as follows: Task The estimated execution speed ratio of the GPU relative to the CPU. This indicates that the GPU executes faster. This indicates that the CPU executes faster; Task Estimated execution time on CPU, scored based on task complexity. (from) Figure 1 The calculation includes a comprehensive complexity score and historical performance profiles of CPU nodes. Task Estimated execution time on GPU, based on task complexity score. Calculate the historical performance profile of CPU nodes.

[0057] Dynamic threshold The calculation formula is as follows: A dynamically adjusted decision threshold is used to determine whether a task should be assigned to the GPU; The baseline threshold is pre-calibrated through offline benchmark testing and is typically set to 1.2. The adjustment coefficient controls the degree to which load differences affect the threshold; its value range is typically [value range missing]. ; Average load of the GPU cluster (normalized utilization). Average load of CPU cluster (normalized utilization). A very small positive number, used to prevent division by zero errors.

[0058] 2) Once a task is assigned to the GPU, the system continuously monitors the GPU status. It uses a diamond-shaped decision box to determine if the GPU is idle and the local queue is below a critical value. If the condition is met (the "Yes" branch in the diagram), the GPU will attempt to steal high-performance tasks from the CPU task queue. Tasks are moved to the GPU; If the conditions are not met (the "No" branch in the diagram), the status quo is maintained and the theft is not performed.

[0059] Before performing a data theft operation, the system needs to conduct a two-stage benefit and cost analysis. Specifically, a benefit assessment diamond-shaped checkbox is used to determine whether the benefit model passes, i.e., whether the performance benefits of migration outweigh the data transfer costs. The theft operation is only performed when the benefits outweigh the costs, ensuring that the performance benefits of migration exceed the costs of data transfer and thus avoiding ineffective migrations.

[0060] The formula for calculating the benefit assessment is as follows: This is the net benefit of performing the task migration. A value greater than 0 indicates that the migration is effective; otherwise, the migration is invalid. This refers to the time saved by executing a task on a GPU compared to a CPU. The calculation formula is: in and The definition is the same as in the calculation of the speedup ratio.

[0061] Migration costs, including data transfer overhead (transferring intermediate task state data from CPU memory to GPU memory) and context switching overhead, are calculated using the following formula: in For the task The size of intermediate state data, Effective bandwidth of the PCIe bus between the CPU and GPU This is a fixed overhead for context switching.

[0062] If the benefit assessment is approved (i.e.) If the task is not selected, the task will be stolen and pulled from the CPU's high-priority queue to the GPU for execution; otherwise, the stealing will be rejected to prevent invalid migration and maintain the original encoding calculation state.

[0063] like Figure 7 The diagram illustrates the complete decision-making and execution process of the system's intelligent recovery when a computing node fails, as described in this invention. This includes: 1) Fault Detection. The system detects a fault in a computing node. Fault detection can be achieved through a heartbeat mechanism; if no response is received from the node for several consecutive heartbeat cycles, the node is determined to be in a faulty state.

[0064] 2) Task Analysis. Identify all tasks affected by the failure and sort them by criticality (e.g., ...). Figure 7 Example: (Task A is marked as "Critical", Task B as "Medium", Task C as "Low"), ensuring that critical path tasks are processed first. When node After being marked as a fault, the system first identifies all those assigned to Task Collection Then, based on the dependency topology graph between tasks, calculate the dependency relationship of each affected task. Latest priority The calculation formula is as follows: in, Affected tasks The criticality score indicates that the higher the recovery priority, the more urgent the task needs to be handled. It is a task In a dependency graph, the depth or out-degree represents the number of subsequent tasks that are blocked. The larger the value, the greater the impact of the task on the overall job completion time. It is a task The estimated remaining computational cost is scored based on the initial task complexity. Estimate the difference between the current progress and the completed progress; These are weighting coefficients used to balance the contribution of dependency resistance and computational cost to criticality, satisfying the following conditions: ,generally Set to 0.6. Set it to 0.4 to prioritize critical path tasks.

[0065] System by Tasks are processed in descending order to ensure that tasks on the critical path are restored first, minimizing the impact of failures on the overall job completion time.

[0066] 3) Cost Assessment. For each task to be recovered, the estimated cost of different recovery strategies is evaluated in parallel. The example in the figure shows a cost comparison of three typical strategies: "Migration" (28ms), "Recovery" (15ms), and "Recalculation" (95ms). Among them, migration refers to switching to a hot standby node to continue execution, which is suitable for scenarios with a small amount of task status data; differential recovery refers to continuing execution from the most recent checkpoint, only migrating incremental data; recalculation refers to abandoning existing results and starting execution from scratch.

[0067] For each affected high-critical task The system constructs a list of candidate recovery nodes from the currently available set of healthy nodes. For each candidate node... and each candidate recovery strategy Estimate its total recovery cost The calculation formula is as follows: in, This refers to the affected tasks Through recovery strategies Restore to candidate node Total recovery cost; This refers to the cost of data migration, which refers to the cost of the task. Intermediate state data from the fault node Migrate to target node The cost is calculated using the following formula: in For the task The size of intermediate state data, The available network bandwidth from the source fault node to the target node; This refers to the cost of recovery from checkpoints, including the overhead of loading checkpoint data from persistent storage and replaying incremental logs, and is only included when a differential recovery strategy is used; This refers to the recalculation cost, i.e., the cost of the task at the candidate node. The estimated execution time from scratch can be scored based on the initial complexity of the task. and nodes Historical performance data estimation; This refers to the shadow switching cost, which is the cost required at the target node during the recovery process. The overhead of reserving resources and establishing a temporary execution environment includes memory pre-allocation, context initialization, etc. This refers to the normalized weighting coefficient of each cost item, used to unify cost items with different dimensions into a comparable range, satisfying... .

[0068] 4) Select strategy and execute recovery. The system selects the strategy with the lowest total recovery cost. Combination, that is, intelligently selecting the recovery strategy with the lowest cost through the decision diamond "selection of the best". Figure 7 The decision diamond splits into three branches, corresponding to the three strategy options: migration, recovery, and recalculation. The selected strategy must satisfy the following constraints: Capacity constraint: target node The predicted free resources (from the dynamic state table) must be able to satisfy the task. The resource requirements, namely: Confidence constraint: target node The prediction confidence level for the predicted idle resources must be higher than a safety threshold. (Usually set to 0.7), that is: To prevent the recovery process from overwhelming the system, all recovery operations are centrally controlled by a global coordination node: Concurrency limit: The number of instances performing state transitions or task recovery simultaneously must not exceed a preset limit (usually set to 20% of the number of cluster nodes). Rate limiting: Sets a rate limit on the use of network bandwidth and node I / O resources to ensure that traffic recovery does not affect the execution of normal foreground tasks; Chunked Incremental Migration: For large task states, chunked incremental migration is supported, allowing tasks to start computation on the target node as soon as some states are in place, further shortening recovery time.

[0069] 5) Update the status table. After recovery is complete, the global coordinating node atomically updates the dynamic status table, marking the failed node as offline and updating the binding relationships between the recovered tasks and the new nodes to ensure the consistency of the entire cluster view. Specific operations include: deleting all records of the failed node in the dynamic status table, adding a new task binding record under the target node for each recovered task, and updating the prediction confidence score of the affected nodes.

[0070] like Figure 8 As shown, the performance comparison chart between the solution described in this invention and the traditional solution includes: 1) This figure compares in detail the traditional static master-slave scheduling method, the fixed rule-based load balancing method, and the multi-core heterogeneous computing resource scheduling method based on dynamic roles and predictive adaptive methods proposed in this invention in terms of system architecture intelligence, resource utilization efficiency, fault recovery capability, and load balancing effect.

[0071] Among them, the system's average coding throughput efficiency is measured, with higher values ​​indicating better coding efficiency. Comparison results show that: In low-dynamic video load scenarios, the video content is less complex. Traditional static master-slave scheduling methods result in some resources being idle due to fixed resource allocation strategies, with an encoding efficiency of only 85%. Fixed rule load balancing methods have improved this by using simple queue length scheduling, reaching 90%. This invention further improves the encoding efficiency to 98% through predictive adaptive GOP segmentation and intelligent load balancing.

[0072] In typical mixed video load scenarios, the complexity of video content fluctuates to a moderate degree. Traditional static master-slave scheduling still maintains an efficiency of 70%; fixed-rule load balancing, unable to dynamically adjust according to content complexity, only improves efficiency to 78%; this invention, through speedup prediction and dynamic threshold-based traffic splitting, significantly improves encoding efficiency to 94%, demonstrating the most obvious advantage.

[0073] In scenarios with high-intensity motion and high-dynamic loads, video content is extremely complex and fluctuates wildly. Traditional static master-slave scheduling methods cannot cope with drastic changes in complexity, resulting in a coding efficiency of only 42%. Fixed-rule load balancing methods improve efficiency to 55% through simple load awareness. This invention, through predictive GOP segmentation, proactive monitoring of dynamic state tables, and cost-aware fault recovery mechanisms, maintains a high coding efficiency of 89% even in high-dynamic scenarios, more than doubling the efficiency of traditional solutions.

[0074] 2) In summary, the video coding task dynamic allocation method proposed in this invention for many-core heterogeneous computing environments has achieved better coding efficiency than traditional methods in three typical scenarios: low dynamic video load, conventional mixed video load, and high-intensity motion / high dynamic load. In particular, the performance advantage is most significant in high complexity and high dynamic scenarios, which verifies the effectiveness and superiority of the method proposed in this invention.

Claims

1. A method for dynamically allocating video coding tasks in a many-core heterogeneous computing environment, characterized in that, Includes the following steps: Step 1: A many-core heterogeneous computing environment is a heterogeneous cluster consisting of multiple physical computing nodes interconnected by a network. At least some of the nodes use many-core architecture processors. Calculate the priority of all nodes in the many-core heterogeneous computing environment and elect a global coordinating node from multiple physical computing nodes based on the node priority. Step 2: Perform predictive adaptive GOP segmentation on the input video stream based on inter-frame motion information and content change information to generate multiple parallel task units; calculate the comprehensive complexity score based on the structure, temporal and spatial characteristics of each parallel task unit, and directly map the score to the initial scheduling priority of the task. Step 3: The global coordination node continuously monitors and maintains the dynamic status table of all physical computing nodes, and the dynamic status table records at least the resource status of each physical computing node; Step 4: The global coordination node, based on the initial scheduling priority obtained in Step 2 and the resource status in the dynamic status table obtained in Step 3, dynamically allocates the multiple parallel task units in Step 2 to physical computing nodes for execution through a load balancing strategy. Step 5: When any physical computing node is detected to have a performance abnormality or failure, the latest priority is obtained by dynamically weighting the task scheduling priority determined in Step 2 and combining it with the system status of the current task with the performance abnormality or failure. The global coordination node immediately marks this physical computing node as unavailable from the dynamic status table, obtains the new resource status, and triggers the rescheduling of all tasks allocated to this physical computing node. The rescheduling re-executes the dynamic task scheduling steps according to the latest priority of the task and the status of the remaining available units in the system to realize the dynamic allocation of video encoding tasks.

2. The method for dynamic allocation of video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 1, the multiple physical computing nodes include multiple CPU nodes and multiple GPU nodes.

3. The method for dynamically allocating video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 1, the priority of all nodes in the many-core heterogeneous computing environment is calculated. Specifically, on each physical computing node, the availability prediction value of the node is generated by taking real-time telemetry data as input and parameters trained by historical telemetry data as a benchmark. The node availability prediction value is then combined with the network topology centrality index and hardware management capability index of the physical computing node according to preset weights to calculate the priority of the node.

4. The method for dynamically allocating video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 2, the inter-frame motion information and content change information specifically include inter-frame motion energy and its trend, motion energy gradient and scene switching information.

5. The method for dynamic allocation of video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 2, predictive adaptive GOP segmentation is performed on the input video stream based on inter-frame motion information and content change information to generate multiple parallel task units, specifically including: The GOP segmentation boundary is determined based on the inter-frame motion energy and its trend, the motion energy gradient, and scene switching information. Multiple parallel task units are then generated based on the GOP segmentation boundary.

6. The method for dynamically allocating video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 3, the resource status includes: availability prediction value, prediction confidence, real-time resource utilization, idle status, and computing role.

7. The method for dynamic allocation of video coding tasks for many-core heterogeneous computing environments according to claim 1, characterized in that, In step 4, the load balancing strategy specifically includes: Based on short-term smooth CPU / GPU utilization metrics, estimate the GPU / CPU speedup ratio S_t for each task; determine whether a task should use CPU or GPU based on S_t and a dynamic threshold; support idle computing units to actively pull tasks from the busy queue, and evaluate whether the benefits cover the costs before migration. If the costs are covered, execute the task; otherwise, do not execute it.