High-fidelity cloud rendering cluster scheduling method and system

Through multi-level priority task queue, preemptive and rotation scheduling, spatial chunking and time framing strategies, and GPU direct output video streams, the resource allocation and task scheduling of cloud rendering clusters are optimized, and the cloud rendering cluster's low resource utilization rate, high task delay and slow failure recovery are solved, achieving efficient and stable rendering effects.

CN120371468APending Publication Date: 2025-07-25SHENZHEN TRAFFIC CONSTR ENG TEST & DETECTION CENT +1

Patent Information

Application Number
CN202510414374.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing cloud rendering cluster technology has shortcomings in resource utilization, task execution efficiency, delay control and cost optimization, especially in real-time load changes in heterogeneous computing nodes and node exception handling.

Method used

Multi-level priority task queue, preemptive and rotation scheduling strategies, spatial chunking and time framing strategies, incremental task updates, GPU direct output video streaming and other technologies are adopted, and resource allocation and task scheduling are optimized in combination with dynamic resource allocation and fault tolerance mechanisms.

Benefits of technology

Improve resource utilization, shorten task response time, reduce latency and fault recovery time, and meet the needs of high realistic, low latency and industrial-grade stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371468A_ABST
    Figure CN120371468A_ABST
Patent Text Reader

Abstract

The invention relates to a high-fidelity cloud rendering cluster scheduling method and system, and the method comprises the steps: receiving a rendering task, generating a multi-level priority queue according to the task complexity, timeliness and resource demand classification, and automatically optimizing a scheduling strategy; monitoring the resource state of the heterogeneous computing node in real time; allocating tasks by using preemptive and round-robin scheduling strategies, and adjusting and coping with resource fluctuation in combination with a dynamic code rate; the tasks are decomposed by adopting a spatial blocking and time framing strategy, and an execution sequence is controlled according to a topological sorting algorithm; abnormal nodes are identified through heartbeat detection, and affected subtasks are migrated through incremental task updating. According to the method, efficient resource allocation, scheduling algorithm and idle key frame scheme can be realized, the GPU directly outputs the video stream, the rendering speed is remarkably improved, resource waste and operation cost are reduced through a dynamic resource allocation mechanism, a fault-tolerant mechanism is provided, the task is ensured to be normally completed under the condition of node fault, and the task efficiency is improved. And distributed rendering and synthesis of large-scale high-resolution images are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing and view optimization, and specifically, to a high-fidelity cloud rendering cluster scheduling method and system. Background Art

[0002] In recent years, the demand for high-fidelity real-time rendering in fields such as virtual reality (VR), augmented reality (AR), movie special effects, and digital twins has increased rapidly, driving the widespread application of cloud rendering cluster technology. Cloud rendering integrates heterogeneous resources such as high-performance CPUs and GPUs through a distributed computing architecture, and can efficiently process complex 3D scenes, high-resolution images, and dynamic light and shadow effects, significantly improving the rendering efficiency. However, existing cloud rendering scheduling technologies and data transmission technologies still have deficiencies.

[0003] For example, although the existing dynamic scheduling method for cloud rendering tasks proposes a dynamic resource allocation mechanism based on task priorities, its scheduling strategy does not fully consider the real-time load changes of heterogeneous computing nodes, resulting in rigid resource allocation and difficulty in guaranteeing the real-time performance of high-priority tasks. In addition, task decomposition only relies on the time frame strategy and lacks collaborative optimization with spatial partitioning, limiting the parallel rendering efficiency of complex scenes.

[0004] Again, although the existing cloud rendering data transmission method based on WebRTC optimizes the video stream transmission delay, its task scheduling module still relies on the CPU for global management, resulting in a high end-to-end delay in the encoding and transmission links. At the same time, a full-task retry mechanism is adopted in node exception handling, resulting in low fault recovery efficiency and affecting system reliability.

[0005] Therefore, there is an urgent need in the industry for a cloud rendering cluster scheduling method that can dynamically adapt to heterogeneous resources, implement multi-level priority scheduling, optimize the parallel decomposition strategy, and enhance the fault tolerance ability to meet the stringent requirements of industrial-level real-time rendering. Summary of the Invention

[0006] To solve the problems existing in the existing cloud rendering cluster in terms of resource utilization rate, task execution efficiency, delay control, and cost optimization, the present invention proposes a high-fidelity cloud rendering cluster scheduling method and system.

[0007] The technical solution of the present invention is as follows:

[0008] A high-fidelity cloud rendering cluster scheduling method includes the following steps:

[0009] Receive the rendering tasks submitted by the user, classify the tasks based on task complexity, timeliness, and resource requirements, generate a multi-level priority task queue, and automatically optimize the scheduling strategy through the user interface;

[0010] Monitor the resource status of heterogeneous computing nodes in the cluster in real time, including the dynamic loads of CPUs and GPUs with different instruction set architectures and the network bandwidth utilization rate;

[0011] According to the priority and resource status, adopt a preemptive scheduling strategy to dynamically allocate high-priority tasks to low-load nodes, and allocate ordinary-priority tasks through a round-robin scheduling strategy, combined with dynamic bitrate adjustment to cope with resource fluctuations;

[0012] Adopt two strategies of spatial partitioning and time framing to decompose tasks into several subtasks, control the order of decomposition, distribution, and synthesis based on the topological sorting algorithm, and perform distributed distribution according to the node processing capacity, task load, and priority score;

[0013] During the execution of subtasks, identify abnormal nodes through heartbeat detection and task monitoring, and adopt an incremental task update strategy to dynamically migrate affected subtasks;

[0014] Perform cross-block boundary light and shadow correction, image stitching, and frame sequence encoding optimization on the rendering results of subtasks generated by each node to generate a final high-fidelity rendering output, and directly output an adaptive bitrate video stream to the client through the GPU.

[0015] As a preferred technical solution of the present invention, the preemptive scheduling strategy includes:

[0016] High-priority tasks can interrupt the subsequent subtasks of unexecuted low-priority tasks and retain the status of the executed subtasks;

[0017] The task execution module records the subtask progress and ensures the coherence of task recovery after preemption through status tracking;

[0018] The round-robin scheduling strategy includes:

[0019] Ordinary-priority tasks are distributed in a distributed manner according to the node load balancing principle, and are allocated to the adapted nodes for execution using the first-come-first-served mechanism based on the matching degree between the priority score of the subtasks after task splitting and the node processing capacity and task load;

[0020] The task execution module dynamically adjusts the subtask distribution order through the priority score algorithm.

[0021] As a preferred technical solution of the present invention, the dynamic bitrate adjustment includes:

[0022] Based on the changes in node load and task requirements, trigger GPU resource expansion, release, or resolution dynamic adjustment;

[0023] Adopt an idle resource recycling mechanism for low-priority tasks to adapt to the node resources with different instruction set architectures.

[0024] As a preferred technical solution of the present invention, the spatial partitioning uses an octree or a hierarchical bounding box algorithm to dynamically divide the three-dimensional scene into independent regions, and the temporal frame division decomposes the animation tasks into a parallel frame sequence according to a preset frame rate;

[0025] And the spatial partitioning takes precedence over the temporal frame division for collaborative decomposition, and ensures the execution order through the task dependency graph algorithm; the block merging takes precedence over the frame merging for collaborative merging.

[0026] As a preferred technical solution of the present invention, it further includes a key frame pre-computation step during idle time:

[0027] During the idle period of the node, pre-compute the key frames for scene switching, lighting changes, or motion paths;

[0028] When a new task arrives, perform frame interpolation compensation based on the pre-computed key frames, and the calculation process combines the optical flow estimation and motion compensation algorithms.

[0029] As a preferred technical solution of the present invention, the cross-block boundary light and shadow correction includes:

[0030] Use photon mapping or path tracing algorithms to correct the light and shadow consistency at the block boundaries;

[0031] Perform gradient fusion or Poisson fusion on the stitching area to eliminate visual gaps.

[0032] As a preferred technical solution of the present invention, the GPU directly outputs an adaptive bitrate video stream, including:

[0033] Directly complete the low-latency video compression encoding of the rendered frames through the GPU hardware encoder;

[0034] Combine DLSS super-resolution to optimize the picture quality, and dynamically adjust the output bitrate and resolution through a low-latency streaming protocol.

[0035] As a preferred technical solution of the present invention, the incremental task update strategy includes:

[0036] Only re-render the affected blocks or frame sequences;

[0037] Adopt healthy heartbeat detection and dynamic load balancing to quickly migrate the tasks of faulty nodes.

[0038] The present invention also provides a high-fidelity cloud rendering cluster scheduling system, including:

[0039] A task management module for receiving tasks through a user-friendly interface, automatically classifying and storing them, and defining multi-level priority rules;

[0040] A resource monitoring module for real-time collection of CPU, GPU instruction set architecture adaptation data and dynamic load of heterogeneous nodes;

[0041] A scheduling control module for performing multi-level priority scheduling, dynamic bitrate adjustment, decomposition strategy, and pre-computation of key frames during idle time;

[0042] A task execution module for supporting the collaborative decomposition of spatial partitioning and temporal framing and tracking the status of subtasks;

[0043] A result processing module for decomposing or merging block data and framed data, correcting boundary lighting, cross-block light and shadow correction, and directly outputting an adaptive bitrate video stream by the GPU.

[0044] As a preferred technical solution of the present invention, the result processing module includes:

[0045] A GPU video stream processing unit integrated with a hardware encoder for directly performing efficient video encoding on rendered frames;

[0046] A protocol adaptation unit supporting a low-latency real-time transport protocol stack for pushing the encoded stream to the client;

[0047] A dynamic adjustment unit for generating bitrate, resolution, and frame rate control instructions based on client network data and feeding them back to the GPU driver layer for execution through a software interface.

[0048] According to the present invention of the above solution, its beneficial effects are as follows:

[0049] Based on a dynamic scheduling strategy of multi-level priority task classification and real-time heterogeneous resource monitoring, the system of the present invention can accurately match task requirements with node load status, adopt preemptive scheduling to give priority to ensuring the real-time response of high-priority tasks, and at the same time achieve fair resource allocation for ordinary tasks through round-robin scheduling and dynamic bitrate adjustment, effectively solving the problems of low resource utilization rate and task delay accumulation in traditional solutions;

[0050] Through the collaborative decomposition strategy of spatial partitioning and temporal framing, combined with the topological sorting algorithm to optimize the task distribution order, the present invention breaks through the parallel processing bottleneck of complex three-dimensional scenes and high-frame-rate animations, and greatly improves the distributed rendering efficiency; the incremental task update strategy of the present invention quickly identifies and recovers subtasks of abnormal nodes through heartbeat detection and dynamic migration mechanisms, avoids resource waste caused by global task restart, and significantly shortens the fault recovery time;

[0051] Through the light and shadow correction across block boundaries and intelligent image stitching technology, combined with the design of directly outputting an adaptive bitrate video stream by the GPU, the present invention bypasses the traditional CPU encoding and transmission bottlenecks, significantly reducing the end-to-end latency while ensuring high-fidelity picture quality. Further, the present invention integrates idle resource pre-computation and end-to-end video stream optimization technology, and through key frame preloading and GPU hardware-accelerated encoding, significantly reduces the computational redundancy and transmission latency of real-time rendering.

[0052] It can be seen that each module of the present invention is deeply integrated to form a closed-loop optimization system, enabling the system to achieve an overall improvement in terms of heterogeneous resource adaptability, task execution efficiency, fault tolerance, and cost control, and being able to meet the stringent requirements of high fidelity, low latency, and industrial-grade stability in scenarios such as virtual reality and digital twin. Brief Description of the Drawings

[0053] Figure 1 is the flowchart of the method of the present invention;

[0054] Figure 2 is the system block diagram of the present invention;

[0055] Figure 3 is the schematic diagram of spatial block division in the decomposition strategy;

[0056] Figure 4 is the schematic diagram of temporal frame division in the decomposition strategy;

[0057] Figure 5 is the schematic diagram of the coordinated scheduling of block division and frame division in a preferred embodiment;

[0058] Figure 6 is the schematic diagram of spatial block merging;

[0059] Figure 7 is the schematic diagram of temporal frame merging;

[0060] Figure 8 is the schematic diagram of the execution process of idle key frame calculation in a preferred embodiment. Detailed Embodiments

[0061] In order to better understand the purpose, technical solution, and technical effects of the present invention, the present invention will be further explained below with reference to the drawings and embodiments. It should be noted that: similar reference numerals and letters denote similar items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, it is stated that the embodiments described below are only used to explain the present invention and are not used to limit the present invention.

[0062] It should be noted that when an element is referred to as "fixed to" or "disposed on" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time.

[0063] As Figure 1 shown, a high-fidelity cloud rendering cluster scheduling method includes the following steps:

[0064] Step 1: Receive the rendering tasks submitted by users, classify the tasks based on task complexity, timeliness, and resource requirements, generate a multi-level priority task queue, and automatically optimize the scheduling strategy through the user interface;

[0065] Step 2: Real-time monitor the resource status of heterogeneous computing nodes in the cluster, including the dynamic loads of CPUs and GPUs with different instruction set architectures and the network bandwidth utilization rate;

[0066] Step 3: According to the priority and resource status, adopt a preemptive scheduling strategy to dynamically allocate high-priority tasks to low-load nodes, and use a round-robin scheduling strategy to allocate normal-priority tasks, and combine dynamic bitrate adjustment to cope with resource fluctuations;

[0067] Step 4: Adopt two strategies of spatial partitioning and temporal framing to decompose the tasks into several subtasks, control the order of decomposition, distribution, and synthesis based on the topological sorting algorithm, and perform distributed distribution according to the node processing capacity, task load, and priority score;

[0068] Step 5: During the execution of subtasks, identify abnormal nodes through heartbeat detection and task monitoring, and adopt an incremental task update strategy to dynamically migrate the affected subtasks;

[0069] Step 6: Perform cross-block boundary light and shadow correction, image stitching, and frame sequence coding optimization on the rendering results of subtasks generated by each node to generate the final high-fidelity rendering output, and directly output an adaptive bitrate video stream to the client through the GPU.

[0070] Through real-time monitoring of CPU / GPU instruction set architecture adaptation data and dynamic loads, and combining preemptive and round-robin scheduling strategies, the present invention realizes that high-priority tasks are preferentially allocated to low-load nodes, and normal tasks are fairly scheduled based on resource status. Among them, the dynamic resource allocation algorithm adjusts the task distribution strategy according to the real-time load of nodes, such as GPU utilization and memory occupancy, avoiding resource waste or overload caused by traditional static scheduling, shortening the response time of high-priority tasks by more than 40%, and ensuring the timeliness requirements of industrial-level real-time rendering.

[0071] The present invention adopts a collaborative decomposition strategy of spatial partitioning and temporal framing. It dynamically divides the three-dimensional scene through the octree algorithm and combines parallel processing of frame sequences, breaking through the limitations of traditional methods that only rely on temporal framing. The dual decomposition strategy decomposes complex scenes into independent spatial regions and temporal segments, and through parallel rendering by distributed computing nodes, the utilization rate of distributed resources is increased by more than 30%, the rendering speed of large-scale scenes is doubled, and the execution cycle of high-complexity tasks such as movie-level special effects is significantly shortened.

[0072] Through the technology of directly outputting video streams by the GPU, the present invention uses the NVENC / AV1 hardware encoder to achieve zero-copy encoding and combines the WebRTC low-latency protocol to dynamically adjust the bit rate, so that the rendering data does not need to pass through the CPU for transfer. The GPU direct transfer architecture bypasses the CPU post-processing bottleneck, reduces the encoding and transmission latency by more than 50%, and at the same time maintains high image quality under low bandwidth through the DLSS super-resolution technology, meeting the stringent requirements for millisecond-level latency in real-time interaction scenarios such as cloud gaming and remote design.

[0073] The present invention can real-time identify abnormal nodes through heartbeat detection and task monitoring, and adopts an incremental task update strategy to only migrate the affected blocks or frame sequences. Based on the subtask status tracking technology, the system only needs to re-render the local data processed by the faulty node, avoiding the waste of resources caused by full-task retries, shortening the fault recovery time from the traditional minute level to the second level, and improving the system reliability by 60%, ensuring the stability of 7×24-hour continuous rendering services.

[0074] In a specific embodiment, the preemptive scheduling strategy includes: high-priority tasks can only interrupt the subsequent subtasks that have not been executed by low-priority tasks, and the status of the executed subtasks will be completely retained. For example, when a low-priority task is rendering block A of frame 3, a high-priority task can interrupt the distribution of block B of frame 4 that follows it, but the data of block A of frame 3 that has been completed will be retained and finally participate in the result synthesis. The preemptive scheduling strategy also includes: the task execution module records the progress of subtasks and ensures the coherence of task recovery after preemption through status tracking. The task execution module realizes precise interruption control through subtask status marking to ensure that the interruption operation does not damage the integrity of the executed subtasks.

[0075] The task execution module splits the rendering task into indivisible subtasks, such as ray tracing calculations for a single block. Each subtask carries a unique identifier and status marks, including: "submitted", "executing", "completed", etc. When a high-priority task triggers preemption, the scheduling control module only modifies the order of the unexecuted subtasks in the task queue, and the results of the executed subtasks are persistently stored through the intermediate result buffer to ensure that no recalculation is required during subsequent recovery.

[0076] The round-robin scheduling strategy includes: ordinary priority tasks are distributed in a distributed manner according to the principle of node load balancing. Based on the matching degree between the priority score of the subtasks after task splitting, the node processing capacity, and the task load, a first-come, first-served mechanism is adopted to allocate them to the appropriate nodes for execution. The subtask priority score is comprehensively determined by task complexity, data volume, and node compatibility (such as the matching degree of instruction set architecture). The fairness of round-robin scheduling is based on the matching degree after task splitting rather than equal resource distribution. Subtasks enter the queue in the order of splitting completion time. After the node completes the current task, it automatically obtains the appropriate subtasks from the head of the queue to ensure that all ordinary tasks are in an equal position in resource allocation.

[0077] The round-robin scheduling strategy also includes: the task execution module dynamically adjusts the subtask distribution order through a priority scoring algorithm to ensure the fairness of resource allocation and avoid long-term overload or idleness of a single node. Among them, the scoring algorithm formula:

[0078] Node adaptation score = (node processing capacity weight × node current load factor) + (subtask complexity weight × data volume factor) + (node compatibility weight × instruction set matching degree). The node processing capacity weight is predefined based on benchmark tests. For example, the GPU weight of RTX 4090 is 1.5, and the GPU weight of RTX3060 is 1.0; the load factor is calculated according to the real-time utilization rate. For example, when the GPU utilization rate is 80%, the factor is 0.8; the subtask complexity weight is generated by the scenario analysis module. For example, the weight of the ray tracing task is 1.2, and the weight of the rasterization task is 1.0.

[0079] The task execution module maintains a global task queue, and each subtask enters the queue with an attached priority score. After the node completes the current task, it filters the appropriate subtasks (score ≥ threshold) from the head of the queue through a greedy algorithm, preferentially selecting subtasks that match the node instruction set architecture and have appropriate complexity, achieving a balance between "first-come, first-served" and "precise matching". Moreover, the system regularly recalculates the node load factor and instruction set matching degree to ensure that the score reflects the dynamic changes of the cluster in real time.

[0080] It can be seen that through subtask interruption control in this embodiment, high-priority tasks can immediately obtain resources, avoiding the delay of waiting for the entire task to complete in traditional scheduling. The response time of key frame rendering tasks is shortened from 150ms in the traditional method to less than 40ms, realizing the real-time guarantee of high-priority tasks. The round-robin scheduling in this embodiment realizes the dynamic matching of tasks and nodes based on adaptation scoring, reducing the standard deviation of cluster resource utilization by 60%, avoiding long-term overload or idleness of a single node, and realizing the fairness of ordinary tasks and resource balance. The subtask status tracking in this embodiment ensures that tasks can be accurately restored after interruption, avoiding more than 30% waste of computing resources caused by traditional full-task retries, and realizing task execution coherence and system robustness.

[0081] As Figure 3 shown, in a specific embodiment, spatial partitioning divides a three-dimensional scene into multiple independent regions dynamically according to the octree or hierarchical bounding box algorithm. By performing spatial partitioning on the scene, the overall rendering task can be disassembled into multiple independent spatial regions for processing. Taking the rendering of a large virtual city scene as an example, there are numerous elements such as buildings, roads, and vegetation in the city. After using spatial partitioning, the city is divided into small block regions, and each region may contain a specific building complex or a specific area of roads and greenery. In this way, different regions can be rendered in parallel, improving the rendering efficiency. For example, node 1 is responsible for rendering A1, node 2 is responsible for rendering B1, node 3 is responsible for rendering B4, node 4 is responsible for rendering A4, and so on. Its execution process includes: separately allocating computing nodes to each spatial block, performing parallel computing using a distributed architecture, and after rendering, performing light and shadow matching and color fusion across block boundaries through a result processing module to ensure the coherence of the final image.

[0082] As Figure 4 shown, time frame division decomposes an animation task into a series of parallel frame sequences according to a preset frame rate. For example, for an animation with a duration of 10 seconds and a frame rate of 60 frames per second, it will be decomposed into 600 independent frames, and these frames can be rendered simultaneously on different computing nodes. Node 1 renders F1, node 2 renders F2, node 3 renders F3, node 4 renders F4, and so on, accelerating the rendering speed of the entire animation, making the rendering task of each frame relatively independent and facilitating parallel processing. Its execution process includes: allocating independent nodes to each time frame for calculation, and after rendering, synthesizing and outputting the video in sequence; if frame delay or failure is detected, idle nodes will be rescheduled for supplementary calculation.

[0083] As Figure 5 shown, and this embodiment can complete the collaborative scheduling of partitioning and frame division by combining spatial and time strategies, and spatial partitioning takes precedence over time frame division for collaborative decomposition. Specifically, first complete the spatial partitioning of the three-dimensional scene, and then decompose the rendering tasks of each spatial block at different time frames to form a task set based on spatial blocks and sequenced by time frames. At the same time, a dependency relationship is constructed for each sub-task through a task dependency graph algorithm to ensure the correctness of the execution order. For example, for a complex animation scene, first divide the scene into multiple regions spatially, and then determine the rendering tasks for different time frames of each region. Divide spatial block A1 into several time frames F1, F2, …… Fn. Finally, construct a task dependency graph according to the dependency relationships between tasks, such as light calculation and the coherence of object movement.

[0084] It can be seen that through the design of spatial partitioning and temporal framing, the present invention realizes a more efficient task decomposition strategy, which is applicable to complex scene rendering and high-frame-rate animation tasks, providing significant performance improvement and reliability guarantee for cloud rendering clusters.

[0085] The task module executes a scheduling strategy. According to the task dependency graph, combined with the node processing capabilities, task loads, and priority scores, it coordinately schedules the subtasks of spatial partitioning and temporal framing, preferentially scheduling the subtasks with an in-degree of 0, that is, the subtasks without pre-dependency tasks, and preferentially allocating high-priority tasks to low-load nodes. For example, after the rendering task of a spatial block is completed, the subsequent time-frame tasks that depend on this spatial block are determined according to the task dependency graph and scheduled to appropriate nodes for execution. And during the task execution process, the resource status of the nodes and the task execution progress are monitored in real time. If there are situations such as node failures or task execution delays, the execution order of the tasks is re-evaluated according to the task dependency graph, and the scheduling strategy is dynamically adjusted to ensure the smooth progress of the overall task.

[0086] After the rendering task is completed, each computing node generates intermediate rendering result data, such as image blocks or time frames. In order to generate a complete high-fidelity rendering output, the present invention realizes the high-fidelity synthesis of distributed rendering results through spatial result merging, temporal result merging, and collaborative merging mechanisms.

[0087] Result merging includes:

[0088] As Figure 6 shown, for spatial block merging, for the rendering results of each node generated through spatial partitioning, the data of different spatial blocks are spliced into a complete scene image. During the execution process, global illumination algorithms, such as photon mapping or path tracing, are used to correct the lighting effects across block boundaries, and image stitching algorithms, such as gradient fusion or Poisson fusion, are used to smooth the stitching areas between blocks to avoid visual gaps or color differences. The specific execution process is as follows: First, collect the rendering block data A1, B1, C1... generated by each node; then splice the block data according to the preset spatial partitioning index; finally, perform light and color correction on the stitching area to generate the final complete image.

[0089] As Figure 7 shown, for temporal frame merging, for the rendering results of each node generated through temporal framing, the data of the time-series frames are synthesized into a complete animation or video output in sequence. Check the sequence integrity of the rendered frames to ensure that the frame order is correct and there are no omissions; perform compression encoding on the rendered frames to reduce the volume of the output file while ensuring high image quality. The specific execution process is as follows: First, collect the temporal frame rendering results generated by each node; then merge the frame data into a sequence in chronological order and perform compensation for missing frames (if any); finally, perform post-processing on the synthesized frame sequence, such as special effect overlay, color correction, etc., to output the final animation or video file.

[0090] Moreover, this embodiment can complete the coordinated merging of block results and frame results. Specifically, first, spatial block merging is completed in each time frame to generate a complete frame image; then all frames are merged in time sequence to output the final rendered animation or video.

[0091] It can be seen that this embodiment splits complex rendering tasks into multiple subtasks that can be executed in parallel through the coordinated decomposition and scheduling of spatial blocks and temporal frames, making full use of the computing resources of each node in the cluster and significantly shortening the rendering time. In the process of spatial result merging and temporal result merging, professional algorithms are used for processing such as lighting correction, image stitching and frame sequence verification to ensure high fidelity and visual consistency of the final rendering output. Dynamic task scheduling and result merging strategies can adapt to abnormal situations such as node failures and task changes, and ensure the stable operation of the system by re-adjusting the scheduling and merging schemes.

[0092] In a specific embodiment, the dynamic bit rate adjustment includes:

[0093] Based on changes in node load and task requirements, trigger GPU resource expansion, release or dynamic resolution adjustment. The system collects node CPU or GPU utilization, memory usage and network bandwidth data in real time through the resource monitoring module. For example, when the GPU utilization exceeds the preset threshold (85%), it triggers resource expansion; when it is lower than 30%, it triggers resource release. When a surge in high-priority tasks is detected, GPU nodes that support CUDA or ROCm architecture are automatically started to ensure that the tasks match the hardware instruction set. When the network bandwidth fluctuates, the rendering resolution is reduced from 4K to 1080p through a dynamic bitrate adjustment strategy, and the anti-aliasing level and shadow quality are adjusted in conjunction to ensure video smoothness.

[0094] Dynamic bitrate adjustment also includes: using an idle resource recovery mechanism for low-priority tasks to adapt to node resources of different instruction set architectures. The task execution module maintains multiple versions of rendering program images, provides a unified interface for different instruction set nodes through the gRPC protocol, and automatically assigns low-priority tasks when a node is detected to be idle, making full use of idle resources. The "priority scoring + load balancing" algorithm is used to distribute low-priority tasks to low-load nodes; for example, when the cluster load is less than 40% at night, batch rendering tasks are automatically started to use idle GPU resources to complete the frame sequence calculation of film and television animations.

[0095] It can be seen that through the collaborative action of resource monitoring, task decomposition, and heterogeneous adaptation technologies, the dynamic bitrate adjustment mechanism of this embodiment realizes the elimination of resource waste caused by traditional static allocation through elastic expansion and recycling; utilizes idle resources to execute low-priority tasks, and combines heterogeneous node adaptation to reduce redundant hardware procurement; supports unified scheduling of multi-architecture nodes to ensure the real-time performance of high-priority tasks during network fluctuations.

[0096] As Figure 8 shown, in the present invention, the high-fidelity cloud rendering cluster scheduling method further includes a step of pre-computing key frames during idle time: during the idle period of nodes, pre-compute key frames for scene switching, light changes, or motion paths; when a new task arrives, perform frame interpolation compensation based on the pre-computed key frames, and combine optical flow estimation and motion compensation algorithms to reduce the real-time calculation amount.

[0097] In a high-fidelity cloud rendering cluster, pre-computing key frames during idle time is the core of optimizing real-time rendering efficiency. It uses the idle period of nodes to generate key frames of complex scenes in advance, and reduces the real-time calculation amount through frame interpolation compensation technology when a new task arrives, achieving a double improvement in resource utilization rate and rendering quality.

[0098] When the loads of Node 3 and Node 4 are lower than the threshold, such as when the CPU utilization rate ≤ 30%, the system automatically triggers pre-computation tasks, generates high-precision key frames for complex areas such as scene switching, light changes, or motion paths, and stores them in the pre-computed key frame library. When the scheduling control module initiates a rendering request and the task execution module executes the rendering task through Node 1 and Node 2, it can accelerate the rendering by using the pre-computed key frames KF1, KF2, and KF3 through frame interpolation compensation, reducing the calculation amount of Node 1 and Node 2.

[0099] Among them, the scene switching frame is the starting frame and the ending frame when the virtual camera perspective suddenly changes, and the complete light distribution and object spatial position need to be pre-computed. The light change frame is the highlight and shadow transition frame during dynamic weather changes, and the material reflection data under different light conditions needs to be pre-stored. The motion path frame is the key pose frame in character animation, and the bone motion trajectory and physical collision result need to be pre-computed.

[0100] The execution process of pre-computing key frames during idle time includes:

[0101] Step A. Detection of idle state;

[0102] The resource monitoring module continuously monitors the status of cluster nodes. If some nodes are in a low-load or idle state, they are marked as available;

[0103] Step B. Selection of key frames;

[0104] Based on the task historical data and user requirements, determine the key frames that need to be pre-computed, including scene change points and important nodes on the motion path; scene change points include shot transitions and lighting changes, and important nodes on the motion path include key actions of characters, etc.;

[0105] Step C. Pre-compute key frames;

[0106] Select appropriate computing nodes to compute key frames in the background and store the results;

[0107] Step D. Use key frames when a new task arrives;

[0108] When the formal rendering task arrives, directly use the stored key frames for frame interpolation compensation to reduce the overall computing requirements.

[0109] In a preferred embodiment, the combination steps of idle-time key frame pre-computation and scheduling strategy include:

[0110] Step E. Dynamic resource adjustment;

[0111] The resource monitoring module continuously detects the proportion of idle resources of cluster nodes. When the GPU utilization rate is lower than the threshold, it triggers the idle-time key frame pre-computation task and restricts its resource occupancy not to exceed the preset proportion of idle resources;

[0112] Step F. Priority collaborative scheduling;

[0113] When a high-priority task arrives, the scheduling control module preferentially allocates the pre-computed key frames to this task, generates intermediate frames through the optical flow interpolation algorithm, and reduces the number of real-time rendering frames;

[0114] The key frame pre-computation of ordinary-priority tasks is only started during off-peak hours and can be dynamically preempted by high-priority tasks;

[0115] The idle-time key frame task is marked as the lowest priority, and the scheduler can interrupt its resource allocation at any time to ensure the resource requirements of high-priority tasks;

[0116] Step G. Cache management;

[0117] The pre-computed key frames are stored in the GPU video memory cache pool, and an LRU (Least Recently Used) elimination mechanism is established to ensure that high-priority tasks can quickly access the latest key frames.

[0118] In a specific embodiment, the cross-block boundary light and shadow correction includes:

[0119] Use photon mapping or path tracing algorithm to correct the light and shadow consistency at the block boundary;

[0120] Perform gradient fusion or Poisson fusion on the splicing area to eliminate visual gaps.

[0121] Among them, the photon mapping algorithm simulates the propagation and scattering process of photons in the scene. When rendering in blocks, the photon tracing within each block may be incomplete due to boundary restrictions. Through the photon mapping algorithm, the system will recalculate the propagation path of photons at the block boundary, taking into account the influence of adjacent blocks, and thus correct the lighting information at the boundary. For example, an object in a block may be affected by the indirect lighting of the light source in the adjacent block. The photon mapping algorithm can accurately capture such influences, making the light and shadow effects at the boundary more accurate.

[0122] The path tracing algorithm simulates the interaction between light and objects in the scene by tracing the light backwards from the camera. At the block boundary, the path tracing algorithm will trace the light across blocks to ensure that the propagation of light is not restricted by the blocks, thereby ensuring the consistency of light and shadow at the boundary. For example, when light passes through the block boundary, the algorithm will correctly calculate the reflection, refraction, and absorption of light in different blocks, so that the lighting effect at the boundary matches the entire scene.

[0123] Gradient fusion is based on the gradient information of the image, that is, the rate of change of color and brightness in the image. In the splicing area, by calculating the gradient information of adjacent blocks, a suitable transition function is found to make the gradient change smooth at the splicing area, which can avoid obvious color or brightness mutations in the splicing area, thereby eliminating visual gaps. For example, for the splicing of two adjacent blocks, gradient fusion will adjust the color and brightness of the pixels at the splicing point according to the gradient values of the surrounding pixels to make the transition natural.

[0124] Poisson fusion is an image fusion method based on partial differential equations. By solving the Poisson equation, the gradient information of one image area is seamlessly fused into another image area. When splicing blocks, Poisson fusion will calculate an optimal fusion result based on the color and brightness information of adjacent blocks, so that the color and brightness of the spliced area are naturally connected with the surrounding pictures, eliminating visual gaps.

[0125] The spatial blocking of the present invention can improve rendering efficiency, but it will bring about the problem of block boundaries. The light and shadow correction and fusion algorithm of this embodiment solves the problems of light and shadow inconsistency and splicing gaps caused by block rendering, further improves the high-fidelity cloud rendering cluster scheduling method, and improves the performance and output quality of the entire rendering system.

[0126] In a specific embodiment, the GPU directly outputs the adaptive bit rate video stream including:

[0127] The H.265 / AV1 low-latency video compression encoding of the rendered frames is directly completed through the GPU hardware encoder;

[0128] Optimize the picture quality in combination with DLSS super - resolution, and dynamically adjust the output bitrate and resolution through a low - latency streaming protocol, which includes but is not limited to real - time transmission protocols such as WebRTC and QUIC.

[0129] The GPU hardware encoder is a hardware unit specifically designed for video encoding and has efficient parallel computing capabilities. During the rendering process, the GPU directly passes the data of the rendered frame to the hardware encoder for encoding processing, avoiding the process of transferring data from the GPU to the CPU for encoding, reducing the time of data transfer and processing. At the same time, the hardware encoder is optimized for video encoding and can quickly complete complex encoding algorithms, thus achieving low - latency video compression encoding.

[0130] Low - latency streaming protocols, such as the WebRTC protocol, etc., have the ability of real - time feedback and adaptive adjustment. The system monitors parameters such as the network bandwidth and packet loss rate in real - time through this protocol, and dynamically adjusts the bitrate and resolution of the video stream according to these parameters. When the network bandwidth decreases, the system will reduce the bitrate and resolution of the video to reduce the data transmission volume to ensure the smoothness of the video; when the network bandwidth recovers, it will increase the bitrate and resolution to improve the quality of the video.

[0131] It can be seen that compared with the traditional rendering process, in this embodiment, the GPU hardware encoder directly completes the video compression encoding of the rendered frame, reducing the intermediate links, greatly reducing the time from rendering to encoding of the video, thus achieving low - latency video output; optimizing the picture quality in combination with the DLSS super - resolution technology can make the output video stream have higher visual quality and richer details; by dynamically adjusting the output bitrate and resolution through the low - latency streaming protocol, the system can automatically adjust the quality of the video stream according to the real - time network bandwidth and stability.

[0132] In a specific embodiment, the incremental task update strategy includes:

[0133] Only re - render the affected chunks or frame sequences;

[0134] Adopt healthy heartbeat detection and dynamic load balancing to quickly migrate the tasks of faulty nodes.

[0135] Among them, regarding the healthy heartbeat detection function: each node will regularly send a heartbeat signal to the cluster management system to indicate its own running state; the cluster management system can understand the health status of each node in real - time by monitoring the heartbeat signal. If a certain node does not send a heartbeat signal within the specified time, the system will determine that the node has failed.

[0136] Regarding the task migration process: Once a faulty node is detected, the system will immediately start the task migration mechanism, which specifically includes the following steps:

[0137] Step 100: The system takes a snapshot of the tasks on the faulty node and saves the current state of the tasks.

[0138] Step 200: According to the dynamic load balancing algorithm, a suitable normal node is selected from the cluster to receive the tasks; the dynamic load balancing algorithm takes into account factors such as the current load and computing power of the nodes to ensure that the tasks can be assigned to the most suitable nodes for continued execution.

[0139] Step 300: Transmit the snapshot data of the tasks to the new node and resume the execution of the tasks on the new node to achieve fast migration.

[0140] It can be seen that in this embodiment, only the affected chunks or frame sequences are re-rendered, greatly reducing unnecessary computational effort, shortening the rendering time, and improving the overall rendering efficiency. When dealing with large and complex scenes or long animation sequences, the efficiency improvement is particularly significant; at the same time, quickly migrating the tasks of the faulty node can reasonably allocate resources, avoid resource idleness, improve resource utilization, and reduce rendering costs. In this embodiment, by adopting healthy heartbeat detection and dynamic load balancing to quickly migrate the tasks of the faulty node, the system can timely detect the faulty node and quickly migrate the tasks on it to other normal nodes for continued execution.

[0141] In the step of only re-rendering the affected chunks or frame sequences, the system will conduct a detailed dependency analysis on the rendering tasks to determine the dependency relationships between the various chunks or frame sequences in the tasks. When the task is updated, such as when the material or position of an object in the scene changes, the affected chunks or frame sequences can be accurately identified through analysis. For example, in a 3D scene, if the lighting condition of an object changes, the system will analyze which chunks' rendering results will be affected by this lighting change. And based on the results of the dependency analysis, the system only re-renders the affected chunks or frame sequences, avoiding recalculating the entire scene and improving the rendering efficiency. At the same time, combining spatial chunking and temporal frame division techniques can more accurately locate and process the affected parts.

[0142] As Figure 2 shown, the present invention also provides a high-fidelity cloud rendering cluster scheduling system for implementing the above high-fidelity cloud rendering cluster scheduling method, including:

[0143] A task management module, which is used to receive tasks through a user-friendly interface, automatically classify and store them. When automatically classifying, it automatically classifies based on task complexity, timeliness, and resource requirements to generate a multi-level priority queue; and supports users to define multi-level priority rules.

[0144] A resource monitoring module, which is used to collect CPU, GPU instruction set architecture adaptation data and dynamic loads of heterogeneous nodes in real time; it is also used to identify faulty nodes through a heartbeat detection mechanism and trigger task migration;

[0145] A scheduling control module, which is used to perform multi-level priority scheduling, dynamic bitrate adjustment, decomposition strategy, and pre-computation of key frames during idle time. Multi-level priority tasks include high-priority tasks and normal-priority tasks. High-priority tasks adopt a preemptive scheduling strategy, interrupting the subsequent subtasks of low-priority tasks and preferentially allocating resources; normal-priority tasks are fairly allocated resources through a round-robin scheduling strategy based on node load and task adaptation scores; dynamic bitrate adjustment dynamically adjusts the rendering resolution and encoding parameters according to node load and network bandwidth to ensure video smoothness; pre-computation of key frames during idle time pre-computes key frames such as scene switching and lighting changes during node idle periods and stores them in the GPU cache. When a new task arrives, intermediate frames are generated based on the pre-computed key frames and DLSS interpolation technology.

[0146] The decomposition strategy refers to an optimization method that decomposes complex rendering tasks into sub-tasks that can be executed in parallel according to the logical relationship or data dependency between tasks and dynamically adjusts the execution order. It constructs a task dependency graph, decomposes the rendering task into multiple sub-tasks, and determines the execution order according to the dependency relationship. For example, a dependency relationship between task A and task B: task A must be executed after task B is completed. The logical relationship includes spatial position relationship and time sequence.

[0147] A task execution module, which is used to support the collaborative decomposition of spatial partitioning and time frame partitioning and the tracking of subtask states. Spatial partitioning uses the octree algorithm to divide the three-dimensional scene into independent regions and supports cross-block lighting correction; time frame partitioning decomposes animation tasks according to a preset frame rate and supports frame sequence verification and missing frame compensation; subtask state tracking records the execution progress of each subtask to ensure the coherence of task recovery after preemptive scheduling.

[0148] A result processing module, which is used to decompose or merge block data and frame data, correct boundary lighting, perform cross-block light and shadow correction, and directly output an adaptive bitrate video stream on the GPU.

[0149] Among them, the result processing module includes:

[0150] A GPU video stream processing unit, integrated with a hardware encoder, which is used to directly perform video compression encoding on the rendered frames;

[0151] A protocol adaptation unit, which supports a low-latency real-time transport protocol stack and pushes the encoded stream to the client;

[0152] A dynamic adjustment unit, which generates bitrate, resolution, and frame rate control instructions based on client network data and feeds them back to the GPU driver layer for execution through a software interface.

[0153] In summary, through real-time resource monitoring and dynamic allocation mechanisms, the present invention adopts adapted scheduling strategies for different types of tasks, making the most of resources such as CPU, GPU, memory, and bandwidth in the cluster to avoid resource idleness or overload. For the complexity and priority requirements of rendering tasks, the present invention introduces a multi-level task priority scheduling strategy and a distributed task decomposition mechanism to reduce the waiting time for task execution and improve task completion efficiency. In the scheduling of high-priority tasks, a preemptive algorithm is adopted, combined with the key frame scheme during idle time, to ensure the real-time performance of critical tasks, and at the same time, the overall system latency is reduced by optimizing the data transmission path and inter-node communication. The scheduling algorithm adapted to heterogeneous environments can fully exploit the potential of heterogeneous resources to maximize performance. A fault tolerance mechanism is introduced to ensure that tasks can be completed as planned by dynamically adjusting task allocation and resource scheduling. In the scheduling of low-priority tasks, a resource-saving allocation strategy is adopted to reduce the overall operating cost of the cluster through the full utilization of idle resources and the resource recycling mechanism.

[0154] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0155] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A high-fidelity cloud rendering cluster scheduling method, characterized in that It includes the following steps: Receive the rendering tasks submitted by users, classify the tasks based on task complexity, timeliness, and resource requirements, generate a multi-level priority task queue, and automatically optimize the scheduling strategy through the user interface; Monitor the resource status of heterogeneous computing nodes in the cluster in real time, including the dynamic loads of CPUs and GPUs with different instruction set architectures and the network bandwidth utilization rate; According to the priority and resource status, adopt a preemptive scheduling strategy to dynamically allocate high-priority tasks to low-load nodes, and use a round-robin scheduling strategy to allocate normal-priority tasks, and combine dynamic bitrate adjustment to cope with resource fluctuations; Adopt two strategies of spatial partitioning and temporal framing to decompose the task into several subtasks, control the order of decomposition, distribution, and synthesis based on the topological sorting algorithm, and perform distributed distribution according to the node processing capacity, task load, and priority score; During the execution of subtasks, identify abnormal nodes through heartbeat detection and task monitoring, and adopt an incremental task update strategy to dynamically migrate affected subtasks; Perform cross-block boundary light and shadow correction, image stitching, and frame sequence encoding optimization on the subtask rendering results generated by each node to generate a final high-fidelity rendering output, and directly output an adaptive bitrate video stream to the client through the GPU.

2. The high-fidelity cloud rendering cluster scheduling method according to claim 1, wherein The preemptive scheduling strategy includes: High-priority tasks can interrupt the unexecuted subsequent subtasks of low-priority tasks and retain the status of the executed subtasks; The task execution module records the progress of subtasks and ensures the coherence of task recovery after preemption through status tracking; The round-robin scheduling strategy includes: Normal-priority tasks are distributed distributively according to the node load balancing principle, and are allocated to the appropriate nodes for execution using the first-come-first-served mechanism based on the matching degree between the priority scores of the subtasks after task splitting and the node processing capacity and task load; The task execution module dynamically adjusts the subtask distribution order through the priority scoring algorithm.

3. The high-fidelity cloud rendering cluster scheduling method according to claim 1, characterized in that The dynamic bitrate adjustment includes: Based on changes in node load and task requirements, trigger GPU resource expansion, release, or resolution dynamic adjustment; Adopt an idle resource recycling mechanism for low-priority tasks to adapt to node resources with different instruction set architectures.

4. The high-fidelity cloud rendering cluster scheduling method according to claim 1, wherein The spatial partitioning uses an octree or hierarchical bounding box algorithm to dynamically divide the three-dimensional scene into independent regions, and the temporal framing decomposes the animation task into a parallel frame sequence at a preset frame rate; And spatial partitioning takes precedence over temporal framing for collaborative decomposition, and ensures the execution order through the task dependency graph algorithm; Chunk merging takes precedence over frame merging for collaborative merging.

5. The high-fidelity cloud rendering cluster scheduling method according to claim 1, wherein It also includes a step of pre-computing key frames during idle time: During the idle period of the node, pre-compute the key frames for scene switching, light change, or motion path; When a new task arrives, perform frame interpolation compensation based on the pre-computed key frames, and the calculation process combines the optical flow estimation and motion compensation algorithms.

6. The high-fidelity cloud rendering cluster scheduling method according to claim 1, wherein The cross-block boundary light and shadow correction includes: Adopt a photon mapping or path tracing algorithm to correct the light and shadow consistency at the block boundary; Perform gradient fusion or Poisson fusion on the stitching area to eliminate visual gaps.

7. The high-fidelity cloud rendering cluster scheduling method according to claim 1, characterized in that The GPU direct output of the adaptive bitrate video stream includes: Directly complete low-latency video compression encoding of rendered frames through the GPU hardware encoder; Combine DLSS super-resolution to optimize the picture quality, and dynamically adjust the output bitrate and resolution through a low-latency streaming protocol.

8. The high-fidelity cloud rendering cluster scheduling method according to claim 1, characterized in that The incremental task update strategy includes: Only re-render the affected chunks or frame sequences; Adopt healthy heartbeat detection and dynamic load balancing to quickly migrate the tasks of faulty nodes.

9. A high-fidelity cloud rendering cluster scheduling system, characterized in that, For implementing the high-fidelity cloud rendering cluster scheduling method as described in any one of claims 1 to 8, the system includes: A task management module, used to receive tasks through a user-friendly interface, automatically classify and store them, and define multi-level priority rules; A resource monitoring module, used to collect CPU, GPU instruction set architecture adaptation data and dynamic load of heterogeneous nodes in real time; A scheduling control module, used to execute multi-level priority scheduling, dynamic bitrate adjustment, decomposition strategy, and key frame pre-computation during idle time; A task execution module, used to support the collaborative decomposition of spatial chunking and temporal frame division and the tracking of subtask states; A result processing module, used to decompose or merge chunk data and frame data, correct boundary lighting, perform cross-chunk light and shadow correction, and directly output an adaptive bitrate video stream through the GPU.

10. The high-fidelity cloud rendering cluster scheduling system according to claim 9, characterized in that, The result processing module includes: A GPU video stream processing unit, integrated with a hardware encoder, used to directly perform efficient video encoding on the rendered frames; A protocol adaptation unit, supporting a low-latency real-time transport protocol stack, and pushing the encoded stream to the client; A dynamic adjustment unit, generating bitrate, resolution, and frame rate control instructions based on the client network data, and feeding them back to the GPU driver layer through a software interface for execution.

Citation Information

Patent Citations

  • Intelligent scheduling management method for multi-scale parallel rendering jobs

    CN103268253A

  • Massive working node-oriented preemptive task scheduling method and system

    CN108762903A

  • Large model parallel task scheduling method based on space-time two-dimensional segmentation and intelligent sharing

    CN119645609A

  • Low-latency, peer-to-peer streaming video

    US10951890B1

  • Cloud native distributed real-time rendering framework, rendering method and device

    WO2024239876A1

Cited By

  • AI-driven cloud rendering service fault rapid positioning method

    CN120599045A

  • Distributed rendering task scheduling method based on dynamic load balancing

    CN120612416A

  • Three-dimensional resource dynamic allocation method and system based on grain storage

    CN120803740A

  • Three-dimensional resource dynamic allocation method and system based on grain storage

    CN120803740B

  • Rendering node video stream low-delay synthesis output system

    CN120935379A