Decoding task scheduling method based on frame feature perception
By dynamically identifying the VPU decoding pressure in real-time communication and using the CPU to perform parallel decoding processing, the problems of tail delay and frame rate drop in high-quality real-time communication are solved, and efficient decoding and stable frame rate are achieved.
Patent Information
- Application Number
- CN202510163046.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to meet the requirements of high frame rate, high resolution and ultra-low latency simultaneously in high-quality real-time communication, resulting in increased tail delay and decreased frame rate.
By dynamically identifying the VPU decoding pressure situation, using the lightweight neural network model to perform decoding pressure estimation, flexibly unloading some decoding tasks to the CPU for parallel processing, reducing decoding time and improving tail delay.
It effectively reduces the tail delay in real-time communication scenarios, while maintaining the stability of high frame rates, especially in high frame rates and high resolution scenarios, which significantly improves decoding efficiency.
Smart Images

Figure CN119966966A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of real-time communication, and in particular relates to a decoding task scheduling method based on frame feature perception. Background Art
[0002] Real-Time Communication (RTC) plays an increasingly important role in today's digital world, especially in video conferencing, cloud gaming, virtual reality (VR), online education, and telemedicine. However, with the popularity of real-time communication applications and the growth of demand, providing high-quality real-time communication experience has become a challenging task.
[0003] High-quality RTC applications have the following characteristics:
[0004] (1) High frame rate: Traditional RTC generally has only 30fps, which puts less pressure on the decoder. High-quality RTC requires 60-120fps, which brings huge challenges to the decoder;
[0005] (2) High resolution: Existing RTC applications default to 720p or even lower, while high-quality RTC should require 1080p or above, or even 2k. The higher the resolution, the longer the decoding time will be;
[0006] (3) Ultra-low latency: For example, video conferencing and cloud gaming both hope to achieve an interaction delay of less than 100 ms.
[0007] The end-to-end delay of RTC applications is mainly composed of video acquisition and encoding at the sender, network transmission, queuing, decoding, and display at the receiver. The increase in transmission delay caused by network fluctuations or the increase in decoding delay caused by decoder performance fluctuations will both lead to an increase in queuing delays, which are important components of the tail delay of RTC applications.
[0008] There are currently two methods to reduce tail latency. The first method is the frame skipping mechanism, which actively manages the frames in the decoding queue at the receiving end, that is, selectively discards some frames without decoding. This method effectively reduces queuing delay. The second method is adaptive frame rate control, which achieves low latency by adaptively coordinating the frame rate at the sending end with the fluctuating network conditions and the decoding queue capacity at the receiving end. However, these solutions will cause a drop in frame rate and cannot meet the high frame rate requirements of high-quality RTC.
[0009] GOP (Group of Pictures) is a sequence of frames organized as a group in a video stream, which determines the prediction relationship between video frames. It is an important part of video coding, and its length and structure directly affect the compression efficiency, decoding performance and application experience of the video. The GOP format commonly used in RTC applications is IPPPP..., that is, the first frame is an I frame, and the subsequent frames are all P frames. Among them, the I frame does not rely on other frames to complete decoding, and the P frame relies on the previous frame to be decoded successfully. Due to such a dependency between frames, for the first method mentioned above, arbitrarily discarding some frames will cause subsequent frames to fail to decode. In order to solve this problem, Scalable Video Coding (SVC) technology is proposed. According to the principle of temporal scalability in SVC, the frames at the top level are not dependent on other frames. These are the frames that can be discarded without causing subsequent frames to fail to decode.
[0010] Video decoding tasks are usually handled by dedicated video processing units (VPUs), such as Qualcomm chips on mobile terminals, Intel chips on PCs (Personal Computers), and NVIDIA chips. According to calculations, when RTC applications are decoding, the usage rate of the VPU is usually high, while the usage rate of the CPU is low. If idle CPU resources can be used to calculate some decoding tasks, the decoding time of a single frame will be effectively reduced, thereby alleviating the problem of decoding queue blocking. This will effectively reduce the tail delay in real-time communication scenarios while maintaining a stable frame rate. Summary of the invention
[0011] In order to overcome the problems existing in the above technologies, the purpose of the present invention is to provide a decoding task scheduling method based on frame feature perception. The present invention dynamically identifies the VPU decoding pressure situation and flexibly unloads part of the decoding tasks to the CPU for parallel processing, thereby reducing the decoding time and improving the tail delay problem.
[0012] The present invention uses a decoding pressure estimation mechanism to determine whether the CPU needs to participate in the decoding task. When the CPU needs to participate in the decoding task, the frame feature perception scheduling mechanism will analyze the frame features in the decoding queue and schedule some decoding calculation tasks reasonably. This will effectively reduce the decoding time of a single frame, thereby alleviating the problem of decoding queue blocking.
[0013] The specific technical solution for achieving the purpose of the present invention is:
[0014] A decoding task scheduling method based on frame feature perception includes the following steps:
[0015] (1) Analyze the characteristic values of the received frame through the decoding pressure estimation mechanism to determine whether the VPU decoder is in a busy state;
[0016] (2) When the prediction result of the decoding pressure estimation mechanism is "H-Latency", mark the current frame and call the frame feature-aware scheduling mechanism;
[0017] (3) The frame feature-aware scheduling mechanism identifies and selects candidate frames for frame-level scheduling based on the dependencies between frames and the feature values of the frames;
[0018] (4) For frames that fail to pass frame-level scheduling, identify and select candidate macroblocks in the frame and perform macroblock-level scheduling;
[0019] (5) Schedule qualified frames or macroblocks to the CPU for parallel decoding processing to reduce tail latency and maintain a stable frame rate.
[0020] Furthermore, the decoding pressure estimation mechanism uses a lightweight neural network model to quickly infer the decoding pressure state as "H-Latency" or "L-Latency" by inputting five characteristic values of frame type, frame size, number of macroblocks, macroblock type and number of frames in the decoding queue.
[0021] Furthermore, the frame level scheduling mechanism specifically includes:
[0022] (1) Candidate frame identification: Using an identification method based on inter-frame dependencies, frames in the decoding queue that are not dependent on other frames are selected as candidate frames;
[0023] (2) Candidate frame selection: By calculating the sum of the decoding time of the candidate frame on the CPU and its decoding time on the VPU, it is determined whether the candidate frame meets the preset conditions.
[0024] Furthermore, the macroblock level scheduling specifically includes:
[0025] (1) Candidate macroblock identification: A method based on intra-frame and inter-frame macroblock dependencies is used to identify macroblocks that can be independently decoded in parallel.
[0026] (2) Candidate macroblock selection: Based on the number of CPU and VPU cores and the principle of spatial locality, select the macroblock that meets the maximum parallelism for decoding.
[0027] The present invention integrates the collaborative work of the decoding pressure estimation mechanism and the frame feature perception scheduling mechanism. Compared with the existing selective frame loss method, the present invention can effectively reduce the tail delay in real-time communication scenarios while maintaining the stability of the frame rate. In particular, for high frame rate and high resolution scenarios, it can effectively improve the decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flow chart of candidate frame identification and selection according to the present invention;
[0029] Figure 2 A flow chart of candidate macroblock identification and selection according to the present invention;
[0030] Figure 3 A schematic diagram for implementing the present invention. DETAILED DESCRIPTION
[0031] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0032] The present invention discloses a decoding task scheduling method based on frame feature perception, and its basic process is as follows: first, the decoding pressure estimation mechanism collects feature value information of the incoming frame and then performs decoding pressure evaluation. If the evaluation result is "H-Latency", the frame feature perception scheduling mechanism is notified. Then, the frames in the decoding queue are analyzed. If the frame meets the conditions of frame-level scheduling, the frame is scheduled to the CPU for decoding. If the frame does not meet the conditions, some macroblocks in the frame are scheduled to the CPU for decoding.
[0033] The decoding pressure estimation mechanism adopts a fast inference model of a lightweight neural network. The model is a fully connected neural network with three layers, including an input layer, a hidden layer and an output layer. It performs inference by inputting five feature values, namely, frame type, frame size, number of macroblocks, macroblock type and number of frames in the decoding queue. The value of the size in the feature is converted into a decimal number and input into the neuron. For example, a frame of 127KB size is formatted as four integers {0,1,2,7}. The same formatting method is also applied to the size of the macroblock. This method can reduce the time cost of the inference stage. It is unnecessary to infer the accurate decoding delay. The present invention simplifies the model by converting the delay inference problem into a simple binary ("busy" or "normal") inference. Specifically, the tracked frames are marked with "H-Latency" and "L-Latency". "H-Latency" indicates that the decoder is in a busy state, while "L-Latency" indicates that the decoder is in a normal state.
[0034] The frame feature-aware scheduling mechanism provides two granularity scheduling schemes, frame-level scheduling and macroblock-level scheduling. Frame-level scheduling is mainly divided into two aspects.
[0035] In a first aspect, a candidate frame identification mechanism is provided, comprising:
[0036] A simple candidate frame identification method based on inter-frame dependency is adopted. The basic idea of this method is to identify frames in the decoding queue that are not dependent on other frames. According to the principle of temporal scalability in scalable video coding technology, the top-level frames are not dependent on other frames. Figure 1 Frames 3 and 7 in the second layer are candidate frames.
[0037] In a second aspect, a candidate frame selection mechanism is provided, including:
[0038] Candidate frames should not be scheduled if they take too long to decode on the CPU. The decoding time on the CPU is ,if Exceeded the frame and subsequent frames The sum of decoding time on VPU , which will increase the delay. Therefore, the candidate frame is only Only then can it be scheduled:
[0039] .
[0040] Macroblock-level scheduling is mainly divided into two aspects.
[0041] In a first aspect, a candidate macroblock identification mechanism is provided, comprising:
[0042] A candidate macroblock identification method based on macroblock dependencies is adopted, including intra-frame macroblocks and inter-frame macroblocks. In an intra-frame macroblock, two independent macroblocks have the same relative position. Starting from the position of the left macroblock, move two positions to the right, and then move one position up to reach the position of the right macroblock. For inter-frame macroblocks, their dependencies span two frames, and the macroblocks in the current frame calculate the position of the dependent macroblock in the previous frame based on their motion vectors as offsets. According to the above two methods, macroblocks that can be processed in parallel between the two frames can be determined, and these are candidate macroblocks.
[0043] In a second aspect, a candidate macroblock selection mechanism is provided, including:
[0044] Using a parallelism-based approach, Figure 2 It shows that during the decoding process, the number of macroblocks that can be processed in parallel at the same time is constantly changing, and the maximum degree of parallelism increases from top to bottom, then reaches a peak state, and finally gradually decreases. The calculation formula for the maximum degree of parallelism is as follows:
[0045]
[0046] in Indicates the number of macroblocks in the width direction of the frame, Indicates the number of macroblocks in the height direction of the frame. After determining the maximum parallelism, the start and end times of macroblock scheduling are selected based on the number of CPU and VPU cores and the principle of spatial locality. Figure 2 For example, it shows that when the number of CPU cores is 2 and the number of VPU cores is 3, the macroblock with dark background is selected.
[0047] like Figure 3 As shown in FIG. 1 , it is an architecture diagram for implementing the present invention. Two new components are added: decoding pressure estimation and frame feature perception scheduling. The implementation is mainly divided into the following steps:
[0048] 1. When the frame receiving end receives a frame of data, the decoding pressure estimation mechanism analyzes the feature value of the frame and inputs the model to predict the decoding pressure.
[0049] 2. If the current frame is "L-Latency", the frame feature perception scheduling will not be awakened, which is consistent with the existing decoding process. Once the current frame is determined to be "H-Latency", it will be marked and the frame feature perception scheduling will be notified.
[0050] 3. Then, start the decoding task scheduling process, find the frames that meet the frame-level scheduling, and then find the macroblocks that meet the macroblock-level scheduling.
[0051] 4. Finally, these decoding tasks that can be processed in parallel are offloaded to the CPU for processing.
[0052] In the above decoding task scheduling process, the scheduling mechanism will analyze the frames in the decoding queue. For each frame, the following information will be obtained: based on the principle of temporal scalability in scalable video coding technology, the layer it is in, the frame it depends on, the number of macroblocks in the frame, and the type of each macroblock. Based on this information, scheduling is performed at two granularities: frame level and macroblock level. The relationship between frame-level scheduling and macroblock-level scheduling is as follows: traverse each frame Fi in the decoding queue D, identify frames suitable for frame-level scheduling, and frames not suitable for frame-level scheduling will be processed using finer-grained scheduling (i.e., macroblock-level scheduling).
[0053] By adopting the above implementation method, the blocking of the decoding queue can be discovered in time and the response can be made more quickly.
[0054] In addition, this implementation requires memory and computational overhead. Memory overhead includes frame level and macroblock level. For frame-level scheduling, the CPU needs to refer to the previous frame to decode the current frame. The required memory space is resolution × 3 × 3 Bytes. Assuming the frame resolution is 1080p, the required memory space size is about 10MB. For macroblock-level scheduling, the CPU needs to refer to the data of several macroblocks to decode the current macroblock, and the required memory space size is about 5KB. Compared with the GB-level memory space of current consumer-grade devices, such memory overhead is completely acceptable. Computational overhead includes decoding pressure estimation model inference and frame feature perception scheduling algorithm. For model inference overhead, the present invention proposes a lightweight neural network that balances accuracy and performance by converting complex delay inference problems into simple binary inference. For scheduling overhead, its algorithm has been simplified. Evaluation results show that the average time taken for inference and scheduling-related calculations per frame does not exceed 500 microseconds, which is significantly lower than the frame processing time (at the 10 millisecond level). Therefore, compared with the time required for frame processing, this part of the overhead is negligible.
Claims
1. A decoding task scheduling method based on frame feature perception, characterized in that: The following steps are involved: (1) Analyze the characteristic values of the received frame through the decoding pressure estimation mechanism to determine whether the VPU decoder is in a busy state; (2) When the prediction result of the decoding pressure estimation mechanism is "H-Latency", mark the current frame and call the frame feature-aware scheduling mechanism; (3) The frame feature-aware scheduling mechanism identifies and selects candidate frames for frame-level scheduling based on the dependencies between frames and the feature values of the frames; (4) For frames that fail to pass frame-level scheduling, identify and select candidate macroblocks in the frame and perform macroblock-level scheduling; (5) Schedule qualified frames or macroblocks to the CPU for parallel decoding processing to reduce tail latency and maintain a stable frame rate.
2. The decoding task scheduling method according to claim 1, characterized in that: The decoding pressure estimation mechanism uses a lightweight neural network model to quickly infer the decoding pressure state as "H-Latency" or "L-Latency" by inputting five characteristic values: frame type, frame size, number of macroblocks, macroblock type, and number of frames in the decoding queue.
3. The decoding task scheduling method according to claim 1, characterized in that: The frame level scheduling mechanism specifically includes: (1) Candidate frame identification: Using an identification method based on inter-frame dependencies, frames in the decoding queue that are not dependent on other frames are selected as candidate frames; (2) Candidate frame selection: By calculating the sum of the decoding time of the candidate frame on the CPU and its decoding time on the VPU, it is determined whether the candidate frame meets the preset conditions.
4. The decoding task scheduling method according to claim 1, characterized in that: The macroblock level scheduling specifically includes: (1) Candidate macroblock identification: A method based on intra-frame and inter-frame macroblock dependencies is used to identify macroblocks that can be independently decoded in parallel. (2) Candidate macroblock selection: Based on the number of CPU and VPU cores and the principle of spatial locality, select the macroblock that meets the maximum parallelism for decoding.
Citation Information
Patent Citations
Video decoding macro-block-grade parallel scheduling method for perceiving calculation complexity
CN105491377A
Visual perception characteristics-combining hierarchical video coding method
US20170085892A1
Cited By
VPU scheduling method and device based on video frame and electronic equipment
CN121173955A