Decoding task scheduling method based on frame feature perception

By dynamically identifying the VPU decoding pressure in real-time communication and using the CPU to perform parallel decoding processing, the problems of tail delay and frame rate drop in high-quality real-time communication are solved, and efficient decoding and stable frame rate are achieved.

CN119966966AInactive Publication Date: 2025-05-09EAST CHINA NORMAL UNIV

Patent Information

Application Number
CN202510163046.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to meet the requirements of high frame rate, high resolution and ultra-low latency simultaneously in high-quality real-time communication, resulting in increased tail delay and decreased frame rate.

Method used

By dynamically identifying the VPU decoding pressure situation, using the lightweight neural network model to perform decoding pressure estimation, flexibly unloading some decoding tasks to the CPU for parallel processing, reducing decoding time and improving tail delay.

Benefits of technology

It effectively reduces the tail delay in real-time communication scenarios, while maintaining the stability of high frame rates, especially in high frame rates and high resolution scenarios, which significantly improves decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966966A_ABST
    Figure CN119966966A_ABST
Patent Text Reader

Abstract

The invention discloses a decoding task scheduling method based on frame feature perception. According to the method, a decoding pressure estimation mechanism and a frame feature perception scheduling mechanism are added in a decoding process. The decoding pressure mechanism adopts a lightweight neural network for flexibly and quickly estimating the decoding pressure. According to the method, each incoming frame is subjected to binary reasoning, and tracked frames are marked by using 'H-Latency' and 'L-Latency'. Only under the condition that 'H-Latency' occurs, a frame feature perception scheduling mechanism can be called to perform decoding task scheduling work. The method comprises the following steps: firstly, analyzing the dependency relationship between frames in a decoding queue and the characteristics of a macro block in each frame, then selecting a part of parallel decoding tasks to be scheduled to a CPU (Central Processing Unit), and delivering the other part of decoding tasks to a video decoder (VPU) for processing according to an original decoding process so as to improve the degree of parallelism, and distributing the decoding tasks at a frame level and a macro block level. According to the method, the tail delay in a real-time communication scene can be effectively reduced, and meanwhile, the stability of the frame rate is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of real-time communication, and in particular relates to a decoding task scheduling method based on frame feature perception. Background Art

[0002] Real-Time Communication (RTC) plays an increasingly important role in today's digital world, especially in video conferencing, cloud gaming, virtual reality (VR), online education, and telemedicine. However, with the popularity of real-time communication applications and the growth of demand, providing high-quality real-time communication experience has become a challenging task.

[0003] High-quality RTC applications have the following characteristics:

[0004] (1) High frame rate: Traditional RTC generally has only 30fps, which puts less pressure on the decoder. High-quality RTC requires 60-120fps, which brings huge challenges to the decoder;

[0005] (2) High resolution: Existing RTC applications default to 720p or even lower, while high-quality RTC should require 1080p or above, or even 2k. The higher the resolution, the longer the decoding time will be;

[0006] (3) Ultra-low latency: For example, video conferencing and cloud gaming both hope to achieve an interaction delay of less than 100 ms.

[0007] The end-to-end delay of RTC applications is mainly composed of video acquisition and encoding at the sender, network transmission, queuing, decoding, and display at the receiver. The increase in transmission delay caused by network fluctuations or the increase in decoding delay caused by decoder performance fluctuations will both lead to an increase in queuing delays, which are important components of the tail delay of RTC applications.

[0008] There are currently two methods to reduce tail latency. The first method is the frame skipping mechanism, which actively manages the frames in the decoding queue at the receiving end, that is, selectively discards some frames without decoding. This method effectively reduces queuing delay. The second method is adaptive frame rate control, which achieves low latency by adaptively coordinating the frame rate at the sending end with the fluctuating network conditions and the decoding queue capacity at the receiving end. However, these solutions will cause a drop in frame rate and cannot meet the high frame rate requirements of high-quality RTC.

[0009] GOP (Group of Pictures) is a sequence of frames organized as a group in a video stream, which determines the prediction relationship between video frames. It is an important part of video coding, and its length and structure directly affect the compression efficiency, decoding performance and application experience of the video. The GOP format commonly used in RTC applications is IPPPP..., that is, the first frame is an I frame, and the subsequent frames are all P frames. Among them, the I frame does not rely on other frames to complete decoding, and the P frame relies on the previous frame to be decoded successfully. Due to such a dependency between frames, for the first method mentioned above, arbitrarily discarding some frames will cause subsequent frames to fail to decode. In order to solve this problem, Scalable Video Coding (SVC) technology is proposed. According to the principle of temporal scalability in SVC, the frames at the top level are not dependent on other frames. These are the frames that can be discarded without causing subsequent frames to fail to decode.

[0010] Video decoding tasks are usually handled by dedicated video processing units (VPUs), such as Qualcomm chips on mobile terminals, Intel chips on PCs (Personal Computers), and NVIDIA chips. According to calculations, when RTC applications are decoding, the usage rate of the VPU is usually high, while the usage rate of the CPU is low. If idle CPU resources can be used to calculate some decoding tasks, the decoding time of a single frame will be effectively reduced, thereby alleviating the problem of decoding queue blocking. This will effectively reduce the tail delay in real-time communication scenarios while maintaining a stable frame rate. Summary of the invention

[0011] In order to overcome the problems existing in the above technologies, the purpose of the present invention is to provide a decoding task scheduling method based on frame feature perception. The present invention dynamically identifies the VPU decoding pressure situation and flexibly unloads part of the decoding tasks to the CPU for parallel processing, thereby reducing the decoding time and improving the tail delay problem.

[0012] The present invention uses a decoding pressure estimation mechanism to determine whether the CPU needs to participate in the decoding task. When the CPU needs to participate in the decoding task, the frame feature perception scheduling mechanism will analyze the frame features in the decoding queue and schedule some decoding calculation tasks reasonably. This will effectively reduce the decoding time of a single frame, thereby alleviating the problem of decoding queue blocking.

[0013] The specific technical solution for achieving the purpose of the present invention is:

[0014] A decoding task scheduling method based on frame feature perception includes the following steps:

[0015] (1) Analyze the characteristic values ​​of the received frame through the decoding pressure estimation mechanism to determine whether the VPU decoder is in a busy state;

[0016] (2) When the prediction result of the decoding pressure estimation mechanism is "H-Latency", mark the current frame and call the frame feature-aware scheduling mechanism;

[0017] (3) The frame feature-aware scheduling mechanism identifies and selects candidate frames for frame-level scheduling based on the dependencies between frames and the feature values ​​of the frames;

[0018] (4) For frames that fail to pass frame-level scheduling, identify and select candidate macroblocks in the frame and perform macroblock-level scheduling;

[0019] (5) Schedule qualified frames or macroblocks to the CPU for parallel decoding processing to reduce tail latency and maintain a stable frame rate.

[0020] Furthermore, the decoding pressure estimation mechanism uses a lightweight neural network model to quickly infer the decoding pressure state as "H-Latency" or "L-Latency" by inputting five characteristic values ​​of frame type, frame size, number of macroblocks, macroblock type and number of frames in the decoding queue.

[0021] Furthermore, the frame level scheduling mechanism specifically includes:

[0022] (1) Candidate frame identification: Using an identification method based on inter-frame dependencies, frames in the decoding queue that are not dependent on other frames are selected as candidate frames;

[0023] (2) Candidate frame selection: By calculating the sum of the decoding time of the candidate frame on the CPU and its decoding time on the VPU, it is determined whether the candidate frame meets the preset conditions.

[0024] Furthermore, the macroblock level scheduling specifically includes:

[0025] (1) Candidate macroblock identification: A method based on intra-frame and inter-frame macroblock dependencies is used to identify macroblocks that can be independently decoded in parallel.

[0026] (2) Candidate macroblock selection: Based on the number of CPU and VPU cores and the principle of spatial locality, select the macroblock that meets the maximum parallelism for decoding.

[0027] The present invention integrates the collaborative work of the decoding pressure estimation mechanism and the frame feature perception scheduling mechanism. Compared with the existing selective frame loss method, the present invention can effectively reduce the tail delay in real-time communication scenarios while maintaining the stability of the frame rate. In particular, for high frame rate and high resolution scenarios, it can effectively improve the decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flow chart of candidate frame identification and selection according to the present invention;

[0029] Figure 2 A flow chart of candidate macroblock identification and selection according to the present invention;

[0030] Figure 3 A schematic diagram for implementing the present invention. DETAILED DESCRIPTION

[0031] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0032] The present invention discloses a decoding task scheduling method based on frame feature perception, and its basic process is as follows: first, the decoding pressure estimation mechanism collects feature value information of the incoming frame and then performs decoding pressure evaluation. If the evaluation result is "H-Latency", the frame feature perception scheduling mechanism is notified. Then, the frames in the decoding queue are analyzed. If the frame meets the conditions of frame-level scheduling, the frame is scheduled to the CPU for decoding. If the frame does not meet the conditions, some macroblocks in the frame are scheduled to the CPU for decoding.

[0033] The decoding pressure estimation mechanism adopts a fast inference model of a lightweight neural network. The model is a fully connected neural network with three layers, including an input layer, a hidden layer and an output layer. It performs inference by inputting five feature values, namely, frame type, frame size, number of macroblocks, macroblock type and number of frames in the decoding queue. The value of the size in the feature is converted into a decimal number and input into the neuron. For example, a frame of 127KB size is formatted as four integers {0,1,2,7}. The same formatting method is also applied to the size of the macroblock. This method can reduce the time cost of the inference stage. It is unnecessary to infer the accurate decoding delay. The present invention simplifies the model by converting the delay inference problem into a simple binary ("busy" or "normal") inference. Specifically, the tracked frames are marked with "H-Latency" and "L-Latency". "H-Latency" indicates that the decoder is in a busy state, while "L-Latency" indicates that the decoder is in a normal state.

[0034] The frame feature-aware scheduling mechanism provides two granularity scheduling schemes, frame-level scheduling and macroblock-level scheduling. Frame-level scheduling is mainly divided into two aspects.

[0035] In a first aspect, a candidate frame identification mechanism is provided, comprising:

[0036] A simple candidate frame identification method based on inter-frame dependency is adopted. The basic idea of ​​this method is to identify frames in the decoding queue that are not dependent on other frames. According to the principle of temporal scalability in scalable video coding technology, the top-level frames are not dependent on other frames. Figure 1 Frames 3 and 7 in the second layer are candidate frames.

[0037] In a second aspect, a candidate frame selection mechanism is provided, including:

[0038] Candidate frames should not be scheduled if they take too long to decode on the CPU. The decoding time on the CPU is ,if Exceeded the frame and subsequent frames The sum of decoding time on VPU , which will increase the delay. Therefore, the candidate frame is only Only then can it be scheduled:

[0039] .

[0040] Macroblock-level scheduling is mainly divided into two aspects.

[0041] In a first aspect, a candidate macroblock identification mechanism is provided, comprising:

[0042] A candidate macroblock identification method based on macroblock dependencies is adopted, including intra-frame macroblocks and inter-frame macroblocks. In an intra-frame macroblock, two independent macroblocks have the same relative position. Starting from the position of the left macroblock, move two positions to the right, and then move one position up to reach the position of the right macroblock. For inter-frame macroblocks, their dependencies span two frames, and the macroblocks in the current frame calculate the position of the dependent macroblock in the previous frame based on their motion vectors as offsets. According to the above two methods, macroblocks that can be processed in parallel between the two frames can be determined, and these are candidate macroblocks.

[0043] In a second aspect, a candidate macroblock selection mechanism is provided, including:

[0044] Using a parallelism-based approach, Figure 2 It shows that during the decoding process, the number of macroblocks that can be processed in parallel at the same time is constantly changing, and the maximum degree of parallelism increases from top to bottom, then reaches a peak state, and finally gradually decreases. The calculation formula for the maximum degree of parallelism is as follows:

[0045]

[0046] in Indicates the number of macroblocks in the width direction of the frame, Indicates the number of macroblocks in the height direction of the frame. After determining the maximum parallelism, the start and end times of macroblock scheduling are selected based on the number of CPU and VPU cores and the principle of spatial locality. Figure 2 For example, it shows that when the number of CPU cores is 2 and the number of VPU cores is 3, the macroblock with dark background is selected.

[0047] like Figure 3 As shown in FIG. 1 , it is an architecture diagram for implementing the present invention. Two new components are added: decoding pressure estimation and frame feature perception scheduling. The implementation is mainly divided into the following steps:

[0048] 1. When the frame receiving end receives a frame of data, the decoding pressure estimation mechanism analyzes the feature value of the frame and inputs the model to predict the decoding pressure.

[0049] 2. If the current frame is "L-Latency", the frame feature perception scheduling will not be awakened, which is consistent with the existing decoding process. Once the current frame is determined to be "H-Latency", it will be marked and the frame feature perception scheduling will be notified.

[0050] 3. Then, start the decoding task scheduling process, find the frames that meet the frame-level scheduling, and then find the macroblocks that meet the macroblock-level scheduling.

[0051] 4. Finally, these decoding tasks that can be processed in parallel are offloaded to the CPU for processing.

[0052] In the above decoding task scheduling process, the scheduling mechanism will analyze the frames in the decoding queue. For each frame, the following information will be obtained: based on the principle of temporal scalability in scalable video coding technology, the layer it is in, the frame it depends on, the number of macroblocks in the frame, and the type of each macroblock. Based on this information, scheduling is performed at two granularities: frame level and macroblock level. The relationship between frame-level scheduling and macroblock-level scheduling is as follows: traverse each frame Fi in the decoding queue D, identify frames suitable for frame-level scheduling, and frames not suitable for frame-level scheduling will be processed using finer-grained scheduling (i.e., macroblock-level scheduling).

[0053] By adopting the above implementation method, the blocking of the decoding queue can be discovered in time and the response can be made more quickly.

[0054] In addition, this implementation requires memory and computational overhead. Memory overhead includes frame level and macroblock level. For frame-level scheduling, the CPU needs to refer to the previous frame to decode the current frame. The required memory space is resolution × 3 × 3 Bytes. Assuming the frame resolution is 1080p, the required memory space size is about 10MB. For macroblock-level scheduling, the CPU needs to refer to the data of several macroblocks to decode the current macroblock, and the required memory space size is about 5KB. Compared with the GB-level memory space of current consumer-grade devices, such memory overhead is completely acceptable. Computational overhead includes decoding pressure estimation model inference and frame feature perception scheduling algorithm. For model inference overhead, the present invention proposes a lightweight neural network that balances accuracy and performance by converting complex delay inference problems into simple binary inference. For scheduling overhead, its algorithm has been simplified. Evaluation results show that the average time taken for inference and scheduling-related calculations per frame does not exceed 500 microseconds, which is significantly lower than the frame processing time (at the 10 millisecond level). Therefore, compared with the time required for frame processing, this part of the overhead is negligible.

Claims

1. A decoding task scheduling method based on frame feature perception, characterized in that: The following steps are involved: (1) Analyze the characteristic values ​​of the received frame through the decoding pressure estimation mechanism to determine whether the VPU decoder is in a busy state; (2) When the prediction result of the decoding pressure estimation mechanism is "H-Latency", mark the current frame and call the frame feature-aware scheduling mechanism; (3) The frame feature-aware scheduling mechanism identifies and selects candidate frames for frame-level scheduling based on the dependencies between frames and the feature values ​​of the frames; (4) For frames that fail to pass frame-level scheduling, identify and select candidate macroblocks in the frame and perform macroblock-level scheduling; (5) Schedule qualified frames or macroblocks to the CPU for parallel decoding processing to reduce tail latency and maintain a stable frame rate.

2. The decoding task scheduling method according to claim 1, characterized in that: The decoding pressure estimation mechanism uses a lightweight neural network model to quickly infer the decoding pressure state as "H-Latency" or "L-Latency" by inputting five characteristic values: frame type, frame size, number of macroblocks, macroblock type, and number of frames in the decoding queue.

3. The decoding task scheduling method according to claim 1, characterized in that: The frame level scheduling mechanism specifically includes: (1) Candidate frame identification: Using an identification method based on inter-frame dependencies, frames in the decoding queue that are not dependent on other frames are selected as candidate frames; (2) Candidate frame selection: By calculating the sum of the decoding time of the candidate frame on the CPU and its decoding time on the VPU, it is determined whether the candidate frame meets the preset conditions.

4. The decoding task scheduling method according to claim 1, characterized in that: The macroblock level scheduling specifically includes: (1) Candidate macroblock identification: A method based on intra-frame and inter-frame macroblock dependencies is used to identify macroblocks that can be independently decoded in parallel. (2) Candidate macroblock selection: Based on the number of CPU and VPU cores and the principle of spatial locality, select the macroblock that meets the maximum parallelism for decoding.

Citation Information

Patent Citations

  • Video decoding macro-block-grade parallel scheduling method for perceiving calculation complexity

    CN105491377A

  • Visual perception characteristics-combining hierarchical video coding method

    US20170085892A1

Cited By

  • VPU scheduling method and device based on video frame and electronic equipment

    CN121173955A